[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118843-en":3,"doc-seo-118843-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118843,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Human-in-the-Loop Mixup","Aligning model representations to humans improves robustness and generalization, yet existing approaches often rely on standard observational data and assume synthetic labels are perceptually aligned. This work examines mixup synthetic data, using its controlled mixing mechanism as a testbed for perceptual–label alignment. The authors design the HILL MixE Suite interfaces, collect 159 participants’ perceptual judgments with uncertainty, and find inconsistent alignment with traditional synthetic labels. These results support reliability gains for downstream models, especially when incorporating human uncertainty, and release judgments via the H-Mix data hub.","Human-in-the-Loop Mixup  \nKatherine M. Collins* 1 Umang Bhatt 1,2 Weiyang Liu 1,3 Vihari Piratla 1 Ilia Sucholutsky4 Bradley Love2,5  \nAdrian Weller 1,2  \n1University of Cambridge  \n2The Alan Turing Institute  \n3Max Planck Institute for Intelligent Systems  \n4Princeton University  \n5University College London  \nAbstract  \nAligning model representations to humans has been found to improve robustness and generalization. However, such methods often focus on standard observational data. Synthetic data is proliferating and powering many advances in machine learning; yet, it is not always clear whether synthetic labels are perceptually aligned to humans – rendering it likely model representations are not human aligned. We focus on the synthetic data used in mixup: a powerful regularizer shown to improve model robustness, generalization, and calibration. We design a comprehensive series of elicitation interfaces, which we release as HILL MixE Suite, and recruit 159 participants to provide perceptual judgments along with their uncertainties, over mixup examples. We find that human perceptions do not consistently align with the labels traditionally used for synthetic points, and begin to demonstrate the applicability of these findings to potentially increase the reliability of downstream models, particularly when incorporating human uncertainty. We release all elicited judgmentsin a new data hub we call H-Mix.  \n1 INTRODUCTION  \nSynthetic data is proliferating, fueled by increasingly powerful generative models, e.g. [Goodfellow et al., 2014a, Dhariwal and Nichol, 2021] . These data are not only consumed directly by people – but, as training predictive models on synthetic data has been found to unlock tremendous advances in machine learning (ML) [Silver et al., 2016, de Melo et al., 2022, Emam et al., 2021, Jordon et al., 2022], synthetic data is increasingly employed to train algorithms serving as engines of many applications humans may in-  \n* [Correspondence to: kmc61@cam.ac.uk](Correspondence to: kmc61@cam.ac.uk)  \nmixup Data  \n{0.1,0.9} λf = 0.1 {0.5,0.5} λf = 0.5 {0.7,0.3} λf = 0.7  \nOriginal Data  \n{1,0} {0,1}  \n(a) How mixup data is constructed  \nAnswer Pool  \nconstruct  \nWhich mixup data best matches the label {0.5,0.5}?  \n(b) Elicitation Setting I: Endorse a synthetic image to match a label  \nconstruct  \nWhat is the label that best matches this mixup data?  \n| \u003Cbr>\u003Cbr>\u003Cbr>Annotator | infer | λf \u003Cbr>Mixing Coefficient\u003Cbr>Confidence |\n| --- | --- | --- |\n\n(c) Elicitation Setting II: Infer the label for a synthetic image  \nFigure 1: Framework overview. A) Synthetic data generating process used in mixup; B) and C) depict elicitation settings: B) participants endorse a synthetic image to match a label, C) participants infer the label for a synthetic image and provide their uncertainty in the corresponding inference.  \nteract with. However, it is not always clear whether human perceptual judgments of synthetically-generated data match the generative process used to create them.  \nAligning networks to match humans’ perceptual inferences could be a way to further ensure model reliability, trustworthiness, downstream performance, and robustness [Nanda et al., 2021, Chen et al., 2022, Fel et al., 2022, Sucholutsky and Griffiths, 2023] . If these data are not aligned with human percepts, then performance potentially could be improved by altering such signals to better match the richness of human judgments: this has proven effective when aligning  \nProceedings of the 39th Conference on Uncertainty in Artificial Intelligence (UAI 2023), PMLR 216:454–464 .  \nmodels with human probabilistic knowledge and perceptual uncertainty [Collins et al., 2022a, Sanders et al., 2022, Sucholutsky et al., 2023] . We argue that one ought to verify whether synthetic data aligns with human perception, and if not, explore whether training with human-relabeled examples improves model performance.  \nIn this work, we take a step in this direction by focusing on m","cbCaiiv4fIIRcjN0","https://ap.wps.com/l/cbCaiiv4fIIRcjN0","pdf",2171018,1,11,"English","en",105,"# Abstract\n# Introduction\n## Motivation: synthetic data and alignment\n## Why focus on mixup\n## Goal and elicitation approach","[{\"question\":\"Why does the paper investigate perceptual alignment for synthetic labels?\",\"answer\":\"It addresses uncertainty about whether human perceptual judgments match the synthetic generative process, since misalignment can undermine reliability and performance improvements expected from synthetic-data training.\"},{\"question\":\"Why is mixup chosen as the focus of the study?\",\"answer\":\"Mixup has a simple, controlled generative mechanism via an explicit mixing coefficient, is widely used and effective as a regularizer, and its linear category boundaries contrast with known nonlinear human perceptual warping.\"},{\"question\":\"What does the study contribute in terms of data collection and tools?\",\"answer\":\"It introduces HILL MixE Suite to elicit perceptual judgments with uncertainties from 159 participants and releases all elicited judgments in a new data hub called H-Mix.\"}]","Human-in-the-Loop Mixup | PDF",1785720576,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"human-in-the-loop-mixup","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/human-in-the-loop-mixup/118843/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does the paper investigate perceptual alignment for synthetic labels?","Question",{"text":75,"@type":76},"It addresses uncertainty about whether human perceptual judgments match the synthetic generative process, since misalignment can undermine reliability and performance improvements expected from synthetic-data training.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is mixup chosen as the focus of the study?",{"text":80,"@type":76},"Mixup has a simple, controlled generative mechanism via an explicit mixing coefficient, is widely used and effective as a regularizer, and its linear category boundaries contrast with known nonlinear human perceptual warping.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the study contribute in terms of data collection and tools?",{"text":84,"@type":76},"It introduces HILL MixE Suite to elicit perceptual judgments with uncertainties from 159 participants and releases all elicited judgments in a new data hub called H-Mix.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]