[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121687-en":3,"doc-seo-121687-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121687,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","Designing Data - Proactive Data Collection and Iteration for Machine Learning","Lack of diversity in data collection leads to significant failures in machine learning applications, while post-collection interventions are time intensive and seldom comprehensive. Designing data introduces an iterative, bias-mitigating approach that connects HCI concepts with ML techniques to help teams evaluate whether datasets reflect real-world variability. The method covers three stages: pre-collection planning, collection monitoring, and data familiarity via out-of-distribution methods. An applied case study shows improved generalization across intersectional groups and supports effective dataset debugging.","Designing Data: Proactive Data Collection and Iteration for  \nMachine Learning  \nAspen Hopkins∗ [dataspen@mit.edu](dataspen@mit.edu)[ ](dataspen@mit.edu)Massachussetts Institute of Technology Cambridge, MA, USA  \nFred Hohman  \n[fredhohman@apple.com](fredhohman@apple.com)[ ](fredhohman@apple.com)Apple Seattle, WA, USA  \nLuca Zappella  \n[lzappella@apple.com](lzappella@apple.com)[ ](lzappella@apple.com)Apple Barcelona, Spain  \narXiv :2301 . 10319v1 [ cs .HC] 24 Jan 2023  \nXavier Suau Cuadros  \n[xsuaucuadros@apple.com](xsuaucuadros@apple.com)[ ](xsuaucuadros@apple.com)Apple Barcelona, Spain  \nDominik Moritz  \n[domoritz@apple.com](domoritz@apple.com)[ ](domoritz@apple.com)Apple Pittsburgh, PA, USA  \nFigure 1: Designing data process compared to conventional machine learning development. Reflexivity ensures appropriate consideration of positionality and expectations is given before prior to collection. Tracking provides insight into unexpected trends during collection. Familiarity facilitates debugging and highlights potentially noisy or underrepresented subpopulations to direct iteration. This figure represents a simplification of the data collection process. The results of Familiarity can also be incorporated into training (after cleaning the dataset) and tracking (after training an initial model).  \nABSTRACT  \nLack of diversity in data collection has caused significant failuresin machine learning (ML) applications. While ML developers perform post-collection interventions, these are time intensive and rarely comprehensive. Thus, new methods to track & manage data collection, iteration, and model training are necessary for evaluating whether datasets reflect real world variability. We present designing data, an iterative, bias mitigating approach to data collection connecting HCI concepts with ML techniques. Our process includes (1) Pre-Collection Planning, to reflexively prompt and document expected data distributions; (2) Collection Monitoring, to systematically encourage sampling diversity; and (3) Data Familiarity, to identify samples that are unfamiliar to a model through Out-of-Distribution (OOD) methods. We instantiate designing data through our own data collection and applied ML case study. We find models trained on “designed” datasets generalize better across intersectional groups than those trained on similarly sized but less  \n∗ Work done at Apple.  \ntargeted datasets, and that data familiarity is effective for debugging datasets.  \n1 INTRODUCTION  \nCurating representative training and testing datasets is fundamental to developing robust, generalizable machine learning (ML) models. However, understanding what is representative for a specific task is an iterative process. ML practitioners need to change data, models, and their associated processes as they become more familiar with their modeling task, as the state of the world evolves, and as ML products are updated or maintained. Iteration directed by this evolving understanding seeks to improve model performance, often editing datasets to ensure desired outcomes. Failure to effectively recognize data quality and coverage needs can lead to biased ML models [53] . Such failures are responsible for the perpetuation—even exacerbation—of systemic power and access differentials and the deployment of inaccessible or defective product experiences. Yet building representative datasets is an arduous undertaking [68, 78] that relies on the efficacy of human-specified data requirements.  \nTo ensure a dataset covers all, or as many, characteristics as possible, specifications must be the result of a comprehensive enumeration of possible categories—an open and hard problem that few have practically grappled with in the context of ML. Further contributing to this difficulty is the realization that it is not enough for the training datasets to be aligned with expected distributions: they must also include enough examples from conceptually harder or less common categories if said catego","cbCais355xeUYnmi","https://ap.wps.com/l/cbCais355xeUYnmi","pdf",8163916,1,18,"English","en",105,"# Introduction\n## Designing data: an iterative approach\n## Pre-Collection Planning\n## Collection Monitoring\n## Data Familiarity via OOD","[{\"question\":\"Why does lack of diversity in data collection cause machine learning failures?\",\"answer\":\"Because datasets may not reflect real-world variability, leading to biased models and inaccessible or defective product experiences. Post-collection fixes are often too slow and rarely comprehensive.\"},{\"question\":\"What are the three stages of the designing data approach?\",\"answer\":\"Pre-Collection Planning prepares and documents expected data distributions, Collection Monitoring encourages sampling diversity, and Data Familiarity identifies unfamiliar samples using out-of-distribution (OOD) methods.\"},{\"question\":\"How does designing data improve model performance across groups?\",\"answer\":\"Models trained on designed datasets generalize better across intersectional groups than models trained on similarly sized but less targeted datasets.\"}]","Designing Data - Proactive Data Collection and Iteration for Machine Learning | PDF",1785806245,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"designing-data-proactive-data-collection-and-iteration-for-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/designing-data-proactive-data-collection-and-iteration-for-machine-learning/121687/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does lack of diversity in data collection cause machine learning failures?","Question",{"text":75,"@type":76},"Because datasets may not reflect real-world variability, leading to biased models and inaccessible or defective product experiences. Post-collection fixes are often too slow and rarely comprehensive.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the three stages of the designing data approach?",{"text":80,"@type":76},"Pre-Collection Planning prepares and documents expected data distributions, Collection Monitoring encourages sampling diversity, and Data Familiarity identifies unfamiliar samples using out-of-distribution (OOD) methods.",{"name":82,"@type":73,"acceptedAnswer":83},"How does designing data improve model performance across groups?",{"text":84,"@type":76},"Models trained on designed datasets generalize better across intersectional groups than models trained on similarly sized but less targeted datasets.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]