[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127198-en":3,"doc-seo-127198-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},127198,549768702563,"Sage","https://ap-avatar.wpscdn.com/avatar/8000c4aa63b76e948b?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786536092046926083",6,"Technology","Preventing Harmful Data Practices by using Participatory Input to Navigate the Machine Learning Multiverse","In light of inherent trade-offs regarding fairness, privacy, interpretability and performance, as well as normative questions, the machine learning (ML) pipeline needs to be made accessible for public input, critical reflection and engagement of diverse stakeholders. This work introduces a participatory approach to gather input from the general public on ML pipeline design. Public input is used to navigate and constrain multiverse decisions during model development and evaluation, democratizing key design choices to better reflect downstream system impacts and combat lazy data practices through iterative implementation on a citizen science platform.","Preventing Harmful Data Practices by using Participatory Input to Navigate the Machine Learning Multiverse  \nJan Simson  \nDepartment of Statistics LMU Munich Munich, Germany  \nMunich Center for Machine Learning (MCML) Munich, Germany [jan.simson@lmu.de](jan.simson@lmu.de)  \nFiona Draxler  \nUniversity of Mannheim Mannheim, Germany [fiona.draxler@uni-mannheim.de](fiona.draxler@uni-mannheim.de)  \nSamuel Mehr  \nSchool of Psychology University of Auckland Auckland, New Zealand Child Study Center Yale University  \nNew Haven, Connecticut, USA [sam@auckland.ac.nz](sam@auckland.ac.nz)  \nChristoph Kern  \nDepartment of Statistics LMU Munich Munich, Germany  \nMunich Center for Machine Learning (MCML) Munich, Germany University of Mannheim Mannheim, Germany [christoph.kern@stat.uni-muenchen.de](christoph.kern@stat.uni-muenchen.de)  \nAbstract  \nIn light of inherent trade-offs regarding fairness, privacy, interpretability and performance, as well as normative questions, the machine learning (ML) pipeline needs to be made accessible for public input, critical reflection and engagement of diverse stakeholders.  \nIn this work, we introduce a participatory approach to gather input from the general public on the design of an ML pipeline. We show how people’s input can be used to navigate and constrainthe multiverse of decisions during both model development and evaluation. We highlight that central design decisions should be democratized rather than “optimized” to acknowledge their critical impact on the system’s output downstream. We describe the iterative development of our approach and its exemplary implementation on a citizen science platform. Our results demonstrate how public participation can inform critical design decisions along the model-building pipeline and combat widespread lazy data practices.  \nCCS Concepts  \n• Computing methodologies → Machine learning; • Humancentered computing → Collaborative and social computing; Human computer interaction (HCI); • Social and professional topics → User characteristics.  \nKeywords  \nParticipatory Design, Machine Learning, Algorithmic Fairness, Multiverse Analysis, Citizen Science, Garden of Forking Paths  \nThis work is licensed under a Creative Commons Attribution 4.0 International License. CHI’25, Yokohama, Japan  \n© 2025 Copyright held by the owner/author(s) .  \nACM ISBN 979-8-4007-1394-1/25/04  \n[https://doi.org/10.1145/3706598.3713482](https://doi.org/10.1145/3706598.3713482)  \nACM Reference Format:  \nJan Simson, Fiona Draxler, Samuel Mehr, and Christoph Kern. 2025. Preventing Harmful Data Practices by using Participatory Input to Navigate the Machine Learning Multiverse. In CHI Conference on Human Factors in Computing Systems (CHI’25), April 26–May 01, 2025, Yokohama, Japan. ACM, New York, NY, USA, 30 pages. [https://doi.org/10.1145/3706598.3713482](https://doi.org/10.1145/3706598.3713482)  \n1 Introduction  \nAlgorithmic decision-making (ADM) fueled by machine learning (ML) algorithms is becoming ubiquitous in many domains, affecting the lives of millions of individuals. Examples include jobseekers that are classified into different risk groups by profiling models [74], refugees that are re-allocated within their host country based on matching algorithms [7] and the denial or approval of healthcare coverage for patients [6, 89, 126] . While such systems are introduced with the aim of improving the effectiveness and efficiency of decision-making, there are also serious concerns that algorithmic decisions can treat the affected individuals unfairly [88, 89, 93] . Fairness implications ofADM ultimately depend on how the underlying models interact with biases and deficits in training data, and thus the design, implementation and evaluation of the ML system is of central concern [17, 117]. Addressing fairness and adverse impacts, therefore, does not only include technical measures but rather needs a broader public discourse where developers, stakeholders and affected individuals meet on an equal ","cbCainIf7GaBQXnp","https://ap.wps.com/l/cbCainIf7GaBQXnp","pdf",1789439,3,1,30,"English","en",105,"# Introduction\n## Public engagement for fair and responsible ML\n## ML multiverse as a garden of forking paths","[{\"question\":\"为什么需要将公共参与引入机器学习（ML）流程？\",\"answer\":\"由于公平、隐私、可解释性与性能之间存在权衡，以及还涉及规范性问题，ML 流程需要接受公众输入与批判性反思，以便多方利益相关者在其部署语境中平等参与设计与评估。\"},{\"question\":\"文中如何利用公众输入来处理“ML multiverse”？\",\"answer\":\"作者提出一种参与式方法，让公众对 ML 管线设计进行输入；这些输入用于在模型开发与评估阶段导航并约束决策多元体，从而影响后续模型选择空间。\"},{\"question\":\"作者如何看待对关键设计决策进行“优化”？\",\"answer\":\"文中强调中心设计决策应当实现民主化而不是仅通过训练数据“优化”，因为其会对系统输出产生关键的下游影响，并且应在多种规范性取舍中进行实质性权衡。\"}]","Preventing Harmful Data Practices by using Participatory Input to Navigate the Machine Learning Multiverse | PDF",1785937454,76,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"preventing-harmful-data-practices-by-using-participatory-input-to-navigate-the-machine-learning-multiverse","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":20},"https://docshare.wps.com/document/technology/",{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/preventing-harmful-data-practices-by-using-participatory-input-to-navigate-the-machine-learning-multiverse/127198/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"为什么需要将公共参与引入机器学习（ML）流程？","Question",{"text":76,"@type":77},"由于公平、隐私、可解释性与性能之间存在权衡，以及还涉及规范性问题，ML 流程需要接受公众输入与批判性反思，以便多方利益相关者在其部署语境中平等参与设计与评估。","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"文中如何利用公众输入来处理“ML multiverse”？",{"text":81,"@type":77},"作者提出一种参与式方法，让公众对 ML 管线设计进行输入；这些输入用于在模型开发与评估阶段导航并约束决策多元体，从而影响后续模型选择空间。",{"name":83,"@type":74,"acceptedAnswer":84},"作者如何看待对关键设计决策进行“优化”？",{"text":85,"@type":77},"文中强调中心设计决策应当实现民主化而不是仅通过训练数据“优化”，因为其会对系统输出产生关键的下游影响，并且应在多种规范性取舍中进行实质性权衡。","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,114,119,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":112,"slug":113},50,"technology",{"id":115,"doc_module":4,"doc_module_name":47,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":120,"doc_module":4,"doc_module_name":47,"category_name":121,"show_sort_weight":22,"slug":122},8,"Research & Report","research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]