[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117429-en":3,"doc-seo-117429-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117429,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Position: Why We Must Rethink Empirical Research in Machine Learning","This position paper warns that a common but incomplete understanding of empirical research in machine learning can produce results that cannot be replicated, making findings unreliable and slowing credible progress. The authors argue that much current ML empirical work is presented as confirmatory research, although it should be treated as exploratory. They also emphasize epistemic and practical limitations of experimentation, linking weak rigor to reduced confidence in applied settings, including high-stakes medical use.","Position: Why We Must Rethink Empirical Research in Machine Learning  \nMoritz Herrmann 1 2 F. Julian D. Lange 1 2 Katharina Eggensperger 3 Giuseppe Casalicchio 4 2 Marcel Wever 5 2 Matthias Feurer 4 2 David Rgamer 4 2 Eyke Hllermeier 5 2 Anne-Laure Boulesteix 1 2 Bernd Bischl 4 2  \nAbstract  \nWe warn against a common but incomplete understanding of empirical research in machine learning that leads to non-replicable results, makes findings unreliable, and threatens to undermine progress in the field. To overcome this alarming situation, we call for more awareness of the plurality of ways of gaining knowledge experimentally but also of some epistemic limitations.  \nIn particular, we argue most current empirical machine learning research is fashioned as confirmatory research while it should rather be considered exploratory.  \n1. The Non-Replicable ML Research Enigma  \nIn his Caltech commencement address “Cargo Cult Science”∗ , 1 Richard Feynman (1974) described how researchers employ practices that conflict with scientific principles to adhere to a certain way of doing things. This position paper warns against similar tendencies in empirical research in machine learning (ML) and calls for a mindset change to address methodological and epistemic challenges of experimentation.  \nThere is ML research that does not replicate. From an empirical scientific perspective, non-replicable research is a fundamental problem. As Karl Popper (1959/2002, p. 66) phrased it: “non-reproducible single occurrences are of no significance to science.”2 Consequently, ML research that  \n1Institute for Medical Information Processing, Biometry, and Epidemiology, Faculty of Medicine, LMU Munich, Munich, Germany 2Munich Center for Machine Learning (MCML), Munich, Germany 3University of T¨ubingen, T¨ubingen, Germany 4Department of Statistics, LMU Munich, Munich, Germany 5Institute of Informatics, LMU Munich, Munich, Germany. Correspondence to: Moritz Herrmann \u003C[moritz.herrmann@lmu.de](moritz.herrmann@lmu.de) > .  \nProceedings of the 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024 . Copyright 2024 by the author(s) .  \n1As our paper contains some jargon, we have included a glossary in the appendix; asterisks (∗) in the text denote covered terms.  \n2Reproducible here does not refer to exact computational reproducibility∗ but generally to arriving at the same scientific conclusions, termed replicability∗ in this paper.  \ndoes not replicate has far-reaching epistemic∗ and practical consequences. From an epistemological∗ point of view, it means that research results are unreliable and, to some extent, it calls into question progress in the field. In practice, it may jeopardize applied empirical researchers’ confidence in experimental results and discourage them from applying ML methods, even though these novel approaches might be beneficial. For example, ML is increasingly being used in the medical domain, and this is often promising in terms of patient benefit. However, there are also examples indicating that applied researchers (are starting to) have concerns about ML being used in this high-stakes area. Consider, for example, this quite drastic warning by Dhiman et al. (2022, p. 2): “Machine learning is often portrayed as offering many advantages [...] . However, these advantages have not yet materialised into patient benefit [...] . Given the increasing concern about the methodological quality and risk of bias of prediction model studies [emphasis added], caution is warranted and the lack of uptake of models in medical practice is not surprising.” That is, if the ML community does not improve rigor in empirical methodological research, we think there may be a risk of a backlash against the use of ML in practice.  \nIn general, there is a growing body of empirical evidence showing that conclusions drawn from experimental results in ML were overly optimistic at the time of publication and could not be replicated in subsequent st","cbCairxdoKgvCRcz","https://ap.wps.com/l/cbCairxdoKgvCRcz","pdf",334880,1,20,"English","en",105,"# The Non-Replicable ML Research Enigma\n## Cargo Cult Science and methodological mindset change\n## Consequences for epistemic reliability and field progress","[{\"question\":\"Why does the paper warn against non-replicable machine learning research?\",\"answer\":\"Non-replicable results are treated as a fundamental scientific problem, leading to unreliable conclusions and questioning progress. The paper highlights both epistemic and practical impacts, including reduced confidence in experimental findings.\"},{\"question\":\"What research posture does the paper argue current empirical ML work follows?\",\"answer\":\"The authors argue that much of current empirical ML research is shaped as confirmatory research. They contend it should instead be considered exploratory.\"},{\"question\":\"How can poor empirical rigor affect real-world adoption of ML, especially in medicine?\",\"answer\":\"Weak methodological quality and risk of bias can trigger caution and discourage the uptake of ML models. The paper notes growing concerns in high-stakes medical contexts, even when patient benefit seems promising in principle.\"}]","Position: Why We Must Rethink Empirical Research in Machine Learning | PDF",1785675830,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"position-why-we-must-rethink-empirical-research-in-machine-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/position-why-we-must-rethink-empirical-research-in-machine-learning/117429/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why does the paper warn against non-replicable machine learning research?","Question",{"text":76,"@type":77},"Non-replicable results are treated as a fundamental scientific problem, leading to unreliable conclusions and questioning progress. The paper highlights both epistemic and practical impacts, including reduced confidence in experimental findings.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What research posture does the paper argue current empirical ML work follows?",{"text":81,"@type":77},"The authors argue that much of current empirical ML research is shaped as confirmatory research. They contend it should instead be considered exploratory.",{"name":83,"@type":74,"acceptedAnswer":84},"How can poor empirical rigor affect real-world adoption of ML, especially in medicine?",{"text":85,"@type":77},"Weak methodological quality and risk of bias can trigger caution and discourage the uptake of ML models. The paper notes growing concerns in high-stakes medical contexts, even when patient benefit seems promising in principle.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":29,"slug":114},6,"Technology","technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":21,"slug":126},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":21,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]