[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118354-en":3,"doc-seo-118354-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118354,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Position - Why We Must Rethink Empirical Research in Machine Learning","This position argues that a common but incomplete understanding of empirical research in machine learning undermines replicability, reliability, and long-term progress. It highlights that much current work is structured as confirmatory research, while the research process should be treated as exploratory. The discussion addresses methodological and epistemic limitations of experimentation and warns that non-replicable findings create practical and epistemological risks, including reduced trust in applied medical use.","Position: Why We Must Rethink Empirical Research in Machine Learning  \nMoritz Herrmann 1 2 F. Julian D. Lange 1 2 Katharina Eggensperger 3 Giuseppe Casalicchio 4 2 Marcel Wever 5 2 Matthias Feurer 4 2 David Rgamer 4 2 Eyke Hllermeier 5 2 Anne-Laure Boulesteix 1 2 Bernd Bischl 4 2  \narXiv :2405 .02200v2 [ cs .LG] 25 May 2024  \nAbstract  \nWe warn against a common but incomplete understanding of empirical research in machine learning that leads to non-replicable results, makes findings unreliable, and threatens to undermine progress in the field. To overcome this alarming situation, we call for more awareness of the plurality of ways of gaining knowledge experimentally but also of some epistemic limitations.  \nIn particular, we argue most current empirical machine learning research is fashioned as confirmatory research while it should rather be considered exploratory.  \n1. The Non-Replicable ML Research Enigma  \nIn his Caltech commencement address “Cargo Cult Science”∗ , 1 Richard Feynman (1974) described how researchers employ practices that conflict with scientific principles to adhere to a certain way of doing things. This position paper warns against similar tendencies in empirical research in machine learning (ML) and calls for a mindset change to address methodological and epistemic challenges of experimentation.  \nThere is ML research that does not replicate. From an empirical scientific perspective, non-replicable research is a fundamental problem. As Karl Popper (1959/2002, p. 66) phrased it: “non-reproducible single occurrences are of no significance to science.”2 Consequently, ML research that  \n1Institute for Medical Information Processing, Biometry, and Epidemiology, Faculty of Medicine, LMU Munich, Munich, Germany 2Munich Center for Machine Learning (MCML), Munich, Germany 3University of T¨ubingen, T¨ubingen, Germany 4Department of Statistics, LMU Munich, Munich, Germany 5Institute of Informatics, LMU Munich, Munich, Germany. Correspondence to: Moritz Herrmann \u003C[moritz.herrmann@lmu.de](moritz.herrmann@lmu.de) > .  \nProceedings of the 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024 . Copyright 2024 by the author(s) .  \n1As our paper contains some jargon, we have included a glossary in the appendix; asterisks (∗) in the text denote covered terms.  \n2Reproducible here does not refer to exact computational reproducibility∗ but generally to arriving at the same scientific conclusions, termed replicability∗ in this paper.  \ndoes not replicate has far-reaching epistemic∗ and practical consequences. From an epistemological∗ point of view, it means that research results are unreliable and, to some extent, it calls into question progress in the field. In practice, it may jeopardize applied empirical researchers’ confidence in experimental results and discourage them from applying ML methods, even though these novel approaches might be beneficial. For example, ML is increasingly being used in the medical domain, and this is often promising in terms of patient benefit. However, there are also examples indicating that applied researchers (are starting to) have concerns about ML being used in this high-stakes area. Consider, for example, this quite drastic warning by Dhiman et al. (2022, p. 2): “Machine learning is often portrayed as offering many advantages [...] . However, these advantages have not yet materialised into patient benefit [...] . Given the increasing concern about the methodological quality and risk of bias of prediction model studies [emphasis added], caution is warranted and the lack of uptake of models in medical practice is not surprising.” That is, if the ML community does not improve rigor in empirical methodological research, we think there may be a risk of a backlash against the use of ML in practice.  \nIn general, there is a growing body of empirical evidence showing that conclusions drawn from experimental results in ML were overly optimistic at the time of publication","cbCaiqo7Dcvo3ddt","https://ap.wps.com/l/cbCaiqo7Dcvo3ddt","pdf",540165,1,20,"English","en",105,"# The Non-Replicable ML Research Enigma\n## From Cargo Cult Science to ML experimentation\n## Consequences of non-replicable findings\n## Evidence of overly optimistic and non-replicated results\n## Questionable research practices and incentives","[{\"question\":\"Why does the document warn that current empirical ML research can lead to non-replicable results?\",\"answer\":\"It argues that many studies rely on an incomplete understanding of empirical research, producing findings that cannot be reliably reproduced and are therefore unreliable.\"},{\"question\":\"What distinction does the document make between confirmatory and exploratory research?\",\"answer\":\"It states that most current empirical machine learning research is treated like confirmatory research, while it should instead be considered exploratory.\"},{\"question\":\"What are the consequences of non-replicable ML research described in the document?\",\"answer\":\"The document describes epistemic and practical consequences: results become unreliable and progress in the field can be questioned, which may also reduce confidence in applied ML outcomes, especially in high-stakes domains like medicine.\"}]","Position - Why We Must Rethink Empirical Research in Machine Learning | PDF",1785683254,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"position-why-we-must-rethink-empirical-research-in-machine-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/position-why-we-must-rethink-empirical-research-in-machine-learning/118354/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why does the document warn that current empirical ML research can lead to non-replicable results?","Question",{"text":76,"@type":77},"It argues that many studies rely on an incomplete understanding of empirical research, producing findings that cannot be reliably reproduced and are therefore unreliable.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What distinction does the document make between confirmatory and exploratory research?",{"text":81,"@type":77},"It states that most current empirical machine learning research is treated like confirmatory research, while it should instead be considered exploratory.",{"name":83,"@type":74,"acceptedAnswer":84},"What are the consequences of non-replicable ML research described in the document?",{"text":85,"@type":77},"The document describes epistemic and practical consequences: results become unreliable and progress in the field can be questioned, which may also reduce confidence in applied ML outcomes, especially in high-stakes domains like medicine.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":29,"slug":114},6,"Technology","technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":21,"slug":126},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":21,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]