[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117239-en":3,"doc-seo-117239-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117239,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Position - Embracing Negative Results in Machine Learning","Publications proposing new machine learning methods are often judged mainly by predictive performance on selected benchmarks, which can distort incentives and waste effort when promising approaches fail to beat existing state of the art. This position paper argues that predictive performance alone is not a reliable indicator of scientific value. It calls for normalizing the publication of negative results and explains why doing so improves the quality and efficiency of the machine learning research community’s scientific output.","Position: Embracing Negative Results in Machine Learning  \nFlorian Karl 1 2 3 Lukas Malte Kemeter 1 Gabriel Dax 1 Paulina Sierak 1  \narXiv :2406 .03980v 1 [ cs .LG] 6 Jun 2024  \nAbstract  \nPublications proposing novel machine learning methods are often primarily rated by exhibited predictive performance on selected problems. In this position paper we argue that predictive performance alone is not a good indicator for the worth of a publication. Using it as such even fosters problems like inef􀀂ciencies of the machine learning research community as a whole and setting wrong incentives for researchers. We therefore put out a call for the publication of “negative” results, which can help alleviate some of these problems and improve the scienti􀀂c output of the machine learning research community. To substantiate our position, we present the advantages of publishing negative results and provide concrete measures for the community to move towards a paradigm where their publication is normalized.  \nNote from the authors: Have some of our publications been rejected due to lacking competitive results and has this been frustrating at times? Yes. However, the following position paper is not a personal vendetta: we truly believe embracing negative results can be an asset for the machine learning research community and want to present an objective deliberation on why. We hope to convince you, the reader, of the same in the following pages and spark discussion as well as change in our community.  \n1. Introduction  \nMachine learning has grown into a prominent research 􀀂eld that has demonstrated large impact on a lot of application domains. The number of machine learning publications has grown exponentially along with the number of  \n1Fraunhofer Institute for Integrated Circuits IIS, Fraunhofer IIS, Nuremberg, Germany 2Ludwig-Maximilians-Universit¨at M¨unchen, Munich, Germany 3Munich Center for Machine Learning, Munich, Germany. Correspondence to: Florian Karl \u003C􀀃o[rian.karl@iis.fraunhofer.de](rian.karl@iis.fraunhofer.de) > .  \nProceedings of the 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024 . Copyright 2024 by the author(s) .  \nactive researchers and funding volume (Maslej et al., 2023; Krenn et al., 2023) . There are many machine learning publications that provide value for the research community: works centered around theory and proofs, benchmarks, survey papers and position papers. However, a large number of machine learning publications examine a (often novel) method and then demonstrate its performance on relevant problems; these are the types of publications we focus onin this work.  \nMachine learning is largely an empirical science: If something works and demonstrates good performance it is often deemed a good result and worthy of publication. On the other hand, if a new method or algorithm is not able to beat the state-of-the-art on a typical benchmark dataset, researchers might quickly abandon their work as it is unlikely to be published. Despite being a somewhat confusing term when it comes to scienti􀀂c results, such outcomes are often deemed to be negative results. In science, the terms positive and negative refer to the postulated null hypothesis, which is then either rejected (positive result) or the results of experimentation do not allow for a rejection (negative result) .  \nDe􀀂nition 1.1. The usual null hypothesis of empirical machine learning is that a proposed method does not exhibit signi􀀂cantly better predictive performance than existing methods on a relevant subset of problems.  \nIn this terminology there is no room for “good” or “bad”results. However, machine learning as an empirical science has developed a strong attachment to predictive performance and it often seems that only very speci􀀂c positive results, those that show that a proposed method beats the state-of-the-art, are considered “good” results. In the context of this paper, when talking about negative results we refer to th","cbCaieiXMNrKv07K","https://ap.wps.com/l/cbCaieiXMNrKv07K","pdf",183743,1,11,"English","en",105,"# Introduction\n## What are negative results in empirical machine learning?\n## Subtypes: novel method negative results and existing method negative results\n## Why embracing negative results matters","[{\"question\":\"Why does the paper argue that predictive performance alone is insufficient?\",\"answer\":\"It states that rating publications only by predictive performance creates inefficiencies and wrong incentives for researchers, and can misrepresent the true scientific worth of a work.\"},{\"question\":\"How does the paper define a negative result in machine learning research?\",\"answer\":\"A negative result occurs when the usual null hypothesis cannot be rejected, meaning the proposed method does not show significantly better predictive performance than existing approaches on relevant problems.\"},{\"question\":\"What are the two subtypes of negative results discussed?\",\"answer\":\"The paper distinguishes novel method negative results (a new method not beating state of the art) from existing method negative results (state-of-the-art methods performing worse than expected, e.g., replication or failure-mode analyses).\"}]","Position - Embracing Negative Results in Machine Learning | PDF",1785674620,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"position-embracing-negative-results-in-machine-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/position-embracing-negative-results-in-machine-learning/117239/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why does the paper argue that predictive performance alone is insufficient?","Question",{"text":76,"@type":77},"It states that rating publications only by predictive performance creates inefficiencies and wrong incentives for researchers, and can misrepresent the true scientific worth of a work.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the paper define a negative result in machine learning research?",{"text":81,"@type":77},"A negative result occurs when the usual null hypothesis cannot be rejected, meaning the proposed method does not show significantly better predictive performance than existing approaches on relevant problems.",{"name":83,"@type":74,"acceptedAnswer":84},"What are the two subtypes of negative results discussed?",{"text":85,"@type":77},"The paper distinguishes novel method negative results (a new method not beating state of the art) from existing method negative results (state-of-the-art methods performing worse than expected, e.g., replication or failure-mode analyses).","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]