[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123005-en":3,"doc-seo-123005-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123005,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Why Machine Learning Models Fail - A Benchmarking Perspective - Dissertation","Machine learning progress has boosted performance in object recognition, language understanding, and related capabilities, driven by models trained directly from data and benchmarks that quantify their performance. Models may appear successful on benchmarks yet fail unpredictably in real-world settings when minor distribution changes like noise, rain, or backgrounds occur. This dissertation links benchmark performance to the desired capability, showing that benchmark design strongly shapes what is measured and that good benchmark scores do not guarantee acquired capability. It studies one-shot object detection, analyzes shortcut learning, and argues for broad evaluation and capability verification.","Why Machine Learning Models Fail: A Benchmarking Perspective  \nDissertation  \nder Mathematisch-Naturwissenschaftlichen Fakult¨at der Eberhard Karls Universit¨at T¨ubingen zur Erlangung des Grades eines Doktors der Naturwissenschaften  \n(Dr. rer. nat.)  \nvorgelegt von  \nClaudio Michaelis  \naus M¨unchen  \nT¨ubingen  \nGedruckt mit Genehmigung der Mathematisch-Naturwissenschaftlichen Fakult¨at der Eberhard Karls Universit¨at T¨ubingen.  \nTag der m¨undlichen Qualiﬁkation: 19.12.2023  \nDekan: Prof. Dr. Thilo Stehle  \n1. Berichterstatter: Prof. Dr. Matthias Bethge  \n2. Berichterstatter: Prof. Dr. Fabian Sinz  \nIch erkl¨are, dass ich die zur Promotion eingereichte Arbeit mit dem Titel: Why Machine Learning Models Fail: A Benchmarking Perspective  \nselbstst¨andig verfasst, nur die angegebenen Quellen und Hilfsmittel benutzt und w¨ortlich oder inhaltlich ¨ubernommene Stellen als solche gekennzeichnet habe. Ich versichere an Eides statt, dass diese Angaben wahr sind und dass ich nichts verschwiegen habe. Mir ist bekannt, dass die falsche Abgabe einer Versicherung an Eides statt mit Freiheitsstrafe bis zu drei Jahren oder mit Geldstrafe bestraft wird.  \nT¨ubingen, den      \nDatum Unterschrift  \nSummary  \nOver the last years, machine performance at object recognition, language understanding and other capabilities that we associate with human intelligence has rapidly improved. One central element of this progress are machine learning models that learn the solution for a task directly from data. The other are benchmarks that use data to quantitatively measure model performance. In combination, they form a virtuous cycle where models can be optimized directly on benchmark performance. But while the resulting models perform very well on their benchmarks, they often fail unexpectedly outside the controlled setting. Innocuous changes such as image noise, rain or the wrong background can lead to wrong predictions. In this dissertation, I argue that to understand these failures, it is necessary to understand the relationship between benchmark performance and the desired capability. To support this argument, I study benchmarks in two ways.  \nIn the ﬁrst part, I investigate how to learn and evaluate a new capability. Therefore, I introduce one-shot object detection and deﬁne di↵erent benchmarks to analyze what makes this task hard for machine learning models and what is needed to solve it. I ﬁnd that CNNs struggle to separate individual objects in cluttered environments, and that one-shot recognition of objects from novel categories can be challenging with real-world objects. I then continue to investigate what makes one-shot generalization diﬃcult in real-world scenes, and identify the number of categories in the training dataset as the central factor. Using this insight, I show that excellent one-shot generalization can be achieved by training on broader datasets. These results highlight how much benchmark design inﬂuences what is measured, and that limitations in benchmarks can be confused for limitations of the models developed with them.  \nIn the second part, I broaden the view and analyze the connection between model failuresin di↵erent areas of machine learning. I ﬁnd that many of these failures can be explained by shortcut learning, models exploiting a mismatch between a benchmark and its associated capability. Shortcut solutions use superﬁcial cues that work very well within the training domain, but are unrelated to the capability. This demonstrates that good benchmarks performance is not suﬃcient to prove that a model acquired the associated capability, and that results have to be interpreted carefully.  \nTaken together, these ﬁndings put in question the common practice of evaluating modelson a single, or at maximum a few, benchmarks. Rather, my results indicate that to anticipate model failures, it is essential to measure broadly. And to avoid them, it is necessary to verify that models acquire the desired capability. This will require in","cbCailQR3NpvH3fM","https://ap.wps.com/l/cbCailQR3NpvH3fM","pdf",13999248,1,147,"English","en",105,"# Summary\n## Relationship between benchmark performance and desired capability\n## Learning and evaluating new capabilities via one-shot object detection\n## Shortcut learning and cautious interpretation of benchmark success\n## Need for broad measurement and verification of acquired capability","[{\"question\":\"Why do machine learning models often fail outside the benchmark setting?\",\"answer\":\"Minor, seemingly harmless changes such as image noise, rain, or an incorrect background can shift inputs away from the controlled benchmark distribution, leading to wrong predictions.\"},{\"question\":\"How does the dissertation connect benchmark performance to the desired capability?\",\"answer\":\"It argues that understanding model failures requires understanding how benchmark performance relates to the underlying capability being targeted, since benchmark design can bias what is measured.\"},{\"question\":\"What is shortcut learning and how does it affect benchmark results?\",\"answer\":\"Shortcut learning occurs when models exploit superficial cues that work well on the benchmark’s training domain but are unrelated to the actual desired capability. This means good benchmark performance is not sufficient evidence that the capability was learned.\"}]","Why Machine Learning Models Fail - A Benchmarking Perspective - Dissertation | PDF",1785814141,370,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"why-machine-learning-models-fail-a-benchmarking-perspective-dissertation","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/why-machine-learning-models-fail-a-benchmarking-perspective-dissertation/123005/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do machine learning models often fail outside the benchmark setting?","Question",{"text":75,"@type":76},"Minor, seemingly harmless changes such as image noise, rain, or an incorrect background can shift inputs away from the controlled benchmark distribution, leading to wrong predictions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the dissertation connect benchmark performance to the desired capability?",{"text":80,"@type":76},"It argues that understanding model failures requires understanding how benchmark performance relates to the underlying capability being targeted, since benchmark design can bias what is measured.",{"name":82,"@type":73,"acceptedAnswer":83},"What is shortcut learning and how does it affect benchmark results?",{"text":84,"@type":76},"Shortcut learning occurs when models exploit superficial cues that work well on the benchmark’s training domain but are unrelated to the actual desired capability. This means good benchmark performance is not sufficient evidence that the capability was learned.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]