[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128255-en":3,"doc-seo-128255-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128255,2336475104736,"Quinn","https://ap-avatar.wpscdn.com/avatar/22000c4c5e0e5b17e70?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786591360781797222",8,"Research & Report","Towards Robust Machine Learning - Benchmarking and Adaptation in Challenging Settings","Neural networks deliver strong results when inputs match training data, but degrade under distribution shift. This thesis addresses the robustness gap for deployed models without requiring additional training or new data collection, including settings such as medical imaging and autonomous driving. A benchmark evaluates test-time adaptation under prolonged, varied shifts, showing many methods initially help yet later degrade; a simple baseline maintains high performance. Mechanistic analysis of entropy-based losses explains accuracy loss under continued optimization and motivates Weighted Flips. The work then extends adaptation ideas to language models for literature recommendation, finding baseline limitations and introducing a retrieval-and-reading agent to improve performance.","Towards Robust Machine Learning: Benchmarking and Adaptation in Challenging Settings  \nDissertation  \nder Mathematisch-Naturwissenschaftlichen Fakultät  \nder Eberhard Karls Universität Tübingen  \nzur Erlangung des Grades eines  \nDoktors der Naturwissenschaften  \n(Dr. rer. nat.)  \nvorgelegt von  \nDipl.-Inform. Ori Press  \naus Petah Tikva, Israel  \nTübingen  \n2025  \nGedruckt mit Genehmigung der Mathematisch-Naturwissenschaftlichen Fakultät der Eberhard Karls Universität Tübingen.  \nTag der mündlichen Qualifikation: 25.07.2025  \nDekan: Prof. Dr. Thilo Stehle  \n1. Berichterstatter/-in: Prof. Dr. Matthias Bethge  \n2. Berichterstatter/-in: Prof. Dr. Seong Joon Oh  \nFor Avia.  \nAbstract  \nNeural networks often excel when their inputs closely match the data on which they were trained, yet they frequently fail when inputs differ even slightly from their training data. This issue, known as distribution shift, remains a significant challenge when deploying machine learning models in practical applications such as medical imaging and autonomous driving. Traditional methods to address distribution shift typically involve additional training or data collection, which may not always be feasible for models already deployed. This thesis explores alternative strategies aimed at enhancing the robustness of already trained models to distribution shifts.  \nThe first part of this work introduces a benchmark specifically designed to evaluate testtime adaptation (TTA) methods under prolonged and varied distribution shifts. Using this benchmark, we demonstrate that while existing TTA techniques initially improve performance, they often lead to performance degradation with extended adaptation. We also propose a simple baseline method capable of consistently outperforming other tested methods, maintaining high performance even throughout prolonged adaptation.  \nBuilding on these insights, the second part analyzes the underlying mechanisms of entropybased loss functions commonly employed in TTA. We show that entropy minimization initially clusters embeddings of similar images together, thus increasing accuracy. However, continued entropy minimization eventually drives input image embeddings further away from training embeddings, thereby reducing accuracy. Leveraging this insight, we propose Weighted Flips (WF), a novel method capable of predicting model accuracy on arbitrary image sets without the need for labeled data.  \nThe final part of this work extends the principles of TTA to language models (LMs), focusing on the task of literature recommendation. We propose a benchmark that evaluates LMs in their ability to infer academic papers when given a short description that references them. Our  \nbenchmark demonstrates that LMs are unable to effectively perform this task. Therefore, we propose a simple agent that allows LMs to search for and read relevant papers, significantly improving their performance.  \nKurzfassung  \nNeuronale Netze erzielen oft hervorragende Ergebnisse, wenn ihre Eingaben den Daten ähneln, auf denen sie trainiert wurden. Sie versagen jedoch häufig, sobald sich die Eingaben auch nur geringfügig von ihren Trainingsdaten unterscheiden. Dieses Problem, bekannt als Distribution Shift (Verteilungsshift), stellt weiterhin eine große Herausforderung dar, wenn maschinelle Lernmodelle in praktischen Anwendungen wie der medizinischen Bildgebung oder dem autonomen Fahren eingesetzt werden. Traditionelle Ansätze zur Bewältigung des Distribution Shifts umfassen typischerweise zusätzliches Training oder die Sammlung neuer Daten, was jedoch nicht immer für bereits eingesetzte Modelle praktikabel ist. Diese Arbeit untersucht daher alternative Strategien, um die Robustheit bereits trainierter Modelle gegenüber Distribution Shifts zu verbessern.  \nIm ersten Teil dieser Arbeit wird ein Benchmark vorgestellt, der speziell zur Bewertung von Testzeit-Adaptionsmethoden (TTA) bei langanhaltenden und vielfältigen Distribution Shiftsentwickelt wurde. Mithilfe d","cbCaiuNz42ItX1f3","https://ap.wps.com/l/cbCaiuNz42ItX1f3","pdf",21365987,3,1,164,"English","en",105,"# Introduction\n## The Rise of Deep Learning\n## Why AI Fails Unexpectedly\n## Towards Reliable AI: Adaptation and Evaluation\n## Thesis Contributions\n# RDumb: A simple approach that q","[{\"question\":\"What problem does this thesis focus on?\",\"answer\":\"It focuses on distribution shift, where neural networks fail when inputs differ from the data used for training. The goal is to improve robustness of already trained models in practical deployments.\"},{\"question\":\"How does the thesis evaluate test-time adaptation methods?\",\"answer\":\"It introduces a benchmark designed to test test-time adaptation under prolonged and varied distribution shifts. The results show common methods can degrade performance with extended adaptation.\"},{\"question\":\"What is Weighted Flips and what insight motivates it?\",\"answer\":\"Entropy minimization initially improves accuracy by clustering embeddings of similar images, but continued minimization eventually pushes embeddings away from training representations and reduces accuracy. Weighted Flips leverages this to predict model accuracy on image sets without labeled data.\"}]","Towards Robust Machine Learning - Benchmarking and Adaptation in Challenging Settings | PDF",1785946265,413,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"towards-robust-machine-learning-benchmarking-and-adaptation-in-challenging-settings","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":20},"https://docshare.wps.com/document/research-report/",{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/towards-robust-machine-learning-benchmarking-and-adaptation-in-challenging-settings/128255/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-25","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does this thesis focus on?","Question",{"text":76,"@type":77},"It focuses on distribution shift, where neural networks fail when inputs differ from the data used for training. The goal is to improve robustness of already trained models in practical deployments.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the thesis evaluate test-time adaptation methods?",{"text":81,"@type":77},"It introduces a benchmark designed to test test-time adaptation under prolonged and varied distribution shifts. The results show common methods can degrade performance with extended adaptation.",{"name":83,"@type":74,"acceptedAnswer":84},"What is Weighted Flips and what insight motivates it?",{"text":85,"@type":77},"Entropy minimization initially improves accuracy by clustering embeddings of similar images, but continued minimization eventually pushes embeddings away from training representations and reduces accuracy. Weighted Flips leverages this to predict model accuracy on image sets without labeled data.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]