[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124739-en":3,"doc-seo-124739-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124739,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Beyond Worst-Case Generalization in Modern Machine Learning - Doctoral Dissertation","Beyond Worst-Case Generalization in Modern Machine Learning investigates why large, over-parameterized models trained on fewer examples than parameters can still generalize well to unseen data. The thesis argues that classical worst-case analysis rarely explains practical success and therefore develops alternatives. It introduces path-norm–based bounds for deep networks with positive homogeneous activations, studies the distribution of test errors to show worst-case failures are extremely rare, and analyzes closeness to Bayes-optimal performance using normalizing flows. It further develops oracle bounds for ensembling and studies a transition at the interpolation threshold.","UC Berkeley  \nUC Berkeley Electronic Theses and Dissertations  \nTitle  \nBeyond Worst-Case Generalization in Modern Machine Learning  \nPermalink  \n[https://escholarship.org/uc/item/62j5g7vd](https://escholarship.org/uc/item/62j5g7vd)  \nAuthor  \nTheisen, Ryan Christopher  \nPublication Date  \n2023  \nPeer reviewed|Thesis/dissertation  \n[eScholarship.org](eScholarship.org) Powered by the California Digital Library  \nUniversity of California  \nBeyond Worst-Case Generalization in Modern Machine Learning  \nby Ryan Christopher Theisen  \nA dissertation submitted in partial satisfaction of the requirements for the degree of Doctor of Philosophy  \nin  \nStatistics  \nin the  \nGraduate Division  \nof the  \nUniversity of California, Berkeley  \nCommittee in charge:  \nProfessor Michael W. Mahoney, Chair Professor Aditya Guntuboyina  \nProfessor Song Mei  \nSummer 2023  \nBeyond Worst-Case Generalization in Modern Machine Learning  \nCopyright 2023  \nby Ryan Christopher Theisen  \n1  \nAbstract  \nBeyond Worst-Case Generalization in Modern Machine Learning  \nby  \nRyan Christopher Theisen  \nDoctor of Philosophy in Statistics  \nUniversity of California, Berkeley  \nProfessor Michael W. Mahoney, Chair  \nThis thesis is concerned with the topic of generalization in large, over-parameterized machine learning systems—that is, how models with many more parameters than the number of examples they are trained on can perform well on new, unseen data. Such systems have become ubiquitous in many modern technologies, achieving unprecedented success across a wide range of domains. Yet, a comprehensive understanding of exactly why modern machine learning models work so well has alluded fundamental understanding. Classical approaches to this question have focused on characterizing how models will perform in the worst-case. However, recent findings have strongly suggested that such an approach is unlikely to yield the desired insight. This thesis is concerned with furthering our understanding of this phenomenon, with an eye towards moving beyond the worst-case. The contents of this thesis are divided into six chapters.  \nIn Chapter 1, we introduce the problem of generalization in machine learning, and briefly overview some of the approaches—both recent and classical—that have been taken to understand it.  \nChapter 2 introduces a novel analyses of deep neural networks with positive homogeneous activation functions. We develop a method to provably sparsify and quantize the parameters of a model by sampling paths through the network. This directly leads to a covering of this class of neural networks, whose size grows with a measure of complexity we call the \"path norm\". Using standard techniques, we can derive new worst-case generalization bounds that improve on previous results appearing in the literature.  \nIn Chapter 3, we take a critical look at the worst-case approach to understanding generalization. To do this, we develop a methodology to compute the full distribution of test errors for interpolating linear classifiers on real-world datasets, and compare this distribution to the performance of the worst-case classifier on the same tasks. We consistently find that, while truly poor, worst-case classifiers indeed exist for these tasks, they are exceedingly rare—so  \n2  \nmuch so that we expect to essentially never encounter them in practice. Moreover, we observe that as models become larger, test errors undergo a concentration around a critical threshold, with almost all classifiers achieving nearly the same error rate. These results suggest that the worst-case approach to generalization is unlikely to describe practical performance of large, over-parameterized models, and that new approaches are needed.  \nIn light of our findings in Chapter 3, in Chapter 4 we ask a complementary question: if modern machine learning systems perform nowhere near the worst-case, how close might their performance be to the best-case? For an arbitrary classification task, we can quantify the ","cbCairWgMaZBPjtf","https://ap.wps.com/l/cbCairWgMaZBPjtf","pdf",8921161,1,127,"English","en",105,"# Abstract\n# Generalization in Machine Learning\n## Positive Homogeneous Deep Neural Networks\n## Worst-Case vs Practical Test Error Distributions\n# Best-Case Proximity and Bayes Error Invariance\n## Normalizing Flows and Bayes-Optimal Evaluation\n# Average-Case Generalization via Ensembling\n## Ensemble Improvement Transition at the Interpolation Threshold\n# Conclusion","[{\"question\":\"What problem does the thesis address about generalization in modern machine learning?\",\"answer\":\"It studies why large, over-parameterized models can perform well on new, unseen data despite having many more parameters than training examples.\"},{\"question\":\"Why does the thesis question classical worst-case approaches?\",\"answer\":\"It finds that worst-case failures exist but are exceedingly rare on real-world datasets, and that test errors concentrate near a threshold as models grow.\"},{\"question\":\"How does the thesis analyze ensembling for deep neural networks?\",\"answer\":\"It derives oracle bounds linking ensemble improvement to disagreement/error ratios, and shows a distinct transition in improvement and disagreement-error at the interpolation threshold.\"}]","Beyond Worst-Case Generalization in Modern Machine Learning - Doctoral Dissertation | PDF",1785894221,320,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"beyond-worst-case-generalization-in-modern-machine-learning-doctoral-dissertation","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/beyond-worst-case-generalization-in-modern-machine-learning-doctoral-dissertation/124739/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the thesis address about generalization in modern machine learning?","Question",{"text":75,"@type":76},"It studies why large, over-parameterized models can perform well on new, unseen data despite having many more parameters than training examples.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why does the thesis question classical worst-case approaches?",{"text":80,"@type":76},"It finds that worst-case failures exist but are exceedingly rare on real-world datasets, and that test errors concentrate near a threshold as models grow.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the thesis analyze ensembling for deep neural networks?",{"text":84,"@type":76},"It derives oracle bounds linking ensemble improvement to disagreement/error ratios, and shows a distinct transition in improvement and disagreement-error at the interpolation threshold.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]