[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118288-en":3,"doc-seo-118288-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118288,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Information-Theoretic Foundations for Machine Learning - arXiv 2407.12288","Machine learning’s rapid progress has often outpaced rigorous theory, leaving practitioners to extrapolate from large empirical studies. This work proposes a mathematically rigorous Bayesian framework grounded in Shannon’s information theory to clarify what lies beyond the “cave” of observed shadows. It characterizes the performance of an optimal Bayesian learner by fundamental information limits, yielding accurate insights across settings from iid sampling to sequential and hierarchical data suitable for meta-learning. It also analyzes misspecified algorithms.","arXiv :2407 . 12288v3 [ stat .ML] 20 Aug 2024  \nInformation-Theoretic Foundations for Machine Learning  \nHong Jun Jeon 1∗ and Benjamin Van Roy2,3  \n1 Department of Computer Science, Stanford University  \n2 Department of Electrical Engineering, Stanford University  \n3 Department of Management Science and Engineering, Stanford University  \nAbstract  \nThe staggering progress of machine learning over the past decade has been a sight to behold. In retrospect, it is both remarkable and unsettling that these milestones were achievable with little to no rigorous theory to guide experimentation. Despite this fact, practitioners have been able to guide their future experimentation via observations from previous large-scale empirical investigations. However, alluding to Plato’s Allegory of the cave, it is likely that the observations which form the field’s notion of reality are but shadows representing fragments of that reality. In this work, we propose a theoretical framework which attempts to answer what exists outside of the cave. To the theorist, we provide a framework which is mathematically rigorous and leaves open many interesting ideas for future exploration. To the practitioner, we provide a framework whose results are very intuitive, general, and which will help form principles to guide future investigations. Concretely, we provide a theoretical framework rooted in Bayesian statistics and Shannon’s information theory which is general enough to unify the analysis of many phenomena in machine learning. Our framework characterizes the performance of an optimal Bayesian learner, which considers the fundamental limits of information. Unlike existing analyses that weaken with increasing data complexity, our theoretical tools provide accurate insights across diverse machine learning settings. Throughout this work, we derive very general theoretical results and apply them to derive insights specific to settings ranging from data which is independently and identically distributed under an unknown distribution, to data which is sequential, to data which exhibits hierarchical structure amenable to meta-learning. We conclude with a section dedicated to characterizing the performance of misspecified algorithms. These results are exciting and particularly relevant as we strive to overcome increasingly difficult machine learning challenges in this endlessly complex world.  \n∗ Correspondence to [hjjeon@stanford.edu](hjjeon@stanford.edu).  \nContents  \n1 Introduction 4  \n2 Related Works 5  \n2.1 Frequentist and Bayesian Statistics ..................................... 5  \n2.2 PAC Learning ................................................. 6  \n2.3 Information Theory ............................................. 6  \n3 A Framework for Learning 7  \n3.1 Probabilistic Framework and Notation ................................... 7  \n3.2 Data Generating Process ........................................... 8  \n3.3 Error ...................................................... 8  \n3.4 Achievable Error ............................................... 9  \nSummary ...................................................... 10  \n4 Requisite Information Theory 11  \n4.1 Entropy .................................................... 11  \n4.2 Conditional Entropy ............................................. 12  \n4.3 Mutual Information ............................................. 13  \n4.4 Differential Entropy ............................................. 13  \n4.5 Requisite Results from Information Theory ................................ 14  \nSummary ...................................................... 17  \n5 Connecting Learning and Information Theory 18  \n5.1 Error is Information ............................................. 18  \n5.2 Characterizing Error via Rate-Distortion Theory ............................. 19  \nSummary ...................................................... 21  \n6 Learning from iid Data 22  \n6.1 Theoretical Results Tailored for iid Data ...................","cbCaiiTR2nicsmC6","https://ap.wps.com/l/cbCaiiTR2nicsmC6","pdf",1363271,1,80,"English","en",105,"# Introduction\n# Related Works\n## Frequentist and Bayesian Statistics\n## PAC Learning\n## Information Theory\n# A Framework for Learning\n## Probabilistic Framework and Notation\n## Data Generating Process\n## Error\n## Achievable Error\n# Requisite Information Theory\n## Entropy\n## Conditional Entropy\n## Mutual Information\n## Differential Entropy\n## Requisite Results from Information Theory\n# Connecting Learning and Information Theory\n## Error is Information\n## Characterizing Error via Rate-Distortion Theory\n# Learning from iid Data\n## Theoretical Results Tailored for iid Data\n## Linear Regression\n## Logistic Regression\n## Deep Neural Networks\n## Nonparametric Learning\n# Learning from Sequences\n## Data Generating Process\n## Binary AR(K) Process\n## Transformer Process","[{\"question\":\"What theoretical framework does the document propose for analyzing machine learning performance?\",\"answer\":\"It proposes a Bayesian framework rooted in Shannon’s information theory, designed to unify analysis across multiple machine learning phenomena under fundamental information limits.\"},{\"question\":\"How does the framework handle different data regimes?\",\"answer\":\"It derives general theoretical results and applies them to iid data, sequential data, and hierarchically structured data that can support meta-learning.\"},{\"question\":\"What does the document conclude about optimal learners and error?\",\"answer\":\"It characterizes the performance of an optimal Bayesian learner by the fundamental limits of information and connects error to information, including characterization via rate-distortion theory.\"}]","Information-Theoretic Foundations for Machine Learning - arXiv 2407.12288 | PDF",1785682819,202,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"information-theoretic-foundations-for-machine-learning-arxiv-240712288","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/information-theoretic-foundations-for-machine-learning-arxiv-240712288/118288/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What theoretical framework does the document propose for analyzing machine learning performance?","Question",{"text":75,"@type":76},"It proposes a Bayesian framework rooted in Shannon’s information theory, designed to unify analysis across multiple machine learning phenomena under fundamental information limits.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the framework handle different data regimes?",{"text":80,"@type":76},"It derives general theoretical results and applies them to iid data, sequential data, and hierarchically structured data that can support meta-learning.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the document conclude about optimal learners and error?",{"text":84,"@type":76},"It characterizes the performance of an optimal Bayesian learner by the fundamental limits of information and connects error to information, including characterization via rate-distortion theory.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":21,"slug":99},"Literature","literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]