[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123597-en":3,"doc-seo-123597-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123597,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","The No Free Lunch Theorem - Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning","No free lunch theorems for supervised learning state that no learner can solve all problems or that all learners obtain the same average accuracy over a uniform distribution of learning problems. This motivates the belief that each problem needs tailored inductive biases. The work argues that while uniformly sampled datasets have high Kolmogorov complexity, real-world data is biased toward low complexity, and neural networks share this preference. It shows domain-specific architectures can compress diverse domains and that pre-trained and random language models generate low-complexity sequences, while supporting automation of model selection when labeled data is scarce or abundant.","The No Free Lunch Theorem, Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning  \nMicah Goldblum * Marc Finzi * Keefer Rowan Andrew Gordon Wilson  \nNew York University  \narXiv :2304 .05366v1 [ cs .LG] 11 Apr 2023  \nAbstract  \nNo free lunch theorems for supervised learning state that no learner can solve all problems or that all learners achieve exactly the same accuracy on average over a uniform distribution on learning problems. Accordingly, these theorems are often referenced in support of the notion that individual problems require specially tailored inductive biases. While virtually all uniformly sampled datasets have high complexity, real-world problems disproportionately generate low-complexity data, and we argue that neural network models share this same preference, formalized using Kolmogorov complexity. Notably, we show that architectures designed for a particular domain, such as computer vision, can compress datasets on a variety of seemingly unrelated domains. Our experiments show that pre-trained and even randomly initialized language models prefer to generate low-complexity sequences. Whereas no free lunch theorems seemingly indicate that individual problems require specialized learners, we explain how tasks that often require human intervention such as picking an appropriately sized model when labeled data is scarce or plentiful can be automated into a single learning algorithm. These observations justify the trend in deep learning of unifying seemingly disparate problems with an increasingly small set of machine learning models.  \n1. Introduction  \nThe problem of justifying inductive reasoning has challenged epistemologists since at least the 1700s (Hume, 1748) . How can we justify our belief that patterns we observed previously are likely to continue into the future without appealing to this same inductive reasoning in a circular fashion? Nonetheless, we adopt inductive reasoning in everyday life whenever we learn from our mistakes or make decisions based on past experience. Likewise, the feasibility of machine learning is entirely dependent on in-  \n*  \nEqual contribution.  \nduction, as models extrapolate from patterns found in previously observed training data to new samples at inference time.  \nMore recently, in the late 1990s, no free lunch theorems emerged from the computer science community as rigorous arguments for the impossibility of induction in contexts seemingly relevant to real machine learning problems (Wolpert, 1996 ; Wolpert & Macready, 1997) . One such no free lunch theorem for supervised learning states that no single learner can achieve high accuracy on every problem (Shalev-Shwartz & Ben-David, 2014) . Another states that, assuming a world where labeling functions of learning problems are drawn from a uniform distribution, every learner is equally good on expectation, achieving the same accuracy as a random guess (Wolpert, 1996) . Such a world would be hostile to inductive reasoning. The assumption that labelings are drawn uniformly ensures that training data is uninformative about unseen samples.  \nIn contrast to this dismal outlook on machine learning, naturally occurring inference problems involve highly structured data, and basic structures are often shared even across seemingly disparate different problems. If we can design learning algorithms with inductive biases that are aligned with this structure, then we may hope to perform inference on a wide range of problems. In this work, we explore the alignment between structure in real-world data and machine learning models through the lens of Kolmogorov complexity.  \nThe Kolmogorov complexity of an output is deﬁned as the length of the shortest program under a ﬁxed language that produces that output. In Section 3, we explain the connection between Kolmogorov complexity and compressibility. We note that virtually all data drawn from a uniform distribution as assumed by the no free lunch theorem of Wolpert (1996) canno","cbCaivw3sm4SOoGH","https://ap.wps.com/l/cbCaivw3sm4SOoGH","pdf",641663,1,21,"English","en",105,"# Introduction\n## Inductive reasoning and learning feasibility\n## No free lunch theorems in supervised learning\n## Real-world structured data and inductive bias\n## Kolmogorov complexity and compressibility\n## Neural networks and low-complexity preference","[{\"question\":\"What do no free lunch theorems imply for supervised learning?\",\"answer\":\"They imply that no learner can achieve high accuracy on every possible problem, and under a uniform distribution of learning labels, all learners have equal expected performance.\"},{\"question\":\"How does Kolmogorov complexity relate to the paper’s argument about inductive bias?\",\"answer\":\"Kolmogorov complexity measures the length of the shortest program generating an output; the paper links this to compressibility and argues that real-world datasets and neural models preferentially produce or select low-complexity outputs.\"},{\"question\":\"What evidence does the paper provide that neural network models prefer low-complexity data?\",\"answer\":\"It reports experiments where both pre-trained and randomly initialized language models favor generation of low-complexity sequences, and vision architectures prefer correct labelings even when inputs lack natural image structure.\"}]","The No Free Lunch Theorem - Kolmogorov Complexity, and the Role of Inductive Biases in Machine Learning | PDF",1785817555,53,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-no-free-lunch-theorem-kolmogorov-complexity-and-the-role-of-inductive-biases-in-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-no-free-lunch-theorem-kolmogorov-complexity-and-the-role-of-inductive-biases-in-machine-learning/123597/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What do no free lunch theorems imply for supervised learning?","Question",{"text":75,"@type":76},"They imply that no learner can achieve high accuracy on every possible problem, and under a uniform distribution of learning labels, all learners have equal expected performance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Kolmogorov complexity relate to the paper’s argument about inductive bias?",{"text":80,"@type":76},"Kolmogorov complexity measures the length of the shortest program generating an output; the paper links this to compressibility and argues that real-world datasets and neural models preferentially produce or select low-complexity outputs.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence does the paper provide that neural network models prefer low-complexity data?",{"text":84,"@type":76},"It reports experiments where both pre-trained and randomly initialized language models favor generation of low-complexity sequences, and vision architectures prefer correct labelings even when inputs lack natural image structure.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]