[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-128840-105":59,"doc-detail-128840-en":130},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":123,"head_meta":125,"extra_data":127,"updated_unix":129},105,"en","position-a-theory-of-deep-learning-must-include-compositional-sparsity","Position: A Theory of Deep Learning Must Include Compositional Sparsity","","Overparametrized deep neural networks (DNNs) succeed in high-dimensional domains where classical shallow models struggle with the curse of dimensionality, yet core principles governing their learning dynamics remain unclear. This position paper argues that success is driven by compositional sparsity: target functions can be built from a small set of constituent functions, each depending only on low-dimensional input subsets. The property is shown to hold for all efficiently Turing-computable functions and is likely present in current learning problems; significant open questions persist for learnability and optimization, motivating a more complete theory of deep learning and intelligence.",{"@graph":69,"@context":122},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/position-a-theory-of-deep-learning-must-include-compositional-sparsity/128840/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/position-a-theory-of-deep-learning-must-include-compositional-sparsity/128840.png","ImageObject",300,407,{"name":92,"@type":93},"Aria","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-19","2026-08-06",true,{"@type":102,"interactionType":103,"userInteractionCount":39},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118],{"name":109,"@type":110,"acceptedAnswer":111},"What key principle does the paper propose for why deep neural networks work well?","Question",{"text":112,"@type":113},"It proposes that DNN success is driven by compositional sparsity in the target function, enabling efficient representation, learning, and generalization in high dimensions.","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"What does compositional sparsity mean in this context?",{"text":117,"@type":113},"Most relevant functions can be composed from a small set of constituent functions, where each constituent relies only on a low-dimensional subset of all inputs.",{"name":119,"@type":110,"acceptedAnswer":120},"Which open theoretical questions remain despite the compositional sparsity explanation?",{"text":121,"@type":113},"The paper notes unresolved issues about learnability and optimization of DNNs, even though compositional sparsity helps explain approximation and generalization.","https://schema.org",{"og:url":83,"og:type":124,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":126,"canonical":83},"index,follow",{"doc_id":128,"site_id":62},128840,1786003820,{"code":4,"msg":5,"data":131},{"doc_id":128,"user_id":132,"nickname":92,"user_avatar":133,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":39,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":139,"language":140,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":67,"update_tm":129,"read_time":144},2336474459895,"https://ap-avatar.wpscdn.com/avatar/22000baeef7a5ed0655?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786071322749376916","POLITECNICO DI TORINO Repository ISTITUZIONALE  \nPosition: A Theory of Deep Learning Must Include Compositional Sparsity  \nOriginal  \nPosition: A Theory of Deep Learning Must Include Compositional Sparsity / Danhofer, David A. ; D'Ascenzo, Davide; Dubach, Rafael; Poggio, Tomaso. -ELETTRONICO. -267:(2025), pp. 81199-81210. ( Forty-second International Conference on Machine Learning Position Vancouver (CAN) July 13th-19th, 2025) .  \nAvailability:  \nThis version is available at: 11583/3001665 since: 2025-10-24T12:01:30Z  \nPublisher: PMLR  \nPublished DOI:  \nTerms of use:  \nThis article is made available under terms and conditions as specified in the corresponding bibliographic description in the repository  \nPublisher copyright  \n(Article begins on next page)  \n21 February 2026  \nPosition: A Theory of Deep Learning Must Include Compositional Sparsity  \nDavid A. Danhofer * 1 2 Davide D’Ascenzo * 1 3 4 Rafael Dubach * 1 5 Tomaso Poggio * 1  \nAbstract  \nOverparametrized Deep Neural Networks (DNNs) have demonstrated remarkable success in a wide variety of domains too high-dimensional for classical shallow networks subject to the curse of dimensionality. However, open questions about fundamental principles, that govern the learning dynamics of DNNs, remain. In this position paper we argue that it is the ability of DNNs to exploit the compositionally sparse structure of the target function driving their success. As such, DNNs can leverage the property that most practically relevant functions can be composed from a small set of constituent functions, each of which relies only on a low-dimensional subset of all inputs. We show that this property is shared by all efficiently Turing-computable functions and is therefore highly likely present in all current learning problems. While some promising theoretical insights on questions concerned with approximation and generalization exist in the setting of compositionally sparse functions, several important questions on the learnability and optimization of DNNs remain. Completing the picture of the role of compositional sparsity in deep learning is essential to a comprehensive theory of artificial – and even general – intelligence.  \n1. Introduction  \nDeep Neural Networks (DNNs) have achieved remarkable breakthroughs across numerous domains, including computer vision, playing games (Silver et al., 2016a), protein structure prediction (Jumper et al., 2021 ; Abramsonet al., 2024), natural language usage, and complex reasoning (Romera-Paredes et al., 2023 ; Trinh et al., 2024 ;  \n*Equal contribution 1 Center for Brains, Minds and Machines (CBMM), MIT, Cambridge, MA, USA 2ETH Zurich, Zurich, Switzerland 3Politecnico di Torino, Torino, Italy 4University of Milan, Milan, Italy 5University of Zurich, Zurich, Switzerland. Correspondence to: Davide D’Ascenzo \u003C[davide.dascenzo@unimi.it](davide.dascenzo@unimi.it) >.  \nProceedings of the 42 nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025 . Copyright 2025 by the author(s) .  \nDeepSeek-AI et al., 2025) . Despite their rapidly growing set of achievements, our understanding of DNNs still lags behind their empirical success. Without deeper theoretical insights, it remains unclear why certain architectures scale so well to high-dimensional tasks, or how to pinpoint the limits of Deep Learning (DL) paradigms.  \nHistorically, many attempts to build intelligent systems relied on logical or rule-based approaches (Newell & Simon, 1976 ; Hayes-Roth et al., 1983) . In contrast, the rise of DNNs from around 2012 onward ushered in learning-based architectures that surpass traditional algorithms in numerous domains. For several years, the focus has been on pushing performance boundaries: from high-dimensional image recognition (AlexNet (Krizhevsky et al., 2012), ResNet (He et al., 2015)) to playing strategic games (AlphaGo (Silver et al., 2016b), AlphaZero (Silver et al., 2018)) . More recently, Large Language Models (LLMs) such as the GPT ","cbCaiiu6zrBozAJ4","https://ap.wps.com/l/cbCaiiu6zrBozAJ4","pdf",329757,13,"English","# Introduction\n## Approximation, Optimization, and Generalization\n## Compositional sparsity and the curse of dimensionality","[{\"question\":\"What key principle does the paper propose for why deep neural networks work well?\",\"answer\":\"It proposes that DNN success is driven by compositional sparsity in the target function, enabling efficient representation, learning, and generalization in high dimensions.\"},{\"question\":\"What does compositional sparsity mean in this context?\",\"answer\":\"Most relevant functions can be composed from a small set of constituent functions, where each constituent relies only on a low-dimensional subset of all inputs.\"},{\"question\":\"Which open theoretical questions remain despite the compositional sparsity explanation?\",\"answer\":\"The paper notes unresolved issues about learnability and optimization of DNNs, even though compositional sparsity helps explain approximation and generalization.\"}]","Position: A Theory of Deep Learning Must Include Compositional Sparsity | PDF",33]