[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128680-en":3,"doc-seo-128680-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128680,962084928432,"Emma Wilson","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Information Theoretic Analysis of Deep Neural Networks - Final Dissertation","The task of learning an objective function that characterizes a deep, non-linear neural network focuses on training all parameters in a regime with large numbers of samples, high input dimensionality, and wide networks. The study uses a teacher-student framework where a teacher network generates labeled data and a student network with the same architecture classifies it. The work performs an information-theoretic analysis by extending known two-layer reductions to a recursive simplification scheme, replacing the last two layers with an equivalent one-layer model until a global one-layer equivalent is obtained.","UNIVERSITÀ DEGLI STUDI DI PADOVA Dipartimento di Fisica e Astronomia 􀀐Galileo Galilei􀀑  \nMaster Degree in Physics of Data  \nFinal Dissertation  \nInformation Theoretic Analysis of Deep Neural Networks  \nInternal thesis supervisor Candidate  \nDr. Michele Allegra Eleonora Bergamin  \nExternal thesis supervisor Prof. Jean Barbier  \nExternal thesis co-supervisors Dr. Francesco Camilli  \nDr. Daria Tieplova  \nAcademic Year 2023/2024  \nAbstract  \nThe task of learning an objective function that characterizes a deep, non-linear neural network is tackled, focusing on training the complete set of network parameters. Our investigation is conducted within a scenario where the number of samples, input dimension, and network width are all notably large. The neural networks under study operate in a teacher-student framework, where the data generated by the teacher network are classi􀀛ed by a student network with an identical architecture. Our main goal is to carry out an information-theoretical analysis of deep neural networks, building upon established results on two-layer networks. Recent conjectures, followed by partial rigorous proofs, show that it is possible to reduce two-layer networks to simpler one-layer networks, commonly referred to as generalized linear models. Remarkably, fundamental information-theoretic quantities such as the mutual information between training data and teacher network weights, as well as the Bayes-optimal generalization error, are well-understood for such simpli􀀛ed networks. Consequently, our strategy involves extending this reduction using a recursive argument. This involves progressively simplifying the network by replacing the last two layers with an equivalent one-layer neural network. The recursion continues until we identify an equivalent one-layer model for the entire network. This recursive approach is expected to provide us with a comprehensive understanding of the network’s behavior and performance.  \nContents  \nAbstract iii  \nList of symbols vii  \n1 Introduction 1  \n2 Framework and background 5  \n2.1 Information theory ..................................... 5  \n2.1.1 Information ..................................... 5  \n2.1.2 Entropy ....................................... 6  \n2.1.3 Joint entropy and conditional entropy ...................... 8  \n2.1.4 Kullback-Leibler divergence ............................ 10  \n2.1.5 Mutual information ................................. 10  \n2.2 Statistical mechanics .................................... 13  \n2.3 Statistical and Bayesian inference ............................. 17  \n2.3.1 Statistical inference ................................. 18  \n2.3.2 Bayesian inference ................................. 18  \n2.3.3 Bayesian inference as a statistical mechanics problem ............. 20  \n2.4 Machine learning ...................................... 23  \n2.4.1 Neural networks .................................. 24  \n2.4.2 Generalized linear models ............................. 26  \n3 Model and setting 29  \n3.1 Teacher-student setup .................................... 29  \n3.2 Model ............................................. 30  \n3.3 Equivalent shallow network ................................ 36  \n3.4 Methods ........................................... 40  \n3.4.1 Stein’s Lemma ................................... 40  \n3.4.2 Nishimori identity ................................. 42  \n3.4.3 Concentration of measure ............................. 42  \n3.4.4 Interpolation method ................................ 44  \n4 Main results 47  \n4.1 Recursion scheme ...................................... 47  \n4.2 Results ............................................ 50  \n4.2.1 Concentration results ................................ 50  \n4.2.2 Free entropy and mutual information results ................... 52  \n4.3 Outline of the proof of Theorem 7 ............................. 56  \n5 Proofs 61  \n5.1 Concentration proofs .................................... 63  \n5.1.1 Function of a sub-Gauss","cbCaiezIlLECS71k","https://ap.wps.com/l/cbCaiezIlLECS71k","pdf",1151043,2,1,117,"English","en",105,"# Abstract\n# List of symbols\n# Introduction\n# Framework and background\n## Information theory\n## Statistical mechanics\n## Statistical and Bayesian inference\n## Machine learning\n# Model and setting\n## Teacher-student setup\n## Model\n## Equivalent shallow network\n## Methods\n## Stein’s Lemma\n## Nishimori identity\n## Concentration of measure\n## Interpolation method\n# Main results\n## Recursion scheme\n## Results\n## Outline of the proof of Theorem 7\n# Proofs\n## Concentration proofs\n## Output kernel properties\n## Approximation Lemma\n## Proof of Theorem 7\n## Proof of Corollary 9\n# Conclusions and future works\n# Bibliography","[{\"question\":\"What scenario and modeling framework does the dissertation use for the neural networks?\",\"answer\":\"It uses a teacher-student setup in which a teacher network generates data that is classified by a student network with the same architecture.\"},{\"question\":\"What is the dissertation’s main information-theoretical goal?\",\"answer\":\"To carry out an information-theoretical analysis of deep neural networks by extending reductions known for two-layer networks to a recursive simplification approach.\"},{\"question\":\"How does the recursive reduction method simplify deep networks?\",\"answer\":\"It repeatedly replaces the last two layers with an equivalent one-layer neural network, continuing until the entire deep network is represented by an equivalent one-layer model.\"}]","Information Theoretic Analysis of Deep Neural Networks - Final Dissertation | PDF",1786002548,295,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"information-theoretic-analysis-of-deep-neural-networks-final-dissertation","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/information-theoretic-analysis-of-deep-neural-networks-final-dissertation/128680/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What scenario and modeling framework does the dissertation use for the neural networks?","Question",{"text":76,"@type":77},"It uses a teacher-student setup in which a teacher network generates data that is classified by a student network with the same architecture.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the dissertation’s main information-theoretical goal?",{"text":81,"@type":77},"To carry out an information-theoretical analysis of deep neural networks by extending reductions known for two-layer networks to a recursive simplification approach.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the recursive reduction method simplify deep networks?",{"text":85,"@type":77},"It repeatedly replaces the last two layers with an equivalent one-layer neural network, continuing until the entire deep network is represented by an equivalent one-layer model.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]