[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128663-en":3,"doc-seo-128663-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128663,962084925782,"Ava Thompson","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Continual Learning - Theoretical and Empirical Analysis of Infinitely Wide Neural Networks","Real-world AI systems must continuously gather, update, accumulate, and use knowledge over time, a process known as continual learning. This work studies a central obstacle: catastrophic forgetting, where learning new tasks can substantially reduce performance on previously learned ones. The analysis focuses on infinite wide (overparametrized) neural networks, leveraging architectural properties to better understand training dynamics. Using the Neural Tangent Kernel under Neural Tangent Parametrisation and Maximal Update Parametrisation, it characterizes the evolution of key quantities and kernels that govern learning.","Università degli Studidi Padova  \nDepartment of Physics and Astronomy “Galileo Galilei”  \nMaster Thesis in Physics Of Data  \nContinual Learning: Theoretical and Empirical Analysys of Infinitely Wide Neural Networks  \nSupervisor Master Candidate  \nProf. Marco Baiesi Alessandro Breccia  \nUniversità degli Studidi Padova  \nCo-supervisor Student ID  \nProf. Thomas Hofmann ँࣿआअआऀं  \nGiulia Lanzillotta Lorenzo Noci ETH Zurich  \nAcademic Year  \nँࣿ ँं - ँࣿ ँः  \nii  \n“E non per un Dio, ma nemmeno per gioco, perchè i ciliegi tornassero in fiore”  \n—Fabrizio De Andrè  \niv  \nAbstract  \nTo handle real-world dynamics, an intelligent system must continuously gather, update, accumulate, and utilise knowledge throughout its existence. This capability, termed continual learning, is essential for AI systems to adapt and evolve over time and to avoid re-training models from scratch in order to perform updates. However, a significant challenge in continual learning is catastrophic forgetting, where acquiring new knowledge often leads to a substantial decline in performance on previously learned tasks. This work aims to look at the origin of the catastrophic forgetting phenomenon under the lens of infinite wide (’overparametrized’) neural networks: thanks to special properties emerging from this architectural choice, it is possible to obtain a deeper knowledge of the training dynamics. Through the analysis of the Neural Tangent Kernel[ऀ] under two different network parametrisations, Neural Tangent Parametrisation(NTP) and Maximal Update Parametrisation(µP)[ँ], we can characterise the evolution of fundamental quantities and kernels governing the training dynamics.  \nvi  \nContents  \nAbstract v  \nList of figures ix  \nList of tables xv  \nListing of acronyms xvii  \nऀ Continual Learning ऀऀ . ऀ Definition ........................................... ऀ  \nऀ . ँ Mathematical Formulation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ँ  \nऀ.ं Scenarios ........................................... ँ  \nऀ.ः Evaluation Metrics ...................................... ं  \nऀ.ः . ऀ Overall Performance ................................ ं  \nऀ.ः.ँ Memory Stability .................................. ं  \nऀ.ः.ं Learning Plasticity ................................. ः  \nऀ.ऄ Different Approaches ..................................... ः  \nऀ.अ Continual Learning Benchmarks ............................... आ  \nँ Neural Tangent Kernel 9  \nँ . ऀ Introduction ......................................... ई  \nँ.ँ Contributions of the NTK .................................. ई  \nँ . ं Theoretical description .................................... ऀࣿ  \nँ .ः Gaussian Processes . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ऀँ  \nँ.ः . ऀ Linearised networks ................................. ऀं  \nँ.ऄ Field Theory Formulation ................................... ऀः  \nं Maximum Update Parametrisation ऀ7 ं . ऀ Dynamical Mean Field Theory (DMFT) ........................... ऀइ  \nं . ऀ . ऀ Introduction .................................... ऀइ  \nं.ऀ.ँ Path Integral Formulation ............................. ऀई  \nं.ऀ.ं Action and Order Parameters definitions ...................... ँऀ  \nं.ऀ.ः Saddle Point Solution ................................ ँँ  \nं.ऀ.ऄ Single Site Stochastic Process ............................ ँऄ  \nं.ऀ.अ DMFT equations .................................. ँअ  \nं.ँ Single Hidden Layer NN ................................... ँआ  \nं.ं Finite width N corrections .................................. ँई  \nः Continual Learning and Parametrisations 3 ऀः . ऀ Continual learning and lazy regime .............................. ंऀ  \nः.ँ Continual learning and rich regime .............................. ंः  \nः.ँ . ऀ Toy model: Single Hidden Layer NN ........................ ंइ  \nः . ं Catastrophic Forgetting ................................... ःࣿ  \nऄ Experiments and Results 43  \nऄ . ऀ Experimental setup ...................................... ःं  \nऄ . ऀ . ऀ Model: si","cbCaihk5kIWm0M9H","https://ap.wps.com/l/cbCaihk5kIWm0M9H","pdf",21446373,2,1,93,"English","en",105,"# Abstract\n# Continual Learning\n## Definition\n## Mathematical Formulation\n## Scenarios\n## Evaluation Metrics\n## Overall Performance\n## Memory Stability\n## Learning Plasticity\n## Different Approaches\n## Continual Learning Benchmarks\n# Neural Tangent Kernel\n## Introduction\n## Contributions of the NTK\n## Theoretical description\n## Gaussian Processes\n## Linearised networks\n## Field Theory Formulation\n# Maximum Update Parametrisation\n## Dynamical Mean Field Theory (DMFT)\n## Single Hidden Layer NN\n## Field Theory Formulation\n# Continual Learning and Parametrisations\n## Continual learning and lazy regime\n## Continual learning and rich regime\n## Toy model: Single Hidden Layer NN\n## Catastrophic Forgetting\n# Experiments and Results\n## Experimental setup\n## Dataset: Permuted MNIST\n## Evaluation metrics and key observables\n## Results\n# Discussions\n## Theoretical alignment\n## Scaling width experimental results\n## Addressing forgetting\n# Conclusions\n# Appendix\n## Additional Figures\n## Algorithmic Implementation\n# References\n# Acknowledgments","[{\"question\":\"What problem does this thesis address in continual learning?\",\"answer\":\"It addresses catastrophic forgetting, where learning new tasks causes a significant performance drop on previously learned tasks.\"},{\"question\":\"How does the thesis study continual learning theory?\",\"answer\":\"It analyzes training dynamics through the Neural Tangent Kernel, comparing two parametrisations that lead to different dynamical characterizations.\"},{\"question\":\"What network setting is emphasized for the theoretical analysis and experiments?\",\"answer\":\"The focus is on infinitely wide (overparametrized) neural networks, including experiments using a single hidden layer setting and Permuted MNIST.\"}]","Continual Learning - Theoretical and Empirical Analysis of Infinitely Wide Neural Networks | PDF",1786002425,234,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"continual-learning-theoretical-and-empirical-analysis-of-infinitely-wide-neural-networks","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/continual-learning-theoretical-and-empirical-analysis-of-infinitely-wide-neural-networks/128663/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does this thesis address in continual learning?","Question",{"text":76,"@type":77},"It addresses catastrophic forgetting, where learning new tasks causes a significant performance drop on previously learned tasks.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the thesis study continual learning theory?",{"text":81,"@type":77},"It analyzes training dynamics through the Neural Tangent Kernel, comparing two parametrisations that lead to different dynamical characterizations.",{"name":83,"@type":74,"acceptedAnswer":84},"What network setting is emphasized for the theoretical analysis and experiments?",{"text":85,"@type":77},"The focus is on infinitely wide (overparametrized) neural networks, including experiments using a single hidden layer setting and Permuted MNIST.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]