[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128569-en":3,"doc-seo-128569-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128569,549768064778,"Finn","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","Generalisation and expressiveness for over-parameterised neural networks","Over-parameterised modern neural networks succeed through expressive power and generalisation capability. Expressive power concerns fitting diverse datasets, while generalisation enables extracting structure from training examples and applying it to unseen data. This thesis studies key challenges behind these properties: it examines why fitting does not guarantee practical expressiveness in deep architectures, and introduces scalable residual mechanisms. It then investigates generalisation without overfitting using information-theoretic and PAC-Bayesian approaches, proposing new learning algorithms and bounds.","Generalisation and expressiveness for over-parameterised neural networks  \nEugenio Clerico  \nMagdalen College Department of Statistics University of Oxford  \nA thesis presented for the degree of Doctor of Philosophy Hilary 2023  \nStatement of Originality  \nI hereby declare that except where specific reference is made to the work of others, the contents of this dissertation are original and have not been submitted in whole or in part for consideration for any other degree or qualification in this or any other university. My personal contributions are as outlined in the authorship forms at the end of each chapter. This dissertation is my own work except as specified in the text, acknowledgements, forms, and papers.  \nEugenio Clerico Hilary 2023  \nAcknowledgements  \nFirst of all, I would like to express my gratitude to my supervisors, George Deligiannidis and Arnaud Doucet. Without their support and guidance, this work would not have been possible. Their mentorship has been invaluable in shaping my research, providing constructive feedback, and offering insightful suggestions.  \nI would like to thank Benjamin Guedj for giving me the opportunity to work alongside him at UCL. A special thanks to Amitis Shidani and Tyler Farghly, who have not only been great collaborators but have also become close friends. Thanks to Bobby He and Soufiane Hayou for the help and motivation they gave me. I also owe much to all the brilliant colleagues and researchers with whom I had the chance to have valuable and insightful interactions (Judith Rousseau, Patrick Rebeschini, Umut Simsekli, Gergely Neu, Badr-Eddine Ch´eriefAbdelladif, and many others) .  \nI am grateful to the Engineering and Physical Sciences Research Council (EPSRC) and Magdalen College, Oxford, for allowing me to pursue my DPhil at Oxford. Also, I would like to thank the administrative and IT staff at the Oxford Statistics Department, who were always there to help whenever I needed their support.  \nThis thesis would not have been possible without the encouragement of my friends, both near and far. Jake Fawkes, Shahine Bouabid, Xi Lin, Dan Manela, Anri Asagumo, Jian Qian, Lorenzo Pacchiardi,  \nValeria Schellino, Romain Fournier, Yiorgos Kalantzis, Julian Thoenniss, Aizhan Shorman, Florent Bonnet, Kam´elia Daudel, Carlo Alfano, Tom Wu, Sam Hall-McMaster, Holly Yeo, and all the other people who shared this journey with me.  \nI will be eternally grateful to Baba and Adrien for their invaluable friendship and exceptional patience.  \nThanks to Marie, Pierre, Filippo and Herbie for their silent support.  \nThank you to my family, who has always been there for me: la Mamma, Bruno, Titti, le Nonne, le Zie, gli Zii, i Cugini e chi pi`u ne ha pi`u ne metta. Finally, I want to thank Sekela for suddenly bringing a burst of joy during the final year of my PhD.  \nAbstract  \nOver-parameterised modern neural networks owe their success to two fundamental properties: expressive power and generalisation capability. The former refers to the model’s ability to fit a large variety of datasets, while the latter enables the network to extrapolate patterns from training examples and apply them to previously unseen data. This thesis addresses a few challenges related to these two key properties.  \nThe fact that over-parameterised networks can fit any data set is not always indicative of their practical expressiveness. This is the object of the first part of this thesis, where we delve into how the input information can get lost when propagating through a deep architecture, and we propose as an easily implementable possible solution the introduction of suitable scaling factors and residual connections.  \nThe second part of this thesis focuses on generalisation. The reason why modern neural networks can generalise well to new data without overfitting, despite being over-parameterised, is an open question that is currently receiving considerable attention in the research community. We explore this subject from inf","cbCaiv4kwyTY4LK3","https://ap.wps.com/l/cbCaiv4kwyTY4LK3","pdf",3415194,1,221,"English","en",105,"# Introduction\n## Supervised learning framework\n## Neural networks\n## Infinite-width limit and Gaussian behaviour\n## Expressiveness\n## Generalisation\n## Contributions\n# Stable and expressive ResNets\n## Mathematical preliminaries\n## Gaussian limit for neural networks\n## Stable ResNets","[{\"question\":\"What properties of over-parameterised neural networks are central to this thesis?\",\"answer\":\"The thesis focuses on expressive power and generalisation capability. Expressive power relates to fitting many datasets, while generalisation refers to applying learned patterns to previously unseen data.\"},{\"question\":\"Why does the thesis argue that fitting any dataset does not automatically ensure practical expressiveness?\",\"answer\":\"It examines how input information can be lost through deep architectures during propagation. Based on this issue, it proposes scaling factors and residual connections to improve effective expressiveness.\"},{\"question\":\"How does the thesis approach generalisation for over-parameterised networks without overfitting?\",\"answer\":\"It explores the open problem from information-theoretic and PAC-Bayesian perspectives. The work includes novel learning algorithms and generalisation bounds designed to explain and formalise this behaviour.\"}]","Generalisation and expressiveness for over-parameterised neural networks | PDF",1786001789,557,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"generalisation-and-expressiveness-for-over-parameterised-neural-networks","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/generalisation-and-expressiveness-for-over-parameterised-neural-networks/128569/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What properties of over-parameterised neural networks are central to this thesis?","Question",{"text":76,"@type":77},"The thesis focuses on expressive power and generalisation capability. Expressive power relates to fitting many datasets, while generalisation refers to applying learned patterns to previously unseen data.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why does the thesis argue that fitting any dataset does not automatically ensure practical expressiveness?",{"text":81,"@type":77},"It examines how input information can be lost through deep architectures during propagation. Based on this issue, it proposes scaling factors and residual connections to improve effective expressiveness.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the thesis approach generalisation for over-parameterised networks without overfitting?",{"text":85,"@type":77},"It explores the open problem from information-theoretic and PAC-Bayesian perspectives. The work includes novel learning algorithms and generalisation bounds designed to explain and formalise this behaviour.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]