[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124844-en":3,"doc-seo-124844-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124844,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","On Kernel and Feature Learning in Neural Networks - Doctor of Philosophy Thesis","Kernel learning and feature learning provide two complementary paradigms for explaining the behaviour of large-scale deep learning systems, inspired by the theory of wide neural networks. The thesis studies relationships and shared themes between the two perspectives through three works: Bayesian views of deep ensembles via kernel learning and Gaussian processes, knowledge distillation using the feature kernel induced by final-layer representations, and self-supervised learning analysis linking eigenvalue decay to the gap between collapsed and whitened features and downstream generalisation under scarce labels.","On Kernel and Feature Learning  \nin Neural Networks  \nA thesis submitted for the degree of Doctor of Philosophy  \nMichaelmas 2022  \n\n| Bobby He\u003Cbr>St Peter’s College | Department of Statistics University of Oxford |\n| --- | --- |\n\nAbstract  \nInspired by the theory of wide neural networks (NNs), kernel learning and feature learning have recently emerged as two paradigms through which we can understand the complex behaviours of large-scale deep learning systems in practice. In the literature, they are often portrayed as two opposing ends of a dichotomy, both with their own strengths and weaknesses: one, kernel learning, draws connections to well-studied machine learning techniques like kernel methods and Gaussian Processes, whereas the other, feature learning, promises to capture more of the rich, but yet unexplained, properties that are unique to NNs.  \nIn this thesis, we present three works studying properties of NNs that combine insights from both perspectives, highlighting not only their differences but also shared similarities. We start by reviewing relevant literature on the theory of deep learning, with a focus on the study of wide NNs. This provides context for a discussion of kernel and feature learning, and against this backdrop, we proceed to describe our contributions. First, we examine the relationship between ensembles of wide NNs and Bayesian inference using connections from kernel learning to Gaussian Processes, and propose a modification that accounts for missing variance at initialisation in NN functions, resulting in a Bayesian interpretation to our trained deep ensembles. Next, we combine kernel and feature learning to demonstrate the suitability of the feature kernel, i.e. the kernel induced by inner products over final layer NN features, as a target for knowledge distillation, where one seeks to use a powerful teacher model to improve the performance of a weaker student model. Finally, we explore the gap between collapsed and whitened features in self-supervised learning, highlighting the decay rate of eigenvalues in the feature kernel as a key quantity that bridges between this gap and impacts downstream generalisation performance, especially in settings with scarce labelled data. We conclude with a discussion, including limitations and future outlook, of our contributions.  \nAcknowledgements  \nI’d like to start by thanking my supervisor Yee Whye, and co-supervisors, Arnaud and George, for their kindness and support during my PhD. To Yee Whye, I owe a great deal for starting me on the long and winding road that lead to much of this thesis, and for his guidance and patience when wrong paths were inevitably taken. I am grateful to Yee Whye also for the encouragement I received throughout my PhD to follow my interests, and for exemplifying qualities like openness, resilience and compassion, that I strive for, both within and outside of my research.  \nAt the Department of Statistics in Oxford, I’d like to thank: Adam, Alan, Bryn, Chris, D´eborah, Edwin, Emile, Emilien, Faaiz, Fran, James, Jean-Francois, Jessie, Jin, Lorenzo, Michael, Natalia, Qinyi, Robert, Sheh and Tyler for their friendship over these past four years. This list wouldn’t be complete without the Japanese van or the countless raspberry buns that came and went along the way. I’d also like to thank the department’s staff, especially Beverley, Joanna, Mark, Stuart and Susan, for their help on any issue of mine, no matter how large nor small.  \nI also owe a lot of gratitude to Mete for his unwavering support whilst allowing me to follow my instincts, and for creating an ideal environment for me that resulted in my time at SRUK being fruitful far beyond my expectations. I want to take the opportunity here to thank Ching-Ling, as it was around now the music started to play again.  \nIt was a real privilege too to intern at DeepMind in London in the final year of my PhD, and for that I am extremely grateful to my host James. I’d like to thank James too","cbCaicbx6xzEuwxk","https://ap.wps.com/l/cbCaicbx6xzEuwxk","pdf",3202045,1,160,"English","en",105,"# Introduction\n## Thesis Outline\n## Omitted work\n# Literature Review\n## Background and Notation\n## Kernel Learning in Neural Networks\n## Feature Learning in Neural Networks","[{\"question\":\"What problem does the thesis address about kernel learning and feature learning?\",\"answer\":\"It examines how kernel learning and feature learning—often framed as opposing paradigms—relate to each other and how they can jointly explain behaviours of wide neural networks.\"},{\"question\":\"How does the thesis connect deep ensembles to Bayesian inference?\",\"answer\":\"It uses connections from kernel learning to Gaussian processes and proposes a modification to account for missing variance at initialisation, yielding a Bayesian interpretation of trained deep ensembles.\"},{\"question\":\"What role does the feature kernel play in knowledge distillation in this thesis?\",\"answer\":\"It shows that the feature kernel induced by inner products over final-layer neural network features is suitable as a target for knowledge distillation from a teacher to a weaker student model.\"}]","On Kernel and Feature Learning in Neural Networks - Doctor of Philosophy Thesis | PDF",1785894957,403,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"on-kernel-and-feature-learning-in-neural-networks-doctor-of-philosophy-thesis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/on-kernel-and-feature-learning-in-neural-networks-doctor-of-philosophy-thesis/124844/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the thesis address about kernel learning and feature learning?","Question",{"text":75,"@type":76},"It examines how kernel learning and feature learning—often framed as opposing paradigms—relate to each other and how they can jointly explain behaviours of wide neural networks.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the thesis connect deep ensembles to Bayesian inference?",{"text":80,"@type":76},"It uses connections from kernel learning to Gaussian processes and proposes a modification to account for missing variance at initialisation, yielding a Bayesian interpretation of trained deep ensembles.",{"name":82,"@type":73,"acceptedAnswer":83},"What role does the feature kernel play in knowledge distillation in this thesis?",{"text":84,"@type":76},"It shows that the feature kernel induced by inner products over final-layer neural network features is suitable as a target for knowledge distillation from a teacher to a weaker student model.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]