[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128794-en":3,"doc-seo-128794-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128794,1099523882367,"Hazel","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Representations in Zero-Shot Meta-Reinforcement Learning - Doctoral Thesis","A doctoral thesis investigates how autonomous AI agents learn sequences of decisions when encountering previously unseen environment situations without explicit state representations. It frames the challenge of sample inefficiency and slow adaptation in reinforcement learning, then studies meta-reinforcement learning for rapid task adaptation while targeting broad generalization. Contributions include stable hypernetwork training for improved generalization, end-to-end objective designs that work effectively with hypernetworks, and permutation-invariant sequence modeling improved by controlled permutation variance.","Representations in Zero-Shot Meta-Reinforcement Learning  \nJacob Beck  \nLinacre College University of Oxford  \nA thesis submitted for the degree of Doctor of Philosophy  \nHilary 2025  \nAbstract  \nA long-standing goal in the field of Artificial Intelligence (AI) is to create autonomous agents capable of making a sequence of decisions to interact with the world. When interacting with its environment, an agent will encounter new situations that it has not previously encountered and for which we cannot provide an explicit representation of the environment. In this case, the agent must learn about its environment through interaction allowing trial and error – a problem well modeled by the field of Reinforcement Learning (RL) . A long-standing goal in reinforcement learning is to create an agent capable of solving a diverse set of learning problems (i.e., broad generalization), and also learn quickly from few samples (i.e., rapid adaptation) . While RL has made remarkable progress in generalization across diverse applications, from Go to nuclear fusion control, it remains sample-inefficient and slow to adapt.  \nIn this thesis, we explore meta-reinforcement learning (meta-RL) as a solution to these challenges. In meta-RL, an agent is directly trained to rapidly adapt over a distribution of tasks. Still, generalization over a broad task distribution in meta-RL remains a significant hurdle. To address this, in Chapter 4, we propose the use of hypernetworks, i.e. , neural networks that generate the weights and biases for other neural networks. In order to train hypernetworks stably, we introduce a novel initialization scheme that enables hypernetworks to improve generalization performance. We explore the implications of hypernetworks further in Chapter 5 . We also show that, surprisingly, when used with our hypernetworks, simple end-to-end objectives can outperform more complex ones. We then explore the role of sequence models in meta-RL in Chapter 6 . We specifically investigate sequence models that are permutation-invariant with respect to historical transitions. This inductive bias is theoretically sufficient for representing the optimal policy and can ameliorate gradient decay. However, we find that introducing controlled permutation variance improves architectural robustness and allows for the representation of sub-optimal policies, which can serve as stepping stones to the optimal solution. Finally, we consider how these methods might differ in practice, by considering an application of meta-supervised learning to protein fitness prediction in Chapter 7 .  \nIn summary, this thesis provides representations that critically improve the capabilities of meta-RL agents. We show that such agents can be trained with hypernetworks, but that doing so requires stable initialization. We show that end-to-end objectives provide effective supervision, but only when using hypernetworks. And, we show that permutation invariance is useful, but only when carefully combined with permutation variance as well. By integrating these insights, we offer a path forward for RL agents capable of both rapid adaptation and broad generalization, advancing a long-standing goal in the field of AI.  \nAcknowledgements  \nTo my advisor, Shimon Whiteson, thank you for your guidance on my research, especially when encouraging me to pivot topics, and for your editorial insight, especially when framing the stories for my papers. You provided plenty of academic space to be independent, but also guidance when needed. A special thanks for editing our many drafts of a monstrously large survey paper, for which a single read was no small feat of stamina and labor.  \nTo my collaborators, Risto Vuorio, Matthew Jackson, Luisa Zintgraf, Tarun Gupta, Zheng Xiong, Evan Liu, Mingfei Sun, and Chelsea Finn, thank you for making my research a reality. To my lab members, Kristian Hartikainen, Clemence Grislain, Nasma Dasser, Benjamin Ellis, Alex Goldie, Alex Zakharov, Bei Peng, Wendelin B","cbCair7mWCVaNCzf","https://ap.wps.com/l/cbCair7mWCVaNCzf","pdf",43339758,1,216,"English","en",105,"# Abstract\n# Chapter 4: Hypernetworks and Initialization for Generalization\n# Chapter 5: Implications of Hypernetworks\n# Chapter 6: Sequence Models and Permutation Invariance\n# Chapter 7: Protein Fitness Prediction via Meta-Supervised Learning","[{\"question\":\"What problem does the thesis address in reinforcement learning?\",\"answer\":\"It addresses reinforcement learning’s sample inefficiency and slow adaptation when agents must handle new situations without explicit environment representations.\"},{\"question\":\"How does the thesis improve meta-reinforcement learning generalization?\",\"answer\":\"It proposes using hypernetworks and introduces a novel initialization scheme to train hypernetworks stably and improve generalization performance.\"},{\"question\":\"What role do sequence models play in the proposed meta-RL methods?\",\"answer\":\"The thesis studies sequence models that are permutation-invariant to historical transitions, and finds that adding controlled permutation variance improves robustness and enables representation of sub-optimal policies.\"}]","Representations in Zero-Shot Meta-Reinforcement Learning - Doctoral Thesis | PDF",1786003472,544,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"representations-in-zero-shot-meta-reinforcement-learning-doctoral-thesis","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/representations-in-zero-shot-meta-reinforcement-learning-doctoral-thesis/128794/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-24","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the thesis address in reinforcement learning?","Question",{"text":76,"@type":77},"It addresses reinforcement learning’s sample inefficiency and slow adaptation when agents must handle new situations without explicit environment representations.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the thesis improve meta-reinforcement learning generalization?",{"text":81,"@type":77},"It proposes using hypernetworks and introduces a novel initialization scheme to train hypernetworks stably and improve generalization performance.",{"name":83,"@type":74,"acceptedAnswer":84},"What role do sequence models play in the proposed meta-RL methods?",{"text":85,"@type":77},"The thesis studies sequence models that are permutation-invariant to historical transitions, and finds that adding controlled permutation variance improves robustness and enables representation of sub-optimal policies.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]