[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121888-en":3,"doc-seo-121888-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121888,4398048949847,"Eliana","https://ap-avatar.wpscdn.com/avatar/400002536579ef2da7f?_k=1778318612642679267",8,"Research & Report","Towards Efficient and Robust Reinforcement Learning via Synthetic Environments and Offline Data","Over the past decade, Deep Reinforcement Learning (RL) has driven advances in sequential decision-making, enabling applications such as superhuman game playing, robotic control, and automated algorithm discovery. Yet deep RL remains sample-inefficient, generalizes poorly beyond the original environment, and often trains unstably. This thesis studies two directions to address these issues: using synthetic environments and data to expand an agent’s experience, and developing principled methods to leverage pre-existing offline datasets to reduce or replace costly online data collection.","Towards Efficient and Robust Reinforcement Learning via Synthetic Environments and Offline Data  \nCong Lu Balliol College  \nUniversity of Oxford  \nA thesis submitted for the degree of Doctor of Philosophy  \nTrinity 2023  \nTo my parents,陆志豹and 姜苏.  \nAcknowledgements  \nFirst and foremost, I’m deeply grateful to my supervisors Mike and Yee Whye for taking a chance on me all those years ago and guiding me through the DPhil process; I will treasure the many fun discussions throughout. I want to particularly thank Mike for showing me a conscientious way to do research, the wonders of being Bayesian, and encouraging me to think about the broader implications of my research. I want to particularly thank Yee Whye for always encouraging me to formulate my thoughts scientifically and delve deeper to understand the underlying phenomena.  \nI am especially fortunate to have spent eight wonderful years in Oxford—one of the gems of this place is the opportunity to meet so many friendly and inspiring people. My interest in reinforcement learning began here with a master’s project with Varun exploring more theoretical questions. After that, I am profoundly indebted to my many fellow students and postdocs who took the time to mentor me and introduced me to many separate fields of machine learning that I still work on today. Being able to work with such close friends throughout the DPhil has been one of the highest privileges I could have imagined.  \nI want to thank Tarun, Christian, and Wendelin for a fun project introducing me to deep and multi-agent RL. I want to thank Luisa for introducing me to meta-RL and also for her perspective and knowledge guiding me through the early days. I want to thank Tim for introducing me to more Bayesian and theoretical reinforcement learning, and also for endless fun discussions on maths and for always providing me with incredibly useful advice. I want to thank Phil and Jack for an awesome journey through research, which this thesis is primarily composed of—Jack particularly for his tirelessness, optimism, and vision for our ideas (also cameos of his dog), and Phil for being my partner in crime in realizing all of those ideas and making that magic happen. I also want to thank my collaborators with whom we had many fun and fruitful discussions including Leo, Max, Kristian, Sebastian, Xingchen, Robin, Vu, Guneet, Anya, Silvia, Gunshi, Karmesh, Matt, and Mikey.  \nI benefited a lot from two internships in industry, showing me how my research might be applied to real-world problems. I spent three lovely months at Microsoft Research  \nin Cambridge working on automated game testing, and I want to particularly thank Raluca, Johan, and Sam for hosting me. I also spent a great seven months working at Waymo Research in Oxford, working on simulation agents for autonomous driving. I’m particularly thankful to Max for hosting me and to Jack, João, Angad, Kyriacos, and Shimon for many useful discussions.  \nMy DPhil was in the EPSRC Centre for Doctoral Training in Autonomous Intelligent Machines and Systems, and I’m extremely fortunate to have been funded through this program. I’m glad to have gone through the journey with my cohort; special thanks as well to Wendy, who went above and beyond to support us all. I am also thankful to Yarin and Shimon for assessing my transfer and confirmation and providing lots of advice throughout. Finally, I am deeply grateful to Professors Jakob Foerster and Tim Rocktäschel for the opportunity to discuss my work during my DPhil viva and for substantial feedback in the process.  \nOn the personal side, I’m grateful to all the friends I made at Balliol College and beyond, far too many to list, and the fun times I had on the committee. In particular, my friends in Block C with whom we went through all the trials and tribulations of the pandemic—without you, I would have surely lost my sanity. Iam grateful to my parents for their unconditional love and support. Finally, I’m extremely grateful to A","cbCaibgT2aladvrH","https://ap.wps.com/l/cbCaibgT2aladvrH","pdf",16684127,1,229,"English","en",105,"# Acknowledgements\n# Originality Statement\n# Abstract\n## Synthetic data and environments for RL\n## Offline-to-online transfer and augmented world models","[{\"question\":\"为什么深度强化学习仍然面临挑战？\",\"answer\":\"尽管深度强化学习在多个领域取得成功，它通常存在样本效率低、对新环境泛化能力差以及训练不稳定等问题。\"},{\"question\":\"本论文如何利用合成环境与合成数据提升强化学习？\",\"answer\":\"论文关注通过生成合成数据与环境来训练强化学习智能体，从而拓展智能体的经验并改善学习效率。\"},{\"question\":\"论文如何利用已有离线数据减少对在线采集的依赖？\",\"answer\":\"论文提出可利用预存数据的原则性技术，用于降低或替代昂贵的在线数据收集需求。\"}]","Towards Efficient and Robust Reinforcement Learning via Synthetic Environments and Offline Data | PDF",1785807504,577,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"towards-efficient-and-robust-reinforcement-learning-via-synthetic-environments-and-offline-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/towards-efficient-and-robust-reinforcement-learning-via-synthetic-environments-and-offline-data/121888/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么深度强化学习仍然面临挑战？","Question",{"text":75,"@type":76},"尽管深度强化学习在多个领域取得成功，它通常存在样本效率低、对新环境泛化能力差以及训练不稳定等问题。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"本论文如何利用合成环境与合成数据提升强化学习？",{"text":80,"@type":76},"论文关注通过生成合成数据与环境来训练强化学习智能体，从而拓展智能体的经验并改善学习效率。",{"name":82,"@type":73,"acceptedAnswer":83},"论文如何利用已有离线数据减少对在线采集的依赖？",{"text":84,"@type":76},"论文提出可利用预存数据的原则性技术，用于降低或替代昂贵的在线数据收集需求。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]