[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128560-en":3,"doc-seo-128560-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128560,549768064622,"Anda","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","Planning with Learned Ignorance-Aware Models - Doctor of Philosophy Dissertation","One aim of artificial intelligence is to build decision-makers that improve from experience through data collected by interacting with environments. World models help agents plan and make counterfactual predictions without additional interaction, but perfect-model planning fails to scale beyond problems where designers can specify accurate dynamics. Learning imperfect models from finite data often creates epistemic uncertainty, where naive planning can become catastrophic under out-of-training distributions. This thesis proposes ignorance-aware agents that plan with learned models using knowledge-equivalent, ignorance-augmented objectives, validated across imitation, social learning, and reinforcement learning on simulated driving, continuous control, video games, and small grid-world benchmarks.","Planning with Learned Ignorance-Aware Models  \nAngelos Filos  \nDepartment of Computer Science  \nUniversity of Oxford  \nThis dissertation is submitted for the degree of  \nDoctor of Philosophy  \nWorcester College Trinity 2022  \nTo my family, Panagiota, Maria, and Paschalis  \n♡  \nAcknowledgments  \nThis thesis would not have been enjoyable or even possible without the support of fantastic individuals and institutions. First and foremost, I would like to express my gratitude to my advisor, Yarin Gal, for providing support and essential feedback throughout the years. Yarin, you have significantly influenced my research while allowing me the time and independence to explore my own path, which has helped me develop as an independent researcher. Moreover, you have introduced me to an array of remarkable researchers and significant opportunities. Your consistent mentoring and encouragement have been invaluable to me. I would also like to thank Sergey Levine, who acted as a co-advisor. Sergey, thank you for always providing helpful advice and insight. I greatly enjoyed our collaboration, even though the COVID-19 lockdown disrupted my visit to your lab. I would like to thank Edward Grefenstette and Jakob Foerster for carefully examining this thesis and providing constructive feedback. I thank Rowan McAllister, Nicholas Rhinehart, Greg Farquhar, Natasha Jaques, Andre Barreto, and Amy Zhang, for being amazing collaborators and mentors. For the work in this thesis, I was lucky to collaborate also with Panos Tigas and Clare Lyle. Thank you for your contributions – I could not have done it without you. I also want to thank the rest of OATML: Sebastian Farquhar, Joost van Amersfoort, Aidan Gomez, Milad Alizadeh, Tim G. J. Rudner, Lewis Smith, Andreas Kirsch, Binxin Ru, Neil Band, Andrew Jesson, Jannik Kossen, Lisa Schut, Gunshi Gupta, and Muhammed T. Razzak, for the interesting discussions and fun times. I would also like to thank all my colleagues at J.P. Morgan: Gregory Sidier, Louis Moussu, Samuel Assefa, Cyrine Chtourou, Mahmoud Mahfouz, Giorgio Vit, Joshua Lockhart, Hans Buehler, Manuela Veloso, Jacobo Roa-Vicens, Tucker Balch. For my colleagues at DeepMind: Andre Barreto, Greg Farquhar, Eszter Vertes, Feryal Behbahani, Zita Marinho, Simon Osindero, Kate Baumli, Matteo Hessel, Hado van Hasselt, David Silver, Diana Borsa, Abram Friesen, Tom Schaul, Ioannis Antonoglou, Wilka Carvalho Carvalho and the RL team, thank you for making me feel so welcome despite the physical distance, and being invested in my internship project. I thank Marilena for assuming and delivering the challenging role of ”partner in DPhil during a pandemic” and for her unwavering support. Last but not least, I want to thank my family: Panagiota, Maria, and Pascalis, for always believing in me and encouraging me to be who I am.  \nWalton Street, Oxford, March 2022  \nDeclaration of originality  \nI, Angelos Filos hereby declare that except where specific reference is made to the work of others, the contents of this dissertation are original and have not been submitted in whole or in part for consideration for any other degree or qualification in this, or any other university. This dissertation is my own work and contains nothing which is the outcome of work done in collaboration with others, except as specified in the text and Acknowledgements.  \nAngelos Filos Trinity 2022  \nAbstract  \nOne of the goals of artificial intelligence research is to create decision-makers (i.e. , agents) that improve from experience (i.e. , data), collected through interaction with an environment. Models of the environment (i.e. , world models) are an explicit way that agents use to represent their knowledge, enabling them to make counterfactual predictions and plans without requiring additional environment interactions. Although agents that plan with a perfect model of the environment have led to impressive demonstrations, e.g., superhuman performance in board games, they are limited to problems t","cbCaij6AhppO6YqP","https://ap.wps.com/l/cbCaij6AhppO6YqP","pdf",23807621,1,190,"English","en",105,"# Acknowledgments\n# Declaration of originality\n# Abstract\n# Ignorance-aware planning with learned world models\n## Motivation and challenges of learned models\n## Knowledge-equivalent objectives to address overoptimisation","[{\"question\":\"Why do learned world models create planning difficulties for agents?\",\"answer\":\"Learned models are generally imperfect due to finite data. Naive planning can exploit model errors under out-of-training distributions, leading to catastrophic failures.\"},{\"question\":\"What is the central approach of this thesis?\",\"answer\":\"The thesis proposes ignorance-aware agents that plan with learned models while quantifying lack of knowledge (ignorance or epistemic uncertainty) through ignorance-augmented objectives called knowledge equivalents.\"},{\"question\":\"How are the proposed ideas evaluated?\",\"answer\":\"The methods are tested in multiple settings including imitation learning from expert demonstrations, social learning from sub-optimal demonstrations, and reinforcement learning with rewards. Evidence is based on simulated autonomous driving, continuous control, video games, and small grid-worlds from pixels.\"}]","Planning with Learned Ignorance-Aware Models - Doctor of Philosophy Dissertation | PDF",1786001754,479,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"planning-with-learned-ignorance-aware-models-doctor-of-philosophy-dissertation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/planning-with-learned-ignorance-aware-models-doctor-of-philosophy-dissertation/128560/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do learned world models create planning difficulties for agents?","Question",{"text":76,"@type":77},"Learned models are generally imperfect due to finite data. Naive planning can exploit model errors under out-of-training distributions, leading to catastrophic failures.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the central approach of this thesis?",{"text":81,"@type":77},"The thesis proposes ignorance-aware agents that plan with learned models while quantifying lack of knowledge (ignorance or epistemic uncertainty) through ignorance-augmented objectives called knowledge equivalents.",{"name":83,"@type":74,"acceptedAnswer":84},"How are the proposed ideas evaluated?",{"text":85,"@type":77},"The methods are tested in multiple settings including imitation learning from expert demonstrations, social learning from sub-optimal demonstrations, and reinforcement learning with rewards. Evidence is based on simulated autonomous driving, continuous control, video games, and small grid-worlds from pixels.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]