[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117632-en":3,"doc-seo-117632-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117632,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Abstraction for Bayesian Reinforcement Learning in Factored POMDPs","Bayesian reinforcement learning addresses the exploration–exploitation trade-off in partially observable Markov decision processes by maintaining a belief over initially unknown dynamics and rewards. While scaling Bayesian reinforcement learning to large problems remains difficult, existing factored and planning approaches often preserve belief-space factors that minimally affect optimal control. This work argues that reinforcement learning prioritizes policy quality over recovering the exact ground-truth model. It proposes integrating abstraction with online planning for factored POMDPs, reducing model size to enable more simulations and improving performance by providing stronger statistics under a fixed simulation budget.","This is an electronic reprint of the original article.  \nThis reprint may differ from the original in pagination and typographic detail.  \nStarre, Rolf A. N. ; Katt, Sammie; Çelikok, Mustafa Mert; Loog, Marco; Oliehoek, Frans A.  \nAbstraction for Bayesian Reinforcement Learning in Factored POMDPs  \nPublished in:  \nTransactions on Machine Learning Research  \nPublished: 01/01/2025  \nDocument Version  \nPublisher's PDF, also known as Version of record  \nPublished under the following license:  \nCC BY  \nPlease cite the original version:  \nStarre, R. A. N. , Katt, S. , Çelikok, M. M. , Loog, M. , & Oliehoek, F. A. (2025) . Abstraction for Bayesian Reinforcement Learning in Factored POMDPs. Transactions on Machine Learning Research, July-2025 . [https://openreview.net/forum?id=HHgdT6m9L9](https://openreview.net/forum?id=HHgdT6m9L9)  \nThis material is protected by copyright and other intellectual property rights, and duplication or sale of all or part of any of the repository collections is not permitted, except that material may be duplicated by you foryour research use or educational purposes in electronic or print form. You must obtain permission for anyother use. Electronic or print copies may not be offered, whether for sale or otherwise to anyone who is not an authorised user.  \nAbstraction for Bayesian Reinforcement Learning in Factored POMDPs  \nRolf A. N. Starre  \nDelft University of Technology  \nSammie Katt  \nAalto University  \nMustafa Mert Çelikok  \nDelft University of Technology  \nMarco Loog  \nRadbout University  \nFrans A. Oliehoek  \nDelft University of Technology  \n[r. a. n.starre@tudelft.nl](r. a. n.starre@tudelft.nl)  \n[sammie.katt@aalto.fi](sammie.katt@aalto.fi)  \n[m. m. celikok@tudelft.nl](m. m. celikok@tudelft.nl)  \n[marco.loog@ru.nl](marco.loog@ru.nl)  \n[f. a. oliehoek@tudelft.nl](f. a. oliehoek@tudelft.nl)  \nReviewed on OpenReview: [https: // openreview. net/ forum? id= HHgdT6m9L9](https: // openreview. net/ forum? id= HHgdT6m9L9)  \nAbstract  \nBayesian reinforcement learning provides an elegant solution to addressing the exploration–exploitation trade-off in Partially Observable Markov Decision Processes (POMDPs) when the environment’s dynamics and reward function are initially unknown. By maintaining a belief over these unknown components and the state, the agent can effectively learn the environment’s dynamics and optimize their policy. However, scaling Bayesian reinforcement learning methods to large problems remains to be a significant challenge. While prior work has leveraged factored models and online sample-based planning to address this issue, these approaches often retain unnecessarily complex models and factors within the belief space that have minimal impact on the optimal policy. While this complexity might be necessary for accurate model learning, in reinforcement learning, the primary objective is not to recover the ground truth model but to optimize the policy for maximizing the expected sum of rewards. Abstraction offers a way to reduce model complexity by removing factors that are less relevant to achieving high rewards. In this work, we propose and analyze the integration of abstraction with online planning in factored POMDPs. Our empirical results demonstrate two key benefits. First, abstraction reduces model size, enabling faster simulations and thus more planning simulations within a fixed runtime. Second, abstraction enhances performance even with a fixed number of simulations due to greater statistical strength. These results underscore the potential of abstraction to improve both the scalability and effectiveness of Bayesian reinforcement learning in factored POMDPs.  \n1 Introduction  \nDeep reinforcement learning methods have achieved significant milestones, such as attaining superhuman performance on Atari games with only 100k frames (Ye et al., 2021), solving highly complex games such as Go (Silver et al., 2016), and achieving high performance in simulated control tasks (Haarnoja et al., 201","cbCaibYO87HHU2Zd","https://ap.wps.com/l/cbCaibYO87HHU2Zd","pdf",2919051,1,35,"English","en",105,"# Abstract\n# 1 Introduction","[{\"question\":\"What problem does Bayesian reinforcement learning solve in POMDPs?\",\"answer\":\"It manages the exploration–exploitation trade-off when environment dynamics and reward functions are initially unknown by maintaining a belief over these elements and the state.\"},{\"question\":\"Why is scaling Bayesian reinforcement learning difficult?\",\"answer\":\"Large problems create significant computational and model-complexity challenges, and prior methods often keep belief factors that do not meaningfully contribute to the optimal policy.\"},{\"question\":\"How does abstraction improve performance in the proposed approach?\",\"answer\":\"Abstraction reduces model size to speed up simulations and enables more planning within a fixed runtime, and it also improves performance under a fixed number of simulations by increasing statistical strength.\"}]","Abstraction for Bayesian Reinforcement Learning in Factored POMDPs | PDF",1785677481,88,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"abstraction-for-bayesian-reinforcement-learning-in-factored-pomdps","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/abstraction-for-bayesian-reinforcement-learning-in-factored-pomdps/117632/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does Bayesian reinforcement learning solve in POMDPs?","Question",{"text":76,"@type":77},"It manages the exploration–exploitation trade-off when environment dynamics and reward functions are initially unknown by maintaining a belief over these elements and the state.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why is scaling Bayesian reinforcement learning difficult?",{"text":81,"@type":77},"Large problems create significant computational and model-complexity challenges, and prior methods often keep belief factors that do not meaningfully contribute to the optimal policy.",{"name":83,"@type":74,"acceptedAnswer":84},"How does abstraction improve performance in the proposed approach?",{"text":85,"@type":77},"Abstraction reduces model size to speed up simulations and enables more planning within a fixed runtime, and it also improves performance under a fixed number of simulations by increasing statistical strength.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]