[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128380-en":3,"doc-seo-128380-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128380,962085564807,"Aurelia","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","Minimax-Bayes Reinforcement Learning - Abstract","Minimax-Bayes reinforcement learning addresses decision making under uncertainty when selecting a prior distribution over Markov decision processes is unclear. The approach models nature as an adversary that chooses the worst-case prior while the agent selects an (adaptive, history-dependent) policy to maximize expected utility. The work studies theoretical and algorithmic properties, identifies when solutions exist, provides regret and value-game analysis, and develops approximate minimax policies with robustness advantages over standard maximum-entropy or uniform priors.","Minimax-Bayes Reinforcement Learning  \nThomas Kleine Buening∗ University of Oslo  \nChristos Dimitrakakis∗ University of Neuchatel  \nHannes Eriksson∗ Zenseact  \nDivya Grover∗ Chalmers University of Technology  \nEmilio Jorge∗ Chalmers University of Technology  \nAbstract  \nWhile the Bayesian decision-theoretic framework offers an elegant solution to the problem of decision making under uncertainty, one question is how to appropriately select the prior distribution. One idea is to employ a worst-case prior.  \nHowever, this is not as easy to specify in sequential decision making as in simple statistical estimation problems. This paper studies (sometimes approximate) minimax-Bayes solutions for various reinforcement learning problems to gain insights into the properties of the corresponding priors and policies. We find that while the worstcase prior depends on the setting, the corresponding minimax policies are more robust than those that assume a standard (i.e. uniform) prior.  \n1 Introduction  \nReinforcement learning is the problem of an agent learning how to act in an unknown environment through interaction and reinforcement. In the standard setting, the learning agent acts in an unknown Markov Decision Process µ, within some class of MDPs M. The agent observes the state st ∈ S of the MDP and selects an action at ∈ A using a policy π . It then observes a reward rt ∈ R and the next state st+1 . The agent’s goal is to maximise utility, defined as the sum of rewards to some horizon T, u = P rt , in expectation, i.e. Eπµ(u), where Eπµ is the expectation under the MDP and policy. Since the true µ is unknown, this optimisation problem is ill-posed. In the Bayesian setting, this conundrum is solved by selecting some subjective prior distribution β over MDPs and  \nProceedings of the 26th International Conference on Artificial Intelligence and Statistics (AISTATS) 2023, Valencia, Spain. PMLR: Volume 206 . Copyright 2023 by the author(s) .  \n∗Authors contributed equally to this work.  \nmaximising Eπβ(u) = RM Eπµ(u)dβ(µ) . Then it remains to compute the optimal adaptive (i.e. history-dependent) policy, something that can be only done approximately in general, due to the fact that the number of adaptive policies increases exponentially with the problem horizon.  \nThe above discussion assumes that the agent has somehow chosen a prior. However, it is not clear how such a prior can be selected from first principles, if we have no domain knowledge, but still want to be robust. The minimax-Bayes idea (Berger, 1985) is to assume that nature selects the worst possible prior β ∗ for the agent, but without knowledge of the agent’s policy. This can be formalised by having nature play the minimising player ina simultaneous-move zero-sum game defined by the expected utility Eπβ(u), where the agent (who maximises) chooses π, and nature (who minimises) chooses β . In simple Bayesian decision problems (e.g. linear regression) the minimax-Bayes problem is well-studied and β∗ sometimes corresponds to a maximum entropy prior. However, in an interactive setting, results are limited to one-shot experiment design (Gr¨unwald and Dawid, 2004), which shows that maximum entropy priors are not the worst-case priors generally.  \nIn reinforcement learning, which can be seen as a sequential generalisation of one-shot experiment design, this problem has not received much attention in the past. Sometimes, the concept of maximum entropy has been used in reinforcement learning as a penalty term on the policy (e.g. Todorov, 2006; Haarnoja et al., 2018; Eysenbach and Levine, 2021) as well as in the context of inverse reinforcement learning (Ziebart, 2010), but an explicit connection to the minimax-Bayes literature has not been made. In preliminary work, Androulakis and Dimitrakakis (2014) analysed variants of the weighted majority algorithm for finding minimax priors in a restricted version of this setting.  \nContributions. In this paper, we study the basic theoretical and al","cbCainr77JbxcNEb","https://ap.wps.com/l/cbCainr77JbxcNEb","pdf",1168211,2,1,17,"English","en",105,"# Introduction\n# Setting\n## Markov Decision Process formulation\n# Regret and optimality relationships\n# Value existence in the Bayesian agent vs. Nature game\n# Algorithms for approximately minimax policies\n## Finite-horizon Bayes-optimal policies\n## Posterior sampling policies\n## Parametrised adaptive policies\n# Related work and conclusions","[{\"question\":\"What problem does the minimax-Bayes framework target in reinforcement learning?\",\"answer\":\"It targets robust decision making when the prior distribution over MDPs is uncertain or cannot be selected reliably from first principles.\"},{\"question\":\"How is the minimax-Bayes interaction between the agent and nature formulated?\",\"answer\":\"It is posed as a simultaneous-move zero-sum game: the agent maximizes expected utility by choosing a policy π, while nature minimizes it by choosing a prior β.\"},{\"question\":\"Why are minimax-Bayes policies described as more robust than standard-prior Bayesian RL policies?\",\"answer\":\"The paper finds that although the worst-case prior depends on the setting, the resulting minimax policies are generally more robust than policies derived under standard maximum-entropy or uniform priors.\"}]","Minimax-Bayes Reinforcement Learning - Abstract | PDF",1785947190,43,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"minimax-bayes-reinforcement-learning-abstract","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/minimax-bayes-reinforcement-learning-abstract/128380/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the minimax-Bayes framework target in reinforcement learning?","Question",{"text":76,"@type":77},"It targets robust decision making when the prior distribution over MDPs is uncertain or cannot be selected reliably from first principles.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is the minimax-Bayes interaction between the agent and nature formulated?",{"text":81,"@type":77},"It is posed as a simultaneous-move zero-sum game: the agent maximizes expected utility by choosing a policy π, while nature minimizes it by choosing a prior β.",{"name":83,"@type":74,"acceptedAnswer":84},"Why are minimax-Bayes policies described as more robust than standard-prior Bayesian RL policies?",{"text":85,"@type":77},"The paper finds that although the worst-case prior depends on the setting, the resulting minimax policies are generally more robust than policies derived under standard maximum-entropy or uniform priors.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]