[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85448-en":3,"doc-seo-85448-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85448,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Beyond Slater’s Condition in Online CMDPs with Stochastic and Adversarial Constraints","Online episodic constrained Markov decision processes (CMDPs) are studied under both stochastic and adversarial constraint feedback. A new algorithm is introduced with stronger guarantees than existing best-of-both-worlds methods: in the stochastic regime it achieves O(√T) regret and constraint violation without requiring Slater’s condition, including cases with no strictly feasible solution. It further controls positive constraint violation. In the adversarial regime it guarantees sublinear constraint violation without Slater’s condition and sublinear α-regret relative to the unconstrained optimum. Synthetic experiments confirm practical effectiveness.","Beyond Slater’s Condition in Online CMDPs with Stochastic and Adversarial Constraints  \nFrancesco Emanuele Stradi [francescoemanuele. stradi@polimi. it](francescoemanuele. stradi@polimi. it)[ ](francescoemanuele. stradi@polimi. it)Politecnico di Milano  \narXiv :2509 .20 1 14v 3 [ cs .LG] 13 Jul 2026  \nEleonora Fidelia Chiefari [eleonorafidelia. chiefari@mail. polimi. it](eleonorafidelia. chiefari@mail. polimi. it)[ ](eleonorafidelia. chiefari@mail. polimi. it)Politecnico di Milano  \nAlberto Marchesi [alberto. marchesi@polimi. it](alberto. marchesi@polimi. it)[ ](alberto. marchesi@polimi. it)Politecnico di Milano  \nMatteo Castiglioni [matteo. castiglioni@polimi. it](matteo. castiglioni@polimi. it)[ ](matteo. castiglioni@polimi. it)Politecnico di Milano  \nNicola Gatti [nicola. gatti@polimi. it](nicola. gatti@polimi. it)[ ](nicola. gatti@polimi. it)Politecnico di Milano  \nJuly 14, 2026  \nAbstract  \nWe study online episodic Constrained Markov Decision Processes (CMDPs) under both stochastic and adversarial constraints. We provide a novel algorithm whose guarantees greatly improve those of the state-of-the-art best-of-both-worlds algorithm introduced by Stradi et al. [2025d] . In the stochastic regime, i.e. , when the constraints are sampled from fixed but unknown distributions, our method achievese  \nO ( √T) regret and constraint violation without relying on Slater’s condition, thereby handling settings where no strictly feasible solution exists. Moreover, we provide guarantees on the stronger notion of positive constraint violation, which does not allow to recover from large violation in the early episodes by playing strictly safe policies. In the adversarial regime, i.e. , when the constraints may change arbitrarily between episodes, our algorithm ensures sublinear constraint violation without Slater’s condition, and achieves sublinear α-regret with respect to the unconstrained optimum, where α is a suitably defined multiplicative approximation factor. We further validate our results through synthetic experiments, showing the practical effectiveness of our algorithm.  \n1 Introduction  \nReinforcement Learning (RL) [Sutton and Barto, 2018] provides a general framework for sequential decisionmaking, where an agent learns to act optimally by interacting with an environment modeled as a Markov Decision Process (MDP) [Puterman, 2014] . While RL has achieved remarkable success in numerous applications, real-world decision-making problems often involve safety and resource constraints that must be respected at every step, leading to the study of Constrained Markov Decision Processes (CMDPs) [Altman, 1999] . CMDPs have been widely employed in safety-critical domains such as autonomous driving [Iseleet al. , 2018 , Wen et al. , 2020], online bidding and advertising [Gummadi et al. , 2012 , Wu et al. , 2018 , He et al. , 2021], and recommendation systems [Singh et al. , 2020], where constraint satisfaction is as crucial as optimizing cumulative reward.  \nCMDPs have been significantly studied within the framework of online learning [Cesa-Bianchi and Lugosi, 2006], where a learner interacts with an environment in a sequential manner and aims to minimize its regret, defined as the difference between the reward attained by the best fixed policy and the learner’s cumulative reward. An algorithm is considered successful if it achieves sublinear regret, meaning that the average regret per round vanishes as the time horizon T grows. Online CMDPs extend this setting by incorporating  \nconstraints on the learner’s behavior, making them a constrained counterpart of classical online learning problems. These algorithms are typically studied under two main assumptions about the environment: in the stochastic setting, rewards (losses) and constraint functions are drawn i.i.d. from an unknown but fixed distribution, while in the adversarial setting they can be chosen arbitrarily by an adversary, potentially depending on past actions.  \nIn the stochastic se","cbCaietAmO2RoE3T","https://ap.wps.com/l/cbCaietAmO2RoE3T","pdf",3593268,3,1,34,"English","en",105,"# Introduction\n## Original Contributions","[{\"question\":\"What problem does the document address in online episodic CMDPs?\",\"answer\":\"It addresses online episodic constrained Markov decision processes where constraints can be either stochastic or adversarial, and the learner must manage reward while controlling constraint violations.\"},{\"question\":\"How does the proposed method improve over best-of-both-worlds algorithms?\",\"answer\":\"It provides improved regret and constraint-violation guarantees compared with prior best-of-both-worlds work, including stronger control of constraint violations without relying on Slater’s condition.\"},{\"question\":\"What guarantees are provided for the stochastic and adversarial regimes?\",\"answer\":\"In the stochastic regime, the method achieves O(√T) regret and constraint violation without Slater’s condition and also bounds positive constraint violation. In the adversarial regime, it ensures sublinear constraint violation without Slater’s condition and achieves sublinear α-regret versus the unconstrained optimum.\"}]",1784203611,86,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"beyond-slaters-condition-in-online-cmdps-with-stochastic-and-adversarial-constraints","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/beyond-slaters-condition-in-online-cmdps-with-stochastic-and-adversarial-constraints/85448/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address in online episodic CMDPs?","Question",{"text":75,"@type":76},"It addresses online episodic constrained Markov decision processes where constraints can be either stochastic or adversarial, and the learner must manage reward while controlling constraint violations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method improve over best-of-both-worlds algorithms?",{"text":80,"@type":76},"It provides improved regret and constraint-violation guarantees compared with prior best-of-both-worlds work, including stronger control of constraint violations without relying on Slater’s condition.",{"name":82,"@type":73,"acceptedAnswer":83},"What guarantees are provided for the stochastic and adversarial regimes?",{"text":84,"@type":76},"In the stochastic regime, the method achieves O(√T) regret and constraint violation without Slater’s condition and also bounds positive constraint violation. In the adversarial regime, it ensures sublinear constraint violation without Slater’s condition and achieves sublinear α-regret versus the unconstrained optimum.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]