[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122142-en":3,"doc-seo-122142-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122142,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Mildly Constrained Evaluation Policy for Offline Reinforcement Learning","Offline reinforcement learning (RL) enforces policy constraints to closely follow the behavior policy, improving the stability of value learning and reducing the risk of selecting out-of-distribution (OOD) actions during test-time inference. Existing methods typically use the same constraint strength for both value estimation and action selection. The proposed Mildly Constrained Evaluation Policy (MCEP) relaxes constraints for inference while keeping more constrained targets for value estimation, and is compatible as a plug-in with prior offline RL methods. Experiments on D4RL MuJoCo locomotion, high-dimensional humanoid, and 16 robotic manipulation tasks show significant gains and further improvements over SOTA, with open-sourced code.","Mildly Constrained Evaluation Policy for Offline Reinforcement Learning  \nLinjie Xu [linjie.xu@qmul. ac.uk](linjie.xu@qmul. ac.uk)  \nQueen Mary University of London  \nZhengyao Jiang [z.jiang@cs.ucl. ac.uk](z.jiang@cs.ucl. ac.uk)  \nUniversity College London  \nJinyu Wang, Lei Song and Jiang Bian {wang.jinyu, [lei.song](lei.song), [jiang.bian}@microsoft. com](jiang.bian}@microsoft. com)  \nMicrosoft Research Asia  \nReviewed on OpenReview: [https: // openreview. net/ forum? id= imAROs79Pb](https: // openreview. net/ forum? id= imAROs79Pb)  \nAbstract  \nOffline reinforcement learning (RL) methodologies enforce constraints on the policy to adhere closely to the behavior policy, thereby stabilizing value learning and mitigating the selection of out-of-distribution (OOD) actions during test time. Conventional approaches apply identical constraints for both value learning and test time inference. However, our findings indicate that the constraints suitable for value estimation may in fact be excessively restrictive for action selection during test time. To address this issue, we propose a Mildly Constrained Evaluation Policy (MCEP) for test time inference with a more constrained target policy for value estimation. Since the target policy has been adopted in various prior approaches, MCEP can be seamlessly integrated with them as a plug-in. We instantiate MCEP based on TD3BC (Fujimoto & Gu, 2021), AWAC (Nair et al., 2020) and DQL (Wang et al. , 2023) algorithms. The empirical results on D4RL MuJoCo locomotion, high-dimensional humanoid and a set of 16 robotic manipulation tasks show that the MCEP brought significant performance improvement on classic offline RL methods and can further improve SOTA methods. The codes are open-sourced at [https://github.com/egg-west/MCEP.git](https://github.com/egg-west/MCEP.git).  \n1 Introduction  \nOffline reinforcement learning (RL) extracts a policy from data that is pre-collected by unknown policies. This setting does not require interactions with the environment thus it is well-suited for tasks where the interaction is costly or risky. Recently, it has been applied to Natural Language Processing (Snell et al. , 2022; Sodhi et al., 2023), e-commerce (Degirmenci & Jones, 2022) and real-world robotics (Kalashnikov et al., 2021; Rafailov et al., 2021; Kumar et al., 2022; Shah et al., 2022; Bhateja et al., 2023) etc. Compared to the standard online setting where the policy gets improved via trial and error, learning with a static offline dataset raises novel challenges. One challenge is the distributional shift between the training data and the data encountered during deployment. To attain stable evaluation performance under the distributional shift, the policy is expected to stay close to the behavior policy. Another challenge is the \"extrapolation error\" (Fujimoto et al., 2019; Kumar et al., 2019) that indicates value estimate error on unseen state-action pairs or Out-Of-Distribution (OOD) actions. Worsely, this error can be amplified with bootstrapping and cause instability of the training, which is also known as deadly-triad (Van Hasselt et al., 2018) . Majorities of model-free approaches tackle these challenges by either constraining the policy to adhere closely to the behavior policy (Wu et al., 2019; Kumar et al., 2019; Fujimoto & Gu, 2021; Wang et al., 2023) or regularising the Q to pessimistic estimation for OOD actions (Kumar et al., 2020; Lyu et al., 2022) . In this work, we focus on policy constraint methods.  \nPolicy constraint methods minimize the disparity between the policy distribution and the behavior distribution. Meanwhile, the strength of policy constraints introduces a tradeoff between stabilizing value estimates and attaining better inference performance. While various policy constraints have been developed to address this tradeoff, it remains a common problem for them that an excessively constrained policy enables stable value estimate but degrades the evaluation performance (Kumar e","cbCaiturwAIidOIP","https://ap.wps.com/l/cbCaiturwAIidOIP","pdf",6043616,1,20,"English","en",105,"# Introduction\n## Offline RL challenges\n## Policy constraint tradeoff\n## Proposed idea and evaluation policy\n## Empirical findings (overview)","[{\"question\":\"What problem does offline reinforcement learning face during deployment?\",\"answer\":\"Offline RL suffers from distributional shift between training data and deployment, and from extrapolation error on unseen state-action pairs or OOD actions, which can destabilize training through bootstrapping.\"},{\"question\":\"How does the proposed Mildly Constrained Evaluation Policy (MCEP) differ from conventional constraint methods?\",\"answer\":\"Conventional approaches apply the same constraints for both value learning and test-time inference, while MCEP uses a more constrained target policy for value estimation and a mildly constrained evaluation policy for test-time action selection.\"},{\"question\":\"What empirical results does MCEP achieve and on which tasks?\",\"answer\":\"On D4RL MuJoCo locomotion, a high-dimensional humanoid, and 16 robotic manipulation tasks, MCEP delivers significant performance improvements on classic offline RL methods and can further improve SOTA methods.\"}]","Mildly Constrained Evaluation Policy for Offline Reinforcement Learning | PDF",1785809047,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mildly-constrained-evaluation-policy-for-offline-reinforcement-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mildly-constrained-evaluation-policy-for-offline-reinforcement-learning/122142/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does offline reinforcement learning face during deployment?","Question",{"text":75,"@type":76},"Offline RL suffers from distributional shift between training data and deployment, and from extrapolation error on unseen state-action pairs or OOD actions, which can destabilize training through bootstrapping.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed Mildly Constrained Evaluation Policy (MCEP) differ from conventional constraint methods?",{"text":80,"@type":76},"Conventional approaches apply the same constraints for both value learning and test-time inference, while MCEP uses a more constrained target policy for value estimation and a mildly constrained evaluation policy for test-time action selection.",{"name":82,"@type":73,"acceptedAnswer":83},"What empirical results does MCEP achieve and on which tasks?",{"text":84,"@type":76},"On D4RL MuJoCo locomotion, a high-dimensional humanoid, and 16 robotic manipulation tasks, MCEP delivers significant performance improvements on classic offline RL methods and can further improve SOTA methods.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]