[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122922-en":3,"doc-seo-122922-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122922,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Learning Optimal Deterministic Policies with Stochastic Policy Gradients - Research Paper","Policy gradient (PG) methods provide a prominent approach for continuous reinforcement learning by learning stochastic parametric (hyper)policies through exploration in action or parameter spaces. However, stochastic controllers can be impractical due to reduced robustness, safety, and traceability, so practitioners often train stochastically and deploy a deterministic version. The paper proposes a framework to analyze this workflow, proves global convergence to the best deterministic policy under (weak) gradient domination, and studies exploration tuning to balance sample complexity and deployed performance, also comparing action- and parameter-based exploration.","Learning Optimal Deterministic Policies with Stochastic Policy Gradients  \nAlessandro Montenegro 1 Marco Mussi 1 Alberto Maria Metelli 1 Matteo Papini 1  \nAbstract  \nPolicy gradient (PG) methods are successful approaches to deal with continuous reinforcement learning (RL) problems. They learn stochastic parametric (hyper)policies by either exploring in the space of actions or in the space of parameters.  \nStochastic controllers, however, are often undesirable from a practical perspective because of their lack of robustness, safety, and traceability. In common practice, stochastic (hyper)policies are learned only to deploy their deterministic version.  \nIn this paper, we make a step towards the theoretical understanding of this practice. After introducing a novel framework for modeling this scenario, we study the global convergence to the best deterministic policy, under (weak) gradient domination assumptions. Then, we illustrate how to tune the exploration level used for learning to optimize the trade-off between the sample complexity and the performance of the deployed deterministic policy.  \nFinally, we quantitatively compare action-based and parameter-based exploration, giving a formal guise to intuitive results.  \n1. Introduction  \nWithin reinforcement learning (RL, Sutton & Barto, 2018) approaches, policy gradient (PG, Deisenroth et al., 2013) algorithms have proved very effective in dealing with realworld control problems. Their advantages include the applicability to continuous state and action spaces (Peters & Schaal, 2006), resilience to sensor and actuator noise (Gravell et al., 2020), robustness to partial observability (Azizzadenesheli et al., 2018), and the possibility of incorporating prior knowledge in the policy design phase (Ghavamzadeh & Engel, 2006), improving explainability (Likmeta et al., 2020) . PG algorithms search directly in the space of parametric policies for the one that maximizes a performance  \n1Politecnico di Milano, Piazza Leonardo da Vinci 32, 20133, Milan, Italy. Correspondence to: Alessandro Montenegro \u003C[alessandro.montenegro@polimi.it](alessandro.montenegro@polimi.it) >.  \nProceedings of the 41 st International Conference on Machine Learning, Vienna, Austria. PMLR 235, 2024 . Copyright 2024 by the author(s) .  \nfunction. Nonetheless, as always in RL, the exploration problem has to be addressed, and practical methods involve injecting noise in the actions or in the parameters. This limits the application of PG methods in many real-world scenarios, such as autonomous driving, industrial plants, and robotic controllers. Indeed, stochastic policies typically do not meet the reliability, safety, and traceability standards of this kind of applications.  \nThe problem of learning deterministic policies has been explicitly addressed in the PG literature by Silver et al. (2014) with their deterministic policy gradient, which spawned very successful deep RL algorithms (Lillicrap et al., 2016 ; Fujimoto et al., 2018) . This approach, however, is affected by several drawbacks, mostly due to its inherent off-policy nature. First, this makes DPG hard to analyze from a theoretical perspective: local convergence guarantees have been established only recently, and only under assumptions that are very demanding for deterministic policies (Xiong et al., 2022) . Furthermore, its practical versions (DDPG, Lillicrap et al., 2016) are known to be very susceptible to hyperparameter tuning.  \nWe study here a simpler and fairly common approach: that of learning stochastic policies with PG algorithms, then deploying the corresponding deterministic version,“switching off” the noise.1 Intuitively, the amount of exploration (e.g., the variance of a Gaussian policy) should be selected wisely. Indeed, the smaller the exploration level, the closer the optimized objective is to that of a deterministic policy. At the same time, with a small exploration, learning can severely slow down and get stuck on bad local optima.  \nPoli","cbCaibllU4842VXA","https://ap.wps.com/l/cbCaibllU4842VXA","pdf",5772337,1,52,"English","en",105,"# Introduction\n## Policy gradient methods and exploration challenges\n## Deterministic policy gradient background and limitations\n## Stochastic-to-deterministic deployment motivation\n## Exploration types: action-based vs parameter-based\n# Main Contributions and Paper Focus","[{\"question\":\"Why do practitioners train stochastic policies and deploy deterministic versions?\",\"answer\":\"Because stochastic controllers often lack robustness, safety, and traceability required in real applications. Training with noise helps learning, then the deterministic version is used for deployment.\"},{\"question\":\"What theoretical result does the paper target?\",\"answer\":\"It establishes global convergence to the best deterministic policy under (weak) gradient domination assumptions, within a new modeling framework for the stochastic-to-deterministic practice.\"},{\"question\":\"How does exploration level affect learning and final deterministic performance?\",\"answer\":\"Smaller exploration makes the optimized objective closer to the deterministic one, but it can slow learning and lead to poor local optima. The paper analyzes how to tune exploration to optimize the sample-complexity vs performance trade-off.\"}]","Learning Optimal Deterministic Policies with Stochastic Policy Gradients - Research Paper | PDF",1785813678,131,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"learning-optimal-deterministic-policies-with-stochastic-policy-gradients-research-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/learning-optimal-deterministic-policies-with-stochastic-policy-gradients-research-paper/122922/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do practitioners train stochastic policies and deploy deterministic versions?","Question",{"text":75,"@type":76},"Because stochastic controllers often lack robustness, safety, and traceability required in real applications. Training with noise helps learning, then the deterministic version is used for deployment.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What theoretical result does the paper target?",{"text":80,"@type":76},"It establishes global convergence to the best deterministic policy under (weak) gradient domination assumptions, within a new modeling framework for the stochastic-to-deterministic practice.",{"name":82,"@type":73,"acceptedAnswer":83},"How does exploration level affect learning and final deterministic performance?",{"text":84,"@type":76},"Smaller exploration makes the optimized objective closer to the deterministic one, but it can slow learning and lead to poor local optima. The paper analyzes how to tune exploration to optimize the sample-complexity vs performance trade-off.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]