[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123387-en":3,"doc-seo-123387-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123387,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Convergence Analysis of Policy Gradient Methods with Dynamic Stochasticity","Policy gradient (PG) methods are widely used in reinforcement learning for continuous control, optimizing stochastic (hyper)policies through exploration in action or parameter space. Many existing convergence results target deterministic policies but assume a fixed stochasticity level chosen using the target suboptimality, which mismatches real practice where stochasticity is often scheduled dynamically. This work develops theoretical guarantees for such dynamic stochasticity by defining PES and analyzing convergence under gradient domination.","Convergence Analysis of Policy Gradient Methods with Dynamic Stochasticity  \nAlessandro Montenegro 1 Marco Mussi 1 Matteo Papini 1 Alberto Maria Metelli 1  \nAbstract  \nPolicy gradient (PG) methods are effective reinforcement learning (RL) approaches, particularly for continuous problems. While they optimize stochastic (hyper)policies via action- or parameter-space exploration, real-world applications often require deterministic policies. Existing PG convergence guarantees to deterministic policies assume a fixed stochasticity in the (hyper)policy, tuned according to the desired final suboptimality, whereas practitioners commonly use a dynamic stochasticity level. This work provides the theoretical foundations for this practice.  \nWe introduce PES, a phase-based method that reduces stochasticity via a deterministic schedule while running PG subroutines with fixed stochasticity in each phase. Under gradient domination assumptions, PES achieves last-iterate convergence to the optimal deterministic policy with a saally,camplealyzextyhe oordemmppϵr5tqic. Ae, itirmoend  \nSL-PG, of jointly learning stochasticity (via an appropriate parameterization) and (hyper)policy parameters. We show that SL-PG also ensures asto th-aptte coimanlvsgechcstw(ithhyaperreolpyϵ3nqly,,br  \nquiring stronger assumptions compared to PES.  \n1. Introduction  \nAmong reinforcement learning (RL, Sutton & Barto, 2018) approaches, policy gradient (PG, Deisenroth et al., 2013) methods achieved significant success in addressing realworld scenarios thanks to their ability to handle continuous state and action spaces (Peters & Schaal, 2006), resilience to sensor and actuator noise (Gravell et al., 2020), and robustness in partially-observable environments (Azizzadenesheliet al., 2018) . Additionally, they enable the incorporation of  \n1Politecnico di Milano, Piazza Leonardo Da Vinci 32, 20133, Milan, Italy. Correspondence to: Alessandro Montenegro \u003C[alessandro.montenegro@polimi.it](alessandro.montenegro@polimi.it) >.  \nProceedings of the 42 nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025 . Copyright 2025 by the author(s) .  \nexpert knowledge in the policy design phase (Ghavamzadeh & Engel, 2006), improving the efficacy, safety, and interpretability of the learned policy (Peters & Schaal, 2008) .  \nPG methods optimize directly over the parameter space of parametric policies in order to improve a performance function (e.g., the expected return) . In RL, addressing the exploration problem is crucial. Agents must try different actions to gather information on long-term outcomes, rather than solely maximizing immediate rewards. In PGs, exploration is typically achieved by injecting noise into either the agent’s actions or the policy parameters. These two exploration strategies are known as action-based (AB) and parameter-based (PB) exploration (Metelli et al., 2018), respectively. In particular, AB exploration, whose prototypical algorithms are REINFORCE (Williams, 1992) and GPOMDP (Baxter & Bartlett, 2001), keeps the exploration at the action level by leveraging stochastic policies (e.g., Gaussian) . Instead, PB approaches, whose prototype is PGPE (Sehnke et al., 2010), explore at the parameter level via stochastic hyperpolicies, used to sample the parameters of an underlying (typically deterministic) policy.  \nFrom a theoretical perspective, significant work has focused on the convergence guarantees of PG methods, particularly for AB exploration (Zhao et al., 2011 ; Papini et al., 2018 ; Yuan et al., 2022 ; Fatkhullin et al., 2023 ; Bhandari & Russo, 2024) . However, these methods produce parameters of stochastic (hyper)policies 1 that often fail to meet reliability, safety, and traceability requirements in real-world applications. The PG literature traditionally addressed learning deterministic policies through deterministic policy gradient algorithms (DPG, Silver et al., 2014), which inspired successful deep RL methods (e.g., DDPG, Li","cbCaifNQthfXJgnX","https://ap.wps.com/l/cbCaifNQthfXJgnX","pdf",1321484,1,47,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why do standard policy gradient convergence results not match real-world practice?\",\"answer\":\"They typically assume a fixed stochasticity level throughout learning, tuned to a target suboptimality. In practice, practitioners often change stochasticity during training, creating a theory-practice gap.\"},{\"question\":\"What is the main idea behind the proposed PES method?\",\"answer\":\"PES uses a phase-based schedule that reduces stochasticity over time while running PG subroutines with fixed stochasticity within each phase.\"},{\"question\":\"What additional approach does the work discuss beyond PES?\",\"answer\":\"It introduces SL-PG, which jointly learns stochasticity (via an appropriate parameterization) and (hyper)policy parameters, but requires stronger assumptions than PES.\"}]","Convergence Analysis of Policy Gradient Methods with Dynamic Stochasticity | PDF",1785816240,118,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"convergence-analysis-of-policy-gradient-methods-with-dynamic-stochasticity","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/convergence-analysis-of-policy-gradient-methods-with-dynamic-stochasticity/123387/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do standard policy gradient convergence results not match real-world practice?","Question",{"text":75,"@type":76},"They typically assume a fixed stochasticity level throughout learning, tuned to a target suboptimality. In practice, practitioners often change stochasticity during training, creating a theory-practice gap.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the main idea behind the proposed PES method?",{"text":80,"@type":76},"PES uses a phase-based schedule that reduces stochasticity over time while running PG subroutines with fixed stochasticity within each phase.",{"name":82,"@type":73,"acceptedAnswer":83},"What additional approach does the work discuss beyond PES?",{"text":84,"@type":76},"It introduces SL-PG, which jointly learns stochasticity (via an appropriate parameterization) and (hyper)policy parameters, but requires stronger assumptions than PES.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]