[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84718-en":3,"doc-seo-84718-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84718,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","Progress and Reliability Oriented Group Policy Optimization for Agentic Reinforcement Learning","Group-based reinforcement learning improves large language model agents on long-horizon interactive tasks by refining policy updates beyond trajectory-level optimization. Step-level group RL enables intermediate-step grouping and comparison, but grouping choices trade off coverage with biased comparisons, or fairness with fragmented singleton-heavy batches. ProGPO (Progress-and Reliability-Oriented Group Policy Optimization) is a learned-critic-free method that preserves exact-prefix action comparison while adding progress-sensitive transition credit from rollout-based state potentials. Semantic expansion plus inverse-variance fusion across history depths yields reliable potentials. Evaluations on ALFWorld and WebShop with Qwen2.5 Instruct show consistent gains under similar compute, with Qwen2.5-3B scalability tests.","Progress-and Reliability-Oriented Group Policy  \nOptimization  \nfor Agentic Reinforcement Learning  \nMingxuan Fan Baidu Inc., China  \nPeiyang Liu Peking University, China  \narXiv :2607 .04242v 1 [ cs .AI ] 5 Jul 2026  \nABSTRACT  \nGroup-based reinforcement learning (RL) has become an effective paradigm for improving large language model agents on long-horizon interactive tasks. To obtain finergrained policy updates than trajectory-level optimization, recent work has moved toward step-level group-based RL, where intermediate steps are grouped and compared within a rollout batch. However, step-level advantage estimation is sensitive to how groups are formed: grouping by broad state keys improves coverage but may compare actions taken under different histories, while enforcing historical consistency yields fairer comparisons at the cost of fragmented groups and missing peer-comparison signal. In this paper, we propose ProGPO (Progress-and Reliability-Oriented Group Policy Optimization), a learned-critic-free method for context-consistent step-level learning. ProGPO keeps exact-prefix action comparison, and complements sparse peer comparisons with transition credit derived from rollout-based state potentials. To estimate these potentials reliably, ProGPO combines semantic expansion with inverse-variance fusion across history depths. We evaluate ProGPO on two challenging agentic tasks, ALFWorld and WebShop, with Qwen2.5-1.5B-Instruct. Results show that ProGPO improves over matched agentic RL baselines under comparable computational overhead, and additional Qwen2.5-3B-Instruct experiments further test the scalability of the proposed method.  \n1 Introduction  \nLarge language models (LLMs) are increasingly deployed as agents that perceive, reason, and actin external environments (Brown et al., 2020; Yao et al., 2023; Liu et al., 2024; Wang et al., 2024) . Representative applications include web navigation (Yao et al., 2022; Deng et al., 2023; Zhou et al., 2024; Zheng et al., 2024), embodied household tasks (Shridhar et al., 2021; Ahn et al., 2022; Driess et al., 2023), and tool-augmented reasoning (Schick et al., 2023; Qin et al., 2024; Jin et al., 2025) . Unlike single-turn generation, these tasks require long-horizon planning, recovery from earlier mistakes, and credit assignment under sparse delayed rewards.  \nReinforcement learning (RL) has therefore become a key post-training paradigm for improving model behavior, from human-preference optimization (Christiano et al., 2017; Ouyang et al., 2022; Bai et al., 2022; Rafailov et al., 2023) to reasoning-oriented RL (DeepSeek-AI et al., 2025) . In particular, group-based methods estimate advantages from multiple sampled responses or trajectories for the same task, avoiding a learned critic while retaining scalable policy optimization. GRPO (Shao et al., 2024) and variants such as DAPO (Yu et al., 2025) compare trajectory-level outcomes within a group, while RLOO (Kool et al., 2019) uses leave-one-out baselines. However, these methods are  \nFigure 1: Motivation for ProGPO. Left: rollout trajectories. Middle: step-level grouping by the current state can mix different histories. Right: context-consistent grouping creates smaller peer groups.  \nprimarily designed for single-turn tasks such as mathematical reasoning and code generation, where each response receives a complete outcome and step-level credit is less ambiguous.  \nIn multi-turn agentic tasks, direct trajectory-wise optimization becomes inefficient. Approaches such as RAGEN (Wang et al., 2025) and Search-R1 (Jin et al., 2025) concatenate the full interaction history into a single sequence, causing the effective context length to grow rapidly with the number of turns. To avoid this context explosion, recent step-level methods such as GiGPO (Feng et al., 2025) group repeated intermediate states and compute within-group advantages, enabling finer-grained updates without per-step extra rollouts.  \nNevertheless, step-level grouping intr","cbCaipKRjWC0b3ko","https://ap.wps.com/l/cbCaipKRjWC0b3ko","pdf",1110204,1,13,"English","en",105,"# Introduction\n## Motivation and background\n## Step-level grouping challenges\n## ProGPO overview and key idea\n## Reliability-aware state potential estimation","[{\"question\":\"What problem does ProGPO address in step-level group policy optimization for agentic RL?\",\"answer\":\"It addresses biased action advantages caused by grouping steps that share only the current observation, and the sparse peer-comparison signal that appears when enforcing strict historical consistency.\"},{\"question\":\"How does ProGPO combine action comparison with progress credit without a learned critic?\",\"answer\":\"ProGPO keeps exact-prefix action comparison within context-consistent groups, then adds transition credit using rollout-based state potentials to capture progress.\"},{\"question\":\"How are the state potentials estimated reliably in ProGPO?\",\"answer\":\"ProGPO applies semantic expansion and uses inverse-variance fusion across different history depths, weighting eligible depths by sample size and return variance to reduce noise and blur.\"}]",1784197826,33,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"progress-and-reliability-oriented-group-policy-optimization-for-agentic-reinforcement-learning","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/progress-and-reliability-oriented-group-policy-optimization-for-agentic-reinforcement-learning/84718/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does ProGPO address in step-level group policy optimization for agentic RL?","Question",{"text":75,"@type":76},"It addresses biased action advantages caused by grouping steps that share only the current observation, and the sparse peer-comparison signal that appears when enforcing strict historical consistency.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does ProGPO combine action comparison with progress credit without a learned critic?",{"text":80,"@type":76},"ProGPO keeps exact-prefix action comparison within context-consistent groups, then adds transition credit using rollout-based state potentials to capture progress.",{"name":82,"@type":73,"acceptedAnswer":83},"How are the state potentials estimated reliably in ProGPO?",{"text":84,"@type":76},"ProGPO applies semantic expansion and uses inverse-variance fusion across different history depths, weighting eligible depths by sample size and return variance to reduce noise and blur.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]