[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83875-en":3,"doc-seo-83875-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83875,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","STAPO: Selective Trajectory-Aware Policy Optimization for LLM Agent Training","Reinforcement Learning (RL) is widely used to train Large Language Model (LLM) agents for long-horizon tasks, yet sparse and delayed rewards can cause trajectory neglect, where agents lose the task goal and interaction history at intermediate steps. Existing Shannon-entropy step supervision conflates state complexity with confidence, weakening decision reliability. Normalized entropy separates uncertainty from state-average behavior to better flag neglect-associated outlier steps. STAPO is proposed as a hierarchical group-based RL framework that selectively optimizes these steps using trajectory-aware rewards and trajectory-independent penalties, improving performance on ALFWorld, WebShop, and Search-Augmented QA.","STAPO: Selective Trajectory-Aware Policy Optimization for  \nLLM Agent Training  \nQiuyi Qi♠♢ * , Tian Liang♠♢ * , Mutian Bao♠♢ * , Jinjian Zhang♢ , Dongnan Liu♢ ,  \nWei Zhou♢ , Linjian Mo♢ , Ming Kong♠†, Jie Liu♣†, Feng Zhang♠ , Qiang Zhu♠†  \n♠ Zhejiang University, ♢ Ant Group, ♣ City University of Hong Kong  \n{qiqiuyi,[zhuq}@zju.edu.cn](zhuq}@zju.edu.cn)  \narXiv :2607 .04963v 1 [ cs .AI] 6 Jul 2026  \nAbstract  \nReinforcement Learning (RL) is the dominant paradigm for training Large Language Model (LLM) agents on long-horizon tasks. However, sparse and delayed rewards often lead to trajectory neglect, in which agents lose focus on the task goal and interaction history at intermediate steps. Prior work has explored step-level supervision using Shannon-entropy–based uncertainty signals, which conflate inherent state complexity with agent confidence and therefore provide unreliable estimates of decision reliability. To address this issue, we propose normalized entropy, which measures confidence deviations relative to an agent’s average behavior under a given state, thereby strengthening the association between low-quality actions and trajectory neglect. Building on this insight, we introduce Selective Trajectory-Aware Policy Optimization (STAPO), a hierarchical groupbased RL framework. STAPO leverages normalized entropy to locate outlier steps associated with trajectory neglect and optimizes them via a joint mechanism of trajectory-aware reward and trajectory-independent penalty, enhancing trajectory awareness while preserving training stability. Extensive experiments on ALFWorld, WebShop, and Search-Augmented QA demonstrate that STAPO achieves state-ofthe-art performance while substantially alleviating trajectory neglect, validating its effectiveness and robustness for agentic tasks.  \n1 Introduction  \nLarge Language Models (LLMs) have evolved from static text generators to autonomous agents  \n* Q. Qi, T. Liang and M. Bao contributed equally to this work.  \n†Q. Zhu, M. Kong and J. Liu are corresponding authors.  \nQ. Zhu is with the College of Artificial Intelligence, Shanghai Institute for Advanced Study, Zhejiang University. M. Kong is with the School of Earth Sciences, Zhejiang University. J. Liu is with the Department of Computer Science, City University of Hong Kong.  \nTime Step (t)  \nFigure 1: Illustration of the differences between Shannon entropy and normalized entropy in locating trajectory neglect.  \ncapable of executing complex, long-horizon tasks, such as embodied control (Shridhar et al., 2021 ; Li et al., 2024) and web navigation (Furuta et al., 2024 ; Zheng et al., 2024 ; Gou et al., 2025) .  \nTo enable long-horizon reasoning in LLM agents, recent work has increasingly adopted Reinforcement Learning (RL) for post-training. In particular, group-based RL algorithms such as RLOO (Kool et al., 2019 ; Ahmadian et al., 2024) and GRPO (Shao et al., 2024) estimate advantages via group sampling, avoiding explicit critic training while improving scalability and training stability. However, training agents for long-horizon tasks remains challenging due to the temporal credit assignment problem induced by sparse and delayed rewards. As illustrated in Figure 1, agents often exhibit trajectory neglect, where the model loses focus on the task goal and interaction history during extended interaction sequences, leading to lowquality actions at certain intermediate steps (see Appendix A for detailed case studies) .  \nWhile recent adaptations like GiGPO (Feng et al., 2025) and RLVMR (Zhang et al., 2025) attempt to mitigate the credit assignment issue by achieving fine-grained step-level advantage estimation, they typically optimize all steps indiscriminately, lacking a selective mechanism to precisely locate and address these outlier steps caused by trajectory neglect. Furthermore, several works have explored leveraging internal uncertainty signals to assist training, including entropy-based regularization and supervision strategies (","cbCaigDwtzXznL4p","https://ap.wps.com/l/cbCaigDwtzXznL4p","pdf",2368327,4,1,22,"English","en",105,"# Introduction\n## Trajectory neglect and credit assignment\n## Limitations of Shannon entropy\n## Normalized entropy metric\n## STAPO framework and selective optimization","[{\"question\":\"What problem does STAPO address in LLM agent training?\",\"answer\":\"STAPO targets trajectory neglect, where an agent loses focus on the task goal and interaction history during extended sequences, leading to low-quality actions at intermediate steps.\"},{\"question\":\"Why is Shannon entropy considered an unreliable signal for locating trajectory neglect?\",\"answer\":\"Shannon entropy conflates inherent state complexity with agent confidence, so high-entropy steps may reflect genuinely complex states rather than low decision reliability.\"},{\"question\":\"How does normalized entropy improve outlier-step detection?\",\"answer\":\"Normalized entropy aggregates actions from the same state across sampled trajectories and measures confidence deviations relative to the state’s average behavior, decoupling state complexity from model confidence.\"}]",1784191151,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"stapo-selective-trajectory-aware-policy-optimization-for-llm-agent-training","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/stapo-selective-trajectory-aware-policy-optimization-for-llm-agent-training/83875/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does STAPO address in LLM agent training?","Question",{"text":75,"@type":76},"STAPO targets trajectory neglect, where an agent loses focus on the task goal and interaction history during extended sequences, leading to low-quality actions at intermediate steps.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is Shannon entropy considered an unreliable signal for locating trajectory neglect?",{"text":80,"@type":76},"Shannon entropy conflates inherent state complexity with agent confidence, so high-entropy steps may reflect genuinely complex states rather than low decision reliability.",{"name":82,"@type":73,"acceptedAnswer":83},"How does normalized entropy improve outlier-step detection?",{"text":84,"@type":76},"Normalized entropy aggregates actions from the same state across sampled trajectories and measures confidence deviations relative to the state’s average behavior, decoupling state complexity from model confidence.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]