[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81697-en":3,"doc-seo-81697-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81697,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","ECHO: Prune To Act, Trace To Learn With Selective Turn Memory in Agentic RL","Long-horizon language agents must repeatedly use tools, collect evidence, and decide under limited context windows. Existing context-management methods compress history but can destroy source-level provenance, so reinforcement learning cannot reliably credit the observations that truly support correct final answers. ECHO introduces selective turn-memory for traceable context reconstruction: it stores each completed turn as source-indexed memory, selects useful records for bounded acting, and routes positive outcome credit to traceable trajectory segments. Experiments on BrowseComp-Plus achieve 43.4% held-out accuracy with efficiency gains and improved zero-shot generalization across multi-objective QA, code generation, and deep information seeking.","arXiv :2606 .3 1650v2 [ cs .LG] 10 Jul 2026  \nECHO: PRUNE TO ACT, TRACE TO LEARN  \nWITH SELECTIVE TURN MEMORY IN AGENTIC RL  \nZijun Xie 1 ,2∗† Binbin Zheng2 ,3∗† Enlei Gong2∗ Jihua Liu2 Yuyang You 1 Lingfeng Liu 1 Jiayao Tang 1 Guanqun Zhao2 Aoqi Hu2 Zeyu Chen2‡  \n1 School of Mathematical Sciences, Peking University 2Baidu Inc.  \n3University of Science and Technology of China [xiezijun714@gmail.com](xiezijun714@gmail.com) [binbinzheng@mail.ustc.edu.cn](binbinzheng@mail.ustc.edu.cn)  \n§ GitHub:xiezijun714-lang/Echo  \nABSTRACT  \nLong-horizon language agents must repeatedly interact with tools, accumulate evidence, and make decisions under bounded context windows. Context-management methods make such rollouts feasible by simplifying past interactions through deletion, folding, or memory editing. However, when useful history is collapsed into compressed states, the reconstructed context may no longer reveal which earlier observations support a successful final answer. This creates amismatch between bounded-context acting and outcome-based reinforcement learning: the policy acts on reconstructed context, while the learner lacks source-level provenance for assigning credit to the evidence that mattered. We propose ECHO, a selective turn-memory framework for traceable context reconstruction in agentic RL. ECHO compresses each completed environment turn into a compact source-indexed memory record, reconstructs bounded policy contexts by selecting useful records, and reuses the selected source indices to route positive outcome credit to the final trajectory segment, reused evidence turns, memory findings, and memory-selection actions. On BrowseComp-Plus, ECHO reaches 43.4% held-out accuracy, outperforming GRPOat 28.9% and the rolling-summary baseline SUPO at 36.1%, while using fewer turns and lower trajectory volume than SUPO. The trained policy also improves zero-shot generalization across multi-objective QA, code generation, and deep information-seeking benchmarks on both dense and MoE backbones.  \n GRPO  SUPO  ECHO  \nAcc Pass@1  \n0.50  \n0.40  \n0.30  \n0.20  \n0.10  \nBrowseComp Plus  \n0 25 50 75 100 115  \nTraining Steps  \nAvg . Turns  \n90  \n60  \n30  \n0  \nTurn Dynamics  \n\n| \u003Cbr> |\n| --- |\n|  |\n\n0 25 50 75 100 115  \nTraining Steps  \nAvg . Trajectories  \n5.0  \nTrajectory Volume  \n4.0  \n3.0  \n2.0  \n1.0  \n0 25 50 75 100 115  \nTraining Steps  \nFigure 1: Held-out accuracy, tool-use turns per rollout, and trajectory volume over training on BrowseComp-Plus with the Qwen3-32B-Instruct backbone for ECHO (purple), GRPO (orange), and SUPO (green) . ECHO traces the upper-left frontier: rising accuracy without the turn and volume growth seen for SUPO.  \n1 INTRODUCTION  \nLarge language models (LLMs) are increasingly deployed as multi-turn agents that interleave reasoning, tool invocation, and environment feedback (Yao et al., 2023 ; Schick et al., 2023) . Reinforcement learning (RL) from verifiable final outcomes has become a central recipe for improving such agents in search, coding, function calling,  \nand deep-research settings (Jin et al., 2025 ; Qian et al., 2025 ; Li et al., 2025 ; Zheng et al., 2025) . As interaction ∗Equal contribution.  \n†This work was done during an internship at Baidu.  \n‡Corresponding author.  \nhorizons grow, however, history management becomes a bottleneck for both acting and learning. The policy must retain useful observations within a bounded context window, while the learner must decide which earlier decisions should receive credit from a sparse final outcome.  \nContext-management methods make long rollouts feasible by simplifying past interactions before the next decision. From the perspective of reconstructed history, these methods differ in what form the completed history becomes: some delete or omit distant turns, some edit an explicit memory state, and others fold long prefixes into compact summaries or internal states. The triggering mechanism may be rule-based or action-based, but this choice is orthogonal to th","cbCaigBLsjqGN9sd","https://ap.wps.com/l/cbCaigBLsjqGN9sd","pdf",757930,3,1,16,"English","en",105,"# Abstract\n# Introduction\n## Motivation: provenance gap in bounded context acting and outcome-based RL\n## Proposed method: ECHO selective turn-memory for traceable context reconstruction","[{\"question\":\"What problem does ECHO address in agentic reinforcement learning?\",\"answer\":\"ECHO targets the provenance gap caused by context compression: the policy can act using reconstructed context, but the learner cannot identify which earlier environment turns should receive credit from sparse final outcomes.\"},{\"question\":\"How does ECHO reconstruct bounded policy context for acting?\",\"answer\":\"ECHO compresses each completed environment tool-use turn into a compact, source-indexed memory record and selects useful records when the context budget becomes binding to rebuild the next bounded context alongside recent interactions.\"},{\"question\":\"How does ECHO improve credit assignment during learning?\",\"answer\":\"ECHO reuses the source-indexed reconstruction trace for learning by keeping only the positive part of trajectory advantage and applying it only to credit tokens that can be traced to the selected evidence turns.\"}]",1784175478,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"echo-prune-to-act-trace-to-learn-with-selective-turn-memory-in-agentic-rl","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/echo-prune-to-act-trace-to-learn-with-selective-turn-memory-in-agentic-rl/81697/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does ECHO address in agentic reinforcement learning?","Question",{"text":75,"@type":76},"ECHO targets the provenance gap caused by context compression: the policy can act using reconstructed context, but the learner cannot identify which earlier environment turns should receive credit from sparse final outcomes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does ECHO reconstruct bounded policy context for acting?",{"text":80,"@type":76},"ECHO compresses each completed environment tool-use turn into a compact, source-indexed memory record and selects useful records when the context budget becomes binding to rebuild the next bounded context alongside recent interactions.",{"name":82,"@type":73,"acceptedAnswer":83},"How does ECHO improve credit assignment during learning?",{"text":84,"@type":76},"ECHO reuses the source-indexed reconstruction trace for learning by keeping only the positive part of trajectory advantage and applying it only to credit tokens that can be traced to the selected evidence turns.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]