[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84397-en":3,"doc-seo-84397-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84397,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Remember When It Matters Proactive Memory Agent for Long Horizon Agents","Long-horizon LLM agents often lose decision-relevant state as trajectories expand, leading to behavioral state decay where requirements, environment facts, prior attempts, diagnoses, and open subgoals no longer reliably shape subsequent actions. The work treats memory as active intervention: a separate memory agent updates a structured memory bank from the recent trajectory and decides whether to inject a concise reminder into the next action step. Plug-and-play results on Terminal-Bench 2.0 and τ2-Bench improve pass@1 across agent strengths, with selective intervention outperforming passive retrieval baselines. An open-weight early step trains Qwen3.5-27B on SETA using SFT and GRPO.","arXiv :2607 .087 16v 1 [ cs .AI] 9 Jul 2026  \nRemember When It Matters: Proactive Memory Agent for Long-Horizon Agents  \nYifan Wu , Lizhu Zhang , Yuhang Zhou , Mingyi Wang , Bo Peng , Serena Li , Xiangjun Fan , Zhuokai Zhao Meta AI  \nIn long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajectories grow, task requirements, environment facts, prior attempts, diagnoses, and open subgoals can be buried in the context window or pushed beyond it, failing to influence decisions when needed. We call this failure mode behavioral state decay. We study memory as an active intervention mechanism rather than passive retrieval. A separate memory agent runs alongside an unmodified action agent, updating a structured memory bank from the recent trajectory and deciding whether to inject a memory-grounded reminder or remain silent. The module is plug-and-play with frontier action agents and existing agent harnesses. Across Terminal-Bench 2.0 and τ2-Bench, it improves pass@1 for both weaker and stronger action agents, with gains of +8 .3 pp on Terminal-Bench and +6 .8 pp on τ 2-Bench. Ablations show that selective intervention outperforms passive bank exposure, always-on injection, advisor-only guidance, and general retrieval. As an early step toward open-weight memory policies, we train Qwen3.5-27B on SETA using SFT and GRPO, improving validation reward and achieving partial transfer to Terminal-Bench.  \nDate: July 10, 2026  \nCorrespondence: Yifan Wu and Zhuokai Zhao at {yfwu, [zhuokai}@meta.com](zhuokai}@meta.com)[ ](zhuokai}@meta.com)Code: [https://github.com/yifannnwu/proactive-memory-agent](https://github.com/yifannnwu/proactive-memory-agent)  \n1 Introduction  \nLLM agents are increasingly evaluated on long-horizon tasks that require many rounds of tool use and environment interaction (Yao et al. , 2023b ; Shinn et al. , 2023 ; Wang et al. , 2024) . These tasks span autonomous command-line execution (Merrill et al. , 2026), multi-step machine-learning engineering (Chan et al. , 2024), and interactive tool use under domain-specific rules (Yao et al. , 2024 ; Barres et al. , 2025) . Because such tasks unfold over many observations, actions, and partial decisions, success depends not only on solving each local problem, but also on preserving information that should constrain future behavior. Yet agents often fail in ways that reveal a breakdown in this state maintenance: they may identify a requirement early in a task and later violate it while fixing an unrelated bug; observe that a command, parameter setting, or implementation path fails and later retry a near-identical variant; or diagnose an error pattern and later treat the same pattern as new. These failures suggest that simply making longer histories available is insufficient. Long-horizon agents need mechanisms for keeping decision-relevant execution state active over time.  \nWe call this failure mode behavioral state decay: during long-horizon execution, information that should shape future actions like task requirements, environment facts, previous attempts, failure diagnoses, intermediate discoveries, and open subgoals stops influencing the agent’s next decision. The information may still be present in the transcript, or may even remain within the model’s context window, but it no longer exerts reliable control over behavior (Liu et al. , 2024) . This distinction is important because many memory approaches focus on whether information can be stored or retrieved, while long-horizon task execution also requires deciding when remembered information should affect the agent’s next action.  \nA natural response is to equip agents with memory. Existing memory systems typically emphasize storing, updating, and retrieving relevant records (Packer et al. , 2023 ; Zhong et al. , 2024 ; Park et al. , 2023 ; Mem0 Team, 2026), which is crucial for personalization, persistent user state, and cross-sessio","cbCaim4FTMQS3Z5z","https://ap.wps.com/l/cbCaim4FTMQS3Z5z","pdf",6882106,7,1,12,"English","en",105,"# Introduction\n## Behavioral State Decay\n## Memory as Active Intervention\n## Evaluation on Terminal-Bench 2.0 and τ2-Bench","[{\"question\":\"What failure mode do the authors call behavioral state decay?\",\"answer\":\"Behavioral state decay describes long-horizon execution where information that should control future actions—task requirements, environment facts, prior attempts, diagnoses, and open subgoals—stops reliably influencing the agent’s next decisions.\"},{\"question\":\"How does the proposed memory agent differ from passive memory systems?\",\"answer\":\"The memory agent runs alongside an unmodified action agent and decides at intervals whether to inject a memory-grounded reminder into the next decision, rather than only storing and retrieving records on demand.\"},{\"question\":\"What evidence shows the approach improves long-horizon performance?\",\"answer\":\"Experiments on Terminal-Bench 2.0 and τ2-Bench show higher pass@1 for both weaker and stronger action agents, with reported gains of about +8.3 pp and +6.8 pp respectively, and ablations indicating selective intervention is best.\"}]",1784195317,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"remember-when-it-matters-proactive-memory-agent-for-long-horizon-agents","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/remember-when-it-matters-proactive-memory-agent-for-long-horizon-agents/84397/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What failure mode do the authors call behavioral state decay?","Question",{"text":76,"@type":77},"Behavioral state decay describes long-horizon execution where information that should control future actions—task requirements, environment facts, prior attempts, diagnoses, and open subgoals—stops reliably influencing the agent’s next decisions.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed memory agent differ from passive memory systems?",{"text":81,"@type":77},"The memory agent runs alongside an unmodified action agent and decides at intervals whether to inject a memory-grounded reminder into the next decision, rather than only storing and retrieving records on demand.",{"name":83,"@type":74,"acceptedAnswer":84},"What evidence shows the approach improves long-horizon performance?",{"text":85,"@type":77},"Experiments on Terminal-Bench 2.0 and τ2-Bench show higher pass@1 for both weaker and stronger action agents, with reported gains of about +8.3 pp and +6.8 pp respectively, and ablations indicating selective intervention is best.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]