[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82235-en":3,"doc-seo-82235-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82235,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Scoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift","Deployed LLM agents depend on agentic context—model-external control text assembled by an operational harness—where the mutable part is a persistent system-level instruction updated from experience. Over long evolution horizons, flat-text maintenance makes verification harder as accumulated instructions interact. Graph-Regularized Agentic Context Evolution (GRACE) represents this persistent instruction as a typed semantic graph and validates updates locally in typed neighborhoods. Accepted graph changes are converted into incremental edits for a deployment text checkpoint. On a telecom agent harness with controlled distribution shifts, GRACE improves strict reliability across replications.","arXiv :2607 .09 175v 1 [ cs .AI] 10 Jul 2026  \nScoped Verification for Reliable Long-Horizon Agentic Context Evolution under Distribution Shift  \nDan C. Hsu  \nRedMind Research, San Francisco, CA, USA National Taiwan University, Taipei, Taiwan  \nLuke Lu  \n[dan@redmindresearch.org](dan@redmindresearch.org)  \n[luke@redmindresearch.org](luke@redmindresearch.org)  \nRedMind Research, San Francisco, CA, USA  \nCode: [https://github.com/RedMind-Research/GRACE](https://github.com/RedMind-Research/GRACE)  \nAbstract  \nDeployed LLM agents rely on agentic context, the model-external textual control content assembled by an operational harness. In this work, the mutable component of that context is a persistent system-level instruction that is updated from operational experience while the model, tools, and harness remain fixed. Over long evolution horizons, flat-text maintenance makes verification increasingly difficult as accumulated instructions grow and interact. We propose Graph-Regularized Agentic Context Evolution (GRACE), which maintains the persistent instruction component as a typed semantic graph and validates proposed updates within the local typed neighborhoods of modified nodes. Accepted graph updates are reconstructed as incremental edits to the textual instruction checkpoint used at deployment. We evaluate GRACE within a fixed telecom agent harness derived from τ 2-bench under a controlled distribution-shift protocol. Across five independent replications, GRACE improves strict reliability, measured by passˆ3, from the Gemini 2.5 Flash zero-shot value of 0 .091 to 0 .673±0 . 136 at the final checkpoint. This exceeds a Gemini 3 .1 Pro zero-shot reference of 0.242 on the same held-out set, while the flat-text HCE baseline finishes at 0.191±0.051 . These results identify two requirements for reliable long-horizon context evolution, a structural substrate that makes verification local and a consolidation mechanism that keeps accumulated instruction content usable.  \nKeywords: agentic context, context evolution, persistent instruction, structural validation, reliability, distribution shift  \n1. Introduction  \nAs large language models have become better at following natural-language instructions, model-external text has become a primary control surface for deployed agents (Khattab et al. , 2024; Zhang et al. , 2026) . In operational agent systems, this text is assembled with task inputs, tool observations, and harness-provided information before each model call (Yao et al. , 2025; Barres et al. , 2025) . We refer to the resulting inference-time input asthe agent context. In the setting studied here, the mutable component of this context is a persistent system-level instruction that specifies the agent’s role, behavioral constraints, procedural guidance, and domain assumptions (Qin et al., 2025; Lee et al., 2024) . When task distributions shift, this persistent instruction can encode stale or incomplete assumptions about the environment it is meant to govern. We study context evolution as the process of revising this persistent context component from operational experience while the model, tools, and harness remain fixed. The evolving artifact may be a playbook (Zhang et al. , 2026), an adaptive memory (Suzgun et al. , 2026), a set of strategic principles (Wu et al. ,  \n© 2026 D.C. Hsu & L. Lu.  \nHsu Lu  \n2025), or a set of structured guidelines (Pei et al. , 2025) . Across these forms, prior work has shown that iterative updates can improve agent behavior (Zhang et al. , 2026; Shinn et al. , 2023) .  \nHowever, over extended evolution horizons, previously accumulated improvements can be undermined by inconsistencies introduced in subsequent steps. Because the persistent instruction is included in the agent context throughout deployment, degradation in this artifact can compromise operational reliability as well as update efficiency. Context collapse during full-document rewriting has been documented as one failure mode in agentic conte","cbCaisrmPDECm4kw","https://ap.wps.com/l/cbCaisrmPDECm4kw","pdf",1577054,1,18,"English","en",105,"# Abstract\n# Introduction\n## Problem: long-horizon context evolution under distribution shift\n## Approach: GRACE with graph-regularized, local verification\n## Experimental setup and evaluation protocol","[{\"question\":\"What is the key mutable component in agentic context studied in the document?\",\"answer\":\"The mutable component is a persistent system-level instruction that defines the agent’s role, constraints, procedures, and domain assumptions and is updated from operational experience.\"},{\"question\":\"Why does reliability degrade during long-horizon context evolution?\",\"answer\":\"Flat-text maintenance causes verification to become increasingly difficult as accumulated instructions interact, potentially introducing inconsistencies that harm operational reliability and update efficiency.\"},{\"question\":\"How does GRACE validate proposed updates to the persistent instruction?\",\"answer\":\"GRACE represents the instruction as a typed semantic graph, then validates each proposed update within local typed neighborhoods around modified nodes; accepted graph updates are translated back into incremental text edits for the deployment checkpoint.\"}]",1784179024,45,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"scoped-verification-for-reliable-long-horizon-agentic-context-evolution-under-distribution-shift","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/scoped-verification-for-reliable-long-horizon-agentic-context-evolution-under-distribution-shift/82235/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the key mutable component in agentic context studied in the document?","Question",{"text":75,"@type":76},"The mutable component is a persistent system-level instruction that defines the agent’s role, constraints, procedures, and domain assumptions and is updated from operational experience.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why does reliability degrade during long-horizon context evolution?",{"text":80,"@type":76},"Flat-text maintenance causes verification to become increasingly difficult as accumulated instructions interact, potentially introducing inconsistencies that harm operational reliability and update efficiency.",{"name":82,"@type":73,"acceptedAnswer":83},"How does GRACE validate proposed updates to the persistent instruction?",{"text":84,"@type":76},"GRACE represents the instruction as a typed semantic graph, then validates each proposed update within local typed neighborhoods around modified nodes; accepted graph updates are translated back into incremental text edits for the deployment checkpoint.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]