[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84900-en":3,"doc-seo-84900-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84900,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","AgentTether Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operations","Large language model agents are increasingly applied to multi-step, stateful tool-use tasks, but production reliability remains limited due to dynamic trajectories where early decisions propagate into later failures and changing external states. Existing remedies provide incomplete coverage: blind retry lacks diagnosis, outcome feedback lacks “where/why,” and self-reflection often lacks grounded evidence. AgentTether provides run-time repair via Transition Units, a dependency-aware Critical Transition Graph, localized diagnosis, and behavior-scoped guidance using Repair Memory, optionally with guarded runtime intervention.","AgentTether: Graph-Guided Diagnosis and Runtime Intervention for Reliable LLM Agent Operations  \nChenyu Zhao 1 , Shenglin Zhang 1,* , Wenwei Gu 1 , Yongqian Sun 1 Dan Pei2 , Chetan Bansal3 , Saravan Rajmohan3 , and Minghua Ma3  \n1 Nankai University, Tianjin, China  \n2 Tsinghua University, Beijing, China  \n3 Microsoft  \n* Corresponding author: Shenglin Zhang.  \narXiv :2607 .06273v 1 [ cs . SE] 7 Jul 2026  \nAbstract—Large language model (LLM) agents are increasingly used for multi-step, stateful tool-use tasks, yet production reliability remains limited. Unlike static software repair, agent repair must recover dynamic trajectories whose early decisions can propagate into later errors and external state changes. Existing automatic remedies address only part of this problem: blind retry adds no diagnosis, outcome feedback says whether a run failed but not where or why, and self-reflection often lacks grounded evidence to prevent the same failure from recurring. We present AgentTether, a run-time repair framework that automates post-run diagnosis and guided recovery without modifying the underlying agent or environment. AgentTether abstracts each run into Transition Units, links them through a dependencyaware Critical Transition Graph, and localizes failure-critical subtrajectories by combining an offline normal-behavior model with a run-local graph detector. It then converts the localized cause into behavior-scoped guidance backed by cross-iteration Repair Memory, and can optionally apply guarded run-time intervention to keep the correction active during re-execution. The same design can be deployed as an offline diagnostic-andguidance tool or as an online repair layer.  \nWe evaluate AgentTether on 261 τ-bench tasks across three domains with Qwen3.7-max, and test cross-model transfer on Banking with GPT-5.4. On the hardest Banking domain, AgentTether repairs 59.04% (49/83) of initially failed Qwen3.7-max tasks and 65.12%(56/86) of initially failed GPT-5.4 tasks. Overall, AgentTether improves repair effectiveness while reducing agent turns and end-to-end approach tokens, suggesting a practical reliability layer that can wrap existing agent deployments, reduce wasted re-execution, and improve recovery without retraining the agent.  \nIndex Terms—LLM agents, agent repair, root cause analysis, graph-guided diagnosis, runtime intervention  \nI. INTRODUCTION  \nLarge language model (LLM) agents now execute multi-step tasks that require planning, tool use, state updates, and iterative interaction with external environments [1]–[7] . As they move from demonstrations to production, reliability becomes the central barrier: single-run success rates remain limited in realistic domains [3], [8]–[14], and even seemingly successful runs may harbor process-level risks such as missing steps, incorrect ordering, tool-protocol violations, unauthorized state changes, or unnecessary actions [15], [16] . Such defects maybe missed by coarse outcome-based evaluation yet impose real operational and compliance costs, as illustrated by deployed chatbots that misstated refund rules or advised legally invalid  \nactions [17], [18] . Improving reliability therefore demands moving beyond final-outcome measurement and blind retry, toward actively recovering the runs that go wrong.  \nWe refer to this corrective process as agent repair: diagnosing and correcting defective agent runs after failure or during execution, without modifying the underlying model or environment. Unlike automated program repair or LLM-based self-debugging, which operate on fixed code artifacts and test oracles [19]–[21], agent repair targets dynamic, stateful, and non-deterministic trajectories whose failures may lack an oracle, originate far upstream, and change external state in ways that are hard to undo. These differences make agent repair more than patch generation and motivate the expert human operator’s repair loop as a reference.  \nFigure 1 illustrates the repair gap on a Banking task ","cbCails7PaSZjhIr","https://ap.wps.com/l/cbCails7PaSZjhIr","pdf",1118062,1,12,"English","en",105,"# Introduction\n## Agent repair and reliability gaps\n## Motivating example: Banking task repair","[{\"question\":\"What problem does AgentTether address in LLM agent reliability?\",\"answer\":\"It targets failures in multi-step, stateful tool-use tasks where early incorrect decisions propagate and where existing approaches do not provide sufficient diagnosis or grounded guidance to prevent recurrence.\"},{\"question\":\"How does AgentTether perform diagnosis and recovery during runtime?\",\"answer\":\"It abstracts each run into Transition Units, links them using a dependency-aware Critical Transition Graph, localizes failure-critical subtrajectories with an offline normal-behavior model and a run-local graph detector, and converts causes into behavior-scoped guidance backed by cross-iteration Repair Memory.\"},{\"question\":\"What outcomes does the evaluation report for AgentTether?\",\"answer\":\"On 261 τ-bench tasks, it repairs 59.04% of initially failed Qwen3.7-max Banking tasks and 65.12% of initially failed GPT-5.4 Banking tasks, while improving repair effectiveness and reducing agent turns and end-to-end approach tokens.\"}]",1784199237,30,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"agenttether-graph-guided-diagnosis-and-runtime-intervention-for-reliable-llm-agent-operations","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/agenttether-graph-guided-diagnosis-and-runtime-intervention-for-reliable-llm-agent-operations/84900/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does AgentTether address in LLM agent reliability?","Question",{"text":74,"@type":75},"It targets failures in multi-step, stateful tool-use tasks where early incorrect decisions propagate and where existing approaches do not provide sufficient diagnosis or grounded guidance to prevent recurrence.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does AgentTether perform diagnosis and recovery during runtime?",{"text":79,"@type":75},"It abstracts each run into Transition Units, links them using a dependency-aware Critical Transition Graph, localizes failure-critical subtrajectories with an offline normal-behavior model and a run-local graph detector, and converts causes into behavior-scoped guidance backed by cross-iteration Repair Memory.",{"name":81,"@type":72,"acceptedAnswer":82},"What outcomes does the evaluation report for AgentTether?",{"text":83,"@type":75},"On 261 τ-bench tasks, it repairs 59.04% of initially failed Qwen3.7-max Banking tasks and 65.12% of initially failed GPT-5.4 Banking tasks, while improving repair effectiveness and reducing agent turns and end-to-end approach tokens.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":28,"slug":120},"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]