[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85039-en":3,"doc-seo-85039-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85039,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Who Broke the System Failure Localization in LLM Based Multi Agent Systems","Large language model (LLM) based multi-agent systems solve complex problems through coordinated reasoning and action, yet their distributed execution makes diagnosing system-level failures difficult. When an run fails, determining the responsible agent and the earliest step where the trajectory becomes irreversibly misdirected is hindered by long-horizon interactions and tightly coupled agent behaviors. The paper introduces AgentLocate, attributing failures to an agent and the earliest decisive step using an LLM judge plus confidence-aware multi-perspective verification and lightweight fine-tuning.","arXiv :2607 .07989v 1 [ cs .CR] 8 Jul 2026  \nWho Broke the System? Failure Localization in LLM-Based Multi-Agent Systems  \nYufei Xia1, Anjun Gao1, Yueyang Quan2, Zhuqing Liu2, Minghong Fang1  \n1University of Louisville, 2University of North Texas  \nAbstract  \nLarge language model (LLM) based multi-agent systems enable complex problem solving through coordinated reasoning and action, but their distributed structure also introduces new challenges in diagnosing systemlevel failures. When an execution fails, identifying which agent is responsible and at what point the trajectory first becomes irreversibly misdirected is difficult due to long-horizon interactions and tightly coupled agent behaviors. In this paper, we study the problem of failure localization in LLM-based multi-agent systems and present AgentLocate, a framework that attributes failures to both a specific agent and the earliest decisive step. AgentLocate combines an LLM-based judging mechanism with multiperspective verification by independent evaluators, whose assessments are aggregated using a confidence-aware strategy. The resulting feedback is further used to adapt the judge through lightweight fine-tuning, improving attribution quality. We evaluate AgentLocate on two complementary benchmarks covering diverse tasks, agent configurations, and trajectory lengths. Experimental results show that AgentLocate consistently outperforms existing failure localization methods in identifying both responsible agents and failure steps, while remaining efficient in terms of token usage and running time.  \n1 Introduction  \nLarge language model (LLM) based multi-agent systems (Li et al., 2024; Hong et al., 2024; Liet al., 2023; Wu et al., 2024; Han et al., 2024; Talebirad & Nadiri, 2023) have recently gained prominence as an effective framework for tackling problems that exceed the capabilities of a single language-model-driven agent. By distributing responsibilities across specialized components, these systems enable richer forms of reasoning, more flexible planning, and coordinated workflows that span diverse domains. They have been applied to areas such as software engineering (Qian et al., 2024; Tufano et al., 2024), scientific discovery (Boiko et al., 2023; Ferraro et al., 2025), information gathering (Shen et al., 2023; Sun et al., 2025), and complex web-based tasks (Zhou et al., 2023; Mialon et al., 2023), where collaborative agent behaviors can yield stronger performance than isolated models. This growing adoption highlights the potential of multi-agent architectures as a foundation for building increasingly capable AI systems.  \nDespite these advantages, the increasing adoption of LLM-based multi-agent systems has also revealed a growing set of reliability concerns. The very features that make multi-agent architectures powerful, such as distributed decision making, role specialization, and toolmediated interactions, also create new pathways for failures to arise (Zhang et al., 2025d; Cemri et al., 2025; Kong et al., 2025) . Unlike single-agent pipelines, where errors typically stem from a localized misprediction, multi-agent systems introduce interdependent behaviors in which a subtle mistake by one component can ripple across subsequent steps, distort shared state, or derail collaborative planning. These intertwined execution paths make it difficult to determine not only when a failure occurs but also which agent or interaction first pushed the system off course. As multi-agent workflows become more intricate and are deployed in increasingly realistic environments, understanding and diagnosing these failure dynamics becomes essential for ensuring dependable operation.  \nAutomatic failure localization in LLM-based multi-agent systems has recently attracted growing attention, with several methods (Zhang et al., 2025c;d; Banerjee et al., 2025; Kong et al., 2025) proposed to attribute failures to a responsible agent and a decisive step. However, our empirical results suggest","cbCailrZDVnmmlDN","https://ap.wps.com/l/cbCailrZDVnmmlDN","pdf",483608,3,1,25,"English","en",105,"# Introduction\n## Motivation and reliability challenges\n## Prior failure localization methods and limitations\n## Contributions","[{\"question\":\"What problem does the paper address in LLM-based multi-agent systems?\",\"answer\":\"It addresses failure localization: identifying which agent is responsible for an execution failure and the earliest step where the trajectory first becomes decisively misdirected.\"},{\"question\":\"How does AgentLocate localize failures in the system?\",\"answer\":\"AgentLocate combines an LLM-based judging mechanism with multi-perspective verification by independent evaluators, aggregates results via a confidence-aware strategy, and uses the feedback to improve the judge through lightweight fine-tuning.\"},{\"question\":\"Why can existing failure localization approaches be unstable in multi-agent settings?\",\"answer\":\"Counterfactual replay can be unstable because changing one step alters later prompts, tool outcomes, and coordination patterns, making the failure-inducing action hard to consistently isolate.\"}]",1784200550,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"who-broke-the-system-failure-localization-in-llm-based-multi-agent-systems","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/who-broke-the-system-failure-localization-in-llm-based-multi-agent-systems/85039/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in LLM-based multi-agent systems?","Question",{"text":75,"@type":76},"It addresses failure localization: identifying which agent is responsible for an execution failure and the earliest step where the trajectory first becomes decisively misdirected.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does AgentLocate localize failures in the system?",{"text":80,"@type":76},"AgentLocate combines an LLM-based judging mechanism with multi-perspective verification by independent evaluators, aggregates results via a confidence-aware strategy, and uses the feedback to improve the judge through lightweight fine-tuning.",{"name":82,"@type":73,"acceptedAnswer":83},"Why can existing failure localization approaches be unstable in multi-agent settings?",{"text":84,"@type":76},"Counterfactual replay can be unstable because changing one step alters later prompts, tool outcomes, and coordination patterns, making the failure-inducing action hard to consistently isolate.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]