[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82995-en":3,"doc-seo-82995-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82995,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Beyond the Leaderboard A Synthesis of Tool Use Planning and Reasoning Failures in Large Language Model Agents","Large language model (LLM) agents are increasingly judged by tool use, multi-step planning, coordination, and long-horizon reasoning. Benchmark improvements often conceal repeatable failure modes reported across diverse evaluations. This synthesis aggregates 27 benchmark, taxonomy, and audit papers (2023–2026) covering 19 benchmarks into a unified cross-cutting taxonomy of limitations. Six failure clusters emerge, including tool/parameter errors, planning constraint failures, long-horizon context degradation, multi-agent coordination issues, safety/security problems, and measurement validity defects.","Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents  \nWael Albayaydh, University of Oxford [wael,[albayaydh@cs.ox.ac.uk](albayaydh@cs.ox.ac.uk)]  \nRui Zhao, University of Oxford [[rui.zhao@cs.ox.ac.uk](rui.zhao@cs.ox.ac.uk)]  \nIvan Flechais, University of Oxford [[ivan.flechais@cs.ox.ac.uk](ivan.flechais@cs.ox.ac.uk)]  \nAbstract  \nLarge language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons. Reported benchmark gains often obscure recurring failure modes documented across otherwise unrelated evaluation efforts. This paper synthesizes 27 benchmark, taxonomy, and audit papers (2023-2026), spanning 19 distinct benchmarks, into a cross-cutting taxonomy of agent limitations. To our knowledge, this is the first synthesis that integrates evidence across tool use, planning, long-horizon reasoning, multi-agent coordination, safety, and measurement validity into a single, unified taxonomy of LLM agent limitations. We identify six failure clusters: (1) tool invocation and parameter-level errors,(2) planning and constraint-satisfaction failures,(3) long-horizon degradation from context accumulation,(4) multi-agent coordination failures, (5) safety and security failures under adversarial or underspecified conditions, and (6) measurement validity problems. The taxonomy was built iteratively by grouping independently reported error categories into themes corresponding to distinct stages of an agent's reasoning-to-action pipeline, and its boundaries are corroborated by several source papers converging on overlapping categories despite unrelated annotation processes. Across clusters we find that failure compounds non-linearly with task length, that sub-skill competence does not reliably compose into end-to-end success, and that additional scaffolding does not uniformly improve reliability. We balance this with an explicit account of where agents have measurably improved—single-turn tool selection, short-horizon web navigation, and narrowly scoped coding tasks all show substantial, credible progress. We consolidate quantitative findings into a categorized comparative table, discuss implications for evaluation methodology and deployment, and argue that interpreting agent progress requires distinguishing genuine capability gains from corrections of earlier measurement error.  \nKeywords: AI agents; tool use; LLM planning; multi-agent systems; benchmark evaluation; long-horizon reasoning; agent safety  \n1. Introduction  \nThe rapid deployment of large language model (LLM) agents—systems that combine LLM reasoning with the ability to call external tools, browse the web, write and execute code, and coordinate with other agents— has been accompanied by an equally rapid proliferation of benchmarks designed to measure their competence. GAIA (Mialon et al., 2024), WebArena (Zhou et al., 2024), AgentBench (Liu et al., 2023), TravelPlanner (Xie et al., 2024), the Berkeley Function-Calling Leaderboard (Patil et al., 2025), and SWEbench (Jimenez et al., 2024) are among the most widely cited examples of this evaluation infrastructure.  \nLeaderboards built on these benchmarks show steady, sometimes dramatic, year-over-year improvement. Reported WebArena success rates for frontier agents rose from approximately 14% at the benchmark's introduction (Zhou et al., 2024) to a reported 61.7% roughly eighteen months later for a specialized enterprise scaffold (Marreed et al., 2025), though even the strongest current agents still trail human performance of approximately 78% . SWE-bench Verified resolution rates for leading coding agents have climbed from approximately 20% in mid-2024 to figures above 50-60% for top systems by 2026 across independent leaderboard reports. Separately, METR's time-horizon analysis—a rigorous, peer-reviewed-  \nadjacent measurement effort that times human experts on ","cbCaieQeE3HUIvWZ","https://ap.wps.com/l/cbCaieQeE3HUIvWZ","pdf",179496,4,1,16,"English","en",105,"# Introduction\n## Agent evaluation benchmarks and leaderboard trends\n## Recurring failure patterns and safety risks\n## Motivation: capability vs measurement validity\n## Contributions and taxonomy overview","[{\"question\":\"What is the document’s main goal?\",\"answer\":\"To synthesize evidence across tool use, planning, long-horizon reasoning, multi-agent coordination, safety, and evaluation measurement validity into a single taxonomy of LLM agent limitations.\"},{\"question\":\"How many sources and benchmarks are synthesized?\",\"answer\":\"The paper synthesizes 27 benchmark, taxonomy, and audit papers from 2023–2026 spanning 19 distinct benchmarks.\"},{\"question\":\"What are the six main clusters of agent failures identified?\",\"answer\":\"Tool invocation and parameter errors; planning and constraint-satisfaction failures; long-horizon degradation from context accumulation; multi-agent coordination failures; safety and security failures; and measurement validity problems.\"}]",1784184523,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"beyond-the-leaderboard-a-synthesis-of-tool-use-planning-and-reasoning-failures-in-large-language-model-agents","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/beyond-the-leaderboard-a-synthesis-of-tool-use-planning-and-reasoning-failures-in-large-language-model-agents/82995/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the document’s main goal?","Question",{"text":75,"@type":76},"To synthesize evidence across tool use, planning, long-horizon reasoning, multi-agent coordination, safety, and evaluation measurement validity into a single taxonomy of LLM agent limitations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How many sources and benchmarks are synthesized?",{"text":80,"@type":76},"The paper synthesizes 27 benchmark, taxonomy, and audit papers from 2023–2026 spanning 19 distinct benchmarks.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the six main clusters of agent failures identified?",{"text":84,"@type":76},"Tool invocation and parameter errors; planning and constraint-satisfaction failures; long-horizon degradation from context accumulation; multi-agent coordination failures; safety and security failures; and measurement validity problems.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]