[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86134-en":3,"doc-seo-86134-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86134,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","AgentCheck: A Reproduce–Intervene–Mitigate Workbench for LLM Agents over MCP","Tool-using LLM agents are often evaluated under the assumption that tools always work, but real deployments face timeouts, stale results, and poisoned tool descriptions that can mislead planners. AgentCheck provides an open-source web workbench that converts an MCP server into an intervention surface: it runs agents on real tools, records responses, injects 12 fault types, and replay-matches cached calls. It enables a reproduce–intervene–confirm loop with deterministic scoring plus an LLM judge validated on human annotations.","AgentCheck: A Reproduce–Intervene–Mitigate Workbench for LLM  \nAgents over MCP  \nAritra Mazumder  \nUniversity of Utah [aritra.mazumder@utah.edu](aritra.mazumder@utah.edu)  \nNusrat Jahan Lia  \nUniversity of Dhaka [bsse1306@iit.du.ac.bd](bsse1306@iit.du.ac.bd)  \narXiv :2607 . 1 1098v 1 [ cs . SE] 13 Jul 2026  \nAbstract  \nTool-using LLM agents are mostly evaluated assuming all tools work. When a tool times out, returns a week-stale value, or has its description poisoned in deployment, the developer needs a controlled way to reproduce the failure, test a fix, and confirm the fix worked before deployment. We present AgentCheck, an open-source web workbench that turnsan MCP server into an intervention surface. AgentCheck runs an agent against its real tools and records every tool response, then re-runs the agent with the response perturbed by a fault (12 types) injector. Matching tool calls are replayed from cache, and later tool calls go live after the agent diverges. This yields areproduce-intervene-confirm loop: the developer toggles a mitigation, re-runs against the identical fault, and sees if the failure goes away. Scoring has two parts: deterministic pass/fail rules, plus an LLM judge for interpretive labels, validated against human annotations. Across five agents, the best passes 105/120 scenariosand the weakest only 77 . The failures are usually silent, confident use of incorrect tool outputs rather than crashes. On the weakest agent, a retry mitigation raises success on timeout error faults from as few as 30% of cases to 100%, whereas stale-data faults remain near 3-4 of  \n10 regardless of the mitigation. AgentCheck makes these failure modes reproducible, comparable, and verifiable before deployment.1  \n1 Introduction  \nTool-using LLM agents are evaluated on a growing set of benchmarks. Mialon et al. (2024) tests general-purpose reasoning with tool calls, Jimenez et al. (2024) evaluates software engineering over real GitHub issues, and Qin et al. (2024) measures API mastery across thousands of endpoints. All three share one assumption that the tools work  \n1Repository: [https://github.com/aritra741/](https://github.com/aritra741/)[ ](https://github.com/aritra741/)AgentCheck. Demonstration: [https://www.youtube.com/](https://www.youtube.com/)[ ](https://www.youtube.com/)[watch?v=h_xmHC-hILU](watch?v=h_xmHC-hILU.)[.](watch?v=h_xmHC-hILU.)  \n(search API returns a result; a file write succeeds) . In real deployments an API may throw a connection timeout or a raw [HTTP error](HTTP error) (Sigdel and Baral, 2026 ; Zhu et al., 2026), a local database may return structurally valid but semantically corrupted stale records (Liu et al., 2026), and a third-party server could publish a poisoned tool description that misleads the agent’s planner into silently uploading private keys (Wang et al., 2026 ; Ye et al., 2026) .  \nRecent work shows that this gap matters. MCPTox (Wang et al., 2026) reports a 72.8% attack-success rate for tool-description poisoning against the most susceptible agent, with fewer than 3% of agents refusing outright. Cemri et al. (2026) formalizes failure modes across system design, inter-agent coordination, output verification and presents failure rates being framework-dependent. This is supported by Roig (2025)’s trace-level analysis which outlines archetypes like overhelpfulness under uncertainty and premature execution. These studies demonstrate that the dominant fault family is a function of system design and alignment choices rather than model scale only. A useful diagnostic instrument must therefore be capable of profiling these errors on custom, target-agent configurations.  \nWe present AgentCheck. AgentCheck treatsa Model Context Protocol (MCP) server (Hou et al., 2025) as that intervention surface. It runs the agent against its real tools, holds every tool response constant except one, perturbs that one with a fault injector, and visualizes the comparison with the clean and faulted runs. The result is arep","cbCaipLH0uCTTYk1","https://ap.wps.com/l/cbCaipLH0uCTTYk1","pdf",4407695,6,1,11,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"What problem does AgentCheck address in evaluating tool-using LLM agents?\",\"answer\":\"AgentCheck targets the gap between benchmark assumptions and real deployments, where tool failures like timeouts, stale outputs, or poisoned tool descriptions can change agent behavior and hide issues. It focuses on controlled reproduction and verification before deployment.\"},{\"question\":\"How does AgentCheck enable the reproduce–intervene–confirm loop?\",\"answer\":\"AgentCheck runs an agent with real MCP tools while recording every tool response, then re-runs the agent with the recorded response perturbed by a fault injector. After applying a mitigation, it re-runs against the identical fault to confirm whether the failure is resolved.\"},{\"question\":\"What scoring methods does AgentCheck use to judge results?\",\"answer\":\"Scoring combines deterministic pass/fail rules and an LLM judge that produces interpretive labels. The judge is validated against human annotations to ensure diagnostic reliability.\"}]",1784208817,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"agentcheck-a-reproduceintervenemitigate-workbench-for-llm-agents-over-mcp","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/agentcheck-a-reproduceintervenemitigate-workbench-for-llm-agents-over-mcp/86134/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does AgentCheck address in evaluating tool-using LLM agents?","Question",{"text":76,"@type":77},"AgentCheck targets the gap between benchmark assumptions and real deployments, where tool failures like timeouts, stale outputs, or poisoned tool descriptions can change agent behavior and hide issues. It focuses on controlled reproduction and verification before deployment.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does AgentCheck enable the reproduce–intervene–confirm loop?",{"text":81,"@type":77},"AgentCheck runs an agent with real MCP tools while recording every tool response, then re-runs the agent with the recorded response perturbed by a fault injector. After applying a mitigation, it re-runs against the identical fault to confirm whether the failure is resolved.",{"name":83,"@type":74,"acceptedAnswer":84},"What scoring methods does AgentCheck use to judge results?",{"text":85,"@type":77},"Scoring combines deterministic pass/fail rules and an LLM judge that produces interpretive labels. The judge is validated against human annotations to ensure diagnostic reliability.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]