[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85799-en":3,"doc-seo-85799-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85799,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","AgentAbstain: Do LLM Agents Know When Not to Act","AgentAbstain evaluates whether tool-using LLM agents can accurately abstain—recognize when not to act—rather than only measuring task success. The benchmark introduces an agent-native taxonomy covering 8 abstention scenarios spanning pre-execution reasoning and runtime discovery. It includes 263 paired tasks across 42 executable sandbox environments, generated via ABSTAINGEN to address manual authoring cost and data contamination risk through automated regeneration and deterministic validation. Results on frontier models show abstention ability is largely independent of general task-solving, highlighting critical post-hoc failure modes.","arXiv :2607 . 10059v 1 [ cs .AI] 11 Jul 2026  \nAGENTABSTAIN: Do LLM Agents Know  \nWhen Not to Act?  \nXun Liu†, Yi Evie Zhang†, Vira Kasprova∗ , Parisa Rabbani∗ , Pardis Sadat Zahraei∗  \nTianyu Zhang∗ , Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran  \nUniversity of Illinois Urbana-Champaign  \nAbstract  \nAgent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain. This gap poses real risks: under ambiguity, conflicting constraints, or tool failures, agents may execute unintended and irreversible actions. To close this gap, we present the first systematic evaluation framework for agentic abstention: the calibrated ability of tool-using LLM agents to recognize when not to act. At its core, AGENTABSTAIN is a paired-task benchmark built on an agent-native taxonomy of 8 abstention scenarios across pre-execution reasoning and runtime discovery. It contains 263 paired tasks across 42 executable sandbox environments, where each pair consists of a should-act task and a shouldabstain variant produced through a controlled perturbation to the instruction, tool, or environment state. Scaling such paired evaluations poses two practical challenges:  \nmanually authoring diverse tasks is expensive, and static benchmarks risk data contamination as models evolve. To address both, we propose ABSTAINGEN, a fully automated pipeline that synthesizes sandbox environments and generates paired tasks end-to-end, validated by deterministic replay and semantic LLM judges. Its scalable design enables on-demand regeneration of fresh task instances, and three independent annotators rate 94–98% of sampled tasks as well-designed.  \nAcross 17 frontier LLMs in 4 agent harnesses, the best agent (Gemini 3.1 Pro) achieves only 59.5% paired accuracy (correct on both the act and abstain sides of each paired task) . More importantly, abstention capability is largely independent of general task-solving capability, indicating that scaling task-solving alone will not close this gap. We identify failure modes such as post-hoc abstention, in which agents execute irreversible actions before recognizing abstention triggers. For instance, an agent may cancel a flight reservation before noticing contradictory rebooking instructions, leaving the user stranded. These findings underscore the need for rigorous abstention evaluation to develop more trustworthy LLM agents.  \nOur code [and dataset are open-sourced at](and dataset are open-sourced at agentabstain.github.io)[ agentabstain.github.io](and dataset are open-sourced at agentabstain.github.io).  \n1 Introduction  \nLarge language model (LLM) agents are increasingly deployed to perform consequential tasks autonomously: booking travel, managing files, executing code, and interacting with APIs on behalf of users [32, 65, 78] . As these agents acquire the ability to commit irreversible actions in real environments, the question becomes: do they know when not to act?  \nConsider a travel-management agent asked to “cancel the outbound flight and rebook a windowseat on the next departure.” A well-calibrated agent should verify rebooking availability before  \ncanceling the existing reservation. Yet in practice, we observe that frontier models cancel the flight †Project lead. ∗ Equal contribution, [listed alphabetically. Correspondence to xunliu@illinois.edu](listed alphabetically. Correspondence to xunliu@illinois.edu), [varunc@illinois.edu](varunc@illinois.edu).  \nPreprint.  \nfirst, discover that no suitable alternative exists, and then inform the user that the rebooking failed, leaving the user stranded with neither the original nor a new reservation. This failure mode combines irreversible side effects with a belated verbal acknowledgment of the problem. It is distinct from, and more dangerous than, a simple wrong answer in a QA setting.  \nAbstention, the ability to recognize when a response or","cbCaiin6CHHBEEZX","https://ap.wps.com/l/cbCaiin6CHHBEEZX","pdf",8420258,7,1,56,"English","en",105,"# Introduction\n## Agentic abstention problem\n## Related work\n## Proposed evaluation framework","[{\"question\":\"What problem does AgentAbstain target in LLM agent deployments?\",\"answer\":\"AgentAbstain targets the ability of tool-using LLM agents to recognize when not to act, especially under ambiguity, conflicting constraints, or tool failures that can lead to unintended irreversible actions.\"},{\"question\":\"How is AgentAbstain structured to measure calibrated abstention?\",\"answer\":\"It uses a paired-task benchmark: each instance is a should-act variant and a should-abstain variant that share the same sandbox but differ by a controlled perturbation to the instruction, tool, or environment state.\"},{\"question\":\"Why can’t existing answer-or-abstain evaluations fully cover tool-using agents?\",\"answer\":\"They evaluate abstention over a single response where environment state cannot be altered, so restraint cannot be verified against tool-call traces, and failures emerging mid-trajectory in sequential tool use are not captured.\"}]",1784206345,141,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"agentabstain-do-llm-agents-know-when-not-to-act","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/agentabstain-do-llm-agents-know-when-not-to-act/85799/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does AgentAbstain target in LLM agent deployments?","Question",{"text":76,"@type":77},"AgentAbstain targets the ability of tool-using LLM agents to recognize when not to act, especially under ambiguity, conflicting constraints, or tool failures that can lead to unintended irreversible actions.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is AgentAbstain structured to measure calibrated abstention?",{"text":81,"@type":77},"It uses a paired-task benchmark: each instance is a should-act variant and a should-abstain variant that share the same sandbox but differ by a controlled perturbation to the instruction, tool, or environment state.",{"name":83,"@type":74,"acceptedAnswer":84},"Why can’t existing answer-or-abstain evaluations fully cover tool-using agents?",{"text":85,"@type":77},"They evaluate abstention over a single response where environment state cannot be altered, so restraint cannot be verified against tool-call traces, and failures emerging mid-trajectory in sequential tool use are not captured.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]