[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83237-en":3,"doc-seo-83237-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83237,962075114101,"Seraphina","https://ap-avatar.wpscdn.com/avatar/e000253a75eb197efd?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780044092746381165",8,"Research & Report","Reason Less Verify More Deterministic Gates Recover a Silent Policy Violation Failure Mode in Tool Using LLM Agents","Tool-using LLM agents can perform actions that violate the very policies they are meant to enforce while still reporting successful completion. In policy-permissive environments, a tool may execute well-formed calls even when a domain policy forbids the resulting state transition, creating silent wrong states such as unverified claims being acted on or records being modified. Study on the 2-bench airline domain finds 78% silent wrong-state failures for a budget agent, and deterministic readonly pre-execution gates significantly improve success by raising full-benchmark success from 29.6% to 42.0%.","Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents  \nVikas Reddy Independent Researcher  \nSumanth Reddy Challaram  \nIndian Institute of Technology Kharagpur  \nKharagpur, India  \nAbhishek Basu  \nMassachusetts Institute of Technology Cambridge, USA  \narXiv :2607 .07405v2 [ cs .AI] 11 Jul 2026  \nAbstract  \nTool-using LLM agents can violate the very policies they are deployed to enforce while appearing to complete the task successfully. In policy-permissive environments, a tool may execute any well-formed call even when the corresponding state transition is forbidden by domain policy. The result is a silent wrong state: the booking is cancelled, the passenger count is changed, or a user claim is acted on without verification, and neither the tool nor the agent’s self-report exposes the violation.  \nWe study this failure mode in the 􀁧2-bench airline domain. On a budget agent, 78% of observed failures are silent wrong-state failures with no tool error, and the aggregate failure rate is reproducible across disjoint seeds rather than reflecting sampling noise. We then evaluate a lightweight intervention: deterministic, readonly pre-execution gates that inspect the proposed tool call and current database state before allowing a write. A four-gate suite raises full-benchmark success from 29.6% to 42.0% on gpt-4o-mini (+12.4pp; paired task-level bootstrap 􀀥 = 0. 0012), and the lift reproduces on a disjoint 15-seed replication set (+12.3pp; 􀀥 = 0. 0008) . A per-gate audit shows this lift is carried almost entirely by one high-precision gate (100% precision over 161 fires), while a second gate has only 5% precision, so gate precision must itself be audited.  \nThe effect is concentrated where the gates actually fire: on the 26/50 firing tasks, success rises by +19.2pp, while movement on the 24 non-firing tasks does not exclude zero. Two negative controls, a self-enforcing retail domain and BFCL, bound the mechanism: gates help when tools are policy-permissive and add little where tools already enforce their own preconditions. Finally, we report suggestive evidence, not a central claim, that the same failure mode persists in a frontier-model harness: gpt-5.2 at default reasoning still attempts policy-violating writes, and the same gate suite improves success from 61.2% to 71.6%(+10.4pp; 􀀥 = 0. 020; n=5, no replication) . The contribution is a bounded evaluation and reliability result: deterministic gates do not guarantee task success, but they can deterministically prevent a known class of silent policy-violating writes at the action boundary.  \nPermission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s) .  \nKDD-ETAAI’26, Jeju, Republic of Korea  \n© 2026 Copyright held by the owner/author(s) .  \nCCS Concepts  \n• Computing methodologies → Artificial intelligence; • Software and its engineering → Software safety; • Security and privacy → Formal methods and theory of security.  \nKeywords  \nLLM agents; tool use; policy compliance; deterministic verification; agent evaluation; runtime enforcement; agent reliability  \nACM Reference Format:  \nVikas Reddy, Sumanth Reddy Challaram, and Abhishek Basu. 2026. Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents. In KDD Workshop on Evaluation and Trustworthiness of Agentic AI (KDD-ETAAI’26), August 2026, Jeju, Republic of Korea. ACM, New York, NY, USA, 7 pages.  \n1 Introduction  \n1.1 A trust problem, not only a capability problem  \nWhen an LLM agent operating a customer-service tool cancels anon-refundable reservation, modifies a passenger coun","cbCaioXSIYDBL4lC","https://ap.wps.com/l/cbCaioXSIYDBL4lC","pdf",465689,3,1,7,"English","en",105,"# Abstract\n# Introduction\n## A trust problem, not only a capability problem\n## A reliability gap, not measurement noise","[{\"question\":\"What causes silent policy-violation failures in tool-using LLM agents?\",\"answer\":\"When tools are policy-permissive, they execute any well-formed tool call even if the domain policy forbids the resulting state transition. The tool may not raise errors, so the agent’s own trace can look successful despite an incorrect state.\"},{\"question\":\"How common are silent wrong-state failures in the airline domain studied?\",\"answer\":\"On the 2-bench airline tasks, 78% of observed failures are silent wrong-state failures where the final database is wrong but no tool error is raised.\"},{\"question\":\"How do deterministic readonly pre-execution gates improve agent reliability?\",\"answer\":\"The gates inspect proposed tool calls and the current database state before allowing any write. A four-gate suite increases full-benchmark success from 29.6% to 42.0%, and the improvement is largely driven by one high-precision gate.\"}]",1784186140,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"reason-less-verify-more-deterministic-gates-recover-a-silent-policy-violation-failure-mode-in-tool-using-llm-agents","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/reason-less-verify-more-deterministic-gates-recover-a-silent-policy-violation-failure-mode-in-tool-using-llm-agents/83237/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What causes silent policy-violation failures in tool-using LLM agents?","Question",{"text":75,"@type":76},"When tools are policy-permissive, they execute any well-formed tool call even if the domain policy forbids the resulting state transition. The tool may not raise errors, so the agent’s own trace can look successful despite an incorrect state.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How common are silent wrong-state failures in the airline domain studied?",{"text":80,"@type":76},"On the 2-bench airline tasks, 78% of observed failures are silent wrong-state failures where the final database is wrong but no tool error is raised.",{"name":82,"@type":73,"acceptedAnswer":83},"How do deterministic readonly pre-execution gates improve agent reliability?",{"text":84,"@type":76},"The gates inspect proposed tool calls and the current database state before allowing any write. A four-gate suite increases full-benchmark success from 29.6% to 42.0%, and the improvement is largely driven by one high-precision gate.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]