[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83767-en":3,"doc-seo-83767-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83767,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","The “I Don’t Know” Filter: Enhancing Agentic Reliability in Function Calling","Language models powering agents have advanced rapidly on function-calling benchmarks, yet training and evaluation metrics often reward positive answers even when uncertainty remains, causing hallucinations. For high-stakes settings that rely on function calls, such hallucinations can lead to severe failures. This work proposes an agent evaluation metric that penalizes incorrect function calls and introduces a lightweight, trainable filter to quantify uncertainty and suppress potentially harmful calls, enabling agents to safely respond with “I don’t know.”","The “I Don’t Know” Filter: Enhancing Agentic Reliability in Function Calling  \nStefan Broecker1 , Mason del Rosario2 , Boris Selitser2 , Thomas Strohmer3  \narXiv :2607 .04034v 1 [ cs . SE] 4 Jul 2026  \nAbstract—The language models that underpin agents have seen a rapid rise in performance on function calling benchmarks. However, the metrics used in the training and evaluation of these models often encourage models to make positive claims even when the answer is uncertain, leading to hallucinations. Such hallucinations can be disastrous when language models are trusted to use function calls to make decisions in high stakes applications. To that end, we propose an agent evaluation metric that takes into account the negative outcomes associated withincorrect function calls. Further, to catch hallucinations before they can cause harm, we propose a lightweight trainable filter that can quantify a language model’s uncertainty and remove potentially harmful function calls. By training that filter to detect and suppress uncertain function calls without modifying the underlying model, we demonstrate a practical path toward agents that know when to say “I don’t know,” a property we argue is essential to production reliability.  \nI. INTRODUCTION  \nAgents are language models (LMs) that combine tools, memory, and reasoning to perform tasks [1] . As agents have improved at tool use, developers in critical industries such as digital infrastructure [2], health care, and public services [3] have begun to trust agents to make decisions in production settings. As an example, errant tool calls in such production environments have resulted in airfares offered at disallowed discounts and eggs purchased at exorbitant prices [4] .  \nA large scale study of agents in production showed a persistent gap in agent performance on benchmarks and in production systems [5] . In that study, reliability was most commonly cited as the bottleneck to agent deployment. A key part of that reliability gap is hallucinations, and given agents’ability to take autonomous action in real-world decision making, hallucinations can cause much more severe and material damage than standalone LMs.  \n[6] makes the claim that hallucinations persist because of conventions in LM evaluations. In most evaluations, an incorrect answer is weighed the same as a non-answer, so guessing when uncertain improves accuracy. However, Bastounis et al. state that this tendency to provide a nonanswer when uncertain is a fundamental roadblock to “trustworthy” (i.e., reliable) AI, and they describe this tendency to hallucinate and the inability to state “I don’t know” as a critical component of the Consistent Reasoning Paradox [7] . In this work, we study “agent reliability” through this lens: the agent’s ability to express uncertainty, i.e., by explicitly  \n1University of California, Davis, Department of Computer Science, [sbroecker@ucdavis.edu](sbroecker@ucdavis.edu)  \n2 Okareo, Inc., {mason, [boris](boris}@okareo.com)[}](boris}@okareo.com)[@okareo.com](boris}@okareo.com)  \n3University of California, Davis, Department of Mathematics, [strohmer@math.ucdavis.edu](strohmer@math.ucdavis.edu)  \nstating “I don’t know”(IDK) . Towards this end, we make the following contributions:  \n• An accuracy metric that penalizes incorrect function calls, emphasizing the danger of an autonomous agent guessing at an answer instead of admitting uncertainty and encouraging IDK responses  \n• A simple filter that uses “whitebox,” “graybox,” and“blackbox” features to identify likely-incorrect function calls (cases of IDK) to reduce the likelihood of hallucinations  \n• A simple method for generating synthetic data that can be used to train that filter, enabling deployment even when labeled function calling data is limited  \n• A study of the filter’s effectiveness in reducing uncertainty across different open source LMs and function calling benchmarks  \nWith these contributions, we demonstrate that explicitly measuring an age","cbCaid82qgbcrNf9","https://ap.wps.com/l/cbCaid82qgbcrNf9","pdf",616120,4,1,6,"English","en",105,"# Introduction\n# Background\n## Uncertainty quantification for LMs\n## Reliable function calling for agents","[{\"question\":\"Why do hallucinations become a reliability risk in agent function calling?\",\"answer\":\"Function-calling agents can take autonomous actions based on uncertain outputs, and evaluation conventions may encourage guessing instead of admitting uncertainty. In high-stakes applications, incorrect function calls can cause serious real-world damage.\"},{\"question\":\"What evaluation contribution does the paper propose?\",\"answer\":\"It proposes an agent evaluation metric that accounts for negative outcomes tied to incorrect function calls, penalizing confident-but-wrong behavior and promoting “I don’t know” responses.\"},{\"question\":\"How does the proposed filter improve reliability without changing the base model?\",\"answer\":\"The work introduces a lightweight trainable filter that estimates a language model’s uncertainty and removes potentially harmful function calls. The filter is trained to suppress uncertain calls without modifying the underlying language model.\"}]",1784190291,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-i-dont-know-filter-enhancing-agentic-reliability-in-function-calling","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/the-i-dont-know-filter-enhancing-agentic-reliability-in-function-calling/83767/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do hallucinations become a reliability risk in agent function calling?","Question",{"text":75,"@type":76},"Function-calling agents can take autonomous actions based on uncertain outputs, and evaluation conventions may encourage guessing instead of admitting uncertainty. In high-stakes applications, incorrect function calls can cause serious real-world damage.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What evaluation contribution does the paper propose?",{"text":80,"@type":76},"It proposes an agent evaluation metric that accounts for negative outcomes tied to incorrect function calls, penalizing confident-but-wrong behavior and promoting “I don’t know” responses.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed filter improve reliability without changing the base model?",{"text":84,"@type":76},"The work introduces a lightweight trainable filter that estimates a language model’s uncertainty and removes potentially harmful function calls. The filter is trained to suppress uncertain calls without modifying the underlying language model.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]