[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85843-en":3,"doc-seo-85843-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85843,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Equal Accuracy, Unequal Evidence: Search APIs as Decision Surfaces for Tool-Using Agents","Search APIs form the core retrieval layer for many tool-using agents and are often used most frequently. Traditional evaluations emphasize answer accuracy from URL, title, and snippet previews, while full-page retrieval remains token-intensive and motivates progressive disclosure. This work frames a commercial search API as a decision surface: the ranked evidence exposed before fetching shapes whether an agent answers, searches again, or spends tokens opening pages. Experiments with a fixed agent over SEALQA-HARD compare Brave, Tavily, and Firecrawl, showing similar accuracy but distinct evidence economies and contradiction-to-gold ratios, making provider choice a retrieval-budget and policy decision.","Equal Accuracy, Unequal Evidence: Search APIs as Decision Surfaces for  \nTool-Using Agents  \nSriram Selvam  \n[selvamsriram@gmail.com](selvamsriram@gmail.com)  \nAnneswa Ghosh  \n[anneswaghosh@gmail.com](anneswaghosh@gmail.com)  \narXiv :2607 . 10 198v 1 [ cs .CL] 11 Jul 2026  \nAbstract  \nSearch APIs are the fundamental retrieval layer for many agents and are often their most frequently used tool. Traditional search APIs provide URLs, titles, and snippets that preview website contents. Because full-page retrieval is token-intensive, agent retrieval architectures increasingly use progressive disclosure:  \nthe agent first sees snippets and then chooses whether to fetch full pages. In such systems, search API performance is often evaluated primarily by answer accuracy. We argue that a commercial search API is better understood as a decision surface: the ranked snippets, URLs, and metadata that determine whether an agent answers immediately, searches again, or spends tokens opening pages. We test this claim with one frozen GPT-5.4 agent, two tools (search_web and fetch_page), and 100 questions from SEALQA-HARD, varying only the search provider (Brave, Tavily, Firecrawl) . A Kimi-K2.6 oracle labels every content element visible to the agent (URL, title, snippet, and fetched page, when fetched), producing 6,869 valid per-URL judgments. We use an audited correct-answer label, semantic_match, which preserves exact matches while accepting harmless formatting and naming variants. Under this measure, the providers remain close (25, 25, 26 / 100), but their evidence economies differ sharply: Brave offers gold-answer-rich snippets, Tavily concentrates gold-supporting URLs at rank 1, and Firecrawl is associated with broader exploration under this fixed agent policy. We also introduce a surface contradiction-to-gold URL ratio, which varies from 0 .92 to 2 .59. Provider choice is therefore a retrieval-budget and policy decision, not merely a recall decision.  \n1 Introduction  \nEarly search-augmented agents often depended primarily on extracted page data provided by search API providers. As complex use cases have evolved,  \nthe agent landscape has shifted toward progressive disclosure: models receive top-n URLs, titles, snippets, and key metadata such as freshness, and then decide whether to fetch full pages. In this new landscape, engineering practice often treats search API providers as replaceable components. If two APIs return plausible top-k URLs and produce similar answer accuracy, the rest of the agent pipeline is assumed to be largely unaffected.  \nThis paper argues that the relevant object is the pre-fetch surface. Before a page is opened, an agent does not see a search index or a corpus; it relies on search API results. That surface is the evidence state on which the agent decides whether to answer, search again, or spend fetch tokens. We call commercial search APIs decision surfaces for tool-using language agents.  \nFor grounded language-model systems, the retrieval API is part of the grounding interface. It determines which evidence is exposed before generation, which contradictions enter context, and how much retrieval budget is spent before an answer is produced. Decision-surface evaluation therefore measures not only retrieval quality, but also the faithfulness and efficiency conditions under which grounded answers are generated.  \nThe distinction matters because final correctness can converge while the internal pipeline diverges. A provider may expose enough snippet evidence for immediate answering. Another may put the relevant URL at rank 1, making a top-result fetch policy effective. A third may expose sparse snippets and be associated with broader exploration under the same policy. These regimes can yield similar aggregate accuracy with different cost, latency, contamination, and failure modes.  \nUnder a controlled protocol, we freeze the answer model, prompt, tools, maximum iterations, judge, and page-fetch backend. The only ex","cbCaiak0rBAp2mvY","https://ap.wps.com/l/cbCaiak0rBAp2mvY","pdf",1087767,4,1,15,"English","en",105,"# Abstract\n# Introduction\n# Related Work","[{\"question\":\"为什么本文将搜索 API 称为“决策面（decision surface）”？\",\"answer\":\"因为在页面被打开之前，代理看到的排序片段、URL与元数据会决定它是立刻作答、再次检索还是消耗代币去抓取页面。评估重点因此不只是检索好坏，还包括决策时的证据状态与效率条件。\"},{\"question\":\"实验中有哪些关键控制变量与变化因素？\",\"answer\":\"本文冻结答案模型、提示、工具、最大迭代次数、评判器以及页面抓取后端。唯一实验条件是搜索 API 提供方在 Brave、Tavily 与 Firecrawl 之间变化。\"},{\"question\":\"结果显示三家搜索提供方的效果差异体现在哪里？\",\"answer\":\"在 audited correct-answer 与 semantic_match 指标下，它们的答案正确性接近（约 25/25/26）。但证据经济差异显著：Brave 更偏向富含“金答案”片段，Tavily 更倾向把支持金答案的 URL 集中到第 1 名，而 Firecrawl 更关联到在固定策略下的更广探索，并伴随不同的表面矛盾到金 URL 比率。\"}]",1784206657,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"equal-accuracy-unequal-evidence-search-apis-as-decision-surfaces-for-tool-using-agents","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/equal-accuracy-unequal-evidence-search-apis-as-decision-surfaces-for-tool-using-agents/85843/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么本文将搜索 API 称为“决策面（decision surface）”？","Question",{"text":75,"@type":76},"因为在页面被打开之前，代理看到的排序片段、URL与元数据会决定它是立刻作答、再次检索还是消耗代币去抓取页面。评估重点因此不只是检索好坏，还包括决策时的证据状态与效率条件。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"实验中有哪些关键控制变量与变化因素？",{"text":80,"@type":76},"本文冻结答案模型、提示、工具、最大迭代次数、评判器以及页面抓取后端。唯一实验条件是搜索 API 提供方在 Brave、Tavily 与 Firecrawl 之间变化。",{"name":82,"@type":73,"acceptedAnswer":83},"结果显示三家搜索提供方的效果差异体现在哪里？",{"text":84,"@type":76},"在 audited correct-answer 与 semantic_match 指标下，它们的答案正确性接近（约 25/25/26）。但证据经济差异显著：Brave 更偏向富含“金答案”片段，Tavily 更倾向把支持金答案的 URL 集中到第 1 名，而 Firecrawl 更关联到在固定策略下的更广探索，并伴随不同的表面矛盾到金 URL 比率。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]