[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82993-en":3,"doc-seo-82993-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82993,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","When Should LLMs Search Counterfactual Supervision for Search Routing","Search-augmented language models can leverage external evidence to overcome gaps in parametric knowledge, yet search can be unhelpful or harmful when it is called on questions the model already answers or when noisy evidence is used instead of correction, clarification, or abstention. This work formulates instance-level search routing as a decision problem: whether search improves task success relative to a no-search execution. Paired no-search vs forced-search outcomes define an oracle over NO SEARCH, SEARCH, and UNSOLVED, training routing policies via supervised fine-tuning and preference optimization.","When Should LLMs Search? Counterfactual Supervision for Search Routing  \nMinho Kim 1 2  \narXiv :2607 .05752v 1 [ cs .CL] 7 Jul 2026  \nAbstract  \nSearch-augmented language models can use external evidence to compensate for limitations in parametric knowledge, but search is not uniformly beneficial: models may call search for questions they can already answer, or rely on noisy evidence when correction, clarification, or abstention would be more appropriate. We formulate this as an instance-level search-routing problem:  \ndeciding whether search is needed to improve task success relative to a no-search execution. To derive supervision, we compare no-search and forced-search outcomes for the same question and construct an oracle over NO   SEARCH, SEARCH, and UNSOLVED based on task-specific success.  \nUsing this oracle as both an evaluation criterion and a learning signal, we train search-routing policies with supervised fine-tuning and preference optimization, improving routing macro-F1 on oracle-eligible examples from 0.7082 to 0.8235 for Gemma E2B and from 0.7053 to 0.8365 for Qwen3.5-4B. Further analysis shows that the learned policies reduce model-specific routing failures: Gemma primarily learns no-search restraint, while Qwen further reduces missed search;  \nresidual UNSOLVED cases reveal heterogeneous bottlenecks involving model capacity, retrieval budget, evidence use, and policy behavior.  \n1. Introduction  \nLanguage models equipped with search tools can use external evidence to answer questions that are difficult to resolve from parametric knowledge alone, including queries about long-tail factual knowledge, recently updated information, or facts that are unlikely to be reliably stored in model parameters (Lewis et al., 2020 ; Izacard et al., 2023 ; Mallen  \n1 Sangmyung University 2DMTLABS. Correspondence to: Minho Kim \u003C[ilw2y123@gmail.com](ilw2y123@gmail.com) >.  \nAccepted at the FAGEN Workshop at the 43 rd International Conference on Machine Learning, Seoul, South Korea, 2026 . Copyright 2026 by the author.  \nFigure 1. Motivation for instance-level search routing. Search can help, be unnecessary, or mislead, motivating instance-level routing.  \net al., 2023) . However, the availability of a search tool does not imply that search should be used for every question. For questions the model can already answer, responding without search may be cheaper and less vulnerable to noisy retrieved evidence. For questions with false premises or missing context, the appropriate response may be to correct the premise, ask for clarification, or abstain, rather than retrieve external evidence (Amayuelas et al., 2024 ; Xie et al., 2026) . In practice, the correct first action depends on the instance: a long-tail factual question may require search to recover evidence, whereas an underspecified question maybe better resolved by clarification than by retrieving generic search results, as illustrated in Figure 1.  \nWe therefore view search-augmented language models not merely as systems that generate search queries, but as systems that must first perform search routing: deciding, for each input question, whether to respond without search or  \ncall a search tool. This first-action decision is important because routing errors arise in two opposing directions. A model that is too reluctant to search may answer directly even when its parametric knowledge is insufficient, producing plausible but incorrect responses. Conversely, a model that relies too heavily on search may invoke retrieval for questions it can already answer, or for questions that search cannot resolve, increasing cost and exposing the final answer to irrelevant or misleading evidence. The objective is not aggregate search frequency, but calling search only when it improves task success.  \nThis framing differs from standard tool-use evaluation. Many tool-calling benchmarks focus on whether a model calls the correct tool with the correct arguments once tool use is known to be app","cbCaiq6tNJXstkrO","https://ap.wps.com/l/cbCaiq6tNJXstkrO","pdf",2977439,4,1,20,"English","en",105,"# Abstract\n# Introduction\n## Motivation and problem framing\n## Research questions\n## Counterfactual traces and routing oracle","[{\"question\":\"What problem does the paper address in search-augmented language models?\",\"answer\":\"It addresses when a model should perform external search versus responding directly, since search can be unnecessary or even misleading for some instances.\"},{\"question\":\"How is the supervision signal constructed for search routing?\",\"answer\":\"For each question, the paper compares a no-search trace with a forced-search trace, then uses their task-specific success to label outcomes with an oracle over NO SEARCH, SEARCH, and UNSOLVED.\"},{\"question\":\"How are routing policies trained and evaluated in the study?\",\"answer\":\"The oracle is used as both an evaluation criterion and a learning signal, training search-routing policies with supervised fine-tuning and preference optimization, which improves routing macro-F1 on oracle-eligible examples.\"}]",1784184513,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"when-should-llms-search-counterfactual-supervision-for-search-routing","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/when-should-llms-search-counterfactual-supervision-for-search-routing/82993/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in search-augmented language models?","Question",{"text":75,"@type":76},"It addresses when a model should perform external search versus responding directly, since search can be unnecessary or even misleading for some instances.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the supervision signal constructed for search routing?",{"text":80,"@type":76},"For each question, the paper compares a no-search trace with a forced-search trace, then uses their task-specific success to label outcomes with an oracle over NO SEARCH, SEARCH, and UNSOLVED.",{"name":82,"@type":73,"acceptedAnswer":83},"How are routing policies trained and evaluated in the study?",{"text":84,"@type":76},"The oracle is used as both an evaluation criterion and a learning signal, training search-routing policies with supervised fine-tuning and preference optimization, which improves routing macro-F1 on oracle-eligible examples.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]