[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82034-en":3,"doc-seo-82034-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},82034,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models","Routing among large language models (LLMs) balances response quality with serving cost, reflecting the gap between deployed routers and a hypothetical per-instance oracle. Test-time resampling can recover selection headroom, yet the guarantee relies on an ideal, label-perfect verifier and an unconstrained oracle budget. This work introduces a budget-aware formulation with an imperfect verifier: allocate a per-query cost between resampling the committed model and rerouting to alternatives to maximize expected correctness. An online resample-or-reroute policy is evaluated via replay experiments and robustness analyses.","Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models  \nTeng-Ruei Chen  \narXiv :2607 .08665v2 [ cs .LG] 10 Jul 2026  \nAbstract—Routing among large language models (LLMs) trades response quality against serving cost, motivated by the reported gap between deployed routers and a per-instance oracle. Recent analysis shows that test-time resampling can recover per-instance selection headroom that no single-commit router captures; however, that guarantee holds only under an idealized oracle equipped with correctness labels and an unconstrained budget, neither of which a deployed system has. To the best of our knowledge, no previous work treats resampling the committed model and rerouting to an alternative model as competing uses of a single per-query cost budget. Therefore, this work formulates budget-aware test-time model selection: given a per-query budget and an imperfect verifier, allocate each unit of budget between resampling and rerouting so that expected correctness is maximized. An online resample-or-reroute (RoR) allocation policy driven by estimated marginal correctness per unit cost is proposed, and its behavior is grounded in the recoverability asymmetry between selection and sampling. Replay experiments on newly regenerated multi-draw correctness tensors from an eleven-model open-weight pool over four benchmarks of differing difficulty show that the proposed RoR policy attains a favorable cost–quality Pareto front relative to single-route, one-commitrouter, budget-aware best-of-K, cascade, and random-allocation baselines for the tested pools, with the largest gains on the most heterogeneous benchmark; an ablation further shows the gains are verifier-gated, shrinking as verifier quality degrades, and robustness replays under a provider price vector and a label-free agreement verifier delineate where the conclusions carry over.  \nIndex Terms—Best-of-K sampling, cost-aware inference, inference efficiency, large language models, model routing, modelselection, test-time compute.  \nI. INTRODUCTION  \nSERVING a query with large language models (LLMs)  \nincreasingly means choosing how to spend inference compute, not just which model to call: inference-time cost and energy are first-class deployment constraints [1], costand QoS-aware resource management for ML services isan established concern of serving infrastructures [2], and the space of candidate models keeps broadening [3] . Model routing promises to cut cost by sending each query to the cheapest model that can answer it, motivated by the large reported gap between a deployed router and a per-instance oracle that, in hindsight, always picks a correct model [4]–[7] .  \nTwo facts complicate this picture. First, under stochastic decoding the per-instance oracle is not a reproducible property: it is built from single draws, so part of the reported gap is single-draw label noise that no single-commit router  \nT.-R. Chen is with the Institute of Bioinformatics and Systems Biology, National Yang Ming Chiao Tung University, Hsinchu 300, Taiwan, and also with Krixvon, Taipei 100, Taiwan (e-mail: [ymchen.bi04g@g2.nctu.edu.tw](ymchen.bi04g@g2.nctu.edu.tw)).  \ncan capture, while the rest is genuine, recoverable specialist advantage [8] . Second, that recoverable component can also be reached without any router at all: test-time resampling (best-of-K on one committed model) provably recovers the selection floor at the oracle’s own budget [8] . The catch is that this guarantee assumes access to correctness labels (a perfect verifier) and spends the oracle’s budget; a real serving system has an imperfect verifier and a fixed cost budget per query.  \nThis work addresses the operational question these results leave open: given a per-query cost budget and an imperfect verifier, should a system resample the model it already committed to, or reroute to a different (possibly more expensive) model? The question is cast as a budgeted correctnessmaximization proble","cbCaig3eLKaBZnUx","https://ap.wps.com/l/cbCaig3eLKaBZnUx","pdf",363611,6,1,10,"English","en",105,"# Introduction\n## Problem setup and motivation\n## Budgeted correctness maximization\n# Proposed method\n## Online resample-or-reroute (RoR) policy\n# Evaluation\n## Replay experiments and baselines\n## Verifier ablation and robustness replays","[{\"question\":\"What problem does the paper address in LLM routing?\",\"answer\":\"It addresses how to choose between spending compute by resampling the already committed model or rerouting to another model when each query has a fixed cost budget and the verifier is imperfect.\"},{\"question\":\"How is budget-aware test-time model selection formulated?\",\"answer\":\"Given a per-query budget and an imperfect verifier, the method allocates budget units between resampling and rerouting to maximize expected correctness.\"},{\"question\":\"What does the evaluation show about the proposed RoR policy?\",\"answer\":\"Replay experiments on regenerated multi-draw correctness tensors across multiple benchmarks report a favorable cost–quality Pareto front versus several baselines, with larger gains on the most heterogeneous benchmark; improvements are gated by verifier quality and diminish as verifier quality degrades.\"}]","Resample or Reroute? Budget-Aware Test-Time Model Selection for Large Language Models | PDF",1784177715,25,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"resample-or-reroute-budget-aware-test-time-model-selection-for-large-language-models","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/resample-or-reroute-budget-aware-test-time-model-selection-for-large-language-models/82034/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What problem does the paper address in LLM routing?","Question",{"text":77,"@type":78},"It addresses how to choose between spending compute by resampling the already committed model or rerouting to another model when each query has a fixed cost budget and the verifier is imperfect.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How is budget-aware test-time model selection formulated?",{"text":82,"@type":78},"Given a per-query budget and an imperfect verifier, the method allocates budget units between resampling and rerouting to maximize expected correctness.",{"name":84,"@type":75,"acceptedAnswer":85},"What does the evaluation show about the proposed RoR policy?",{"text":86,"@type":78},"Replay experiments on regenerated multi-draw correctness tensors across multiple benchmarks report a favorable cost–quality Pareto front versus several baselines, with larger gains on the most heterogeneous benchmark; improvements are gated by verifier quality and diminish as verifier quality degrades.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,116,121,124,129,132,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":22,"slug":134},"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":108,"slug":138},19,"General","general"]