[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-151943-en":3,"doc-seo-151943-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},151943,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","ROC-n-reroll - How verifier imperfection affects test-time scaling","Test-time scaling improves language model performance by spending extra compute during inference, and many methods rely on a verifier to enable resampling such as Best-of-N and Rejection Sampling. This work develops theory for the impact of verifier imperfection, proving instance-level accuracy is determined by the geometry of the verifier’s ROC curve. Experiments with Qwen and LLaMA verifiers on GSM8K and MATH500 confirm that for concave ROC curves RS beats BoN at fixed compute, while both match in the infinite-compute limit. The paper also shows high-compute behavior cannot be reliably inferred from low-compute observations.","arXiv :2507 . 12399v3 [ cs .LG] 17 Aug 2026  \nROC-n-reroll: How verifier imperfection affects test-time scaling  \nFlorian E. Dorner 1,2,3,4 , Yatong Chen*3,4 , Andr´e F. Cruz*3,4 , and Fanny Yang 1  \n1 ETH Z¨urich, 2 Max Planck ETH Center for Learning Systems, 3 Max Planck Institute for  \nIntelligent Systems, T¨ubingen, 4 T¨ubingen AI Center  \nAugust 18, 2026  \nAbstract  \nTest-time scaling aims to improve language model performance by leveraging additional compute during inference. Many works have empirically studied techniques such as Best-of-N (BoN) and Rejection Sampling (RS) that make use of a verifier to enable test-time scaling. However, to date there is little theoretical understanding of how verifier imperfection affects performance — a gap we address in this work. Specifically, we prove that the instance-level accuracy of these methods is precisely characterized by the geometry of the verifier’s ROC curve. Our theory has two important takeaways, confirmed by experiments with Qwen and LLama models on GSM8K and MATH500 . First, for any query with a concave verifier ROC curve, RS outperforms BoN for fixed compute, while both methods converge to the same accuracy in the infinite-compute limit. Second, it is generally impossible to predict the highcompute performance of either method based on observations in the low-compute regime.  \n1 Introduction  \nJust as further scaling up large language model (LLM) pre-training started to show diminishing returns, OpenAI released o1, vastly improving upon the state-of-the-art on many challenging benchmarks [1] . Instead of spending more compute on pre-training, o1 was the first flagship LLM to prominently improve performance by spending additional compute at test-time. Since then, interest in test-time scaling has exploded [2–8] .  \nThere are two broad approaches to test-time scaling: resampling and “reasoning”. Both approaches typically use a verifier—a scoring mechanism that evaluates the quality or correctness of an LLM’s outputs—but at different stages of the pipeline. Resampling methods employ a verifier at test-time to filter or rank candidate responses after they are generated [9] . In contrast, reasoning methods employ a verifier to modify how the LLM generates outputs, usually increasing output quality at the cost of increased response length. Most prominently, the verifier can be used as a reward for post-training with reinforcement learning (RL) [3] .  \nIn practice, both test-time scaling approaches have primarily been successful in domains where a reliable oracle verifier can be implemented—e.g., coding using unit tests and math using ground-truth numerical solutions. As such, previous theoretical analysis has focused on the scaling behavior of pass@N, the probability that at least one of the N candidate responses is correct [10, 11] . In most domains, however, access to a perfectly accurate verifier is not realistic. Errors may slip through: insecure code can pass static tests [12], and flawed reasoning can arrive at the correct numerical answer [13] . More broadly, there has been an increasing interest in using another language model as a verifier [14, 15], an approach that can be applied to any domain, but has been shown to have far from perfect accuracy [16, 17] .  \n* Equal contribution.  \nTrue Pos itive Rate  \n1.00 0.75 0.50 0.25  \n0.00  \nVerifier ROC  \n0.0 0.5 1.0  \nFalse Positive Rate  \nAccuracy  \n0.8  \n0.6  \n0.4  \n0.2  \nRejection Sampling  \n5 10 15  \nNumber of Generations  \n0.8  \n0.6  \n0.4  \n0.2  \nBest-of-N  \n\n|  | |  | \u003Cbr>|  |\n| --- | --- | --- | --- | --- |\n| | |  |  |  |\n|  |  | | Q 32B |  |\n| |  | \u003Cbr>| Q 14B\u003Cbr>Q 4B |  |\n| |  |  | Predicted |  |\n|  |  |  |  |  |\n|  |  |  |  |  |\n\n0 10 20 30  \nNumber of Generations  \nFigure 1: Empirical performance (markers) of RS (middle) and BoN (right) on GSM8K test question 58, overlaid with theoretical predictions (lines) . Different verifiers scale similarly at first, but then diverge. RS matches BoN accuracy, using less","cbCaij7sH6pi8CMd","https://ap.wps.com/l/cbCaij7sH6pi8CMd","pdf",5172165,1,47,"English","en",105,"# Introduction\n## Two approaches to test-time scaling\n## Verifier reliability and theoretical gap\n# Main contributions\n## ROC-curve characterization of accuracy\n## RS vs BoN under concave ROC curves\n## Empirical validation on GSM8K and MATH500","[{\"question\":\"Can high-compute performance be predicted from low-compute observations?\",\"answer\":\"No; the paper argues that extrapolation from the low-compute regime to high-compute performance is generally impossible for these methods.\"}]","ROC-n-reroll - How verifier imperfection affects test-time scaling | PDF",1787853257,118,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"roc-n-reroll-how-verifier-imperfection-affects-test-time-scaling","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/roc-n-reroll-how-verifier-imperfection-affects-test-time-scaling/151943/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-04","2026-08-27",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"Can high-compute performance be predicted from low-compute observations?","Question",{"text":76,"@type":77},"No; the paper argues that extrapolation from the low-compute regime to high-compute performance is generally impossible for these methods.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":85},[86,90,94,98,103,108,113,116,121,124,128],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":99,"doc_module":4,"doc_module_name":46,"category_name":100,"show_sort_weight":101,"slug":102},5,"Comic",60,"comic",{"id":104,"doc_module":4,"doc_module_name":46,"category_name":105,"show_sort_weight":106,"slug":107},6,"Technology",50,"technology",{"id":109,"doc_module":4,"doc_module_name":46,"category_name":110,"show_sort_weight":111,"slug":112},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":114,"slug":115},30,"research-report",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},9,"Religion & Spirituality",20,"religion-spirituality",{"id":119,"doc_module":4,"doc_module_name":46,"category_name":122,"show_sort_weight":119,"slug":123},"World Cup","world-cup",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":125,"slug":127},10,"Lifestyle","lifestyle",{"id":129,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":99,"slug":131},19,"General","general"]