[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-151944-en":3,"doc-seo-151944-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},151944,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","ROC-N-REROLL - HOW VERIFIER IMPERFECTION AFFECTS TEST-TIME SCALING - Paper","Test-time scaling improves language model performance by using extra compute during inference. Many approaches rely on a verifier via Best-of-N sampling and Rejection Sampling, yet verifier imperfection lacks a clear theoretical account. This work proves that instance-level accuracy is determined by the geometry of the verifier’s ROC curve. Experiments with Qwen and LLaMA models on GSM8K and MATH500 confirm: for concave ROC curves, RS beats BoN at fixed compute and both meet as compute grows large.","ROC-N-REROLL: HOW VERIFIER IMPERFECTION AFFECTS TEST-TIME SCALING  \nFlorian E. Dorner 1,2,3,4 , Yatong Chen* 3,4 , Andr F. Cruz* 3,4 , and Fanny Yang 1  \n1ETH Z¨urich, 2Max Planck ETH Center for Learning Systems, 3Max Planck Institute for Intelligent Systems, T¨ubingen, 4T¨ubingen AI Center  \nABSTRACT  \nTest-time scaling aims to improve language model performance by leveraging additional compute during inference. Many works have empirically studied techniques such as Best-of-N (BoN) and Rejection Sampling (RS) that make use of averifier to enable test-time scaling. However, to date there is little theoretical understanding of how verifier imperfection affects performance—a gap we address in this work. Specifically, we prove that the instance-level accuracy of these methods is precisely characterized by the geometry of the verifier’s ROC curve. Our theory has two important takeaways, confirmed by experiments with Qwen and LLama models on GSM8K and MATH500 . First, for any query with a concave verifier ROC curve, RS outperforms BoN for fixed compute, while both methods converge to the same accuracy in the infinite-compute limit. Second, it is generally impossible to predict the high-compute performance of either method based on observations in the low-compute regime.  \n1 INTRODUCTION  \nJust as further scaling up large language model (LLM) pre-training started to show diminishing returns, OpenAI released o1, vastly improving upon the state-of-the-art on many challenging benchmarks (OpenAI, 2024) . Instead of spending more compute on pre-training, o1 was the first flagship LLM to prominently improve performance by spending additional compute at test-time. Since then, interest in test-time scaling has exploded (Muennighoff et al., 2025; Guo et al., 2025; Kimi et al., 2025; Qu et al., 2025; Aggarwal and Welleck, 2025; Zaremba et al., 2025; Kavukcuoglu, 2025) .  \nThere are two broad approaches to test-time scaling: resampling and “reasoning”. Both approaches typically use a verifier—a scoring mechanism that evaluates the quality or correctness of an LLM’soutputs—but at different stages of the pipeline. Resampling methods employ a verifier at testtime to filter or rank candidate responses after they are generated (Cobbe et al., 2021) . In contrast, reasoning methods employ a verifier to modify how the LLM generates outputs, usually increasing output quality at the cost of increased response length. Most prominently, the verifier can be used asa reward for post-training with reinforcement learning (RL) (Guo et al., 2025) .  \nIn practice, both test-time scaling approaches have primarily been successful in domains where a reliable oracle verifier can be implemented—e.g., coding using unit tests and math using groundtruth numerical solutions. As such, previous theoretical analysis has focused on the scaling behavior of pass@N, the probability that at least one of the N candidate responses is correct (Brown et al., 2024; Schaeffer et al., 2025) . In most domains, however, access to a perfectly accurate verifier is not realistic. errors may slip through: insecure code can pass static tests (Zhou et al., 2024), and flawed reasoning can arrive at the correct numerical answer (Petrov et al., 2025) . More broadly, there has been an increasing interest in using another language model as a verifier (Huang et al., 2025a; Songet al., 2025), an approach that can be applied to any domain, but has been shown to have far from perfect accuracy (Bavaresco et al., 2025; Dorner et al., 2025) .  \nDespite growing interest in verifier-based test-time scaling, the relationship between scaling behavior and the properties of imperfect verifiers remains poorly understood. This work addresses this  \n* Equal contribution. Code: [https://github.com/socialfoundations/roc-n-reroll](https://github.com/socialfoundations/roc-n-reroll)  \nTrue Pos itive Rate  \n1.00 0.75 0.50 0.25  \n0.00  \nVerifier ROC  \n0.0 0.5 1.0  \nFalse Positive Rate  \nAccuracy  \n0.8  \n0.6  \n0.4  \n0.2  ","cbCaisxSA2SFPRBT","https://ap.wps.com/l/cbCaisxSA2SFPRBT","pdf",787096,1,42,"English","en",105,"# Introduction\n## Test-time scaling approaches\n## Role and limits of imperfect verifiers\n## Contributions overview","[{\"question\":\"What does ROC-N-REROLL study about verifier imperfection in test-time scaling?\",\"answer\":\"It analyzes how the accuracy of verifier-based methods changes when the verifier is imperfect, showing that instance-level performance is governed by the verifier ROC geometry.\"},{\"question\":\"How do Rejection Sampling (RS) and Best-of-N (BoN) relate through the verifier ROC curve?\",\"answer\":\"For a fixed query, both methods’ accuracy depends on the generator’s initial accuracy and the verifier’s ROC curve, not on finer implementation details.\"},{\"question\":\"Why can’t high-compute performance be reliably predicted from low-compute behavior?\",\"answer\":\"The paper shows extrapolation from early scaling observations is generally impossible, because different verifiers can diverge at higher numbers of test-time samples even when low-compute accuracy matches.\"}]","ROC-N-REROLL - HOW VERIFIER IMPERFECTION AFFECTS TEST-TIME SCALING - Paper | PDF",1787853259,106,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"roc-n-reroll-how-verifier-imperfection-affects-test-time-scaling-paper","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/roc-n-reroll-how-verifier-imperfection-affects-test-time-scaling-paper/151944/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-09-04","2026-08-27",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What does ROC-N-REROLL study about verifier imperfection in test-time scaling?","Question",{"text":76,"@type":77},"It analyzes how the accuracy of verifier-based methods changes when the verifier is imperfect, showing that instance-level performance is governed by the verifier ROC geometry.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How do Rejection Sampling (RS) and Best-of-N (BoN) relate through the verifier ROC curve?",{"text":81,"@type":77},"For a fixed query, both methods’ accuracy depends on the generator’s initial accuracy and the verifier’s ROC curve, not on finer implementation details.",{"name":83,"@type":74,"acceptedAnswer":84},"Why can’t high-compute performance be reliably predicted from low-compute behavior?",{"text":85,"@type":77},"The paper shows extrapolation from early scaling observations is generally impossible, because different verifiers can diverge at higher numbers of test-time samples even when low-compute accuracy matches.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]