[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84818-en":3,"doc-seo-84818-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84818,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","LLM-as-a-Verifier: A General-Purpose Verification Framework","The work introduces LLM-as-a-Verifier, a general-purpose verification framework that treats correctness determination as a new scaling axis for large language models. It replaces discrete judge scores with a probabilistic expectation over scoring-token logits to produce continuous feedback and reduce tie rates. Verification can scale across score granularity, repeated evaluation, and criteria decomposition, improving calibration and accuracy via variance and complexity reduction. A cost-efficient ranking algorithm selects the best candidate using verifier preference probabilities. Results span coding, robotics, and medical domains, and fine-grained signals support progress tracking and dense-reward RL for SAC/GRPO.","arXiv :2607 .0539 1v2 [ cs .AI ] 7 Jul 2026  \n\n| LLM-as-a-Verifier: A General-Purpose Verification Framework\u003Cbr>Jacky Kwok1 , Shulu Li2 , Pranav Atreya2 , Yuejiang Liu1 , Yixing Jiang1 Chelsea Finn1 , Marco Pavone1 ,3 , Ion Stoica2 , Azalia Mirhoseini1\u003Cbr>1 Stanford University 2UC Berkeley 3NVIDIA Research |\n| --- |\n|  |\n| Figure 1: Overall Performance Results. Our proposed framework, LLM-as-a-Verifier, achieves stateof-the-art performance across coding, robotics, and medical domains: Terminal-Bench V2 (86.5%), SWE-Bench Verified (78.2%), RoboRewardBench (87.4%), and MedAgentBench (73.3%) .\u003Cbr>Abstract: Scaling pre-training, post-training, and test-time compute have become the central paradigms for improving the capabilities of large language models (LLMs) . In this work, we identify verification—the ability to determine the correctness of a solution—as a new scaling axis. To unlock this and demonstrate its effectiveness, we introduce LLM-as-a-Verifier, a general-purpose verification framework that provides finegrained feedback for agentic tasks without requiring additional training. Unlike standard LM judges that prompt LLMs to produce discrete scores for candidate solutions, LLM-as-a-Verifier computes the expectation over the distribution of scoring token logits to generate continuous scores. This probabilistic formulation substantially reduces tie rates when comparing complex solutions and enables verification to scale along multiple dimensions: (1) score granularity,(2) repeated evaluation, and (3) criteria decomposition. In particular, we show that scaling the scoring granularity leads to better separation between positive and negative solutions, resulting in more calibrated comparisons. Moreover, scaling repeated evaluation and criteria decomposition consistently leads to additional gains in verification accuracy through variance and complexity reduction. To make verification scaling practical, we further introduce a cost-efficient ranking algorithm for selecting the best solution among candidates using the preference probabilities derived from the verifier’s continuous scores. LLM-as-a-Verifier is effective across coding, robotics, and medical domains. It achieves state-of-the-art performance on Terminal-Bench V2 (86.5%), SWE-Bench Verified (78.2%), RoboRewardBench (87.4%), and MedAgentBench (73.3%) . Beyond verification, the fine-grained signals from LLM-as-a-Verifier can also serve as a proxy for estimating task progress. We build extensions for Claude Code and Codex, enabling developers to monitor and improve their own agentic systems. Finally, we show that LLM-as-a-Verifier can be used as a dense reward signal for RL, improving the sample efficiency of SAC and GRPO on robotics and mathematical reasoning benchmarks. |\n\n [llm-as-a-verifier.com](llm-as-a-verifier.com)  llm-as-a-verifier  claude-code-extension  \nCorresponding author: Jacky Kwok \u003C[jackykwok@stanford. edu](jackykwok@stanford. edu) >  \nFigure 2: Multiple modalities, many applications, one unified verification framework. We present LLM-as-a-Verifier, a general-purpose framework that provides fine-grained feedback for any modality without requiring additional training. By leveraging the full distribution of scoring-token logits, our method captures evaluation uncertainty and enables verification to scale along three dimensions: score granularity, repeated evaluation, and criteria decomposition. The resulting fine-grained feedback can be used for test-time scaling, progress tracking, and reinforcement learning.  \n1. Introduction  \nRecent advances in large language models (LLMs) have established scaling as a central paradigm for improving their capabilities. Performance has been driven by scaling along multiple axes, including pre-training data and compute, post-training optimization, and test-time inference [1–3] .  \nHowever, while generation has benefited significantly from these scaling paradigms, verification—the ability to determine the quality or correct","cbCaihcafzor8Y5d","https://ap.wps.com/l/cbCaihcafzor8Y5d","pdf",4485476,1,31,"English","en",105,"# Introduction\n## Verification as a new scaling axis\n## Probabilistic continuous scoring\n## Scaling dimensions for verification\n## Cost-efficient ranking and applications","[{\"question\":\"What problem does LLM-as-a-Verifier address?\",\"answer\":\"It targets the limited scalability of verification—deciding whether a solution is correct—compared with the well-studied scaling of generation.\"},{\"question\":\"How does LLM-as-a-Verifier differ from standard language-model judges?\",\"answer\":\"Instead of producing discrete scores, it computes the expectation over scoring-token logits to generate continuous verification scores.\"},{\"question\":\"Which domains and evaluation benchmarks show the framework’s effectiveness?\",\"answer\":\"The framework reports state-of-the-art performance on Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench, spanning coding, robotics, and medical tasks.\"}]",1784198484,78,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"llm-as-a-verifier-a-general-purpose-verification-framework","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/llm-as-a-verifier-a-general-purpose-verification-framework/84818/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does LLM-as-a-Verifier address?","Question",{"text":75,"@type":76},"It targets the limited scalability of verification—deciding whether a solution is correct—compared with the well-studied scaling of generation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does LLM-as-a-Verifier differ from standard language-model judges?",{"text":80,"@type":76},"Instead of producing discrete scores, it computes the expectation over scoring-token logits to generate continuous verification scores.",{"name":82,"@type":73,"acceptedAnswer":83},"Which domains and evaluation benchmarks show the framework’s effectiveness?",{"text":84,"@type":76},"The framework reports state-of-the-art performance on Terminal-Bench V2, SWE-Bench Verified, RoboRewardBench, and MedAgentBench, spanning coding, robotics, and medical tasks.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]