[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83415-en":3,"doc-seo-83415-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83415,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Do You Need a Frontier Model as a Citation Verifier Benchmarking Rubric LLMs for Deep Research Source Attribution","Reinforcement learning increasingly uses an LLM judge to score rubric criteria, where the judge effectively becomes the reward model. Trusted use requires knowing the minimum judge capability and how much bias it introduces in citation quality for deep-research systems. The study benchmarks citation-quality scoring as a structured rubric task over attribution–citation pairs, evaluating source relevance and factual support with an LLM plus deterministic link accessibility checks. On an adversarial long-form benchmark, 8 off-the-shelf judge models are compared against gold labels across 1,248 rubric decisions including hard adjudication cases. Results show cheaper judges can be competitive, and directional bias in scalar F1 can be reinforced by RL loops, so judge calibration is necessary and need not require the most expensive model.","arXiv :2607 .08700v 1 [ cs .CL] 9 Jul 2026  \nDo You Need a Frontier Model as a Citation Verifier? Benchmarking Rubric LLMs for Deep-Research Source Attribution  \nEthan Leung* Elias Lumer Corey Feld Austin Huber Vamse Kumar Subbiah Kevin Paul  \nCommercial Technology and Innovation Office, PricewaterhouseCoopers, U.S.  \nAbstract  \nReinforcement learning increasingly relies on an LLM judge to score each rubric criterion, and that judge acts as the reward model during training. Before such a signal can be trusted, we need to know how capable the judge must be and how biased it is. We study this calibration question for citation quality in deep-research systems, where a search-grounded LLM must support each claim it writes with a cited source. Citation quality is a structured rubric task in which each attribution-citation pair is judged along two dimensions that require an LLM, source relevance and factual support. On an adversarial long-form benchmark, we score  \n8 off-the-shelf LLM judges from 3 model families against gold labels over 1,248 rubric decisions, all of which were human-reviewed and 378 of which were hard cases adjudicated from judge disagreements. Cheaper judges remain competitive across both dimensions, with GPT-5-mini attaining the strongest source-relevance pass-class F1 at 0.908 (κ=0.636), while on factual support the judges are statistically indistinguishable (overlapping confidence intervals), so no single model dominates. At comparable F1, the judges still differ substantially in pass-rate drift, false positive rate, and false negative rate. Scalar F1 obscures this directional bias, yet it is exactly what a downstream reinforcement learning loop would reinforce. Calibrating the judge is therefore a prerequisite for using citation rubrics as reward signals, and our results show that this calibration does not require the most expensive available model.  \n1 Introduction  \nReinforcement learning with verifiable rewards (RLVR) has become the dominant post-training recipe for tasks where output quality can be automatically checked [Tyagi et al., 2026, Mahmoud et al., 2026, DeepSeek-AI, 2025, Lambert et al., 2024a] . In domains where no programmatic verifier exists (medical advice, scientific writing, instruction following), the community has converged on prompt-specific rubrics with weighted criteria, each scored by an LLM judge whose aggregate judgment serves as the scalar reward [Rezaei et al., 2026, Mahmoud et al., 2026, Liu et al., 2023] . This convergence makes explicit a connection that is easy to overlook. The model that scores each criterion acts as the reward model, so designing the training rubric amounts to designing theoperationalized goal. These scoring calls are also a verification bottleneck at scale [Trivedy et al. , 2026], which raises a practical question about how capable a judge needs to be before the reward signal degrades.  \nCitation quality in deep-research and other search-augmented systems is a strong test case for this question. These systems produce long-form answers where each factual claim is supported by a retrieved, cited source, and faithful citation generation of this kind is an active RLVR training target [Nakano et al., 2022, Menick et al., 2022, Asai et al., 2024, Gao et al., 2023b, Bohnet  \net al., 2022] . How well a judge evaluates citations also depends on retrieval quality, since the harness and retrieval strategy shape which sources get cited [Sen et al., 2026] . The quality of each citation decomposes into criteria applied per attribution-citation pair. Two of them require an LLM judgment, whether the source is topically relevant to the claim and whether it factually supports it, while a third, whether the cited URL is accessible, is a deterministic check. Each LLMjudged criterion is an independent rubric judgment structurally similar to a process reward model (PRM) step [Lightman et al., 2024], a per-step judgment used as an intermediate reward signal, and the task is narro","cbCaisVj0oPrXtot","https://ap.wps.com/l/cbCaisVj0oPrXtot","pdf",372202,4,1,17,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why is judge capability and bias important for citation verifier rubrics in deep-research systems?\",\"answer\":\"In RLVR setups, the LLM judge’s scalar rubric scores function as the reward model during training. If the judge lacks capability or introduces bias, the reward signal can degrade or reinforce systematic errors.\"},{\"question\":\"How is citation quality evaluated in the paper’s rubric?\",\"answer\":\"Citation quality is decomposed into rubric criteria per attribution–citation pair: source relevance and factual support require LLM judgments, while link accessibility is checked deterministically.\"},{\"question\":\"Do the results show that frontier models are necessary as citation verifiers?\",\"answer\":\"No. The study finds cheaper judges remain competitive, with GPT-5-mini achieving the strongest source-relevance pass-class F1, while factual-support performance differences are statistically indistinguishable across judges. Calibration is shown to be required, but it does not mandate the most expensive model.\"}]",1784187419,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"do-you-need-a-frontier-model-as-a-citation-verifier-benchmarking-rubric-llms-for-deep-research-source-attribution","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/do-you-need-a-frontier-model-as-a-citation-verifier-benchmarking-rubric-llms-for-deep-research-source-attribution/83415/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is judge capability and bias important for citation verifier rubrics in deep-research systems?","Question",{"text":75,"@type":76},"In RLVR setups, the LLM judge’s scalar rubric scores function as the reward model during training. If the judge lacks capability or introduces bias, the reward signal can degrade or reinforce systematic errors.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is citation quality evaluated in the paper’s rubric?",{"text":80,"@type":76},"Citation quality is decomposed into rubric criteria per attribution–citation pair: source relevance and factual support require LLM judgments, while link accessibility is checked deterministically.",{"name":82,"@type":73,"acceptedAnswer":83},"Do the results show that frontier models are necessary as citation verifiers?",{"text":84,"@type":76},"No. The study finds cheaper judges remain competitive, with GPT-5-mini achieving the strongest source-relevance pass-class F1, while factual-support performance differences are statistically indistinguishable across judges. Calibration is shown to be required, but it does not mandate the most expensive model.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]