[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85047-en":3,"doc-seo-85047-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85047,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals","LLM-as-judge assumes that consistency—judge agreement or a model’s self-samples—indicates correctness, but agreement is not accuracy: models can concur due to shared bias, memorized heuristics, option-position priors, or systematic hallucination. A large cross-runner audit (265,000 samples across GPQA Diamond and AIME) finds agreement a positive yet weak predictor (ρ 0.20–0.59) that varies by regime, best for unsaturated mid-tier models and compute allocation, but worst for an over-confident frontier model. Deidentified per-run rows are released.","When LLMs Agree, Are They Right?  \nAuditing Self-Consistency and Cross-Model Agreement as Confidence  \nSignals  \nKaihua Ding*  \nUniversity of Pennsylvania  \n[dkaihua@upenn.edu](dkaihua@upenn.edu)  \narXiv :2607 .08065v 1 [ cs .AI] 9 Jul 2026  \nAbstract  \nLLM-as-judge (Zheng et al., 2023) is increasingly the default for evaluating AI systems in enterprise pipelines, often scaled to ensembles (Verga et al., 2024) or ‘mixture-of-experts’(Shazeer et al., 2017) panels of judges. These systems share a key assumption: that consistency—agreement among judges, or among a model’s own samples—indicates correctness. We show this assumption is unreliable. Agreement is not accuracy: a model can agree with itself, and different models can agree with eachother, out of shared bias, a memorized heuristic, or an option-position prior rather than truth. We ask when agreement is nonetheless a usable proxy, in a large-scale cross-runner study:  \n53 runners drew K=50 samples for assigned overlapping cases across comparisons of model tier, prompting, and scale on GPQA Diamond and AIME—265,000 samples. Using majoritycorrectness as the deployment label and a hierarchical runner-clustered bootstrap, agreement is a positive but weak predictor (ρ 0.20– 0.59, all positive under item-clustered resampling) whose usefulness is regime-dependent: best for unsaturated mid-tier models and for allocating compute, and worst—over-confident yet no more accurate—for the most consistent frontier model (agreement ≥ 0.8 on 77% of GPQA case-result entries, 48% of those wrong) . An exploratory cross-family check on three Claude tiers shows the same frontier over-confidence, with confident errors recurring across providers above a marginalpreserving null. Self-consistency is thus a conditional proxy for correctness, not a standalone confidence score. We publicly release the deidentified per-run rows and answer distributions.  \n*The author’s prior work spans output-based error estimation (Ding, 2018) and AI-system evaluation and assessment design (Ding, 2025a) .  \n1 Introduction  \nLarge language models are increasingly used to evaluate AI systems—as LLM-as-judge, and in ensemble or multi-judge (‘mixture-of-experts’) panels now common in enterprise evaluation pipelines (Zheng et al., 2023) . These pipelines rest on a pervasive intuition: that agreement signals correctness—if several judges concur, or if a model returns the same answer across many stochastic samples, the answer is trusted, and if they disagree it is treated as a guess. But agreement-with-itself isnot the same as being right: a model can repeat an answer because it is genuinely certain, or because every sample passes through the same memorized heuristic, shared misconception, option-position prior, or systematic hallucination. This is the selfreferential analogue of the self-preference bias documented when models judge their own outputs (Zheng et al., 2023 ; Panickssery et al., 2024)—agreement with oneself can encode shared bias rather than correctness. Self-consistency decoding operationalizes this by returning the majority answer over K samples (Wang et al., 2023), building on chain-of-thought prompting (Wei et al., 2022) . The same agreement signal is now reused well beyond decoding—as a confidence estimate for costbased routing and cascades (Chen et al., 2023), for adaptive sample-budget allocation (Aggarwal et al., 2023), and for selective prediction and abstention (Geifman and El-Yaniv, 2017 ; Kamath et al., 2020) .  \nThat self-consistency raises accuracy via majority voting is well established, as is the broader finding that elicited LLM confidence is often miscalibrated and can even worsen after post-training (Kadavath et al., 2022 ; Lin et al., 2022 ; Tian et al., 2023 ; OpenAI, 2023) . Less systematically examined is how reliably the agreement signal itself functions as the confidence proxy that routing and abstention systems already assume it to be—and when it fails—under a controlled, cross-replicate","cbCaiuxAQxzqWkxu","https://ap.wps.com/l/cbCaiuxAQxzqWkxu","pdf",315128,1,10,"English","en",105,"# Introduction\n# Setup\n## Data provenance","[{\"question\":\"What additional checks and resources does the study provide?\",\"answer\":\"An option-shuffle control shows part of the “confidence” is positional, and a cross-family check across multiple Claude tiers indicates recurring confident errors shared across providers. The study also releases a deidentified, schema-validated dataset and analysis pipeline.\"}]",1784200613,25,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"when-llms-agree-are-they-right-auditing-self-consistency-and-cross-model-agreement-as-confidence-signals","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/when-llms-agree-are-they-right-auditing-self-consistency-and-cross-model-agreement-as-confidence-signals/85047/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What additional checks and resources does the study provide?","Question",{"text":75,"@type":76},"An option-shuffle control shows part of the “confidence” is positional, and a cross-family check across multiple Claude tiers indicates recurring confident errors shared across providers. The study also releases a deidentified, schema-validated dataset and analysis pipeline.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,126],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":45,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":21,"slug":125},"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]