[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84379-en":3,"doc-seo-84379-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84379,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","When the Judge Changes, So Does the Measurement Auditing LLM-as-Judge Reliability","An LLM-as-judge score can change even when candidate responses remain fixed, because the evaluator itself is part of the measurement instrument. The study treats this evaluator-replacement ambiguity as a measurement-validity issue and tests two real-world upgrade routes across four datasets. Results show judge upgrades are not interchangeable: only Qwen3 1.7B→4B yields a robust adjacent gain. Stronger judges reduce but cannot eliminate position and verbosity bias, and reliability under juries or debate requires auditable protocol logs, dataset slices, and bias probes.","When the Judge Changes, So Does the Measurement: Auditing  \nLLM-as-Judge Reliability  \nZongyou Yang  \nDyson School of Design Engineering, Imperial College London London, United Kingdom [zy2926@ic.ac.uk](zy2926@ic.ac.uk)  \nYinghan Hou  \nDepartment of Electrical and Electronic Engineering, Imperial College London London, United Kingdom [yh24@ic.ac.uk](yh24@ic.ac.uk)  \nXiaokun Yang∗ School of Electronic Information, Nanchang Institute of Technology Nanchang, China [yangxk@bupt.cn](yangxk@bupt.cn)  \narXiv :2607 .08535v 1 [ cs .CL] 9 Jul 2026  \nAbstract  \nAn LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed. We treat this evaluator-replacement ambiguity as a measurementvalidity problem. Across four judgment datasets, we compare two upgrade paths available in practice: scaling Qwen3 dense judges from 1.7B to 32B parameters and moving across MiniMax M2–M2.7 released APIs. The main pattern is that judge upgrades are not interchangeable: only Qwen3 1.7B→4B gives a robust adjacent gain, while MiniMax adjacent releases do not. Stronger judges reduce but do not remove position and verbosity bias. Repeated-sample juries add little when errors are correlated. Structured debate can move decisions substantially, but without parser and fallback logs those shifts cannot be attributed to deliberation. We argue that LLM-as-judge reports should include dataset slices, bias probes, error-dependence estimates, and protocol audit trails.  \nCCS Concepts  \n• Computing methodologies → Natural language processing; Machine learning.  \nKeywords  \nLLM-as-judge, automatic evaluation, model scaling, evaluation reliability, bias, jury aggregation  \n1 Introduction  \nLLM-as-judge evaluation is now often used as a measuring instrument for model quality. The difficulty is that the instrument is itself a model. When a system’s score changes after replacing the evaluator, the change is ambiguous: it may indicate that the new judge is more capable, that it is biased differently, that it fails on a different slice of the benchmark, or that the evaluation pipeline parsed and aggregated its outputs differently. This ambiguity is not a minor implementation detail. It determines whether an LLM-as-judge score can be interpreted as evidence about the candidate systems at all.  \nWe call this the evaluator-replacement ambiguity: when the measured preference outcome changes after replacing the judge, the source of the change is not identifiable from accuracy alone. This paper studies the ambiguity through two observable interventions available in evaluation practice. The first is a parameter-scaling decision, represented by Qwen3 dense judges from 1.7B to 32B parameters [21]. The second is a released-model upgrade path, represented by MiniMax M2–M2.7 APIs evaluated as released [15] .  \n∗ Corresponding author.  \nThese axes are evidence sources rather than the paper’s object of explanation. The MiniMax M2 report documents the released series and its agent-oriented training pipeline, but it does not make our API sequence a controlled ablation; consequently, this paper does not make a causal claim about MiniMax internals.  \nThis formulation leads to three research questions. RQ1 asks whether judge reliability improves similarly along a parameter axis and a released-model upgrade path. If evaluator capability is the dominant factor, both axes should show consistently positive adjacent gains across datasets; if reliability is workload-dependent, gains should vary by dataset and significance should not transfer uniformly. RQ2 asks whether higher aggregate accuracy also reduces known judge biases. Higher-capability judges should be less bias-sensitive, but nonzero flip rates would show that capability does not eliminate measurement artifacts. RQ3 asks whether protocol-level upgrades, such as juries and debate, change reliability beyond single-judge scaling. Jury gains should be limited when error correlation is high, whi","cbCaibjlQVnG2wea","https://ap.wps.com/l/cbCaibjlQVnG2wea","pdf",557109,5,1,6,"English","en",105,"# 1 Introduction\n# 2 Related Work","[{\"question\":\"What problem does the paper call evaluator-replacement ambiguity?\",\"answer\":\"It refers to cases where the measured preference outcome changes after replacing the judge, making the cause of the score change unclear from accuracy alone.\"},{\"question\":\"Which evaluator upgrade paths are compared in the study?\",\"answer\":\"The paper compares parameter scaling using Qwen3 dense judges (1.7B to 32B) and released-model upgrades using MiniMax M2–M2.7 APIs evaluated as released.\"},{\"question\":\"Do stronger judges fully remove scoring bias and measurement artifacts?\",\"answer\":\"Stronger judges reduce but do not eliminate position and verbosity bias; some artifacts remain observable in the evaluation outcomes.\"}]",1784195207,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"when-the-judge-changes-so-does-the-measurement-auditing-llm-as-judge-reliability","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/when-the-judge-changes-so-does-the-measurement-auditing-llm-as-judge-reliability/84379/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper call evaluator-replacement ambiguity?","Question",{"text":76,"@type":77},"It refers to cases where the measured preference outcome changes after replacing the judge, making the cause of the score change unclear from accuracy alone.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which evaluator upgrade paths are compared in the study?",{"text":81,"@type":77},"The paper compares parameter scaling using Qwen3 dense judges (1.7B to 32B) and released-model upgrades using MiniMax M2–M2.7 APIs evaluated as released.",{"name":83,"@type":74,"acceptedAnswer":84},"Do stronger judges fully remove scoring bias and measurement artifacts?",{"text":85,"@type":77},"Stronger judges reduce but do not eliminate position and verbosity bias; some artifacts remain observable in the evaluation outcomes.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]