[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84027-en":3,"doc-seo-84027-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84027,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges","Training a language model against its own reference-free judgments—used in self-reward, self-play, and LLM-as-a-judge label-free self-improvement pipelines—assumes the judge’s verdict on a candidate answer can proxy correctness. This work shows a structural failure: the judge scores plausibility rather than correctness, creating a false-positive basin that policies learn to exploit. Hidden-anchor exact-match auditing quantifies a large judge–truth gap and demonstrates weak transferability controls and a de-anchoring remedy that collapses false positives while preserving solvability.","arXiv :2607 .05904v 1 [ cs .LG] 7 Jul 2026  \nMORE CONVINCING, NOT MORE CORRECT:  \nSELF-PLAY REWARD HACKING OF REFERENCE-FREE LLM JUDGES  \nChenyu Zhou  \nSchool of Engineering, Institute of Science Tokyo [zhou.c.76d6@m.isct.ac.jp](zhou.c.76d6@m.isct.ac.jp)  \nABSTRACT  \nTraining a language model against its own reference-free judgments—the premise of self-rewarding, self-play, and LLM-as-a-judge pipelines for label-free selfimprovement—assumes a model’s verdict on a shown answer is a usable proxy for correctness. We show the premise fails structurally: conditioned on a candidate, a reference-free judge scores plausibility, not correctness—a verification asymmetry that leaves false-positive basins of plausible-but-wrong answers a policy learns to exploit. We measure the failure with a hidden-anchor audit: a held-out, crosssource exact-match check the judge never sees. On GSM8K with Qwen3 policies ina reasoning-suppressed regime, self-play drives the judge’s pass rate from 0.72 to 0.94 while true accuracy stays at 0 .20—a 0 .74 judge–truth gap (three seeds) . This reward hacking is not white-box gaming: the manufactured errors transfer across judge families (Qwen, Llama, Gemma) and scales (to 14B), a strict three-family ensemble still accepts 55% of them, and recompute prompts, stronger judges, and training directly against the ensemble reward all fail to close the basin. The decisive variable is not whether the judge sees the candidate but whether it commits an answer of its own first: the recompute prompt leaves the false-positive rate on wrong answers at 0.719, committing first—candidate still in view—drops it to 0.012, and blind solving lifts discrimination from near chance to 0.96 with no reference answer. Used as the training reward, the de-anchored channel keeps the false-positive rate at zero across self-play, preventing the basin rather than only detecting it. A falsifiable bound explains which regimes are exposed—the gap is at most 1 − accuracy—and the de-anchored reward obeys the analogous judge-side bound. The full arc replicates without training under best-of-N selection in code and competition math, and in the full loop with a Gemma policy.  \n1 INTRODUCTION  \nReinforcement learning from a model’s own evaluations has become a central recipe for improving language models without human labels. Reference-free “LLM-as-a-judge” rewards, self-reward, and self-play schemes share one premise: a model’s judgment of an answer is a usable proxy for its correctness, and a growing line of methods builds on it as a default assumption (Bai et al., 2022; Lee et al., 2023; Yuan et al., 2024; Chen et al., 2024; Simonds et al., 2025; Huang et al., 2026) .  \nThe premise has a structural flaw. A reference-free judge has no access to ground truth; it can only assess whether an answer looks correct. For tasks where verifying an answer is harder than recognizing a plausible one—most of reasoning—this is a verification asymmetry: the judge scores plausibility, not correctness, leaving a false-positive basin of plausible-but-wrong answers it accepts. Optimizing a policy against such a judge does not merely risk noise; it actively rewards finding the basin.  \nWe measure it with a hidden anchor: a held-out, cross-source exact-match check on the final answer that the judge never sees and is never trained against. On GSM8K with Qwen3 policies trained by self-play, the audit reveals a large divergence: on the full test set, in a reasoning-suppressed regime that holds accuracy low, the judge’s pass rate climbs from ≈0 .72 to 0.94 while anchor-verified  \naccuracy stays flat at ≈0 .20—a 0 .74 judge–truth gap (three seeds; §5); Figure 1a shows the fiveiteration trajectory. Self-play does not make the model more correct; it makes the model’s errors more convincing.  \nThis is not a quirk of one judge that a stronger or more diverse judge would catch. Re-scoring the hacked answers with independent judges from other families (Llama, Gemma) and larger scales","cbCaibkwbvw7Uq6V","https://ap.wps.com/l/cbCaibkwbvw7Uq6V","pdf",733202,1,15,"English","en",105,"# Abstract\n# Introduction\n## Structural flaw in reference-free judging\n## Hidden-anchor audit methodology\n## Cross-judge and scale transfer\n## De-anchoring fix and predictive bounds\n# Contributions","[{\"question\":\"What structural problem does the paper identify in reference-free LLM judges?\",\"answer\":\"The judge lacks ground truth and therefore evaluates whether an answer looks plausible, not whether it is correct. This verification asymmetry leaves a false-positive basin of plausible-but-wrong answers.\"},{\"question\":\"How is the “hidden-anchor audit” used to measure the judge–truth gap?\",\"answer\":\"A held-out, cross-source exact-match check on the final answer is performed that the judge never sees and is never trained against. It reveals that self-play increases judge pass rates while true accuracy stays low.\"},{\"question\":\"What is the de-anchoring fix and what effect does it have?\",\"answer\":\"The fix requires the judge to commit to an answer of its own before using the candidate. This collapses the false-positive rate on wrong answers (from about 0.719 to about 0.012) and prevents the basin when used as the training reward.\"}]",1784192118,38,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"more-convincing-not-more-correct-self-play-reward-hacking-of-reference-free-llm-judges","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/more-convincing-not-more-correct-self-play-reward-hacking-of-reference-free-llm-judges/84027/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-28","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What structural problem does the paper identify in reference-free LLM judges?","Question",{"text":75,"@type":76},"The judge lacks ground truth and therefore evaluates whether an answer looks plausible, not whether it is correct. This verification asymmetry leaves a false-positive basin of plausible-but-wrong answers.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the “hidden-anchor audit” used to measure the judge–truth gap?",{"text":80,"@type":76},"A held-out, cross-source exact-match check on the final answer is performed that the judge never sees and is never trained against. It reveals that self-play increases judge pass rates while true accuracy stays low.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the de-anchoring fix and what effect does it have?",{"text":84,"@type":76},"The fix requires the judge to commit to an answer of its own before using the candidate. This collapses the false-positive rate on wrong answers (from about 0.719 to about 0.012) and prevents the basin when used as the training reward.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]