[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86112-en":3,"doc-seo-86112-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86112,1374391974468,"Eden","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","When the Reward Suite Is Leaky: A Preregistered Causal Contrast of Natural Verifier False Positives in RLVR","Deployed RLVR code test suites contain natural, persistent, asymmetric false positives where the same wrong program reliably passes the same weak tests, unlike the symmetric resampled noise assumed by prior noise-robustness work. A preregistered two-arm causal contrast trains GRPO on identical MBPP tasks and settings, using the original (leaky) versus MBPP+ extra tests (hardened) as rewards. Across families, average held-out effects are bounded and false-positive mass tracks static leakiness audits, yet human adjudication finds a large residual of genuinely wrong code rewarded.","arXiv :2607 . 1 1022v 1 [ cs .LG] 13 Jul 2026  \nWHEN THE REWARD SUITE IS LEAKY: A PREREGISTERED CAUSAL CONTRAST OF NATURAL VERIFIER FALSE POSITIVES IN RLVR  \nChuyifei Zhang  \nBeijing Jiaotong University [24222058@bjtu.edu.cn](24222058@bjtu.edu.cn)  \nABSTRACT  \nThe test suites used as RLVR rewards for code have natural false positives: pertask, persistent, asymmetric errors that accept the same wrong programs everytime they appear, unlike the symmetric or resampled noise assumed by existing noise-robustness analyses. We run a preregistered two-arm causal contrast on a deployed suite: GRPO on identical MBPP tasks, seeds, and compute (Qwen2.5-Coder-1.5B-Instruct, 5 seeds × 400 steps), rewarded by the original MBPP tests (leaky) versus the MBPP+ extra tests (hardened; extra-tests-only, §3.1) . Two further families replicate the design under a preregistration frozen before their data existed: deepseek-coder-1.3b-instruct, registered as the decision arm, and Llama- 3.2-1B-Instruct. Claims below are tagged [C] (confirmatory: preregistered) or [E](exploratory) . [C] The average held-out effect (scored by the extra-test suite throughout) is bounded: non-inferior under a preregistered 1.5-pt margin (gap 0.20 pt, one-sided 95% upper bound 0.75 pt), with the upper bound on the same side of the margin under all six estimators in all three families. [C] Rewarded false-positive mass tracks a cheap static leakiness audit computed before training (Spearman 0.80 raw, 0.79 under difficulty control), and the registered train-side test puts the leak-stratum FP share +43.8 pt above clean tasks (CI90 [43.05, 44.57];  \nthe FP definition is classifier-free, so these numbers are invariant under all later audit revisions) . [E] Auditing every rewarded FP under signed, human-adjudicated rules finds a large residual of verified genuinely wrong code: 47.57% recordweighted (task-cluster bootstrap-95 [36.4, 60.5]; estimator variants within two points) . The reward paid for real bugs, not merely suite artifacts. Both replication families reproduce a large residual share—45.37% and 62.78%—but the shares replicate “large,” not one number. [E] On statically leaky held-out tasks the leakyarm improves less (family A) . This is an association only, its specification-search correction prices one axis of freedom (a lower bound), and it does not replicate cleanly across families: one reversed-direction candidate, one underpowered positive. [E] Mechanism evidence is consistent with selection of pre-existing error modes rather than learned exploitation. FP incidence sits at its long-run share from step 0 and does not grow within our horizon (flat in two families, declining in the third); the largest channel shows 55/55 distinct wrong programs and no hacking signatures; and untrained base models already produce the same wrong outputs under the leaky filter. The 8.37-pt reward inflation shows no held-out counterpart within the bound reported above, though parameter-level sharpening below behavioral resolution is not excluded. In an exploratory follow-up outside these confirmatory claims, we turn the same instrument on the frontier judges themselves: on their own false positives they self-assess only weakly, a deconfounded same-author test is unresolved, and even the highest-scoring reader we probe stays far below its score on a weaker policy’s errors—two subjects on MBPP, licensing nothing about frontier models in general (§5) . The practical boundary we map: a cheap static audit locates where FP mass will sit before training—exposure, not final damage. Hardening the reward—swapping the leaky tests for the far larger extra-test suite—removes the measurement inflation, though at this scale it buys little capability.  \n1 INTRODUCTION  \nReinforcement learning from verifiable rewards (RLVR) trains code models against test suites and treats a pass as ground truth (Lambert et al., 2024; DeepSeek-AI, 2025) . Deployed suites do not earn that trust: a recent audit of two public co","cbCairg2fWadTo8j","https://ap.wps.com/l/cbCairg2fWadTo8j","pdf",1670669,4,1,37,"English","en",105,"# Introduction\n## Motivation and problem definition\n## Preregistered two-arm causal contrast design\n## Static leakiness audit approach","[{\"question\":\"What does the paper mean by “leaky” reward suites in RLVR?\",\"answer\":\"A deployed suite accepts wrong solutions consistently, rewarding the same incorrect programs across rollouts. These errors are per-task and persistent, creating a directional reward artifact rather than random label noise.\"},{\"question\":\"How is the causal test set up to measure the impact of leaky versus hardened rewards?\",\"answer\":\"The paper runs a preregistered two-arm experiment using GRPO to train the same model on identical MBPP tasks, seeds, and compute. The only difference is the reward test suite: the original MBPP tests (leaky) versus the MBPP+ extra tests (hardened).\"},{\"question\":\"What does the study find about false positives and the true quality of rewarded solutions?\",\"answer\":\"False-positive mass correlates strongly with a pre-training static leakiness audit, but auditing rewarded false positives with human rules reveals a substantial residual share of genuinely wrong code. The reward therefore pays not just for suite artifacts, but for real bugs.\"}]",1784208595,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"when-the-reward-suite-is-leaky-a-preregistered-causal-contrast-of-natural-verifier-false-positives-in-rlvr","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/when-the-reward-suite-is-leaky-a-preregistered-causal-contrast-of-natural-verifier-false-positives-in-rlvr/86112/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does the paper mean by “leaky” reward suites in RLVR?","Question",{"text":75,"@type":76},"A deployed suite accepts wrong solutions consistently, rewarding the same incorrect programs across rollouts. These errors are per-task and persistent, creating a directional reward artifact rather than random label noise.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the causal test set up to measure the impact of leaky versus hardened rewards?",{"text":80,"@type":76},"The paper runs a preregistered two-arm experiment using GRPO to train the same model on identical MBPP tasks, seeds, and compute. The only difference is the reward test suite: the original MBPP tests (leaky) versus the MBPP+ extra tests (hardened).",{"name":82,"@type":73,"acceptedAnswer":83},"What does the study find about false positives and the true quality of rewarded solutions?",{"text":84,"@type":76},"False-positive mass correlates strongly with a pre-training static leakiness audit, but auditing rewarded false positives with human rules reveals a substantial residual share of genuinely wrong code. The reward therefore pays not just for suite artifacts, but for real bugs.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]