[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83285-en":3,"doc-seo-83285-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83285,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Agon Competitive Cross-Model RL with Implicit Rival Grading of Reasoning","Reinforcement learning from verifiable rewards sharpens reasoning models, yet typical methods grade only final answers and never score the reasoning trace, which can encourage verbosity over genuine thinking on hard problems. Agon introduces competitive cross-model RL: two alternating models act as each other’s graders, rewarding out-solving the rival using implicit judging during training without process labels or reward-model supervision. On DeepMath with Qwen3, Agon doubles GRPO pass@1 and improves ordering across tasks and model families.","arXiv :2607 .07690v 1 [ cs .LG] 8 Jul 2026  \nAgon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning  \nVladislav Beliaev  \nIndependent Researcher  \n[belyaev. vladislav. nw@gmail. com](belyaev. vladislav. nw@gmail. com)  \n[thinkdense. ai](thinkdense. ai)  \nAbstract  \nReinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today’s reasoning models, yet it grades only the final answer. On hard problems this trains models to write more rather than to think better, since the trace itself is never graded and no label for good thinking exists. We introduce Agon, which makes two competing models each other’s graders. Both attempt the same problem; in alternating roles, one drafts a solution and the other reads it while solving, and each is rewarded for out-solving the other. To win, a model must out-reason a rival that has seen its work, so reasoning is judged implicitly during training, with no process labels and no reward model. Because both models are optimized, each faces a progressively stronger rival, which single-model RL cannot provide. The two need only be comparably strong and behaviorally different. At inference the pair deploys as it trains, a twostage cascade in which one model drafts and the other answers after reading the draft. On the hard split of DeepMath with Qwen3, this doubles GRPO’s pass@1, roughly eight times the gain of an untrained Mixture-of-Agents pass over the same base. The ordering replicates on competitive-programming code and across model families (Qwen3.5, Gemma 4) . For now the models talk in text; the next step is to let them reason together in latent space.  \n1 Introduction  \nReinforcement learning from verifiable rewards has become the standard tool for sharpening the reasoning of large language models (LLMs) on tasks such as mathematics and code [Guo et al., 2025, Shao et al. , 2024] . The dominant recipe is on-policy and outcome-based: sample a group of rollouts from the current policy, score each by whether its final answer is correct, and update with group-relative advantages, as in GRPO [Shao et al., 2024] . The reward is attached to the answer; the chain of thought that produced it is never graded. On problems the model can already partly solve this is benign, but on hard problems it creates a strong incentive to write more: a longer chain affords more chances to stumble onto the right answer, so additional text is the cheapest way to raise expected reward. The trace inflates with hedging and backtracking (“hmm,”“wait,”“let me reconsider”), and accuracy grows far slower than length (accuracy rises only modestly while length inflates by an order of magnitude [Aggarwal & Welleck, 2025]) . The model ends up producing more reasoning per problem without producing better reasoning per token.  \nThe obvious fix, grading the reasoning directly, is intractable. There is no ground truth for “good thinking”: no label says which step was the insight and which was filler, and a learned process reward model is expensive, brittle, and itself unverifiable [Lightman et al., 2023] . So the trace stays unscored, and the length pathology persists. We ask instead: can a second model supply the missing signal? If a different policy attempts the same problem and we reward each model for out-reasoning the other, then the trace is graded implicitly by the rival, without any process labels. A step that leads to a win is reinforced; filler that the opponent exploits is punished.  \nThe grader must be a different model, and the game must be competitive. Post-training RL is, by construction, self-improvement: the policy is optimized on signal it generates itself, which reinforces the very blind spots that created its errors: a model auditing its own work tends to plateau [Huang et al., 2024] . A second model with different failure modes breaks this closed loop. But merely having two models is not  \nDeepMath-hard pass@1 (%)  \n80  \n60  \n40  \n20  \n0  \n\n|  |  |  |  |  |  |  |  |  |  | ","cbCaijKPllQotWA1","https://ap.wps.com/l/cbCaijKPllQotWA1","pdf",688865,2,1,15,"English","en",105,"# Abstract\n# Introduction\n## Limitations of answer-grading\n## Implicit rival-grading via competition\n## Competitive multi-model training setup\n## Experimental results and related comparisons","[{\"question\":\"Why does conventional verifiable-reward RL often produce longer reasoning traces without better thinking?\",\"answer\":\"Because rewards attach to the final answer while the reasoning trace is ungraded, harder problems incentivize writing more text to stumble on a correct outcome. Accuracy tends to improve much more slowly than length.\"},{\"question\":\"How does Agon grade reasoning implicitly during training?\",\"answer\":\"Agon trains two models competitively: one drafts a solution while the other reads it and is rewarded for solving correctly and beating the opponent. The rival’s success implicitly grades the usefulness of the trace without explicit process labels.\"},{\"question\":\"What makes Agon different from self-consistency or mixture-of-agents approaches?\",\"answer\":\"Self-consistency or agent mixtures typically rely on agreement or averaging, which can regress toward consensus and dilute quality when models differ. Agon instead uses direct competition with comparable-strength but behaviorally different models under an out-reasoning objective.\"}]",1784186495,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"agon-competitive-cross-model-rl-with-implicit-rival-grading-of-reasoning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/agon-competitive-cross-model-rl-with-implicit-rival-grading-of-reasoning/83285/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does conventional verifiable-reward RL often produce longer reasoning traces without better thinking?","Question",{"text":75,"@type":76},"Because rewards attach to the final answer while the reasoning trace is ungraded, harder problems incentivize writing more text to stumble on a correct outcome. Accuracy tends to improve much more slowly than length.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Agon grade reasoning implicitly during training?",{"text":80,"@type":76},"Agon trains two models competitively: one drafts a solution while the other reads it and is rewarded for solving correctly and beating the opponent. The rival’s success implicitly grades the usefulness of the trace without explicit process labels.",{"name":82,"@type":73,"acceptedAnswer":83},"What makes Agon different from self-consistency or mixture-of-agents approaches?",{"text":84,"@type":76},"Self-consistency or agent mixtures typically rely on agreement or averaging, which can regress toward consensus and dilute quality when models differ. Agon instead uses direct competition with comparable-strength but behaviorally different models under an out-reasoning objective.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]