[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84260-en":3,"doc-seo-84260-105":30,"detail-sidebar-cat-0-en-105":88},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84260,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems","Group Relative Policy Optimization (GRPO) provides little learning signal on the hardest problems: when none of a rollout group succeeds, group-relative advantages collapse to zero and that problem contributes no gradient. Adaptive Trace Prefix Control makes a hard problem temporarily easier by prepending a correct reference prefix, turning success rate into a tunable knob via prefix length. AdaPrefix-GRPO uses a feedback controller to keep realized success rate near 50% during training (where GRPO’s gradient signal is maximal), then anneals assistance away so deployment solves without a prefix. On hard math, it more than doubles GRPO accuracy at matched FLOPs and reduces trace length, with larger gains on smaller models.","arXiv :2607 .07674v 1 [ cs .LG] 8 Jul 2026  \nMax Out GRPO Signal:  \nAdaptive Trace Prefix Control for Hard Reasoning Problems  \nVladislav Beliaev  \nIndependent Researcher  \n[belyaev. vladislav. nw@gmail. com](belyaev. vladislav. nw@gmail. com)  \n[thinkdense. ai](thinkdense. ai)  \nAbstract  \nGroup Relative Policy Optimization (GRPO) stalls on a model’s hardest problems: when no rollout ina group succeeds, the group-relative advantages vanish and the problem contributes no gradient, wasting the frontier examples we most want to learn from. Prepending a correct prefix of a reference solution raises the success rate, making prefix length a continuous knob on difficulty. Concurrent methods set the knob once; AdaPrefix-GRPO turns it into a feedback controller: throughout training it adjusts how much of the solution each problem gets, holding its success rate near 50%, where GRPO’s gradient signal is largest, then withdraws the assistance entirely, so the deployed model solves problems unaided. On hard math, at matched training FLOPs, it more than doubles GRPO’s accuracy on held-out problems from the training distribution for a 0.6B model (2.1×), with 1.6 × on Qwen3-1.7B and 1.7 × on AIME, while roughly halving trace length. The method is implemented in data preparation plus a loss mask on prefix tokens; the trainer is otherwise stock. The smaller the model, the larger the gain.  \n1 Introduction  \nReinforcement learning from verifiable rewards is now the standard tool for improving the reasoning of large language models (LLMs) on tasks such as mathematics and code [Guo et al., 2025, Lambert et al., 2024] . Building on policy-gradient methods [Schulman et al., 2017, Ahmadian et al., 2024] and RLHF [Ouyang et al., 2022], the dominant recipe is on-policy: sample a group of G rollouts from the current policy, score each by an outcome reward, and update using group-relative advantages, as in GRPO [Shao et al., 2024] . This recipe has a structural blind spot. On a hard problem, one the model almost never solves, all G rollouts fail, every reward is identical, the group-relative advantage is exactly zero, and the gradient contribution of that problem is zero [Yu et al., 2025, Zhang et al., 2025] . The model burns sampling compute on the problem and learns nothing from it. Worse, these are frequently the problems we most want to learn: the ones at the frontier of the model’s ability.  \nCommon remedies work around this rather than fix it. Oversampling and dynamic filtering, as in DAPO [Yu et al., 2025], discard degenerate all-correct/all-wrong groups but cannot manufacture a correct rollout where the model has none. Warm-starting RL with supervised fine-tuning on reference traces (or rejection-sampled ones, STaR/RAFT-style [Zelikman et al., 2022, Dong et al., 2023]) injects competence but induces entropy collapse that hurts subsequent exploration [Chu et al., 2025, Cui et al., 2025] . A natural idea is instead to make hard problems temporarily easier so they yield a signal, then remove the assistance. If we prepend a correct prefix of a solution and ask the model to complete it, the conditional success rate rises (empirically near-monotonically) with prefix length: with no prefix the model solves almost nothing, with along enough prefix it solves almost everything. Prefix length is therefore a continuous dial on difficulty, and for most reachable success rates k/G there is a prefix length that achieves it. This observation underlies aline of concurrent work [Setlur et al., 2026, Qu et al., 2026, Huang et al., 2025] .  \nThe central question this paper addresses is what success rate to aim for, and how to hit it, i.e. how to max out the GRPO signal on hard problems. Concurrent methods fix a small set of prefix lengths per problem, chosen once from the base model so that the conditioned base accuracy is “reasonable” [Setlur et al., 2026] . We argue this leaves most of the available signal on the table for two reasons. First, the GRPO  \nFigure 1:","cbCaiqwI50T8uTIn","https://ap.wps.com/l/cbCaiqwI50T8uTIn","pdf",586158,5,1,13,"English","en",105,"# Abstract\n# Introduction\n## Problem: GRPO stalls on unsolved hard cases\n## Limitations of common remedies\n## Core idea: prefix length as a difficulty dial\n## Research question and contribution: AdaPrefix-GRPO controller\n# Method overview","[{\"question\":\"How does AdaPrefix-GRPO decide what success rate to target during training?\",\"answer\":\"It measures realized batch-mean k/G and adjusts a single global base prefix length toward a target k/G using a root-finding (secant or binary search) update, keeping success rate near the regime where GRPO’s gradient signal is largest (around 50%).\"},{\"question\":\"What ensures the deployed model does not rely on prefixes?\",\"answer\":\"All prefixes are annealed to zero during training so the deployed policy never sees a prefix at test time, withdrawing the assistance after the controller has extracted training signal.\"}]",1784194430,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":83,"head_meta":85,"extra_data":87,"updated_unix":28},"max-out-grpo-signal-adaptive-trace-prefix-control-for-hard-reasoning-problems","",{"@graph":36,"@context":82},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/max-out-grpo-signal-adaptive-trace-prefix-control-for-hard-reasoning-problems/84260/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78],{"name":73,"@type":74,"acceptedAnswer":75},"How does AdaPrefix-GRPO decide what success rate to target during training?","Question",{"text":76,"@type":77},"It measures realized batch-mean k/G and adjusts a single global base prefix length toward a target k/G using a root-finding (secant or binary search) update, keeping success rate near the regime where GRPO’s gradient signal is largest (around 50%).","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What ensures the deployed model does not rely on prefixes?",{"text":81,"@type":77},"All prefixes are annealed to zero during training so the deployed policy never sees a prefix at test time, withdrawing the assistance after the controller has extracted training signal.","https://schema.org",{"og:url":52,"og:type":84,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":86,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":89},[90,94,98,102,106,111,116,119,124,127,131],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Comic",60,"comic",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},6,"Technology",50,"technology",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":117,"slug":118},30,"research-report",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},9,"Religion & Spirituality",20,"religion-spirituality",{"id":122,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":122,"slug":126},"World Cup","world-cup",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":128,"slug":130},10,"Lifestyle","lifestyle",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":20,"slug":134},19,"General","general"]