[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82468-en":3,"doc-seo-82468-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82468,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity","Three popular training methods for verifiable language-model reasoning—Group Relative Policy Optimization (GRPO), Dr. GRPO, and Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO)—all act through a single quantity: the standard deviation of group rewards produced by sampling multiple responses for one prompt. Disagreement among verifier marks drives the learning signal, vanishing for unanimous outcomes and maximizing near even splits. For binary rewards, the paper proves the group-standard-deviation identity: the update magnitude equals the gradient length scaled by σ, giving a precise account of learning strength, difficulty bias, and sampling behavior, confirmed on Big-Math and controlled experiments.","arXiv :2607 .00152v1 [ cs .LG] 30 Jun 2026  \nGRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity  \nYong Yi Bay* Kathleen A. Yearick*  \nPhD, University of Illinois at Urbana-Champaign  \nAB S T RAC T  \nThree of the most popular methods for training language models to reason look like three different tricks. They are not. All three adjust a single number: standard deviation, reflecting how much a prompt’s sampled answers disagree. When such a model is trained, it answers each problem many times, and an automatic checker marks every answer right or wrong. The standard deviation of those marks measures the disagreement: largest when the answers split evenly between right and wrong, and zero when they all agree. Group Relative Policy Optimization (GRPO) divides by this number, GRPO Done Right (Dr. GRPO) drops the division, and Decoupled Clip and Dynamic Sampling Policy Optimization (DAPO) discards the groups where it is zero. Each is presented as its own fix, yet this paper proves they are three settings of one dial. That dial is not cosmetic: for right-or-wrong rewards, the disagreement is exactly the size of the training update, the group-standard-deviation identity. A split group teaches the most, while a unanimous group teaches nothing and falls silent. The same result says which problems deserve the most weight and how many tries each one needs. This paper confirms the intuition on a large real difficulty dataset (Big-Math) and in a controlled training run. What looks like a harmless normalization step  \nis the dial that decides where learning happens and how strongly.  \nKeywords GRPO · reward normalization · silent groups · difficulty bias · group size · dynamic sampling · RLVR · LLM reasoning  \n1 Introduction  \nTeaching a language model to reason starts with a step that looks wasteful: the same prompt is answered many times over, on purpose. This happens during reinforcement learning, not when a user chats with the model. The model produces a group of candidate answers, an external verifier marks each one as correct or incorrect, and the optimizer updates the model so that rewarded answer paths become more likely. The repeated answers are not redundant outputs for a user; they are measurements of the model’s current uncertainty on that prompt.  \nA simple example captures the mechanism. Suppose a model attempts the same math problem eight times. If all eight attempts are wrong, there is no successful attempt to imitate. If all eight are right, there is no failed attempt to move away from. The useful training case is mixed: some attempts are right and some are wrong. Only then can the training rule compare the two sides. In this sense, a prompt teaches through its within-group disagreement.  \nGroup Relative Policy Optimization [GRPO; 1, 2], the workhorse of current verifiable reasoning training, is built around this comparison. GRPO operates in the reinforcement learning with verifiable rewards setting  \n* Equal contribution. Correspondence: {yongyibay, [kallie.a.yearick}@gmail.com](kallie.a.yearick}@gmail.com).  \n(RLVR), where an automatic checker returns a reward, usually 1 for a correct final answer and 0 for an incorrect one. The model supplies the candidate answers; the verifier supplies the rewards; GRPO supplies the rule that converts those rewards into advantages; the optimizer changes the model parameters.  \nFigure 1: The training-time loop studied in this paper. The trainer samples one prompt many times to compare its correct and incorrect attempts, and the group reward standard deviation σ measures whether they disagree.  \nThe mean subtraction inside GRPO needs little controversy: subtracting any action-independent baseline preserves the policy gradient while reducing variance [3] . The contested step is the next one, division by the group standard deviation. It is often dismissed as a normalization detail, yet Liu et al. [4] identify it as the source of a question-level","cbCaid3LgkvfhPvU","https://ap.wps.com/l/cbCaid3LgkvfhPvU","pdf",474208,2,1,18,"English","en",105,"# Introduction\n## Reinforcement learning with verifiable rewards\n## The lens: group reward standard deviation σ\n## Finite-group analysis and the identity","[{\"question\":\"Why does this training approach sample multiple answers for the same prompt?\",\"answer\":\"It measures the model’s uncertainty by comparing correct versus incorrect candidates produced within a group. The verifier assigns rewards (typically 1 or 0), enabling the optimizer to update using within-group disagreement.\"},{\"question\":\"What is the key quantity that unifies GRPO, Dr. GRPO, and DAPO?\",\"answer\":\"The group reward standard deviation σ. GRPO divides by σ, Dr. GRPO omits that division, and DAPO discards groups where σ equals zero.\"},{\"question\":\"How does the paper relate σ to the strength and direction of the policy update?\",\"answer\":\"For binary rewards, the update decomposes into a direction term comparing mean scores of correct versus incorrect rollouts and a multiplier σ that vanishes for unanimous groups and peaks for balanced splits. This yields the group-standard-deviation identity tying σ to the gradient magnitude.\"}]",1784180697,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"grpo-dr-grpo-and-dapo-are-three-operations-on-one-number-the-group-standard-deviation-identity","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/grpo-dr-grpo-and-dapo-are-three-operations-on-one-number-the-group-standard-deviation-identity/82468/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-18","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does this training approach sample multiple answers for the same prompt?","Question",{"text":75,"@type":76},"It measures the model’s uncertainty by comparing correct versus incorrect candidates produced within a group. The verifier assigns rewards (typically 1 or 0), enabling the optimizer to update using within-group disagreement.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the key quantity that unifies GRPO, Dr. GRPO, and DAPO?",{"text":80,"@type":76},"The group reward standard deviation σ. GRPO divides by σ, Dr. GRPO omits that division, and DAPO discards groups where σ equals zero.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper relate σ to the strength and direction of the policy update?",{"text":84,"@type":76},"For binary rewards, the update decomposes into a direction term comparing mean scores of correct versus incorrect rollouts and a multiplier σ that vanishes for unanimous groups and peaks for balanced splits. This yields the group-standard-deviation identity tying σ to the gradient magnitude.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]