[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83807-en":3,"doc-seo-83807-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83807,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","dOPSD On-Policy Self-Distillation for Diffusion Language Models","Diffusion large language models generate text by iterative denoising of a masked sequence, offering parallel decoding but making post-training reasoning difficult. Off-policy supervised fine-tuning suffers exposure bias, while reinforcement learning relies on sparse, sequence-level rewards and is hard to apply without tractable likelihoods. On-policy self-distillation uses a privileged-information teacher for dense token supervision, yet privilege typically depends on instance-specific references not available at inference, weakening student gains. dOPSD derives teacher privilege from the student’s own denoising trajectory, using later decoded steps for token evaluation and improving reasoning and code generation on Dream and LaDA.","dOPSD: On-Policy Self-Distillation for Diffusion Language Models  \nPhuong Tuan Dat 1 , Qi Li 1 , Xinchao Wang 1 ∗  \n1National University of Singapore  \n{phuongtuandat, [liqi}@u.nus.edu](liqi}@u.nus.edu), [xinchao@nus.edu.sg](xinchao@nus.edu.sg)  \n[https://github.com/tuandattt/dOPSD](https://github.com/tuandattt/dOPSD)  \narXiv :2607 .04428v 1 [ cs .CL] 5 Jul 2026  \nAbstract  \nDiffusion large language models (dLLMs) generate text by iteratively denoising a masked sequence, offering a parallel alternative to autoregressive models, but eliciting strong reasoning through post-training remains difficult: supervised fine-tuning is off-policy and suffers from exposure bias, while reinforcement learning gives only sparse, sequence-level rewards and is hard to apply without tractable sequence likelihoods. On-policy self-distillation (OPSD) offers a promising alternative, using one model as both student and teacher to provide dense, token-level, on-policy supervision, but its effectiveness hinges on giving the teacher privileged information (PI) - typically an instance-specific ground-truth reference unavailable at inference - so the student ends up distilling a weak PI-free consensus policy that yields little improvement on dLLM reasoning. We introduce dOPSD, which instead derives the teacher’s privilege directly from the student’sown denoising trajectory, evaluating masked positions using later, more-decoded steps of that same trajectory rather than an external label, so the teacher’s advantage emerges from the model’s own decoding process; on Dream and LLaDA, dOPSD improves both in-domain math reasoning and out-ofdomain code generation, outperforming supervised and onpolicy baselines.  \n1 Introduction  \nDiffusion large language models (dLLMs) have recently emerged as a competitive, non-autoregressive alternative to standard left-to-right language models (Yu, Li, and Wang 2025; Song et al. 2025) . Rather than generating one token at a time, a dLLM begins from a fully masked sequence and produces text through an iterative denoising process: at each step it predicts the clean tokens at the masked positions and commits the most confident ones, progressively filling in the sequence. This paradigm scales to billions of parameters and rivals autoregressive models on general language tasks, while offering parallel decoding and bidirectional context. Yet eliciting strong reasoning from dLLMs through posttraining remains challenging. Supervised fine-tuning on reference solutions is off-policy and suffers from exposure bias, while reinforcement learning with verifiable rewards (RLVR)(Wen et al. 2025; Yang et al. 2026) is costly, supplies only  \n∗Corresponding author.  \na sparse sequence-level reward, and is itself hard to adapt to diffusion models that lack a tractable sequence likelihood. This motivates a post-training signal that is simultaneously dense, on-policy, and free of external teachers or reward models.  \nOn-policy self-distillation (OPSD) was recently proposed for autoregressive LLMs as exactly such a signal (Zhao et al. 2026) . A single model plays two roles that differ only in their conditioning context: a student that sees only the problem, and a teacher that is additionally conditioned on PI, a groundtruth reference solution. The student generates an on-policy rollout, and is trained to match the teacher’s dense, per-token distribution along that rollout, needing neither a larger external teacher nor a reward model. The teacher’s strength, however, comes entirely from an external, instance-specific label that the student never sees at inference. Recent analysis shows this is more than a practical nuisance: unable to condition on the PI, OPSD ends up optimizing a weak, PImarginalized “consensus” of the per-problem teachers rather than any single strong teacher (Zhu et al. 2026). The key, then, is not to discard the privileged teacher, but to source its privilege from something the model itself produces, keeping the dense, on-polic","cbCairTSPHkF1LDj","https://ap.wps.com/l/cbCairTSPHkF1LDj","pdf",794157,7,1,12,"English","en",105,"# Abstract\n# Introduction\n## Motivation: Challenges in Post-Training for dLLMs\n## Prior Work: On-Policy Self-Distillation (OPSD)\n## Key Observation: Privilege from Diffusion Decoding Trajectory\n## Proposed Method: dOPSD","[{\"question\":\"What problem does dOPSD address in diffusion language models?\",\"answer\":\"dOPSD targets the difficulty of eliciting strong reasoning from diffusion language models during post-training when dense, on-policy supervision is hard to obtain.\"},{\"question\":\"Why does standard OPSD underperform for diffusion-based reasoning?\",\"answer\":\"OPSD’s teacher relies on privileged information from instance-specific ground-truth references unavailable at inference, causing the student to distill a weak PI-free consensus policy.\"},{\"question\":\"How does dOPSD obtain privileged information without external labels?\",\"answer\":\"dOPSD derives the teacher’s privilege directly from the student’s own denoising trajectory by evaluating masked positions using later, more-decoded steps of the same trajectory.\"}]",1784190544,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"dopsd-on-policy-self-distillation-for-diffusion-language-models","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/dopsd-on-policy-self-distillation-for-diffusion-language-models/83807/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does dOPSD address in diffusion language models?","Question",{"text":76,"@type":77},"dOPSD targets the difficulty of eliciting strong reasoning from diffusion language models during post-training when dense, on-policy supervision is hard to obtain.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why does standard OPSD underperform for diffusion-based reasoning?",{"text":81,"@type":77},"OPSD’s teacher relies on privileged information from instance-specific ground-truth references unavailable at inference, causing the student to distill a weak PI-free consensus policy.",{"name":83,"@type":74,"acceptedAnswer":84},"How does dOPSD obtain privileged information without external labels?",{"text":85,"@type":77},"dOPSD derives the teacher’s privilege directly from the student’s own denoising trajectory by evaluating masked positions using later, more-decoded steps of the same trajectory.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]