[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86039-en":3,"doc-seo-86039-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86039,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation","On-Policy Self-Distillation (OPSD) improves and aligns Large Language Models, but it can degrade downstream performance in complex reasoning. This work studies the failure mode and defines “Thinking Collapse” as a sharp drop in native intermediate reasoning, measured by epistemic-token density (ET per 1k). Entropy-based gradient masking and token-level target analysis show the collapse is triggered by aggressive teacher gradients at high-student-entropy decision forks. An Adaptive Dual-Perspective OPSD (AD-OPSD) moderates the objective via asymmetrical pointwise divergence gating, preserving native thinking capacity while retaining error correction, achieving up to +4.1% absolute accuracy improvements.","Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation  \nKeqin Peng 1 , Chen Li 1 , Yuanxin Ouyang 1 , Yancheng Yuan2∗, Liang Ding3 *  \n1Beihang University 2Hong Kong Polytechnic University 3Alibaba Group  \n[keqin.peng@buaa.edu.cn](keqin.peng@buaa.edu.cn) [liangding.liam@gmail.com](liangding.liam@gmail.com)  \narXiv :2607 . 10805v 1 [ cs .CL] 12 Jul 2026  \nAbstract  \nOn-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs) . However, in complex reasoning tasks, OPSD paradoxically degrades downstream performance. In this paper, we systematically investigate this pathology and identify a severe optimization trap we define as Thinking Collapse—a sharp decline in the model’s native intermediate reasoning behavior, measured by epistemic-token density (ET per 1k) . Through entropy-based gradient masking and token-level target analysis, we show that this collapse is triggered by aggressive teacher gradients at high-student-entropy decision forks, where student epistemic tokens are frequently suppressed into teacher non-epistemic targets and are highly concentrated in high pointwise student-teacher divergence regions. To resolve this optimization pathology, we propose Adaptive Dual-Perspective OPSD (AD-OPSD), a robust control framework that dynamically moderates the self-distillation objective. AD-OPSD selectively anchors high-suppression-risk sandboxed tokens toa reference prior derived from the frozen base model via an asymmetrical pointwise divergence gate, preserving native thinking capacity while retaining OPSD’s error-correcting power. Extensive experiments across competitive mathematical benchmarks show that AD-OPSD improves over standard OPSD by up to +4.1% absolute average accuracy across diverse model scales and datasets. Further analysis demonstrates that AD-OPSD mitigates thinking collapse and generalizes robustly to different post-training paradigms.  \n1 Introduction  \nRecently, on-policy self-distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (Zhao et al.,  \n* Corresponding Authors.  \n2026 ; Shenfeld et al., 2026 ; Hübotter et al., 2026) . By distilling ground-truth-conditioned, on-policy target distributions of a teacher model into the student policy over its own sampled rollouts, OPSD avoids the exposure bias of off-policy imitation and the sparse-reward bottleneck of reinforcement learning. This dense, token-level supervision has yielded remarkable success in general alignment and preference-tuning tasks.  \nHowever, when applied to complex reasoning tasks, OPSD paradoxically degrades downstream performance (Kim et al., 2026b ; Kaur et al., 2026) . Prior studies attribute this degradation to the suppression of epistemic verbalizations (Kim et al., 2026b) or related “fork suppression” at highentropy decision points (Kaur et al., 2026) . Yet, these works primarily document empirical symptoms; the underlying mechanics of where and why this suppression occurs remain unclear, leaving how to resolve this optimization trap without losing OPSD’s error-correcting power an open challenge. In this paper, we systematically investigate this pathology, identifying a severe optimization trap we define as Thinking Collapse—a precipitous decline in the model’s native intermediate reasoning steps, measured as epistemic tokens per 1K generated tokens (ET per 1k) . To understand its mechanics, inspired by the findings of Wang et al.(2026) and Xu et al. (2026) that reasoning updatesand critical decision forks are heavily concentrated on small sets of high-entropy tokens, we propose an entropy-based gradient masking diagnostic experiment to locate the spatial boundary of thinking collapse. We find that masking gradients on highstudent-entropy tokens substantially recovers thinking density, but also exposes a correction trade-off.  \nSpecifically, masking the self-distillation gradients of only the top 20%","cbCaiaro4Zb3h4zZ","https://ap.wps.com/l/cbCaiaro4Zb3h4zZ","pdf",424544,3,1,16,"English","en",105,"# Abstract\n# 1 Introduction","[{\"question\":\"What problem does the paper address in On-Policy Self-Distillation (OPSD)?\",\"answer\":\"OPSD can paradoxically degrade downstream performance on complex reasoning tasks. The paper focuses on identifying and explaining the underlying optimization pathology called “Thinking Collapse.”\"},{\"question\":\"How is “Thinking Collapse” measured in this work?\",\"answer\":\"Thinking Collapse is defined as a precipitous decline in native intermediate reasoning steps, measured by epistemic-token density (ET per 1k).\"},{\"question\":\"What approach does AD-OPSD use to mitigate Thinking Collapse?\",\"answer\":\"AD-OPSD uses an adaptive, asymmetrical soft-gating mechanism that anchors high-suppression-risk tokens to a reference prior derived from the frozen base model via a sigmoid pointwise KL gate, preserving native exploratory reasoning while keeping corrective gradients.\"}]",1784208018,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"diagnosing-and-mitigating-thinking-collapse-in-on-policy-self-distillation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/diagnosing-and-mitigating-thinking-collapse-in-on-policy-self-distillation/86039/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in On-Policy Self-Distillation (OPSD)?","Question",{"text":75,"@type":76},"OPSD can paradoxically degrade downstream performance on complex reasoning tasks. The paper focuses on identifying and explaining the underlying optimization pathology called “Thinking Collapse.”","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is “Thinking Collapse” measured in this work?",{"text":80,"@type":76},"Thinking Collapse is defined as a precipitous decline in native intermediate reasoning steps, measured by epistemic-token density (ET per 1k).",{"name":82,"@type":73,"acceptedAnswer":83},"What approach does AD-OPSD use to mitigate Thinking Collapse?",{"text":84,"@type":76},"AD-OPSD uses an adaptive, asymmetrical soft-gating mechanism that anchors high-suppression-risk tokens to a reference prior derived from the frozen base model via a sigmoid pointwise KL gate, preserving native exploratory reasoning while keeping corrective gradients.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]