[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85489-en":3,"doc-seo-85489-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85489,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Stable On-Policy Distillation through Adaptive Target Reformulation","Knowledge distillation (KD) transfers capabilities from large language models to smaller students, but conventional supervised KD can create a training–inference distribution mismatch. On-policy KD learns from student-generated outputs, yet optimization may become unstable when the novice policy diverges sharply from the expert teacher, leading to pathological gradients or diversity collapse. Veto stabilizes objective-level training by reformulating the target distribution in logit space, using a tunable β to veto harmful gradients and balance decisiveness and diversity. Experiments on reasoning and generation tasks show consistent gains over baselines, with code provided.","Stable On-Policy Distillation through Adaptive Target Reformulation  \nIjun Jang Jewon Yeom Juan Yeo Hyunggyu Lim Taesup Kim†  \nGraduate School of Data Science, Seoul National University {ijun0824, jewon0908, juanyeo, sjrksjek, [taesup.kim}@snu.ac.kr](taesup.kim}@snu.ac.kr)  \narXiv :2601 .07 155v 3 [ cs .LG] 11 Jul 2026  \nAbstract  \nKnowledge distillation (KD) is a widely adopted technique for transferring knowledge from large language models to smaller student models. However, conventional supervised KD often suffers from a distribution mismatch between training and inference. While on-policy KD approaches attempt to mitigate this issue by learning directly from student-generated outputs, they frequently encounter training instabilities because the distributional gap between the novice student and the expert teacher is often too wide to bridge directly. These challenges manifest as pathological gradients in forward KL objectives or diversity collapse in reverse KL regimes. To address these limitations, we propose Veto, an objective-level reformulation that constructs a geometric bridge in the logit space. Unlike prior methods that mix data samples, Veto creates an intermediate target distribution that promotes alignment between the teacher and the student. By introducing a tunable parameter β , Veto serves asan Adaptive Gradient Veto that stabilizes optimization by suppressing harmful gradientson low-confidence tokens, while simultaneously acting as a Decisiveness Knob to balance reward-driven performance with output diversity. Experiments across reasoning and generation tasks demonstrate that Veto consistently outperforms supervised fine-tuning and existing baselines. The code is available at [https://github.com/jjun-0824/Veto](https://github.com/jjun-0824/Veto).  \n1 Introduction  \nKnowledge distillation (KD) (Bucilu et al., 2006 ; Hinton et al., 2015) is widely used for transferring capabilities from proprietary models to efficient open-source counterparts and facilitating model self-improvement (Xu et al., 2024) . Beyond model compression, KD has recently become an important mechanism for improving large language  \n†Corresponding author.  \nmodels (LLMs) through self-training and alignment, where a strong teacher supervises a weaker or partially trained student. However, traditional supervised KD (Sanh et al., 2019 ; Kim and Rush, 2016) suffers from exposure bias (Ranzato et al., 2015 ; Bengio et al., 2015): the mismatch between teacher-provided trajectories and the student’s selfgenerated outputs leads to degraded performance in autoregressive tasks (Zhang et al., 2019 ; Arora et al., 2022), especially in long-horizon generation and reasoning settings where errors compound overtime.  \nOn-policy knowledge distillation addresses this limitation by aligning training with inference-time behavior, learning directly from student-generated outputs (Agarwal et al., 2024 ; Lin et al., 2020) . This enables the student to receive feedback on its own predictions and adapt within regions of the probability space it is likely to visit at test time. Building on this idea, recent methods attempt to bridge the teacher-student gap using reverse KL divergences (Gu et al., 2023 ; Wen et al., 2023) or interleaved sampling strategies (Xu et al., 2025) . While these approaches try to bridge the gap by mixing teacher and student tokens at the data level, they largely overlook a critical challenge: the stability of the optimization objective itself. Even with mixed data, forcing a novice student to match an expert’s sharp distribution creates a steep optimization cliff.  \nIn practice, early-stage student policies are highly noisy and often assign near-zero probability to teacher-preferred tokens. As illustrated in Table 1, this leads to unreliable teacher feedback and pathological gradients. Standard forward KL objectives suffer from gradient explosion on such ignorant tokens (Agarwal et al., 2024), while reverse KL objectives, although numerically","cbCaipH9VCP3uiJP","https://ap.wps.com/l/cbCaipH9VCP3uiJP","pdf",722124,2,1,11,"English","en",105,"# Abstract\n# 1 Introduction\n# Gradient Forward KL\n# Veto\n# Mode Collapsing (Reverse KL)\n# Figure 1: Overview of stable on-policy knowledge distillation","[{\"question\":\"What problem does this work address in on-policy knowledge distillation?\",\"answer\":\"Conventional on-policy KD can become unstable because the distribution gap between an early-stage student and the expert teacher is often too large, causing pathological gradients (forward KL) or diversity collapse (reverse KL).\"},{\"question\":\"How does Veto improve stability?\",\"answer\":\"Veto performs an objective-level reformulation by constructing an intermediate target distribution in logit space that bridges teacher and student agreement regions, suppressing harmful updates on low-confidence tokens.\"},{\"question\":\"What role does the parameter β play in Veto?\",\"answer\":\"β acts both as an Adaptive Gradient Veto to prevent gradient explosion in forward KL settings and as a Decisiveness Knob that balances output decisiveness with distributional diversity in reverse KL regimes.\"}]",1784203982,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"stable-on-policy-distillation-through-adaptive-target-reformulation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/stable-on-policy-distillation-through-adaptive-target-reformulation/85489/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does this work address in on-policy knowledge distillation?","Question",{"text":75,"@type":76},"Conventional on-policy KD can become unstable because the distribution gap between an early-stage student and the expert teacher is often too large, causing pathological gradients (forward KL) or diversity collapse (reverse KL).","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Veto improve stability?",{"text":80,"@type":76},"Veto performs an objective-level reformulation by constructing an intermediate target distribution in logit space that bridges teacher and student agreement regions, suppressing harmful updates on low-confidence tokens.",{"name":82,"@type":73,"acceptedAnswer":83},"What role does the parameter β play in Veto?",{"text":84,"@type":76},"β acts both as an Adaptive Gradient Veto to prevent gradient explosion in forward KL settings and as a Decisiveness Knob that balances output decisiveness with distributional diversity in reverse KL regimes.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]