[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82850-en":3,"doc-seo-82850-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82850,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Turning Off-Policy Tokens On-Policy: A Plug-in Approach for Improving LLM Alignment","Reinforcement learning post-training for large language models follows a rollout-then-update workflow, producing unavoidable off-policy training data. Importance sampling corrects this mismatch but becomes unstable as token-level correction ratios compound, leading to exploding variance. Selective Importance Sampling (SIS) transfers off-policy tokens into on-policy tokens using a rejection-sampling inspired token acceptance test, assigning unit weight to accepted tokens while retaining standard IS for rejected ones. SIS is plug-in only modifying importance ratios, with negligible overhead, and is theoretically shown to reduce token- and sequence-level off-policy gradient mismatch. Experiments on dense and MoE LLMs across math and agent benchmarks show consistent objective gains and improved robustness under off-policy regimes.","arXiv :2607 .04728v 1 [ cs .CL] 6 Jul 2026  \nTURNING OFF-POLICY TOKENS ON-POLICY: A PLUGIN APPROACH FOR IMPROVING LLM ALIGNMENT  \nYu Li 1 ,∗ , Xiuyu Li 1 ,∗ ,‡, Mingyang Yi 1 ,†,  \nJiaxing Wang2 , zhangliangxu2 , Zhaolong Xing2 , Zhen Chen2  \n[1](1 Renmin University of China 2 JD.com)[ Renmin University of China](1 Renmin University of China 2 JD.com)[ 2](1 Renmin University of China 2 JD.com)[ JD.com](1 Renmin University of China 2 JD.com)[ ](1 Renmin University of China 2 JD.com){liyu0929,[yimingyang](yimingyang}@ruc.edu.cn)[}](yimingyang}@ruc.edu.cn)[@ruc.edu.cn](yimingyang}@ruc.edu.cn)  \nABSTRACT  \nReinforcement learning (RL) post-training for large language models (LLMs) follows a efficient paradigm of “rollout then update”, which inevitably results in offpolicy training data. To resolve this, Importance sampling (IS) is proposed, while the token-level ratios compound over long sequences, causing severe variance exploded. A natural idea is “transferring” these off-policy token into on-policy token, so that the importance scores for correction are unnecessary. Following this idea, we propose Selective Importance Sampling (SIS), which is inspired by rejection sampling. Concretely, SIS implements by viewing off-policy model as proposal distribution, and implement a token-level rejection test: accepted tokens are viewed as on-policy, so that receive unit importance score, while rejected tokens retain the standard IS correction. Our proposed SIS is theoretically proved reducing the gap between token-level and sequence-level off-policy gradient estimators. The SIS acts as a plug-in that only modifies the importance ratio in the policy loss, adding negligible wall-clock overhead, and can be combine with avast vary of RL post-training algorithms. Experiments on dense and MoE LLMs across math and agent benchmarks show that SIS consistently improves all objectives, while providing substantially stronger robustness under off-policy data.  \n1 INTRODUCTION  \nReinforcement learning (RL) (Sutton et al., 1998) has become the dominant paradigm for posttraining large language models, driving substantial progress in reasoning (Guo et al., 2025; Chenet al., 2025), code generation (Jiang et al., 2025), tool use (Li et al., 2026b), and broader capabilities (Li et al., 2026a;c; Zhang et al., 2026a;b; Tu et al., 2026) . The RL process follows an order of “rollout then update”. However, owing to the considerations of efficiencies in sample reuse (Noukhovitchet al., 2025), asynchronous auto-regressive inference (Fu et al., 2026), researchers inevitably use off-policy data from stale policies to update model. The resulting distribution mismatch biases gradient estimates and degrades training stability (Ma et al., 2025), making effective off-policy gradient correction a fundamental challenge for RL-based post-training methods.  \nImportance sampling (IS) (Tokdar & Kass, 2010) offers a principled correction to off-policy gradient by multiplying a correction ratio, but its variance can grow rapidly in long-horizon reasoning (Metelli et al., 2020), and is sensitive to some outliers (Schulman et al., 2017) . To mitigate this, mainstream methods (Schulman et al., 2017; Shao et al., 2024; Yu et al., 2025) rely on hard clipping. Although effective for stabilization, hard clipping suppresses gradients from highly off-policy samples and loses useful learning signal. Some recent methods (Zheng et al., 2025b; Chen et al., 2025; Gao et al., 2025a) mitigate this issue by soft clipping, but remain heuristic.  \nInstead of working on clipping method, we address this problem by “transferring” the off-policy tokens into on-policy tokens, which fundamentally resolve the off-policy problem. Our method, Selec-  \n∗Equal contribution.  \n†Corresponding author.  \n‡[Work done during an internship at JD.com](Work done during an internship at JD.com).  \ntive Importance Sampling (SIS), a simple and effective method conduct this transfer via acceptancerejection sampling (Robert &","cbCaiiJ20UcE9Gyy","https://ap.wps.com/l/cbCaiiJ20UcE9Gyy","pdf",3190948,3,1,23,"English","en",105,"# Abstract\n# Introduction\n## Off-policy training challenge in RL post-training\n## Importance sampling and variance/clipping limitations\n## SIS method and token-level transfer idea\n## Contributions\n# Background\n## Off-Policy Policy Gradient","[{\"question\":\"What evidence supports SIS effectiveness and practicality?\",\"answer\":\"SIS provides consistent gains across dense and MoE LLMs, improves robustness under challenging off-policy regimes, and acts as a plug-in that only modifies the importance ratio with negligible wall-clock overhead, with ablations showing low sensitivity to hyperparameters.\"}]",1784183409,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"turning-off-policy-tokens-on-policy-a-plug-in-approach-for-improving-llm-alignment","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/turning-off-policy-tokens-on-policy-a-plug-in-approach-for-improving-llm-alignment/82850/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What evidence supports SIS effectiveness and practicality?","Question",{"text":75,"@type":76},"SIS provides consistent gains across dense and MoE LLMs, improves robustness under challenging off-policy regimes, and acts as a plug-in that only modifies the importance ratio with negligible wall-clock overhead, with ablations showing low sensitivity to hyperparameters.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]