[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83432-en":3,"doc-seo-83432-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83432,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators","OPSD-V introduces an on-policy self-distillation approach for post-training few-step autoregressive (AR) video diffusion models. Prior DMD-style few-step AR generators generate long videos with low latency but accumulate errors and weaken motion dynamics. OPSD-V trains with real long-video data as temporal context, providing dense trajectory-level supervision. The student follows exact inference-time rollout, while the teacher uses AR-consistent temporal caching to generate corrective targets, improving long-horizon quality and motion without altering sampling or inference-time cache.","arXiv :2607 .08766v 1 [ cs .CV] 9 Jul 2026  \nOPSD-V  \nOn-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators  \nHongyu Liu1,2 , Chun Wang1,2 , Feng Gao1,†, Xuanhua He1,2 ,  \nYue Ma2 , Ziyu Wan3 , Yong Zhang1,‡, Xiaoming Wei1 , Qifeng Chen2,†  \n1Meituan 2HKUST 3 City University of Hong Kong †Corresponding authors ‡Project lead [hliudq@connect.ust.hk](hliudq@connect.ust.hk)  \nWe propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators, often obtained through DMD-style distillation, can generate long videos with low latency, but still suffer from error accumulation and weakened motion dynamics during long autoregressive rollout. OPSD-V aims to further reduce long-horizon degradation and improve motion dynamics while preserving the original few-step inference path. Our key idea is to introduce real long-video data as temporal context during training and use it to provide dense trajectory-level supervision. Compared with relying only on a short-clip teacher distribution, real long videos offer a richer and cleaner target distribution for supervising long AR rollouts. Specifically, the student follows the exact inferencetime rollout, generating each chunk conditioned on its own previously generated KV cache. In parallel, the teacher is evaluated at the same student-visited denoising states, but uses a cleaner AR-consistent temporal cache in which older history can be replaced by real-video context. To maintain autoregressive consistency and prevent the teacher from becoming a fully teacher-forced oracle, both branches share an initial real-video prefix, and the teacher keeps its most recent cache chunk generated by the model itself. This design provides dense denoising-level corrective targets under on-policy AR cache dynamics, without changing the sampler, number of denoising steps, or inference-time cache mechanism. We apply OPSD-V to representative fewstep AR video models, including Self-Forcing and LongLive. Experiments show consistent improvements in visual quality, motion dynamics, and VBenchLong scores. In a user study with 10 participants comparing 20 video pairs, OPSD-V is preferred over the base models in 66.0% of overall-preference judgments (82.5% excluding ties), demonstrating the effectiveness of on-policy self-distillation with real long-video context for long-horizon AR video generation.  \nProject page: [https://meigen-ai.github.io/OPSD-V](https://meigen-ai.github.io/OPSD-V)  \nCode: MeiGen-AI/OPSD-V  \n1 Introduction  \nVideo generation has rapidly evolved from short, offline clip synthesis to large-scale video foundation models capable of producing high-resolution, text-aligned, and temporally coherent videos. Modern video generation models are largely built upon diffusion transformers (DiTs) (Peebles & Xie, 2022), whose scalability has enabled substantial progress in visual fidelity, motion quality, prompt following, and long-video generation (Yang et al., 2024 ; Kong et al., 2024 ; Wan Team, 2025 ; Meituan LongCat Team et al., 2025 ; HaCohen et al., 2026 ; Brooks et al., 2024) . Despite these advances, most large-scale video foundation models are still primarily designed for offline generation, where an entire video is synthesized after sampling rather than produced continuously during user interaction.  \nIn contrast to offline video generation, real-time video generation requires a model to produce content sequentially with low viewing latency. This requirement naturally favors autoregressive (AR) video models, where each new frame or chunk is generated conditioned on previous outputs, and transformer KV caches are reused to maintain historical context efficiently. To make AR generation practical, recent methods combine causal video modeling with few-step distillation, which compresses slow diffusion or flow models into efficient few-step generators (Yin et al., 2024b ;a; Wang e","cbCaitn4y2Cw6Edg","https://ap.wps.com/l/cbCaitn4y2Cw6Edg","pdf",19289975,4,1,17,"English","en",105,"# Introduction\n## Problem: long-horizon degradation in few-step AR video generation\n## Prior approaches and their limitations\n## OPSD-V core idea: on-policy self-distillation with real long-video context\n## Training design: on-policy rollout and AR-consistent teacher caching\n## Experiments and results\n## User study evaluation","[{\"question\":\"What problem does OPSD-V target in few-step autoregressive video generation?\",\"answer\":\"It addresses long-horizon degradation, including error accumulation and weakened motion dynamics during autoregressive rollout.\"},{\"question\":\"How does OPSD-V use long-video data differently from previous distillation methods?\",\"answer\":\"OPSD-V introduces real long-video data as temporal context and uses it to provide dense trajectory-level supervision, rather than relying only on short-clip teacher distributions.\"},{\"question\":\"Does OPSD-V change the sampler, number of denoising steps, or inference-time cache mechanism?\",\"answer\":\"No. The method provides dense denoising-level corrective targets under on-policy AR cache dynamics without changing the sampler, denoising steps, or inference-time cache behavior.\"}]",1784187633,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"on-policy-self-distillation-for-post-training-few-step-autoregressive-video-generators","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/on-policy-self-distillation-for-post-training-few-step-autoregressive-video-generators/83432/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does OPSD-V target in few-step autoregressive video generation?","Question",{"text":75,"@type":76},"It addresses long-horizon degradation, including error accumulation and weakened motion dynamics during autoregressive rollout.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does OPSD-V use long-video data differently from previous distillation methods?",{"text":80,"@type":76},"OPSD-V introduces real long-video data as temporal context and uses it to provide dense trajectory-level supervision, rather than relying only on short-clip teacher distributions.",{"name":82,"@type":73,"acceptedAnswer":83},"Does OPSD-V change the sampler, number of denoising steps, or inference-time cache mechanism?",{"text":84,"@type":76},"No. The method provides dense denoising-level corrective targets under on-policy AR cache dynamics without changing the sampler, denoising steps, or inference-time cache behavior.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]