[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81586-en":3,"doc-seo-81586-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81586,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Transition Matching Distillation for Fast Video Generation","Large video diffusion and flow models enable high-quality video generation, yet their multi-step sampling is too slow for real-time interactive applications. Transition Matching Distillation (TMD) distills video diffusion models into efficient few-step generators by matching the teacher’s multi-step denoising trajectory with a compact probability transition process. Each transition is implemented as a lightweight conditional flow using a hierarchical student with a main backbone and a flow head, trained via distribution matching. Experiments on Wan text-to-video models show strong speed–quality trade-offs, outperforming prior distilled approaches under similar inference costs.","arXiv :2601 .0988 1v2 [ cs .CV] 9 Jul 2026  \nTransition Matching Distillation for Fast Video Generation  \nWeili Nie¹∗ , Julius Berner¹∗ , Nanye Ma², Chao Liu¹, Saining Xie², Arash Vahdat¹  \n¹NVIDIA ²NYU ∗ equal contribution  \nFigure 1 . Generated examples from TMD. Four frames of 5s 480p videos generated from two text prompts using our TMD method (distilled from Wan2 . 1 14B T2V) with two different (effective) number of function evaluations (NFE) .  \nAbstract  \nLarge video diffusion and flow models have achieved remarkable success in high-quality video generation, but their use in real-time interactive applications remains limited due to their inefficient multi-step sampling process. In this work, we present Transition Matching Distillation (TMD), a novel framework for distilling video diffusion models into efficient few-step generators. The central idea of TMD is to match the multi-step denoising trajectory of a diffusion model with a few-step probability transition process, where each transition is modeled as a lightweight conditional flow. To enable efficient distillation, we decompose the original diffusion backbone into two components: (1) a main backbone, comprising the majority of early layers, that extracts semantic representations at each outer transition step; and (2) a flow head, consisting of the last few layers, that lever-  \nages these representations to perform multiple inner flow updates. Given a pretrained video flow model, we first introduce a flow head to the model, and adapt it into a conditional flow map. We then apply distribution matching distillation to the student model with flow head rollout in each transition step. Extensive experiments on distilling Wan2 .1 1.3B and 14B text-to-video models demonstrate that TMD provides a flexible and strong trade-off between generation speed and visual quality. In particular, TMD outperforms existing distilled models under comparable inference costs in terms of visual fidelity and prompt adherence.  \n1. Introduction  \nRecent progresses in large-scale diffusion models [19, 45] have significantly advanced the frontier of video generation [1, 5, 15, 25, 49, 57] . Open-sourced models (such as  \nHunyuanVideo [25], Wan [49] and Cosmos [1]) and commercial text-to-video (T2V) systems (such as Sora, Veo and Kling) demonstrate remarkable capabilities in synthesizing coherent and photorealistic videos from text prompts. Despite their success, sampling inefficiency remains a central bottleneck. Standard diffusion models rely on a multistep denoising process, often requiring hundreds of iterative steps, to progressively transform noise into realistic outputs [25, 49] . This iterative nature leads to high inference latency and computational cost, rendering large diffusion models impractical for interactive applications such as realtime video generation, content editing, or world modeling for agent training. Accelerating diffusion sampling without sacrificing visual quality becomes a key open challenge.  \nA growing body of research has explored diffusion distillation to compress long denoising trajectories into a small number of inference steps. Existing approaches can be broadly categorized into two families: (1) trajectory-based distillation, which includes knowledge distillation [36, 40] and consistency models [16, 17, 35, 46] that directly regress the teacher’s denoising trajectories; and (2) distributionbased distillation, encompassing adversarial [41, 42] and variational score distillation [58, 59, 70] methods that align the student and teacher distributions. These techniques can reduce the sampling process to as few as one or two steps in the image domain. However, extending them to video diffusion models presents unique challenges. Videos exhibit high spatiotemporal dimensionality and complex interframe dependencies, making it difficult to preserve both global motion coherence and fine-grained spatial details during distillation. Most existing methods treat the diffusion ","cbCaipEuoQKzeE08","https://ap.wps.com/l/cbCaipEuoQKzeE08","pdf",18384255,3,1,26,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"What problem does Transition Matching Distillation (TMD) address in video generation?\",\"answer\":\"TMD targets the sampling inefficiency of large video diffusion models, which typically require many iterative steps and therefore cause high inference latency for interactive use cases.\"},{\"question\":\"How does TMD reduce the number of sampling steps while maintaining quality?\",\"answer\":\"TMD replaces the many-step denoising trajectory with a compact few-step probability transition process, where each transition captures distribution evolution across separated noise levels.\"},{\"question\":\"What is the role of the hierarchical student architecture in TMD?\",\"answer\":\"The student is decomposed into a main backbone that extracts semantic representations at each outer transition step and a flow head that performs inner refinement through multiple lightweight flow updates.\"}]",1784174521,66,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"transition-matching-distillation-for-fast-video-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/transition-matching-distillation-for-fast-video-generation/81586/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does Transition Matching Distillation (TMD) address in video generation?","Question",{"text":75,"@type":76},"TMD targets the sampling inefficiency of large video diffusion models, which typically require many iterative steps and therefore cause high inference latency for interactive use cases.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does TMD reduce the number of sampling steps while maintaining quality?",{"text":80,"@type":76},"TMD replaces the many-step denoising trajectory with a compact few-step probability transition process, where each transition captures distribution evolution across separated noise levels.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the role of the hierarchical student architecture in TMD?",{"text":84,"@type":76},"The student is decomposed into a main backbone that extracts semantic representations at each outer transition step and a flow head that performs inner refinement through multiple lightweight flow updates.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]