[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82413-en":3,"doc-seo-82413-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82413,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Wan-Dancer Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation","Generating long-duration, high-definition, rhythm-synchronized dance videos directly from music remains difficult because diffusion-based video models typically cannot reliably go beyond about 20 seconds. Existing methods based on intermediate skeletons or end-to-end synthesis often suffer from temporal drift, identity inconsistency, and repetitive motion when extended. Wan-Dancer introduces a hierarchical global-to-local framework that plans keyframes using full-track musical context and refines local motion for long-range coherence, validated by stable 720p/30fps videos over one minute across multiple genres with audio and text conditioning.","arXiv :2607 .09581v1 [ cs .CV] 10 Jul 2026  \nWan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation  \nMingyang Huang, Peng Zhang, Li Hu, Guangyuan Wang, Bang Zhang  \nTongyi Lab, Alibaba Group  \n{hongcan.hmy, futian.zp, hooks.hl, yixuan.wgy, [zhangbang.zb}@alibaba-inc.com](zhangbang.zb}@alibaba-inc.com)  \nAbstract. Generating long-duration, high-definition, and rhythmically synchronized dance videos directly from music remains a significant challenge, primarily due to the temporal constraints of current diffusion models, which typically fail beyond 20 seconds. Existing approaches, whether they rely on intermediate 3D skeletons or on end-to-end video synthesis, suffer from temporal drift, identity inconsistency, and repetitive motion patterns when extended to longer horizons. To address these limitations, we propose a novel hierarchical framework for minute-scale coherent music-to-dance generation. Our method decouples the process into global keyframe planning and local temporal refinement, leveraging full-track musical context to ensure long-range coherence. Key innovations include dynamic frame rate adaptation via time-mapped RoPE embeddings for precise alignment, an optical-flow-based loss function to enhance motion continuity, and motion-speed control to preserve high-fidelity details during rapid movements. Extensive experiments demonstrate that our framework surpasses the conventional duration barrier, generating stable, 720p/30fps videos exceeding one minute with superior temporal stability. Furthermore, the model exhibits robust versatility across five distinct dance genres, conditioned on both audio and textual prompts, establishing a new state-of-the-art in coherent, long-form dance video synthesis.  \nKeywords: Video Generation · Music-to-Dance · Global-to-Local · Minutescale  \n1 Introduction  \nRecent advances in diffusion-based video generation have significantly elevated the visual fidelity and motion realism of short-form content. However, current state-of-the-art models, such as Wan [29], HunyuanVideo [33], and Seedance [23], remain fundamentally constrained to temporal windows of 5-15 seconds. This limitation stems from the quadratic computational complexity of self-attention mechanisms over long sequences and the inherent difficulty in maintaining longrange temporal coherence without catastrophic forgetting. While existing systems attempt to extend generation duration through segment-wise stitching or  \n2 Mingyang, H., Peng, Z., Li, H., Guangyuan, W., Bang, Z.  \nsliding-window denoising strategies, these approaches frequently introduce visible artifacts, including temporal drift [7,35], scene inconsistency, and progressive identity degradation as the video length increases.  \nWithin the specific domain of music-to-dance generation, research has predominantly evolved along two distinct paradigms, each facing significant limitations. The first, known as the music-to-motion pipeline, involves predicting 3D skeletal trajectories from audio inputs followed by a separate rendering stage [11, 15, 16, 18, 19, 25, 27, 30, 31] . While this decoupled approach offers explicit control over pose dynamics, it is inherently constrained by the brevity of generated clips (typically under 20 seconds) and often suffers from rendering artifacts and insufficient motion richness. Conversely, the second paradigm employs end-to-end methods [31,36] that inherit the temporal constraints of general video diffusion systems. These approaches struggle to generate coherent long sequences, frequently exhibiting identity flickering, degraded spatial resolution, and repetitive motion patterns due to an inadequate modeling of long-horizon rhythmic structures.  \nAddressing this core objective requires establishing a robust mapping between the auditory modality and the visual modality, demanding models that can capture fine-grained cross-modal correlations—aligning rhythm, tempo, and emotional cues with complex spati","cbCaic4tVV703bPQ","https://ap.wps.com/l/cbCaic4tVV703bPQ","pdf",8519482,2,1,17,"English","en",105,"# Introduction\n## Background and limitations of current video diffusion\n## Two paradigms in music-to-dance generation\n## Motivation: cross-modal mapping and long-horizon coherence\n## Proposed hierarchical global-to-local framework","[{\"question\":\"Why is minute-scale music-to-dance video generation difficult with current diffusion models?\",\"answer\":\"Temporal constraints in diffusion-based video models typically fail beyond around 20 seconds, making long-range rhythm and coherence hard to maintain.\"},{\"question\":\"What problem do existing music-to-dance methods face when extending video length?\",\"answer\":\"They often introduce temporal drift, scene inconsistency, identity degradation, and repetitive motion patterns due to insufficient long-horizon modeling.\"},{\"question\":\"How does the proposed Wan-Dancer framework improve temporal coherence over long sequences?\",\"answer\":\"It decouples global keyframe planning from local temporal refinement, using full-track musical context to preserve long-range rhythm and reduce discontinuities.\"}]",1784180194,43,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"wan-dancer-hierarchical-framework-for-minute-scale-coherent-music-to-dance-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/wan-dancer-hierarchical-framework-for-minute-scale-coherent-music-to-dance-generation/82413/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is minute-scale music-to-dance video generation difficult with current diffusion models?","Question",{"text":75,"@type":76},"Temporal constraints in diffusion-based video models typically fail beyond around 20 seconds, making long-range rhythm and coherence hard to maintain.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem do existing music-to-dance methods face when extending video length?",{"text":80,"@type":76},"They often introduce temporal drift, scene inconsistency, identity degradation, and repetitive motion patterns due to insufficient long-horizon modeling.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the proposed Wan-Dancer framework improve temporal coherence over long sequences?",{"text":84,"@type":76},"It decouples global keyframe planning from local temporal refinement, using full-track musical context to preserve long-range rhythm and reduce discontinuities.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]