[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85956-en":3,"doc-seo-85956-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85956,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Dance to Music Generation Leveraging Pre-training with Unpaired Data and Contrastive Alignment","Dance-to-music generation supports choreography assistance and automatic accompaniment by synchronizing body motion with sound over time. Using human joint positions as motion representation enables explicit control, lightweight processing, privacy preservation, and straightforward integration with motion capture or pose estimation systems. A key obstacle is limited high-quality paired dance–music data because accurate synchronization and rights clearance are costly. The proposed dance-conditioned framework leverages unpaired and paired data via pretrained unimodal encoders, beat-guided contrastive alignment, and ControlNet-style conditioning on a text-to-audio diffusion model, improving alignment and audio quality on AIST++.","Dance to Music Generation leveraging Pre-training with Unpaired  \ndata and Contrastive Alignment  \nRyota Kimura 1 ,2 Sangheon Park 1 ,3 Natalia Polouliakh 1 Taketo Akama 1  \n1 Sony Computer Science Laboratories 2 Keio University 3 Georgia Institute of Technology  \n[ryokimu-2000@keio.jp](ryokimu-2000@keio.jp) , [sangheon@gatech.edu](sangheon@gatech.edu) , [nata@sony.csl.co.jp](nata@sony.csl.co.jp) , [taketo.akama@gmail.com](taketo.akama@gmail.com)  \narXiv :2607 . 10537v 1 [ cs . SD] 12 Jul 2026  \nAbstract  \nDance-to-music generation is a promising task for applications such as choreography support and automatic accompaniment, where temporal coordination between body movement and sound is essential. In particular, using human joint positions as the motion representation is attractive because they explicitly capture body dynamics while being lightweight, privacy-preserving, and easy to integrate with motion capture and poseestimation pipelines. A central challenge in this setting, however, is the scarcity of high-quality paired dance–music data, since collecting accurately synchronized pairs is costly and often constrained by copyright and performance rights. This makes it difficult to train end-to-end models solely from paired data. To address this issue, we propose a dance-conditioned music generation framework that efficiently exploits both unpaired and paired data. Our method combines pretrained unimodal encoders for motion and music, beat-guided contrastive pretraining to align their feature spaces, and a ControlNet-style conditioning module on top of a pretrained text-to-audio diffusion model. Experiments on AIST++ demonstrate that the proposed techniques improve both dance–music alignment and audio quality, as confirmed by quantitative and qualitative evaluations. Compared to a state-of-the-art method, our approach achieves superior dance alignment performance and competitive audio quality. Code is available at [https:](https:)//[github.com/kmraven/AudioLDM-ControlNet](github.com/kmraven/AudioLDM-ControlNet).  \n1 Introduction  \nDance-to-music generation, which generates music from dance motion, is a useful task in settings where temporal consistency between body movement and sound is important, such as choreography support and automatic accompaniment. In particular, using joint positions as dance information has important significance. Joint positions are a representation that allows information related to body movement to be handled relatively explicitly while separating it from appearancerelated factors contained in video, and thus they have advantages from the viewpoints of clarifying the conditioning target and improving controllability. In addition, because joint positions are lightweight and highly anonymous, they are easy to handle in practical ap-  \nplications and are also well suited for integration with motion capture, pose estimation, and various sensorprocessing pipelines. Therefore, dance-to-music generation using joint positions has unique importance both methodologically and practically. For this problem setting, several prior studies have already been conducted, mainly focusing on architectural design based on autoregressive models or diffusion models [1–7] .  \nOn the other hand, a central challenge in this problem setting is the lack of high-quality paired dance– music data. Collecting datasets that accurately synchronize dance motion data, such as time series of joint positions, with music is costly. It requires securing performers, preparing recording environments, annotating the data, and verifying synchronization. In addition, copyright issues related to music and rights associated with performances and choreography pose major barriers. Furthermore, because dance covers a wide range of styles, it is not easy to construct a large-scale paired dataset that sufficiently includes rare styles and individualized bodily expressions. As a result, compared with single-modal data such as motion-only or musiconly","cbCaistX32t4Oehd","https://ap.wps.com/l/cbCaistX32t4Oehd","pdf",1304902,4,1,"English","en",105,"# Abstract\n# Introduction\n## Problem: scarcity of paired dance–music data\n## Related work and motivation\n## Proposed framework and key contributions","[{\"question\":\"Why is dance-to-music generation challenging in practice?\",\"answer\":\"High-quality paired dance–music data is scarce because accurately synchronized motion and audio collection is costly and constrained by copyright and performance rights. This limits end-to-end learning from paired data alone.\"},{\"question\":\"What motion representation does the method use, and what are its benefits?\",\"answer\":\"The method uses human joint positions as the motion representation. It captures body dynamics explicitly, remains lightweight and anonymous, preserves privacy, and is easier to integrate with motion capture and pose estimation pipelines.\"},{\"question\":\"How does the proposed approach use both unpaired and paired data?\",\"answer\":\"It employs pretrained unimodal encoders for motion and music, aligns their feature spaces with beat-guided contrastive pretraining, and adds a ControlNet-style conditioning module on top of a pretrained text-to-audio diffusion model.\"}]",1784207367,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"dance-to-music-generation-leveraging-pre-training-with-unpaired-data-and-contrastive-alignment","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":20},"https://docshare.wps.com/document/dance-to-music-generation-leveraging-pre-training-with-unpaired-data-and-contrastive-alignment/85956/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why is dance-to-music generation challenging in practice?","Question",{"text":74,"@type":75},"High-quality paired dance–music data is scarce because accurately synchronized motion and audio collection is costly and constrained by copyright and performance rights. This limits end-to-end learning from paired data alone.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What motion representation does the method use, and what are its benefits?",{"text":79,"@type":75},"The method uses human joint positions as the motion representation. It captures body dynamics explicitly, remains lightweight and anonymous, preserves privacy, and is easier to integrate with motion capture and pose estimation pipelines.",{"name":81,"@type":72,"acceptedAnswer":82},"How does the proposed approach use both unpaired and paired data?",{"text":83,"@type":75},"It employs pretrained unimodal encoders for motion and music, aligns their feature spaces with beat-guided contrastive pretraining, and adds a ControlNet-style conditioning module on top of a pretrained text-to-audio diffusion model.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]