[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86300-en":3,"doc-seo-86300-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86300,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Qwen Music Technical Report","Qwen-Music is a large-scale music generation model designed to produce highly musical, high-fidelity songs with complete vocal singing. It supports text-to-music creation from descriptions, lyrics, and attributes, as well as reference-audio cover generation with controllable style and vocal characteristics. The system splits work into semantic composition (tokenized at 25 Hz and modeled by Qwen-Music-LLM) and acoustic rendering via generative stereo synthesis, improving detail and waveform naturalness. Trained on 5M+ hours of multilingual music, it achieves state-of-the-art results on multiple objective metrics and strong human preference.","arXiv :2607 . 1 1699v 1 [ cs . SD] 13 Jul 2026  \nQwen-Music Technical Report  \nQwen Team  \nAbstract  \nIn this report, we introduce Qwen-Music, a powerful music generation model capable of producing highly musical and high-fidelity songs with complete vocal singing. QwenMusic supports two core tasks: Text to Music Generation, which create entirely new songs from text descriptions, lyrics, and musical attributes, and Cover Song Generation, which reinterprets existing songs with different styles and vocal characteristics. Architecturally, Qwen-Music integrates three core components: Qwen-Music-Tokenizer, Qwen-Music-LLM, and Qwen-Music-Render. Qwen-Music-Tokenizer compresses audio into a 25 Hz single-codebook stream of Music Semantic Tokens that preserve semantic and melodic information for LLM prediction. Based on these tokens, QwenMusic-LLM performs autoregressive music semantic modeling, with a key novelty being a melody-token-based chain-of-thought (Melody-CoT) mechanism that plans melodies before full-song generation, improving creativity, musicality, structural coherence, and reference-audio-based melody cloning. To overcome the fidelity limitations of discrete semantic tokens, Qwen-Music-Render performs generative stereo rendering, enriching acoustic details and producing high-fidelity stereo waveforms. Finally, we train QwenMusic-LLM on more than 5 million hours of multilingual music data covering hundreds of languages. We first apply quality-aware pre-training curriculum, then use progressive post-training, comprising supervised initialization, offline DPO, and online GSPO, to further improve musicality and instruction-following ability. Across 600 Chinese and English prompts, Qwen-Music achieves state-of-the-art results in 13 of 16 objective musicality and audio-quality metrics. Professional evaluators also prefer Qwen-Music over leading proprietary systems. For cover song generation, Qwen-Music preserves reference melodies more accurately than Suno V5.5, Suno V5, and MiniMax Cover on the AI-generated reference set, and outperforms MiniMax Cover on most metrics in the real-world popular-song reference set.  \nFigure 1: Qwen-Music is a powerful and controllable music generation model capable of producing songs with complete vocal singing. It generates complete songs from text descriptions, lyrics, and musical attributes, and uses reference audio for cover song generation with different styles and vocal characteristics.  \n1 Introduction  \nMusic generation has recently emerged as a key challenge in generative modeling, aiming to synthesize music that is musically coherent, acoustically natural, and semantically aligned with user intent (Agostinelli et al., 2023; Schneider et al., 2023; Copet et al., 2024; Lam et al., 2024; Majumder et al., 2024; Huang et al., 2023; Chen et al., 2024) . Unlike general audio generation, song generation requires the joint modeling of lyrics, melody, rhythm, vocal performance, instrumentation, and musical structure over minutes (Yang et al., 2026; Lei et al., 2026b;a; Liu et al., 2025; Gong et al., 2025; 2026; Yuan et al., 2025; Ning et al., 2025; Jiang et al., 2025; Lei et al., 2024) . A useful system should support open-ended song generation from text descriptions and lyrics while allowing users to control musical attributes such as genre, mood, instrumentation, and vocal timbre. It should also reinterpret existing songs by preserving a reference melody while changing style or vocal characteristics. Despite rapid progress in large-scale audio generation models, generating songs with stable melodic development, clear lyric articulation, realistic singing voices, and natural instrumental accompaniment remains challenging.  \nThe central technical challenge is the mismatch between semantic composition and acoustic rendering. At the semantic level, a model must plan lyrics, vocal melody, section structure, repetition, and stylistic progression over long horizons. At the acoustic level, it must rend","cbCaiibmR1aMlJh3","https://ap.wps.com/l/cbCaiibmR1aMlJh3","pdf",5843029,10,1,23,"English","en",105,"# Abstract\n# Introduction\n## Problem and Motivation\n## Key System Design\n# Core Principles\n## Compact Semantic Composition\n## Acoustic Rendering\n# Experiments and Results\n## Human Preference Evaluation\n## Cover Song Generation Performance","[{\"question\":\"Qwen-Music支持哪些核心任务？\",\"answer\":\"Qwen-Music支持两类任务：Text to Music Generation可根据文本描述、歌词和音乐属性生成全新歌曲；Cover Song Generation可在给定参考音频的基础上以不同风格与人声特征重新诠释现有歌曲。\"},{\"question\":\"Qwen-Music的整体架构如何划分语义与音频渲染？\",\"answer\":\"架构由Qwen-Music-Tokenizer、Qwen-Music-LLM和Qwen-Music-Render组成。Tokenizer将音频压缩为25 Hz的音乐语义令牌流，LLM基于令牌进行自回归语义建模并以Melody-CoT规划旋律，Render通过生成式立体声渲染补足音频细节并输出高保真双声道波形。\"},{\"question\":\"Qwen-Music在质量与人类偏好评测中表现如何？\",\"answer\":\"在600个中文与英文提示下，Qwen-Music在16项客观音乐性与音频质量指标中的13项达到最新水平；专业评测者也更倾向于Qwen-Music而非领先的商用系统。同时，在翻唱任务中，它在参考旋律保留准确性与多项指标上优于对比模型，并在真实世界流行歌曲参考集的大多数指标上表现更好。\"}]",1784210323,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"qwen-music-technical-report","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/qwen-music-technical-report/86300/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Qwen-Music支持哪些核心任务？","Question",{"text":76,"@type":77},"Qwen-Music支持两类任务：Text to Music Generation可根据文本描述、歌词和音乐属性生成全新歌曲；Cover Song Generation可在给定参考音频的基础上以不同风格与人声特征重新诠释现有歌曲。","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Qwen-Music的整体架构如何划分语义与音频渲染？",{"text":81,"@type":77},"架构由Qwen-Music-Tokenizer、Qwen-Music-LLM和Qwen-Music-Render组成。Tokenizer将音频压缩为25 Hz的音乐语义令牌流，LLM基于令牌进行自回归语义建模并以Melody-CoT规划旋律，Render通过生成式立体声渲染补足音频细节并输出高保真双声道波形。",{"name":83,"@type":74,"acceptedAnswer":84},"Qwen-Music在质量与人类偏好评测中表现如何？",{"text":85,"@type":77},"在600个中文与英文提示下，Qwen-Music在16项客观音乐性与音频质量指标中的13项达到最新水平；专业评测者也更倾向于Qwen-Music而非领先的商用系统。同时，在翻唱任务中，它在参考旋律保留准确性与多项指标上优于对比模型，并在真实世界流行歌曲参考集的大多数指标上表现更好。","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":20,"slug":134},"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]