[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83170-en":3,"doc-seo-83170-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83170,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","MMGenre Benchmarking Singing Voice Synthesis across Multiple Musical Genres","Singing voice synthesis (SVS) has advanced quickly, yet generalization across musical genres remains insufficiently studied, because existing benchmarks are dominated by pop and cannot support systematic genre-level diagnosis. MMGenre introduces a multi-genre SVS benchmark covering 10 major genres and 26 subgenres, backed by an automatic pipeline that builds genre-aligned music scores using text-to-music generation. Evaluations show synthesized vocals have weak genre separability, and zero-shot adaptation brings only marginal gains, while lightweight genre-specific continued training substantially improves genre alignment. MMGenre provides a standardized framework and highlights core challenges.","MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres  \nWenhao Feng  1, Yuxun Tang  1, Jiatong Shi  2, Qin Jin  1 ,∗∗  \n1 AIM3 Lab, Renmin University of China, China  \n2 Carnegie Mellon University, United States  \n[wenhaofeng@ruc.edu.cn](wenhaofeng@ruc.edu.cn) , [tangyuxun@ruc.edu.cn](tangyuxun@ruc.edu.cn) , [jiatongs@cs.cmu.edu](jiatongs@cs.cmu.edu) , [qjin@ruc.edu.cn](qjin@ruc.edu.cn)  \narXiv :2607 .06986v 1 [ cs . SD] 8 Jul 2026  \nAbstract  \nSinging voice synthesis (SVS) has progressed rapidly, yet its ability to generalize across diverse musical genres remains underexplored. Existing benchmarks are heavily biased toward pop music, limiting systematic analysis of genre-dependent behavior. We introduce MMGenre, a benchmark for multi-genre SVS diagnosis, supported by an automatic pipeline for constructing genre-aligned music scores. MMGenre spans 10 major genres and 26 subgenres, enabling comprehensive analysis of genre-aware synthesis. Extensive evaluation of representative SVS models reveals limited genre discrimination: synthesized vocals across genres exhibit highly similar acoustic characteristics and weak separability. While zero-shot genre adaptation yields only marginal improvements, lightweight genrespecific continued training leads to substantial gains. MMGenre provides a standardized framework for multi-genre SVS evaluation and exposes critical challenges in achieving genre-aware singing voice synthesis.  \nIndex Terms: singing voice synthesis, genre-aware evaluation, benchmark  \n1. Introduction  \nSinging voice synthesis (SVS) aims to generate expressive singing voices directly from symbolic music scores and has achieved substantial progress with recent neural and diffusionbased models [1–4] . Prior research has largely focused on improving naturalness [5–7], expressiveness [8], and controllability [9, 10] of synthesized singing voices.  \nHowever, singing performance is inherently shaped by musical genre, a high-level semantic attribute that is immediately recognizable to human listeners. Genres such as pop, rock, jazz, and classical music exhibit systematic differences in vocal timbre, articulation, phrasing, rhythmic emphasis, and expressive conventions [11] . In real-world music production and consumption, genres form stable perceptual categories that shape listener evaluation and preference [12, 13] . Despite its perceptual salience and practical importance, genre has rarely been treated as a first-class evaluation dimension in SVS research. A fundamental yet underexplored question remains: how well do current SVS models generalize across musical genres?  \nThis gap is largely due to data and evaluation limitations. Public SVS datasets are overwhelmingly dominated by pop music (e.g., M4Singer [14], Opencpop [15], ACEOpencpop [16]), offering limited coverage of other genres and making large-scale systematic genre-level analysis infeasible. Existing style-aware SVS studies primarily target lower-level  \nattributes, such as emotion, tempo, pitch range, or specific **indicates the corresponding author.  \nsinging techniques [9, 10, 17, 18], rather than modeling genre asa holistic musical condition. Consequently, it remains unclear whether current SVS systems truly capture genre-specific characteristics or simply reproduce surface-level acoustic patterns learned from biased data.  \nTo overcome these data limitations, we turn to recent advances in text-to-music (T2M) generation [19–23] . Modern T2M systems enable explicit genre specification through natural language prompts. Compared to real-world recordings, T2M-generated music is easier to scale, poses fewer copyright constraints, and can be systematically diversified across genres, making it a practical and scalable resource for constructing genre-balanced evaluation data.  \nLeveraging these advances, we propose an automatic pipeline that uses T2M models to generate genre-aligned music scores, enabling scalable construction of multi-genre data.","cbCailfuQI3OO5cF","https://ap.wps.com/l/cbCailfuQI3OO5cF","pdf",904976,3,1,6,"English","en",105,"# Introduction\n## Gap in existing SVS benchmarks\n## Data and evaluation limitations\n## Proposed MMGenre and pipeline\n## Key evaluation findings","[{\"question\":\"What is MMGenre and what problem does it address?\",\"answer\":\"MMGenre is a benchmark designed to evaluate how well singing voice synthesis models generalize across multiple musical genres. It addresses the lack of systematic genre coverage in existing SVS benchmarks, which are largely biased toward pop music.\"},{\"question\":\"How are genre-aligned music scores constructed for MMGenre?\",\"answer\":\"MMGenre uses an automatic pipeline that leverages text-to-music (T2M) generation to produce genre-aligned music scores at scale, enabling genre-balanced evaluation data.\"},{\"question\":\"What do evaluations reveal about genre awareness in current SVS models?\",\"answer\":\"Synthesized vocals across genres show highly similar acoustic characteristics and weak separability, even when conditioned on genre-specific scores. Zero-shot genre adaptation yields only marginal improvements, whereas lightweight genre-specific continued training delivers substantial gains.\"}]",1784185728,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mmgenre-benchmarking-singing-voice-synthesis-across-multiple-musical-genres","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/mmgenre-benchmarking-singing-voice-synthesis-across-multiple-musical-genres/83170/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is MMGenre and what problem does it address?","Question",{"text":75,"@type":76},"MMGenre is a benchmark designed to evaluate how well singing voice synthesis models generalize across multiple musical genres. It addresses the lack of systematic genre coverage in existing SVS benchmarks, which are largely biased toward pop music.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How are genre-aligned music scores constructed for MMGenre?",{"text":80,"@type":76},"MMGenre uses an automatic pipeline that leverages text-to-music (T2M) generation to produce genre-aligned music scores at scale, enabling genre-balanced evaluation data.",{"name":82,"@type":73,"acceptedAnswer":83},"What do evaluations reveal about genre awareness in current SVS models?",{"text":84,"@type":76},"Synthesized vocals across genres show highly similar acoustic characteristics and weak separability, even when conditioned on genre-specific scores. Zero-shot genre adaptation yields only marginal improvements, whereas lightweight genre-specific continued training delivers substantial gains.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]