[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86283-en":3,"doc-seo-86283-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86283,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Teaching Speech Enhancement Models to Sing: Domain Adaptation from Speech Enhancement to Singing Voice Separation","State-of-the-art speech enhancement models rely on large labeled datasets, while singing voice separation typically faces limited training data. This work reframes singing voice separation as domain adaptation from speech enhancement to singing voice separation. Two adaptation strategies are evaluated: full fine-tuning and parameter-efficient fine-tuning with LoRA on both discriminative and generative models. Adaptation improves separation by 0.29–1.8 dB SDR; full fine-tuning achieves best SVS but harms enhancement via catastrophic forgetting. LoRA reaches competitive SVS while preserving speech enhancement capability with only 6–12% extra parameters and the generative model generalizes better to an unseen test set.","TEACHING SPEECH ENHANCEMENT MODELS TO SING: DOMAIN ADAPTATION FROM SPEECH ENHANCEMENT TO SINGING VOICE  \nSEPARATION  \nPaul A. Bereuter1, Mark D. Plumbley2, Alois Sontacchi 1  \n1Institute of Electronic Music and Acoustics, University of Music and Performing Arts, Graz, Austria  \n2Department of Informatics, King’s College London, London, United Kingdom  \narXiv :2607 . 1 1630v 1 [ cs . SD] 13 Jul 2026  \nABSTRACT  \nState-of-the-art speech enhancement models benefit from large-scale labeled datasets, whereas singing voice separation models suffer from limited available training data. To address this limitation, we formulate singing voice separation as domain adaptation from speech enhancement to singing voice separation. We investigate two fine-tuning strategies: full fine-tuning and parameter-efficient fine-tuning using Low-Rank Adaptation (LoRA) on a discriminative and a generative model. Models with either adaptation strategy outperform the same architectures trained from scratch by 0.29-1.8 dBin Signal-to-Distortion-Ratio. Full fine-tuning yields the highest singing voice separation performance, but catastrophic forgetting degrades speech enhancement performance. LoRA fine-tuning achieves competitive singing voice separation performance while preserving the original speech enhancement capability with only 6-12% additional parameters compared to the base speech enhancement model. Furthermore, the generative model shows improved generalization to an unseen test set. The results demonstrate that adapting pretrained speech enhancement models is an effective strategy for training singing voice separation models in data-scarce scenarios.  \nIndex Terms— singing voice separation, speech enhancement, domain adaptation, low rank adaption  \n1. INTRODUCTION  \nClassical speech enhancement (SE) has primarily addressed acoustic degradations such as environmental noise and reverberation, as tackled in organized challenges such as the Deep Noise Suppression (DNS) Challenge [1–3] . The recent transition to more universal enhancement scenarios, such as the Universality, Robustness, and Generalizability for Speech Enhancement (URGENT) Challenge [4], reflects a broader shift towards models capable of handling diverse interference conditions beyond traditional denoising tasks. Enabled by increasingly large-scale training datasets covering a wide range of degradations, these developments have led to the emergence of generative and hybrid discriminative–generative enhancement architectures [5–8] . Broadening the term speech enhancement, these models  \nThe computing infrastructure used in this work was funded by the digital research infrastructure project “Interactive Audiovisual Digital Twins of Performance Venues”, supported by the Austrian Federal Ministry of Education, Science and Research (BMFWF) . This work was supported by the Engineering and Physical Sciences Research Council (EPSRC) [grant numbers EP/T019751/1, EP/Y028805/1, UKRI397] . For the purpose of open access, the authors have applied a Creative Commons Attribution (CC BY) license to any Author Accepted Manuscript version arising.  \nnot only enable separation of speech from diverse interference conditions but also restoration of degraded speech structures. While subsets of large-scale SE datasets include singing voice as part of the target vocal signals, the associated interference conditions typically remain dominated by acoustic distortions such as environmental noise or reverberation, rather than structured musical accompaniment.  \nIn contrast, singing voice separation (SVS) models are commonly trained independently using music datasets such as MUSDB- 18-HQ [9] and MoisesDB [10], which provide isolated vocal and accompaniment stems for supervised training under realistic musical interference conditions. However, if combined, these datasets remain limited to approximately 35 hours of labeled data, whereas the recent URGENT challenge corpus [4] amounts to roughly 700 hours of audio. Di","cbCaih1DKGD7UkjJ","https://ap.wps.com/l/cbCaih1DKGD7UkjJ","pdf",238364,6,1,5,"English","en",105,"# Abstract\n# Introduction\n# Method\n## Tasks related to separation of vocal signals","[{\"question\":\"What problem does the paper address in training singing voice separation models?\",\"answer\":\"It addresses the gap that singing voice separation usually has limited labeled training data compared with speech enhancement, which benefits from large-scale labeled datasets.\"},{\"question\":\"How do the authors perform domain adaptation from speech enhancement to singing voice separation?\",\"answer\":\"They reformulate SVS as domain adaptation and evaluate two strategies: full fine-tuning of pretrained SE models and parameter-efficient fine-tuning using LoRA.\"},{\"question\":\"What is the main trade-off between full fine-tuning and LoRA fine-tuning?\",\"answer\":\"Full fine-tuning yields the highest SVS performance but can degrade speech enhancement due to catastrophic forgetting, while LoRA preserves the original speech enhancement capability with only 6–12% additional parameters while still improving SVS.\"}]",1784210027,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"teaching-speech-enhancement-models-to-sing-domain-adaptation-from-speech-enhancement-to-singing-voice-separation","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/teaching-speech-enhancement-models-to-sing-domain-adaptation-from-speech-enhancement-to-singing-voice-separation/86283/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address in training singing voice separation models?","Question",{"text":76,"@type":77},"It addresses the gap that singing voice separation usually has limited labeled training data compared with speech enhancement, which benefits from large-scale labeled datasets.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How do the authors perform domain adaptation from speech enhancement to singing voice separation?",{"text":81,"@type":77},"They reformulate SVS as domain adaptation and evaluate two strategies: full fine-tuning of pretrained SE models and parameter-efficient fine-tuning using LoRA.",{"name":83,"@type":74,"acceptedAnswer":84},"What is the main trade-off between full fine-tuning and LoRA fine-tuning?",{"text":85,"@type":77},"Full fine-tuning yields the highest SVS performance but can degrade speech enhancement due to catastrophic forgetting, while LoRA preserves the original speech enhancement capability with only 6–12% additional parameters while still improving SVS.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]