[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84881-en":3,"doc-seo-84881-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84881,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling","Single-channel speech separation remains difficult for real-world deployment due to source permutation ambiguity, stochastic sampling variability in generative models, and the operational burden of long-recording inference with chunk-wise processing. A conditional flow-matching approach is introduced to yield an ordered two-source output conditioned on the mixture. A frozen speaker encoder provides source ordering in training and enables biometric best-of-N candidate selection plus chunk-level channel alignment. Libri2Mix evaluation uses SI-SDR, PESQ, and ESTOI, with downstream ASR (cpWER) and speaker verification (EER).","Flow Matching-Based Speech Source Separation with Best-of-N Biometric  \nSampling  \nAnastasia Zorkina 1 Alexandr Anikin 1 Nikita Khmelev 1 2 Anastasiya Korenevskaya 1 Sergey Novoselov 1 2 Vladimir Volokhov 1 2 Maxim Korenevsky 2 Yuriy Matveev 1  \narXiv :2607 .06088v 1 [ cs . SD] 7 Jul 2026  \nAbstract  \nSingle-channel speech separation remains challenging for real-world deployment due to source permutation ambiguity, sampling variability of generative models, and the difficulty of processing long recordings with chunk-wise inference.  \nWe address these issues with a conditional flowmatching-based method that produces an ordered two-source output conditioned on the mixture. A frozen speaker encoder defines the source order during training and is reused at inference for biometric best-of-N candidate selection and chunklevel channel alignment. We evaluate separation quality on Libri2Mix benchmark using SI-SDR, PESQ, and ESTOI, and measure downstream impact using cpWER for automatic speech recognition and EER for speaker verification. The results show that the proposed Transformer U-Net variant is competitive with strong baselines in objective separation metrics and achieves the lowest downstream automatic speech recognition and speaker verification error rates in all evaluated settings.  \n1. Introduction  \nSpeech source separation, also known as the “cocktail party problem”, is a fundamental task in audio processing and speech technologies (Li et al., 2025 ; Araki et al., 2025 ; Wang & Luo, 2025) . It aims to extract individual speech signals from a mixture, benefiting systems such as automatic speech recognition (ASR), speaker recognition (SR), intelligent voice assistants, and others. Despite rapid progress driven by deep learning, speech separation remains challenging due to acoustic complexity, the ill-posed nature of singlechannel inversion, varying overlap patterns, permutation ambiguity, and high speaker variability.  \n1ITMO University, Speech Processing Group, Russia 2 Speech Technology Center Ltd., R&D department, Russia. Correspondence to: Anastasia Zorkina \u003C[zorkina@speechpro.com](zorkina@speechpro.com) >.  \nModern deterministic speech source separation systems based on Transformer and Conformer architectures achieve SI-SDRi above 24 dB on the WSJ0-2mix benchmark (Zhao et al., 2024 ; Shin et al., 2024) . Generative counterparts, including hybrid schemes (Lutati et al., 2024 ; Wang et al., 2024) and fully generative diffusion or flow matching models (Scheibler et al., 2023 ; Dong et al., 2025 ; Scheibler et al., 2025), offer new modeling paradigms that improve perceptual naturalness at the cost of sampling variance and higher computation.  \nDeploying these systems in practice faces additional hurdles: non-causality, processing of long recordings, downstream integration, and computational efficiency. In this work, we propose a practical speech separation system based on conditional flow matching, built on generative speech enhancement models from the NVIDIA NeMo Toolkit (Juki et al., 2024 ; Ku et al., 2025) . We formulate two-speaker separation as a conditional generation of a structured waveform that contains both separated sources. Besides, we introduce a best-of-N biometric sampling procedure that selects the most speaker-disentangled generation among multiple stochastic candidates.  \nThe main contributions of this work are as follows: 1) A flow-matching-based speech separation method adapted from generative speech enhancement. 2) A best-of-N biometric criterion for inference-time candidate selection.  \n3) Long-form processing via chunk-wise generation combined with speaker-recognition-based channel tracking. 4) Evaluation of downstream ASR and SR performance after use of our speech source separation system.  \n2. Background  \nSpeech source separation. Formally, the task is to recover K signals s 1 , . . . , sK from a mixture m = Pk sk. Permutation-invariant training (PIT) (Yu et al., 2017 ; Kolbæk et al., 2017) resolves","cbCaitoHKOPAAUFe","https://ap.wps.com/l/cbCaitoHKOPAAUFe","pdf",2032139,1,6,"English","en",105,"# Introduction\n# Background\n## Speech source separation\n## Speaker recognition\n## Best-of-N sampling","[{\"question\":\"What problem does single-channel speech source separation face in real-world use?\",\"answer\":\"It is limited by source permutation ambiguity, sampling variability from generative models, and the difficulty of handling long recordings via chunk-wise inference.\"},{\"question\":\"How does the proposed method determine an ordered two-source output?\",\"answer\":\"It uses a conditional flow-matching model where a frozen speaker encoder defines the source order during training and is reused at inference for candidate selection and channel alignment.\"},{\"question\":\"How is the approach evaluated and what downstream effects are measured?\",\"answer\":\"Separation quality is measured on Libri2Mix with SI-SDR, PESQ, and ESTOI, while downstream impact is assessed using cpWER for ASR and EER for speaker verification.\"}]",1784198990,15,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"flow-matching-based-speech-source-separation-with-best-of-n-biometric-sampling","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/flow-matching-based-speech-source-separation-with-best-of-n-biometric-sampling/84881/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does single-channel speech source separation face in real-world use?","Question",{"text":75,"@type":76},"It is limited by source permutation ambiguity, sampling variability from generative models, and the difficulty of handling long recordings via chunk-wise inference.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method determine an ordered two-source output?",{"text":80,"@type":76},"It uses a conditional flow-matching model where a frozen speaker encoder defines the source order during training and is reused at inference for candidate selection and channel alignment.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the approach evaluated and what downstream effects are measured?",{"text":84,"@type":76},"Separation quality is measured on Libri2Mix with SI-SDR, PESQ, and ESTOI, while downstream impact is assessed using cpWER for ASR and EER for speaker verification.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]