[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82022-en":3,"doc-seo-82022-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},82022,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","PS4：代理监督联合训练用于真实目标说话人提取","Training target speaker extraction (TSE) for real conversational mixtures remains difficult because large-scale supervision corpora and clean target-speaker references are unavailable. PS4 presents a proxy-supervised training framework tailored for real-world overlaps, combining a 71,771-sample corpus from four public datasets (Chinese and English) and a joint optimization strategy. Fine-tuning starts from a public BSRNN checkpoint while updating only the separator. Four differentiable objectives—ASR cross-entropy, speaker similarity, frame-level VAD, and perceptual audio quality—are jointly optimized. On REAL-T, PS4 ranks 2nd overall with the best speaker similarity and timing F1.","PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction  \nWanyi Ning 1 ,2 , Wei Zhou 1 , Yingpeng Li 1 , Yinshang Guo3 , Haitao Qian 1 , Yiming Cheng 1  \n1 Yijiahe AI, Nanjing, China 2 Tianjin University, Tianjin, China 3 Nanjing University, Nanjing, China  \n[ningwanyi@126.com](ningwanyi@126.com)  \narXiv :2607 .08 1 1 1v 1 [ cs . SD] 9 Jul 2026  \nAbstract—Training target speaker extraction (TSE) models for real conversational mixtures remains challenging because large-scale training corpora and clean target speech for supervision are unavailable. We present PS4, a proxy-supervised training framework for TSE in real conversational mixtures, with two main contributions. First, we construct a large-scale corpus of 71,771 training samples derived from four public datasets, covering both Chinese and English scenarios. Each sample contains an overlapping speech mixture, per-speaker enrollment audio, a ground-truth transcript, and frame-level voice activity labels. Second, we propose a proxy-supervised joint training strategy that fine-tunes a BSRNN-based TSE model using four complementary differentiable objectives: ASR crossentropy, speaker similarity, frame-level voice activity detection, and perceptual audio quality. Starting from a publicly available pre-trained checkpoint, only the BSRNN separator is updated during fine-tuning. On the REAL-T challenge1 leaderboard, PS4 ranks 2nd overall, achieving the best speaker similarity and timing F1 among all submitted systems.  \nIndex Terms—target speaker extraction, proxy supervision, joint training, real conversational speech  \nI. INTRODUCTION  \nTarget speaker extraction (TSE) aims to isolate the speech of a specific speaker from a multi-talker mixture given a short enrollment utterance as a reference [1]–[5] . It has attracted increasing interest as a core component of personalized speech interfaces, meeting transcription, and assistive hearing systems. State-of-the-art TSE models [6]–[10] have achieved impressive results on standard benchmarks such as VoxCeleb [11], WSJ0-2mix [12] and LibriMix [13] . However, these benchmarks are constructed by artificially mixing clean singlespeaker recordings. In contrast, real conversational recordings exhibit substantially different characteristics: reverberation, background noise, device-specific distortions, and natural turntaking patterns that produce irregular overlap durations and speaker ratios [14]–[19] . Because clean reference signals for individual speakers are unavailable in real-world recordings, it is not straightforward to apply conventional signal-level supervision like SI-SNR loss [6],[20] to train TSE models on such data. As a result, existing TSE systems are predominantly trained on simulated mixtures and may degrade when deployed in real conversational scenarios.  \nTo bridge the gap between simulated benchmarks and realworld deployment, REAL-T [19] was introduced as the first benchmark for evaluating TSE systems on real conversational  \n1[https://real-tse.github.io/challenge/](https://real-tse.github.io/challenge/)  \nmixtures. It provides carefully curated real multi-talker recordings with corresponding speaker enrollment utterances and evaluation annotations, enabling systematic assessment of TSE models under realistic acoustic conditions. However, REALT is solely an evaluation benchmark and does not provide a matched training corpus or clean target speech for supervision. Consequently, existing TSE models, including the official REAL-T baseline, still rely on simulated mixtures for training [1], [2], [9], [19] . How to effectively leverage large-scale real conversational recordings for TSE training therefore remains an open challenge.  \nIn this paper, we present PS4, a proxy-supervised training framework for target speaker extraction from real conversational recordings. Instead of relying on unavailable clean target speech, PS4 leverages multiple proxy supervision signals that can be obtained directly from real","cbCaiqzNui4kFqOo","https://ap.wps.com/l/cbCaiqzNui4kFqOo","pdf",980542,5,1,4,"English","en",105,"# Introduction\n## Real conversational evaluation and the training gap\n## PS4 contributions\n# REAL-PS4 corpus","[{\"question\":\"Why is training target speaker extraction on real conversational mixtures challenging?\",\"answer\":\"Real conversational recordings lack clean target-speaker reference speech and large-scale matched supervision corpora. Their acoustic conditions include reverberation, noise, device distortions, and natural overlapping/turn-taking patterns, making conventional signal-level supervision hard to apply.\"},{\"question\":\"How does PS4 construct its training corpus?\",\"answer\":\"PS4 builds a proxy-supervised corpus (REAL-PS4) by reformatting four public datasets into a unified REAL-T compatible format. The corpus contains 71,771 training samples spanning Chinese and English scenarios, with overlapping mixtures, enrollment utterances, transcripts, and voice activity annotations.\"},{\"question\":\"What is the proxy-supervised joint training strategy in PS4?\",\"answer\":\"PS4 fine-tunes a BSRNN-based TSE model using four complementary differentiable objectives: ASR cross-entropy, speaker similarity, frame-level voice activity detection, and perceptual audio quality. During fine-tuning, only the BSRNN separator is updated starting from a public pre-trained checkpoint.\"}]","PS4：代理监督联合训练用于真实目标说话人提取 | PDF",1784177631,10,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"ps4-proxy-supervised-joint-training-for-real-target-speaker-extraction","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":22},"https://docshare.wps.com/document/ps4-proxy-supervised-joint-training-for-real-target-speaker-extraction/82022/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-29","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is training target speaker extraction on real conversational mixtures challenging?","Question",{"text":76,"@type":77},"Real conversational recordings lack clean target-speaker reference speech and large-scale matched supervision corpora. Their acoustic conditions include reverberation, noise, device distortions, and natural overlapping/turn-taking patterns, making conventional signal-level supervision hard to apply.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does PS4 construct its training corpus?",{"text":81,"@type":77},"PS4 builds a proxy-supervised corpus (REAL-PS4) by reformatting four public datasets into a unified REAL-T compatible format. The corpus contains 71,771 training samples spanning Chinese and English scenarios, with overlapping mixtures, enrollment utterances, transcripts, and voice activity annotations.",{"name":83,"@type":74,"acceptedAnswer":84},"What is the proxy-supervised joint training strategy in PS4?",{"text":85,"@type":77},"PS4 fine-tunes a BSRNN-based TSE model using four complementary differentiable objectives: ASR cross-entropy, speaker similarity, frame-level voice activity detection, and perceptual audio quality. During fine-tuning, only the BSRNN separator is updated starting from a public pre-trained checkpoint.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":22,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":47,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":47,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":30,"doc_module":4,"doc_module_name":47,"category_name":132,"show_sort_weight":30,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":47,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]