[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86473-en":3,"doc-seo-86473-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86473,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Breaking the Quality–Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization","Generative streaming models for Target Speaker Extraction (TSE) often suffer a quality–intelligibility trade-off: maximizing perceptual audio quality can harm speech intelligibility, while optimizing for intelligibility can degrade audio quality. The document shows this trade-off stems from an inappropriate choice of optimization anchor, where direct optimization against audio quality metrics triggers catastrophic reward hacking that erases pronunciation-critical content. It proposes enriched local spectro-temporal modeling using a larger Conformer convolution kernel and a WavLM-anchored DPO fine-tuning strategy that resists hacking via deep acoustic feature anchoring. Under 560 ms streaming chunks, the method yields a 10.9% relative intelligibility improvement (WER 0.138→0.123) with marginal gains in audio quality and speaker similarity.","arXiv :2607 . 10 19 1v 1 [ cs . SD] 11 Jul 2026  \nBreaking the Quality–Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization  \nShuhai Peng 1 ∗ , Jinjiang Liu 1 ∗ , Hui Lu2 , Liyang Chen 1 , Guiping Zhong3 , Jiakui Li3 , Shiyin Kang3 , Zhiyong Wu 1†  \n1 Tsinghua University,  \n2 The Chinese University of Hong Kong,  \n3 SenseTime  \nAbstract. Generative streaming models for Target Speaker Extraction (TSE) commonly exhibit a quality–intelligibility trade-off, wherein naive optimization for perceptual audio quality tends to degrade speech intelligibility, and conversely. We reveal that this trade-off arises not from the constraints of streaming architectures, but from an inappropriate choice of optimization anchor. Directly optimizing against audio quality metrics induces catastrophic reward hacking, where content critical to pronunciation and intelligibility is systematically erased to maximize a proxy score. To break this bottleneck, we propose two complementary improvements: an enlarged Conformer convolution kernel for richer local spectro-temporal modeling, and WavLM-anchored Direct Preference Optimization (DPO) fine-tuning strategy. DPO preference pairs are ranked by WavLM cosine similarity, a deep acoustic feature encoding both phonetic structure and speaker identity, providing an optimization anchor that resists hacking. Under a 560 ms streaming chunk size, the proposed method achieves a 10.9% relative intelligibility improvement (word error rate: 0.138 →0.123), with marginal simultaneous gains in audio quality and speaker similarity.  \nKeywords: Direct Preference Optimization · Streaming TSE · Quality– Intelligibility Trade-off · Deep Feature Anchoring · Reward Hacking.  \n1 Introduction  \nTarget Speaker Extraction (TSE) aims to isolate the speech of a designated speaker from a complex acoustic mixture of interfering speakers and background noise [1] . Unlike blind source separation, which treats all concurrent sources symmetrically, TSE leverages auxiliary reference cues, typically a short enrollment utterance of the target speaker, to focus on a single acoustic target. This capability underpins critical real-world applications including teleconferencing systems, voice-controlled assistants, and automatic speech recognition in multitalker environments.  \n∗ Equal contribution.  \n† Corresponding author.  \n2 S. Peng et al.  \nFor decades, the TSE domain was dominated by discriminative approaches such as SpEx+ [2] and WeSep [3], which estimate time-frequency masks or filters to suppress interference. While computationally efficient, these methods introduce processing artifacts and struggle to reconstruct missing spectral details. More recently, generative models have catalyzed a paradigm shift: by treating speech extraction as a conditional generation task, language-model-based generative architectures such as TSELM-L [4] and LauraTSE [5] achieve substantially higher audio fidelity compared to discriminative counterparts. However, these generative models rely on global bidirectional context that is fundamentally incompatible with real-time streaming deployments, where only past and current information is available.  \nThe streaming-aware StarTSE [6] established that autoregressive generative backbones could be adapted for streaming scenarios through chunk-wise interleaved splicing. However, adapting these models to streaming TSE exposes a critical tension intrinsic to real-time generation: the quality–intelligibility tradeoff. A generative policy tuned strictly for acoustic smoothness will suppress the high-frequency bursts of stop consonants and fricatives, yielding high perceptual scores but severe intelligibility collapse. Conversely, optimizing purely for WER degrades perceptual audio quality. In the streaming setting, word error rate (WER) serves as the primary deployment criterion for intelligibility: it reflects the intelligibility and accuracy with which the mode","cbCaisPkADiWaUia","https://ap.wps.com/l/cbCaisPkADiWaUia","pdf",879827,4,1,14,"English","en",105,"# Introduction\n## Target Speaker Extraction and Streaming Constraints\n## Quality–Intelligibility Trade-off and Reward Hacking\n## Preference-based Fine-tuning with DPO\n## Comparing DPO Ranking Anchors\n# Main Contributions","[{\"question\":\"What causes the quality–intelligibility trade-off in streaming TSE according to the document?\",\"answer\":\"The document attributes the trade-off to an inappropriate optimization anchor: optimizing directly for audio quality metrics leads to catastrophic reward hacking that removes pronunciation-critical information needed for intelligibility.\"},{\"question\":\"How does the proposed method improve streaming target speaker extraction performance?\",\"answer\":\"It combines a larger Conformer convolution kernel for richer local spectro-temporal modeling with a WavLM-anchored Direct Preference Optimization (DPO) fine-tuning strategy that uses deep acoustic feature similarity as a robust anchor.\"},{\"question\":\"What improvement does the method achieve under a 560 ms streaming chunk size?\",\"answer\":\"It achieves a 10.9% relative intelligibility improvement, reducing word error rate from 0.138 to 0.123, with marginal simultaneous gains in audio quality and speaker similarity.\"}]",1784211938,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"breaking-the-qualityintelligibility-trade-off-in-streaming-target-speaker-extraction-via-deep-feature-anchored-preference-optimization","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/breaking-the-qualityintelligibility-trade-off-in-streaming-target-speaker-extraction-via-deep-feature-anchored-preference-optimization/86473/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What causes the quality–intelligibility trade-off in streaming TSE according to the document?","Question",{"text":75,"@type":76},"The document attributes the trade-off to an inappropriate optimization anchor: optimizing directly for audio quality metrics leads to catastrophic reward hacking that removes pronunciation-critical information needed for intelligibility.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method improve streaming target speaker extraction performance?",{"text":80,"@type":76},"It combines a larger Conformer convolution kernel for richer local spectro-temporal modeling with a WavLM-anchored Direct Preference Optimization (DPO) fine-tuning strategy that uses deep acoustic feature similarity as a robust anchor.",{"name":82,"@type":73,"acceptedAnswer":83},"What improvement does the method achieve under a 560 ms streaming chunk size?",{"text":84,"@type":76},"It achieves a 10.9% relative intelligibility improvement, reducing word error rate from 0.138 to 0.123, with marginal simultaneous gains in audio quality and speaker similarity.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]