[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84775-en":3,"doc-seo-84775-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84775,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Towards Robust Uncertainty-Aware Speaker Modeling","Speaker embeddings compress frame-level acoustic evidence for speaker recognition, but uncertainty-aware methods often misestimate uncertainty and become miscalibrated when speech conditions change across domains. This work proposes a robust uncertainty modeling framework that improves both estimation and adaptation. An Inter- and Intra-Speaker-Aware Uncertainty Softmax jointly learns uncertainty using inter-speaker separability and intra-speaker variability. A Uncertainty-Calibrated Domain Adaptation (UCDA) method performs lightweight label-free calibration. Experiments on in-domain and cross-domain benchmarks show consistent gains in uncertainty reliability and speaker recognition robustness.","Towards Robust Uncertainty-Aware Speaker  \nModeling  \nJunjie Li 1 , Yang Xiao2 , Kong Aik Lee 1  \n1Department of Electrical and Electronic Engineering, The Hong Kong Polytechnic University, Hong Kong SAR  \n2 The University of Melbourne, Australia  \narXiv :2607 .04937v 1 [ cs . SD] 6 Jul 2026  \nAbstract—Speaker embeddings aggregate frame-level acoustic features into compact representations for speaker recognition. Recent uncertainty-aware speaker modeling approaches further characterize the reliability of speaker embeddings by estimating their associated uncertainty. However, existing methods often suffer from inaccurate uncertainty estimation and uncertainty miscalibration under domain shifts. To address these challenges, we propose a robust uncertainty modeling framework from both estimation and adaptation perspectives. Specifically, we introduce an Inter- and Intra-Speaker-Aware Uncertainty Softmax that incorporates both inter-speaker separability and intra-speaker variability into uncertainty learning, enabling uncertainty estimates to better capture the reliability of speaker embeddings. Furthermore, we propose an Uncertainty-Calibrated Domain Adaptation (UCDA) framework to mitigate uncertainty miscalibration caused by domain mismatch. Extensive experiments on both in-domain and cross-domain benchmarks demonstrate that the proposed approach consistently improves uncertainty reliability and speaker recognition robustness.  \nIndex Terms—speaker verification, cross-domain, uncertainty estimation  \nI. INTRODUCTION  \nSpeaker recognition is widely used in biometric authentication [1], personalized human–machine interaction [2], and intelligent surveillance [3] . Modern systems typically extract fixed-dimensional speaker embeddings from variable-length utterances for similarity-based scoring [4] . However, in realworld scenarios, speech is frequently corrupted by background noise, reverberation, and channel mismatch [4], introducing severe frame-level uncertainty. Conventional pooling methods, such as average pooling [5]–[7] and attention-based pooling [8]–[13], rely on deterministic weighting strategies and fail to account for frame reliability, resulting in degraded embedding quality under unconstrained conditions.  \nTo mitigate this, uncertainty modeling represents embeddings as Gaussian distributions [14]–[27], where the mean encodes speaker identity and the covariance captures estimation uncertainty to down-weight unreliable frames during pooling and scoring. Recently, the uncertainty-aware additive angular margin softmax (UAAM-Softmax) loss [22] was introduced to jointly optimize speaker discrimination and uncertainty estimation. Although effective, UAAM-Softmax relies primarily on inter-speaker separability as a supervisory signal, while the intrinsic intra-speaker variability of speaker embeddings isnot explicitly considered. As a result, the learned uncertainty estimates may not fully reflect the underlying variability and reliability of speaker embeddings. Moreover, uncertainty  \nestimation is highly sensitive to domain mismatch. Acoustic and environmental variations across datasets induce substantial distribution shifts [4], [22], which can lead to uncertainty miscalibration, where the estimated uncertainty is no longer well aligned with the actual reliability of speaker embeddings, thereby degrading cross-domain performance [28] .  \nTo address the above limitations, we propose a unified framework for robust uncertainty modeling in speaker recognition. First, we propose an Inter- and Intra-SpeakerAware Uncertainty Softmax that incorporates both interspeaker relationships and intra-speaker variability into uncertainty learning. By exploiting complementary supervisory signals, the proposed objective enables uncertainty estimates to better reflect the underlying variability and reliability of speaker embeddings. Second, we introduce an UncertaintyCalibrated Domain Adaptation (UCDA) framework that improves the robustness o","cbCaioqWeRUG64cf","https://ap.wps.com/l/cbCaioqWeRUG64cf","pdf",4428006,3,1,"English","en",105,"# I. Introduction\n## Speaker recognition use cases and embedding uncertainty\n## Gaussian uncertainty modeling and limitations of existing losses\n## Proposed robust framework: ISI-Aware Softmax and UCDA\n# II. Background: Uncertainty-Aware Model\n## Linear-Gaussian formulation for posterior aggregation\n## Mean/variance propagation through BN and FC layers\n# III. Uncertainty-Aware Softmax","[{\"question\":\"Why do uncertainty-aware speaker modeling methods struggle under domain shift?\",\"answer\":\"Acoustic and environmental differences across datasets create distribution shifts, causing uncertainty miscalibration where estimated uncertainty no longer matches embedding reliability. This degrades cross-domain speaker recognition.\"},{\"question\":\"What is the Inter- and Intra-Speaker-Aware Uncertainty Softmax?\",\"answer\":\"It is a Softmax objective that incorporates both inter-speaker separability and intra-speaker variability into uncertainty learning. This helps uncertainty estimates better reflect the reliability and variability of speaker embeddings.\"},{\"question\":\"How does UCDA improve robustness of uncertainty estimation?\",\"answer\":\"UCDA performs lightweight, label-free adaptation by updating only the uncertainty estimation module. It encourages target-domain uncertainty distributions to move toward a source-domain prior to improve calibration while preserving discriminative information.\"}]",1784198162,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"towards-robust-uncertainty-aware-speaker-modeling","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":20},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/towards-robust-uncertainty-aware-speaker-modeling/84775/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"Why do uncertainty-aware speaker modeling methods struggle under domain shift?","Question",{"text":74,"@type":75},"Acoustic and environmental differences across datasets create distribution shifts, causing uncertainty miscalibration where estimated uncertainty no longer matches embedding reliability. This degrades cross-domain speaker recognition.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"What is the Inter- and Intra-Speaker-Aware Uncertainty Softmax?",{"text":79,"@type":75},"It is a Softmax objective that incorporates both inter-speaker separability and intra-speaker variability into uncertainty learning. This helps uncertainty estimates better reflect the reliability and variability of speaker embeddings.",{"name":81,"@type":72,"acceptedAnswer":82},"How does UCDA improve robustness of uncertainty estimation?",{"text":83,"@type":75},"UCDA performs lightweight, label-free adaptation by updating only the uncertainty estimation module. It encourages target-domain uncertainty distributions to move toward a source-domain prior to improve calibration while preserving discriminative information.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]