[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83323-en":3,"doc-seo-83323-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83323,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","COALA: Robust Contextualized Speech-Augmented Language Modeling for ASR via Contrastive Regularizer and Biasing Score Estimation","Contextual biasing integrates external knowledge into automatic speech recognition (ASR) to improve domain-specific entity recognition. This paper proposes COALA (COntextualized ASR Leveraging biAsing scoring), a robust framework for enhancing speech-augmented language models (SLMs) in challenging multi-entity settings. Because SLMs have context-window limits, COALA identifies relevant entities from large biasing lists by mapping SLM latent representations into a discriminative space that scores audio–entity match strength. Experiments on LibriSpeech show consistent gains across biasing-list scales using MPD-Loss and DPD-Loss.","COALA: Robust Contextualized Speech-augmented Language Modeling for ASR via Contrastive Regularizer and Biasing Score Estimation  \nJhih-Rong Guo, Bi-Cheng Yan, Tien-Hong Lo, Berlin Chen  \nNational Taiwan Normal University, Taiwan  \n{jhihrong, 80847001s, teinhonglo, [berlin](berlin}@ntnu.edu.tw)[}](berlin}@ntnu.edu.tw)[@ntnu.edu.tw](berlin}@ntnu.edu.tw)  \narXiv :2607 .08 1 17v 1 [ cs .CL] 9 Jul 2026  \nAbstract  \nContextual biasing seeks to integrate external knowledge into automatic speech recognition (ASR) systems to accurately recognize domain-specific entities. In this paper, we propose COALA1 (COntextualized ASR Leveraging biAsing scoring), a robust framework designed to enhance speech-augmented language models (SLMs) in complex multi-entity scenarios. Considering the inherent context-window limitations of SLMs, identifying relevant target entities from a large-scale biasing list is crucial for effective recognition. To this end, COALA maps SLM latent representations into a specialized discriminative space to quantify the matching intensity between audio segments and candidate entities. Furthermore, we address the training collapse in prior study when handling multi-target utterances—where multiple rare words co-occur. Experimental results on the LibriSpeech benchmark demonstrate that COALA consistently achieves superior contextual biasing performance across various biasing list scales.  \nIndex Terms: speech recognition, contextual biasing, biasing scoring, contrastive learning  \n1. Introduction  \nDriven by the monolithic nature and streamlined training process, end-to-end (E2E) automatic speech recognition (ASR) systems [1, 2, 3, 4] have gained widespread attention across both academia and industry. More recently, as witnessed by the remarkable success of large language models (LLMs) in the natural language processing community, efforts in ASR have pivoted toward leveraging LLMs for speech recognition and audio reasoning tasks. Through the synergy of a speech encoder anda language model, speech-augmented language models (SLMs) serve as a promising paradigm for ASR [5, 6, 7, 8, 9] .  \nHowever, SLMs still struggle with uncommon, domainspecific entities such as contact names, proper nouns, and other named entities. Accordingly, contextual biasing techniques have been widely investigated to incorporate external knowledge into ASR systems, steering ASR hypotheses toward the accurate recognition of target entities and exerting a profound impact on modern speech applications.. Existing literature on contextual biasing can be broadly categorized into two active strands of research, namely, inference-time and trainingtime biasing, based on the stage of external knowledge integration. Inference-time biasing approaches, such as shallow fusion [10, 11] and on-the-fly rescoring [12, 13], typically construct ngram finite state transducers (FSTs) from a curated knowledge dataset, which are then utilized to selectively bias the search  \n1The source code and experimental setup for this study are available at [https://github.com/Guo0911/COALA](https://github.com/Guo0911/COALA).  \nFigure 1: Distribution of utterances by number of target entities on the LibriSpeech corpus.  \nspace and dynamically boost the emission probability of target entities during the decoding stage, often triggered by specific prefixes (e.g., ’call’ or ’play’) . On a separate front, trainingtime biasing methods, such as attention-based biasing adapters [14, 15] and trie-based pointer generators [16, 17], which are trained to align acoustic features with contextual information via location-aware attention or neural shortcuts, effectively internalizing the recognition ability to prioritize target entities from custom biasing lists in an E2E manner.  \nDespite the progress in inference-time and training-time biasing techniques, seamlessly integrating these methods into the evolving architectures of SLMs remains a formidable challenge. As the scale of the biasing list expand","cbCais1eCxPs12WS","https://ap.wps.com/l/cbCais1eCxPs12WS","pdf",1596651,3,1,5,"English","en",105,"# Introduction\n# Methods\n## Biasing Scoring\n## Training Objectives\n# Experiments","[{\"question\":\"What problem does COALA address in speech-augmented language models (SLMs)?\",\"answer\":\"COALA targets the difficulty SLMs face with uncommon, domain-specific named entities in multi-entity scenarios, especially under limited context windows and interference from many distractor entities.\"},{\"question\":\"How does COALA choose relevant entities from a large biasing list?\",\"answer\":\"COALA maps SLM latent representations into a specialized discriminative space to quantify the matching intensity between audio segments and candidate entities.\"},{\"question\":\"What training objectives does COALA introduce to handle multi-target utterances?\",\"answer\":\"COALA introduces MPD-Loss (Multi-Positive Discriminative Loss) and DPD-Loss (Decoupled Point-wise Discriminative Loss) to resolve gradient conflicts and support stable training when multiple rare words co-occur.\"}]",1784186732,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"coala-robust-contextualized-speech-augmented-language-modeling-for-asr-via-contrastive-regularizer-and-biasing-score-estimation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/coala-robust-contextualized-speech-augmented-language-modeling-for-asr-via-contrastive-regularizer-and-biasing-score-estimation/83323/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does COALA address in speech-augmented language models (SLMs)?","Question",{"text":75,"@type":76},"COALA targets the difficulty SLMs face with uncommon, domain-specific named entities in multi-entity scenarios, especially under limited context windows and interference from many distractor entities.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does COALA choose relevant entities from a large biasing list?",{"text":80,"@type":76},"COALA maps SLM latent representations into a specialized discriminative space to quantify the matching intensity between audio segments and candidate entities.",{"name":82,"@type":73,"acceptedAnswer":83},"What training objectives does COALA introduce to handle multi-target utterances?",{"text":84,"@type":76},"COALA introduces MPD-Loss (Multi-Positive Discriminative Loss) and DPD-Loss (Decoupled Point-wise Discriminative Loss) to resolve gradient conflicts and support stable training when multiple rare words co-occur.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]