[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82142-en":3,"doc-seo-82142-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82142,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition","Large language model (LLM)-based audio-visual speech recognition (LLM-AVSR) uses complementary audio and visual cues to improve robustness in adverse acoustics. Existing methods often project independently pretrained audio/visual encoders into an LLM space and fuse them as soft prompts, but they rarely correct cross-modal representational discrepancy. This work introduces an optimal transport (OT) semantic alignment framework that couples modality features to LLM linguistic embeddings, yielding probabilistic soft pseudo-labels for contrastive learning. Experiments on LRS3-TED show consistent gains and state-of-the-art results in clean and noisy conditions.","Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition*  \nXugang Lu 1 , Peng Shen 1 , Yu Tsao2 , Hisashi Kawai 1  \n1. National Institute of Information and Communications Technology, Kyoto, Japan  \n2. Research Center for Information Technology Innovation, Academia Sinica, Taiwan  \narXiv :2607 .0900 1v 1 [ cs . SD] 10 Jul 2026  \nAbstract—Large language model (LLM)-based audio-visual speech recognition (LLM-AVSR) has recently demonstrated strong robustness in adverse acoustic environments by leveraging complementary audio and visual information. Existing approaches typically employ independently pretrained acoustic and visual encoders, whose outputs are projected and fused as soft prompts to condition an LLM for speech recognition. However, most methods perform multimodal fusion without explicitly addressing the representational discrepancy between audio, visual and text modalities, potentially limiting the effectiveness of crossmodal integration. In this paper, we propose an optimal transport (OT)-based semantic alignment framework for LLM-AVSR. The proposed method explicitly bridges the modality gap by aligning the acoustic and visual representations with reference to the linguistic embedding space of the LLM before multimodal fusion. Specifically, OT is used to estimate probabilistic coupling matrices that characterize structured correspondences between modalityspecific features and linguistic embeddings. The resulting OT couplings are further utilized as soft pseudo-labels to supervise contrastive learning, encouraging the extraction of semantically coherent and cross-modal consistent audio-visual representations. By anchoring multimodal features to the linguistic space of the LLM, the proposed framework facilitates more effective multimodal fusion and decoding. We implement the proposed framework using a Whisper-based acoustic encoder, an AVHuBERT-based visual encoder, and a LLaMA3.2-3B decoder. Experiments conducted on the LRS3-TED benchmark demonstrate consistent improvements over strong baselines and achieve stateof-the-art performance under both clean and noisy evaluation conditions across a wide range of signal-to-noise ratios (SNRs).  \nIndex Terms—Audio-visual ASR, optimal transport, feature fusion.  \nI. INTRODUCTION  \nAudio-visual speech recognition (AVSR) has attracted increasing attention due to its robustness in adverse acoustic environments. By leveraging complementary visual cues, such as modality-invariant features [1], modality temporal dynamics [2], and lip movements [3], AVSR systems can substantially alleviate performance degradation caused by background noise [3]–[6] . Recent advances in AVSR have shown that the effective integration of audio and visual information within Conformer-based encoder-decoder architectures, often combined with hybrid CTC/Attention objectives [7], can significantly improve recognition performance [3], [4] .  \nWith the rapid development of large language models (LLMs), recent studies have extended AVSR to LLM-based speech recognition frameworks (LLM-AVSR) [5], [8]–[10] . These systems typically employ pretrained modality-specific encoders to extract acoustic and visual representations, which  \nare subsequently projected into the embedding space of an LLM and used as soft prompts for autoregressive transcription generation. Benefiting from the strong contextual modeling and linguistic reasoning capabilities of LLMs, such approaches have demonstrated improved robustness and generalization.  \nHowever, visual speech recognition (VSR) inherently suffers from ambiguity because multiple phonemes may correspond to identical or highly similar lip shapes [6], [11] . This ambiguity becomes even more severe in continuous speech due to co-articulation effects, where the same viseme may correspond to different phonetic realizations depending on the surrounding context [6], [11] . Although acoustic information can partially compensate for this limitation, speech","cbCaiirF1g4LuGSZ","https://ap.wps.com/l/cbCaiirF1g4LuGSZ","pdf",403480,1,7,"English","en",105,"# Introduction\n## Background: AVSR and LLM-AVSR\n## Challenges: visual ambiguity and representational mismatch\n## Motivation: semantic alignment beyond temporal synchronization","[{\"question\":\"What limitation do most existing LLM-based audio-visual speech recognition methods have?\",\"answer\":\"They typically fuse projected audio and visual representations without explicitly addressing representational discrepancy between the modalities, which can weaken cross-modal integration and semantic grounding.\"},{\"question\":\"How does the proposed optimal transport (OT) approach improve semantic alignment?\",\"answer\":\"OT estimates probabilistic coupling matrices that relate acoustic/visual features to reference linguistic embeddings in the LLM space, producing soft pseudo-labels that supervise contrastive learning for cross-modal consistency.\"},{\"question\":\"What model components and evaluation results are reported?\",\"answer\":\"The framework uses a Whisper-based acoustic encoder, an AVHuBERT-based visual encoder, and a LLaMA3.2-3B decoder, and it achieves consistent improvements and state-of-the-art performance on the LRS3-TED benchmark under both clean and noisy conditions across a range of SNRs.\"}]",1784178423,18,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"optimal-transport-based-semantic-alignment-for-llm-based-audio-visual-speech-recognition","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/optimal-transport-based-semantic-alignment-for-llm-based-audio-visual-speech-recognition/82142/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What limitation do most existing LLM-based audio-visual speech recognition methods have?","Question",{"text":75,"@type":76},"They typically fuse projected audio and visual representations without explicitly addressing representational discrepancy between the modalities, which can weaken cross-modal integration and semantic grounding.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed optimal transport (OT) approach improve semantic alignment?",{"text":80,"@type":76},"OT estimates probabilistic coupling matrices that relate acoustic/visual features to reference linguistic embeddings in the LLM space, producing soft pseudo-labels that supervise contrastive learning for cross-modal consistency.",{"name":82,"@type":73,"acceptedAnswer":83},"What model components and evaluation results are reported?",{"text":84,"@type":76},"The framework uses a Whisper-based acoustic encoder, an AVHuBERT-based visual encoder, and a LLaMA3.2-3B decoder, and it achieves consistent improvements and state-of-the-art performance on the LRS3-TED benchmark under both clean and noisy conditions across a range of SNRs.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]