[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84130-en":3,"doc-seo-84130-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84130,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Hierarchical Acoustic Semantic Modeling Modality Separation and Semantic Coherence for Full Duplex SLMs","Developing seamless, high-performance, native intelligent full-duplex Spoken Language Models (SLMs) remains a critical challenge, because modality interference degrades knowledge and damages semantic integrity, making interactions feel unnatural. Fine-grained analysis of optimization dynamics shows the root cause: inherent gradient conflicts between acoustic and semantic modeling when both modalities share a deep parameter space. Lychee-FD mitigates this via hierarchical parameter separation and a semantic alignment channel, improving speech intelligence and full-duplex fluency without sacrificing inference efficiency.","arXiv :2607 .06540v2 [ cs .CL] 9 Jul 2026  \n Lychee  \nHierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs  \nZhenyu Liu1,2 , Xuanyu Zhang1,2 , Yunxin Li1,2 , Qixun Teng1 , Shenyuan Jiang1 , Haolan Chen, Minjun Zhao, Fanbo Meng, Yu Xu, Yancheng He, Baotian Hu1,2, 􀀀 Haizhou Li2,3 , Min Zhang1,2  \n1 School of Computer Science and Technology, Harbin Institute of Technology, Shenzhen, China  \n2 Center for Language, Intelligence and Machines, Shenzhen Loop Area Institute, Shenzhen, China  \n3 School of Artificial Intelligence, The Chinese University of Hong Kong, Shenzhen, China  \n [https://huggingface.co/HIT-TMG/Lychee-FD](https://huggingface.co/HIT-TMG/Lychee-FD)  \n [https://github.com/HITsz-TMG/Lychee-FD](https://github.com/HITsz-TMG/Lychee-FD)  \n [https://hitsz-tmg.github.io/Lychee-FD](https://hitsz-tmg.github.io/Lychee-FD)  \nCorrespondence: [liuzhenyuhit@gmail.com](liuzhenyuhit@gmail.com), [25s151197@stu.hit.edu.cn](25s151197@stu.hit.edu.cn), [liyx@hit.edu.cn](liyx@hit.edu.cn)  \n[hubaotian@hit.edu.cn](hubaotian@hit.edu.cn), [zhangmin2021@hit.edu.cn](zhangmin2021@hit.edu.cn), [haizhouli@cuhk.edu.cn](haizhouli@cuhk.edu.cn)  \nAbstract  \nDeveloping seamless, high-performance, native intelligent full-duplex Spoken Language Models (SLMs) remains a critical challenge and long-standing goal for the speech and NLP community. Despite notable progress, recent endeavors are fundamentally constrained by severe modality interference, which causes substantial knowledge degradation and compromises semantic integrity—ultimately making full-duplex SLMs feel unnatural and unintelligent. In this paper, through an exhaustive fine-grained analysis of model optimization dynamics, we uncover the root cause of such performance degradation, revealing that modality interference arises from inherent gradient conflicts between acoustic and semantic modeling when the two modalities are forced to share a deep parameter space.  \nGuided by this key insight, we introduce Lychee-FD, a native end-to-end full-duplex framework designed to mitigate modality interference. Importantly, we propose a hierarchical parameter separation strategy that decouples conflicting modalities in deep layers while preserving cross-modality coherence via a dedicated semantic alignment channel. Extensive experiments on multiple full-duplex benchmarks demonstrate that our method significantly advances the state of the art, yielding substantial improvements in both speech intelligence (+7.4% on Spoken QA) and full-duplex interaction fluidity (+28.5% on FullDuplexBench 1.5) without compromising inference efficiency. To the best of our knowledge, this work is the first to achieve two key advances: 1) uncovering and elucidating the root cause of modality interference in full-duplex SLMs, and 2) designing an elegant hierarchical model together with a practical solution for seamless, high-performance, native intelligent full-duplex SLMs.  \n1 Introduction  \nThe rapid evolution of Large Language Models (LLMs) has fundamentally reshaped our daily lives, establishing them as ubiquitous assistants capable of complex reasoning and instruction following. Within this landscape, Spoken Language Models (SLMs) represent a significant paradigm shift from text-based to voice-based interaction. Despite recent advancements in Omni-modal models capable of seamless voice interaction (Hurst et al., 2024 ; Zhan et al., 2024 ; Li et al., 2025b ; Kaplan et al., 2020 ; Hoffmann et al., 2022 ; Xu et al., 2023 ; Gao et al., 2025), a critical disparity remains between artificial agents and authentic human conversation.  \n∗􀀀 Corresponding author.  \nCurrently, most voice interactions are constrained to a rigid half-duplex mode, where the system strictly alternates between listening and speaking in a sequential manner. In contrast, authentic human conversation is inherently full-duplex, requiring the ability to continuously process incoming audio streams while concurren","cbCaineb0yP1Hy46","https://ap.wps.com/l/cbCaineb0yP1Hy46","pdf",1899424,1,22,"English","en",105,"# Introduction\n## Full-duplex SLMs and the half-duplex gap\n## Modality interference and optimization bottlenecks","[{\"question\":\"What problem prevents native full-duplex Spoken Language Models from feeling natural and intelligent?\",\"answer\":\"Severe modality interference between acoustic and semantic modeling degrades knowledge and compromises semantic integrity, leading to unnatural and unintelligent behavior.\"},{\"question\":\"What is the root cause of modality interference in full-duplex SLMs?\",\"answer\":\"It stems from inherent gradient conflicts between the acoustic and semantic modeling objectives when the two modalities are forced to share a deep parameter space.\"},{\"question\":\"How does Lychee-FD address modality interference to improve performance?\",\"answer\":\"Lychee-FD uses hierarchical parameter separation to decouple conflicting modalities in deep layers, while a dedicated semantic alignment channel preserves cross-modality coherence.\"}]",1784193187,55,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"hierarchical-acoustic-semantic-modeling-modality-separation-and-semantic-coherence-for-full-duplex-slms","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/hierarchical-acoustic-semantic-modeling-modality-separation-and-semantic-coherence-for-full-duplex-slms/84130/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":11},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem prevents native full-duplex Spoken Language Models from feeling natural and intelligent?","Question",{"text":75,"@type":76},"Severe modality interference between acoustic and semantic modeling degrades knowledge and compromises semantic integrity, leading to unnatural and unintelligent behavior.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the root cause of modality interference in full-duplex SLMs?",{"text":80,"@type":76},"It stems from inherent gradient conflicts between the acoustic and semantic modeling objectives when the two modalities are forced to share a deep parameter space.",{"name":82,"@type":73,"acceptedAnswer":83},"How does Lychee-FD address modality interference to improve performance?",{"text":84,"@type":76},"Lychee-FD uses hierarchical parameter separation to decouple conflicting modalities in deep layers, while a dedicated semantic alignment channel preserves cross-modality coherence.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]