[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85490-en":3,"doc-seo-85490-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85490,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","TagSpeech End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding","TagSpeech introduces a unified LLM-based framework that performs multi-speaker ASR and diarization with fine-grained temporal grounding. The method combines decoupled semantic and speaker streams trained via Serialized Output Training to learn turn-taking dynamics, with an interleaved time anchor mechanism that provides precise timestamp prediction and synchronizes semantic understanding and speaker tracking. Compared with speaker-attributed ASR or implicit diarization, TagSpeech explicitly models who spoke what and when end-to-end, improving DER on AMI and AliMeeting, including under overlapping speech, while using parameter-efficient training by freezing the LLM backbone.","TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding  \nMingyue Huo 1 and Yiwen Shao2 and Yuheng Zhang 1  \n1University of Illinois Urbana-Champaign 2Johns Hopkins University  \n[mhuo5@illinois.edu](mhuo5@illinois.edu) , [yshao18@jhu.edu](yshao18@jhu.edu) , [yuhengz2@illinois.edu](yuhengz2@illinois.edu)  \narXiv :2601 .06896v2 [ ee ss .AS] 13 Jul 2026  \nAbstract  \nWe present TagSpeech, a unified LLM-based framework that utilizes Temporal Anchor Grounding for joint multi-speaker ASR and diarization. The framework is built on two key designs: (1) decoupled semantic and speaker streams fine-tuned via Serialized Output Training (SOT) to learn turn-taking dynamics; and  \n(2) an interleaved time anchor mechanism that not only supports fine-grained timestamp prediction but also acts as a synchronization signal between semantic understanding and speaker tracking. Compared to previous works that primarily focus on speaker-attributed ASR or implicit diarization, TagSpeech addresses the challenge of fine-grained speaker–content alignment and explicitly models who spoke what and when in an end-to-end manner. Experiments on AMI and AliMeeting benchmarks demonstrate that our method achieves consistent improvements in Diarization Error Rate (DER) over strong end-to-end baselines, including Qwen-Omni and Gemini, particularly in handling complex speech overlaps. Moreover, TagSpeech employs a parameter-efficient training paradigm in which the LLM backbone is frozen and only lightweight projectors are trained, resulting in strong performance with low computational cost 1.  \n1 Introduction  \nHuman communication frequently involves multiple speakers, overlapping speech, and rapid turntaking. Consequently, applications like meeting transcription require not only recognizing what was said, but also identifying \"who spoke what and when. \" This motivates a unified framework  \nfor joint automatic speech recognition (ASR) and speaker diarization with high temporal precision, as illustrated in Figure 1.  \n1Code, model and demo are available at [https://github](https://github). com/AudenAI/Auden/tree/main/examples/tagspeech  \nScenario: Multi-Speaker Meeting  \n* illustration generated by Google Gemini  \nFigure 1: Comparison between a conventional cascaded pipeline (top) and our end-to-end framework TagSpeech (bottom) for multi-speaker speech processing.  \nTask Definition and Disambiguation. With the emergence of Large-Audio Language Models (LALMs), LLM-based architectures have become a natural choice for unified multi-speaker modeling, enabling contextual reasoning in challenging conditions and speaker-attributed generation within a single decoding process.  \nHowever, a critical ambiguity persists in recent literature regarding what constitutes “joint ASRand diarization.” As summarized in Table 1, several methods, including SpeakerLM (Yin et al., 2025) and JEDIS-LLM (Shi et al., 2025), term their task\"diarization\" yet primarily address who said what, omitting explicit start and end timestamps. Although effective for speaker-attributed ASR, such  \n\n| Recent Work on Multi-speaker ASR & Diarization | Use LLM | Transcription | Speaker | Timestamp |\n| --- | --- | --- | --- | --- |\n| Pyannote (Plaquet and Bredin, 2023), Cheng et al. (2025) | No | ✗ | ✓ | ✓ |\n| Sortformer(Park et al., 2025); Meta-CAT (Wang et al., 2025b); DNCASR (Zheng et al., 2025) | No | ✓ | ✓ | ✗ |\n| SpeakerLM (Yin et al., 2025); JEDIS-LLM (Shi et al., 2025) | Yes | ✓ | ✓ | ✗ |\n| MT-LLM (Meng et al., 2025) | Yes | ✓ | ✗ | ✗ |\n| DiarizationLM (Wang et al., 2024) * post-processing only | Yes | ✗ | ✗ | ✗ |\n| \u003Cbr>TagSpeech (Ours) | \u003Cbr>Yes | \u003Cbr>✓ | \u003Cbr>✓ | \u003Cbr>✓ |\n\nTable 1: Comparison of recent works by explicit outputs. While several approaches are described as joint ASR and diarization systems, diarization is often realized as speaker-attributed transcription without explicit timestamps. Our work addresses all three aspects explicitly.  \nomission prevents eva","cbCaiv3aONKcNbbu","https://ap.wps.com/l/cbCaiv3aONKcNbbu","pdf",837076,1,16,"English","en",105,"# Abstract\n# Introduction\n## Task Definition and Disambiguation\n## Challenge\n## Contributions","[{\"question\":\"How does TagSpeech compare to prior joint ASR/diarization approaches?\",\"answer\":\"Earlier methods often omit explicit start/end timestamps, whereas TagSpeech explicitly predicts transcription (what), speaker labels (who), and precise timestamps (when) end-to-end, leading to DER improvements especially in overlapping speech.\"}]",1784203986,40,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"tagspeech-end-to-end-multi-speaker-asr-and-diarization-with-fine-grained-temporal-grounding","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/tagspeech-end-to-end-multi-speaker-asr-and-diarization-with-fine-grained-temporal-grounding/85490/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How does TagSpeech compare to prior joint ASR/diarization approaches?","Question",{"text":75,"@type":76},"Earlier methods often omit explicit start/end timestamps, whereas TagSpeech explicitly predicts transcription (what), speaker labels (who), and precise timestamps (when) end-to-end, leading to DER improvements especially in overlapping speech.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,111,114,119,122,126],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":45,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":45,"category_name":109,"show_sort_weight":28,"slug":110},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":45,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]