[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84327-en":3,"doc-seo-84327-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":20,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84327,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech","A system for Task 1 of the MLC-SLM 2026 Challenge addresses multilingual two-speaker conversational speech by tightly coupling speaker diarization with an adapted Qwen3-ASR-1.7B recognizer. The diarization front end performs VAD, subsegment generation, CAMPPlus speaker embeddings, two-speaker spectral clustering, and RTTM-based segmentation, then groups segments by language or region for decoding. ASR adaptation uses supervised full fine-tuning, LoRA with three-pipeline TTS synthetic augmentation, and GRPO reinforcement learning with WER/CER rewards and penalties. tcpMER reaches 23.70 on development and 17.97 on evaluation.","Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker  \nConversational Speech  \nHao Wu 1,∗, RongQi Han 1,∗, Zhen Wang 1, Wei Liang2, Wei Xu 1  \n1 Shanghai Qi Zhi Institute, Shanghai, China  \n2Megatronix (Beijing) Technology Co., Ltd  \n[wuhao@sqz.ac.cn](wuhao@sqz.ac.cn)  \narXiv :2607 .08208v 1 [ cs .CL] 9 Jul 2026  \nAbstract  \nThis paper describes our self-designed system for Task 1 of the MLC-SLM 2026 Challenge for multilingual two-speaker conversational speech. The system combines a modular speaker diarization front end with a challenge-adapted Qwen3-ASR-1.7B recognizer. The diarization front end performs voice activity detection, subsegment generation, CAMPPlus speaker embedding extraction, two-speaker spectral clustering, and RTTMbased audio segmentation. The resulting speaker-attributed segments are grouped by language or region and decoded by the adapted ASR model. For ASR adaptation, we ﬁrst perform supervised full ﬁne-tuning on the ofﬁcial training data, then apply LoRA ﬁne-tuning with synthetic speech generated by a threepipeline TTS-based synthetic speech augmentation framework, and ﬁnally reﬁne the model using GRPO reinforcement learning with rewards based on WER/CER and penalties for hallucination, repetition, and length deviation. On the ofﬁcial development set, the full system achieves an average tcpMER of 23.70, reducing the error rate by 6.83 absolute points relative to the released Qwen-ASR-1.7B performance. On the ﬁnal evaluation set, the system achieves an average tcpMER of 17.97 . Ablation results show that supervised ﬁne-tuning provides the largest gain, while synthetic-speech LoRA adaptation and reinforcement learning further improve robustness.  \nKey Words: multilingual speech recognition, speaker diarization, speech language model, LoRA, TTS synthetic speech augmentation  \n1. Introduction  \nTranscribing two-speaker conversational recordings into speaker-attributed text is a demanding multilingual task: a single system must simultaneously decide who spoke when, segment the audio accordingly, and recognize what was said in each language. The difﬁculty is compounded by short turns, overlapping speech, frequent speaker switches, and the heterogeneous orthographies of the 21 language/region conditions covered by the MLC-SLM 2026 Challenge Task 1 . Because diarization errors are inherited by the recognizer and recognition errors are inherited by the ﬁnal score, treating speaker diarization (SD) and automatic speech recognition (ASR) as independent modules leaves substantial accuracy on the table; a carefully coupled SD+ASR pipeline is therefore central to a competitive system.  \nThe challenge evaluates submissions with the timeconstrained permutation mixed error rate (tcpMER), in which a language-adaptive mixed error rate (MER) is ﬁrst computed per language—word error rate for space-separated scripts and character error rate for Japanese, Korean, and Thai—and then  \n∗Hao Wu and RongQi Han contributed equally to this work.  \nminimized over speaker permutations under time constraints, with diarization quality reported separately through the diarization error rate (DER) . The 2026 Task 1 setting is deﬁned by the ofﬁcial challenge page as multilingual conversational speech diarization and recognition [1] .  \nOur SD ASR system, denoted SQZ-Qwen-ASR-1.7B, targets Task 1 only. On the diarization side, we build a modular front-end within the 3D-Speaker framework [2] that performs voice activity detection (VAD), CAMPPlus speaker embedding extraction, and spectral clustering with the speaker count ﬁxed to two, followed by RTTM-driven audio cutting. On the recognition side, we start from the Qwen3-ASR-1.7B speech LLM [3] and adapt it to the MLC-SLM task setting through a threestage recipe: full supervised ﬁne-tuning (SFT) on the ofﬁcial training set, low-rank adaptation (LoRA) [4] on synthetic speech, and Group Relative Policy Optimization (GRPO) reinforcement learning that penalizes hallucinated and repeti","cbCaitKp9C0g9jZb","https://ap.wps.com/l/cbCaitKp9C0g9jZb","pdf",154638,5,1,"English","en",105,"# Introduction\n# System framework and training strategy\n## Stage-wise inference pipeline\n## Speaker diarization","[{\"question\":\"How does the system combine diarization and ASR for better accuracy?\",\"answer\":\"It treats speaker diarization and ASR as a coupled pipeline: diarization outputs speaker-attributed segments that guide how audio is segmented and transcribed by the adapted Qwen-ASR model.\"},{\"question\":\"What are the main steps in the diarization front end?\",\"answer\":\"It performs VAD, generates short subsegments, extracts CAMPPlus speaker embeddings, runs two-speaker spectral clustering with the speaker count fixed to two, and produces RTTM-based audio segmentation.\"},{\"question\":\"How is the ASR model adapted during training?\",\"answer\":\"Training follows three stages: supervised full fine-tuning on the official data, LoRA fine-tuning using synthetic speech from a three-pipeline TTS augmentation framework, and GRPO reinforcement learning using rewards tied to WER/CER and penalties for hallucination, repetition, and length deviation.\"}]",1784194846,13,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"diarization-guided-qwen-asr-adaptation-for-multilingual-two-speaker-conversational-speech","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/diarization-guided-qwen-asr-adaptation-for-multilingual-two-speaker-conversational-speech/84327/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How does the system combine diarization and ASR for better accuracy?","Question",{"text":75,"@type":76},"It treats speaker diarization and ASR as a coupled pipeline: diarization outputs speaker-attributed segments that guide how audio is segmented and transcribed by the adapted Qwen-ASR model.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the main steps in the diarization front end?",{"text":80,"@type":76},"It performs VAD, generates short subsegments, extracts CAMPPlus speaker embeddings, runs two-speaker spectral clustering with the speaker count fixed to two, and produces RTTM-based audio segmentation.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the ASR model adapted during training?",{"text":84,"@type":76},"Training follows three stages: supervised full fine-tuning on the official data, LoRA fine-tuning using synthetic speech from a three-pipeline TTS augmentation framework, and GRPO reinforcement learning using rewards tied to WER/CER and penalties for hallucination, repetition, and length deviation.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]