[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82917-en":3,"doc-seo-82917-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82917,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Progressive Refinement An Iterative Pseudo Labeling Approach for Mandarin English Code Switching ASR","An iterative pseudo-labeling training approach improves Mandarin-English code-switching ASR by exploiting unlabeled speech to reduce dependence on limited CS-labeled data. The method proceeds through three phases: pseudo-label generation from a large unlabeled corpus, two-stage bilingual model training combining pretraining with pseudolabeled data and fine-tuning on supervised CS plus monolingual utterances, and iterative refinements that improve handling of dynamic, spontaneous language alternations. Experiments report notable Mix Error Rate reductions on SEAME devman (6.35%) and devsge (8.29%).","PROGRESSIVE REFINEMENT: AN ITERATIVE PSEUDO-LABELING APPROACH FOR  \nMANDARIN-ENGLISH CODE-SWITCHING ASR  \nQu Yang 1 ,2 ,∗ , Cakra Wardhana 1, Tim Ng 1  \n1 Apple, Singapore 2 National University of Singapore, Singapore  \narXiv :2607 .05224v 1 [ cs .CL] 6 Jul 2026  \nABSTRACT  \nCode-switching (CS), alternating languages within the same utterance, poses significant challenges for automatic speech recognition (ASR) due to limited CS training data. This paper applies an iterative pseudo-labeling training approach to CS-ASR for the first time, demonstrating its effectiveness in leveraging unlabeled data to improve CS-ASR performance. The approach comprises three phases: pseudo-label generation, two-stage bilingual model training, and iterative improvements. It begins by generating pseudo-labels from a large unlabeled corpus, creating a semi-supervised dataset. This dataset supports a two-stage training framework where the model is pre-trained and then fine-tuned on supervised CS data. Iterative refinements further enhance the model’s accuracy in handling complex CS scenarios. Our approach significantly advances CS-ASR systems, achieving notable Mix Error Rate (MER) reductions on SEAME’s devman (6.35%) and devsge (8.29%) subsets.  \nIndex Terms— Speech recognition, Code-switching, Pseudolabeling, Semi-supervised learning  \n1. INTRODUCTION  \nCode-switching (CS) is common in multilingual societies like Southeast Asia [1] . This linguistic phenomenon poses unique challenges for automatic speech recognition (ASR) systems [2, 3, 4], as it introduces complex variations and unpredictable switches between languages that conventional ASR models struggle to handle. Improving ASR for code-switching speech is therefore crucial to enhancing user experience in multilingual environments, particularly for voice assistant technologies [5] .  \nCode-switching ASR has gained attention recently, with researchers exploring various approaches to address its complexities. Studies have focused on language-specific architectures, such as joint CTC-attention with language identification (LID) [6, 7, 8], and transformer-based multi-encoder-decoder frameworks [9] . Selfsupervised pre-training with multilingual data has also shown promise for improving Mandarin-English CS-ASR [10, 11] . The dual-encoder transformer network [12], which employs encoders pre-trained in Mandarin and English to extract language-specific features, has inspired further developments in this framework [13, 14, 15] . However, the scarcity of labeled code-switching data remains a major obstacle to advancing ASR performance for mixed-language speech.  \nIn this work, we apply an iterative pseudo-labeling training approach to English-Mandarin code-switching ASR for the first time, demonstrating its effectiveness in improving CS-ASR performance. While pseudo-labeling has been widely explored in ASR research [16, 17, 18], our approach differs in key aspects. First, those existing pseudo-labeling methods focus on improving monolingual or  \n∗ Work done during an internship at Apple. [quyang@u.nus.edu](quyang@u.nus.edu)  \ncode-mixing [19] ASR where lexical items from different languages appear in the same utterance but without frequent grammatical switching. In contrast, our method specifically addresses codeswitching speech which involves more dynamic and spontaneous alternation between languages. Second, our approach leverages unlabeled data that we assume to contain code-switching scenarios, whereas prior methods typically rely on monolingual or code-mixing datasets that lack such code-switching contexts. Finally, our initial ASR model (M0) combines Mandarin and English monolingual data during training, providing it with a robust understanding of both languages from the outset. This bilingual initialization sets our method apart from previous pseudo-labeling works.  \nOur proposed method consists of three phases: pseudo-label generation, two-stage bilingual model training (comprising pretraining an","cbCaihVgkEfjvWVc","https://ap.wps.com/l/cbCaihVgkEfjvWVc","pdf",867358,3,1,5,"English","en",105,"# Abstract\n# Introduction\n# Iterative Pseudo-Labeling Training\n## CTC+Attention Model","[{\"question\":\"Why does code-switching make ASR difficult for Mandarin-English speech?\",\"answer\":\"Code-switching introduces unpredictable and complex alternations between languages within the same utterance, creating variations that conventional ASR models handle poorly when CS training data is limited.\"},{\"question\":\"What are the three phases of the proposed iterative pseudo-labeling approach?\",\"answer\":\"The pipeline includes pseudo-label generation from a large unlabeled corpus, two-stage bilingual model training (pretraining then fine-tuning), and iterative improvements that refine the model using newly generated pseudo-labels.\"},{\"question\":\"How is the model trained during two-stage bilingual training?\",\"answer\":\"It is first pre-trained on pseudolabeled data to learn general patterns, then fine-tuned on a smaller supervised dataset containing both monolingual and code-switching utterances to improve CS performance.\"}]",1784183940,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"progressive-refinement-an-iterative-pseudo-labeling-approach-for-mandarin-english-code-switching-asr","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/progressive-refinement-an-iterative-pseudo-labeling-approach-for-mandarin-english-code-switching-asr/82917/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does code-switching make ASR difficult for Mandarin-English speech?","Question",{"text":75,"@type":76},"Code-switching introduces unpredictable and complex alternations between languages within the same utterance, creating variations that conventional ASR models handle poorly when CS training data is limited.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What are the three phases of the proposed iterative pseudo-labeling approach?",{"text":80,"@type":76},"The pipeline includes pseudo-label generation from a large unlabeled corpus, two-stage bilingual model training (pretraining then fine-tuning), and iterative improvements that refine the model using newly generated pseudo-labels.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the model trained during two-stage bilingual training?",{"text":84,"@type":76},"It is first pre-trained on pseudolabeled data to learn general patterns, then fine-tuned on a smaller supervised dataset containing both monolingual and code-switching utterances to improve CS performance.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]