[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84666-en":3,"doc-seo-84666-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84666,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Reinforcement Learning for Data-Efficient Code-Switched ASR","Audio-language models can be prompted for code-switched speech, yet their decoding is not optimized for language switching and often fails at code-switch boundaries. The work introduces Reinforcement Learning with Verifiable Rewards (RLVR) to adapt audio-language models to code-switched ASR using group relative policy optimization. Rewards combine an error-rate objective and a script fidelity penalty, plus a two-pass draft-and-refinement decoding procedure. Experiments on Qwen2-Audio across 10 language pairs show that RLVR with 10% data matches full-data LoRA supervised fine-tuning, with gains largest for typologically distant pairs and zero-shot transfer to human-recorded corpora.","Reinforcement Learning for Data-Efficient Code-Switched ASR  \nZiwei Ye  1 ,∗∗, Peter Vickers  2  \n1 Independent Researcher  \n2 Spotify Canada  \n[zxy1677@rit.edu](zxy1677@rit.edu) , [pvickers@spotify.com](pvickers@spotify.com)  \narXiv :2607 .02757v 1 [ cs .CL] 2 Jul 2026  \nAbstract  \nAudio-language models can be prompted for code-switched speech, but their decoding is not optimized for code-switching and often fails at language boundaries. We propose a practical reinforcement learning with verifiable rewards recipe for dataefficient adaptation of audio-language models to code-switched ASR using group relative policy optimization, combining an error rate reward with a script fidelity reward that penalizes wrong writing systems and a two-pass draft-and-refinement procedure. Using Qwen2-Audio as a reproducible testbed across 10 language pairs, training on only TTS code-switched speech, we show that RLVR with 10% of the data matches LoRA supervised fine-tuning trained on the full dataset, with the largest gains on typologically distant pairs. The error rate reward eliminates translation errors while the script fidelity reward separately reduces script contamination without degradation. These gains transfer zero-shot to a human-recorded code-switching corpus.  \nIndex Terms: reinforcement learning, code-switching, automatic speech recognition, reward design  \n1. Introduction  \nCode-switching is a natural communicative behavior for the majority of the world’s multilingual population [1, 2] . Despite its prevalence, code-switched automatic speech recognition (CS-ASR) remains challenging due to the need to model rapidly shifting acoustic and linguistic cues, compounded by the scarcity of labeled code-switched training data [3, 4] . Since large-scale speech models are typically pre-trained on monolingual corpora [5, 6, 7], data-efficient adaptation methods are needed to endow speech-augmented large language models (speech-LLMs) with robustness to code-switching.  \nInstruction-following speech-LLMs can ingest language context via prompting for flexible output control, yet their autoregressive decoding is trained with token-level cross-entropy, which does not directly optimize sequence-level metrics such as CER and is subject to exposure bias [8] . These issues can be amplified at code-switch boundaries, where language confusion and script hallucination have been observed [9] .  \nWe frame CS-ASR as a verifiable reward problem: given reference transcripts, we can compute automatic rewards and use RL to directly optimize sequence-level transcription quality. Rather than targeting state-of-the-art performance, we use an established, publicly available speech-LLM (Qwen2-Audio) as a controlled testbed to study how reward design and data efficiency interact for code-switched adaptation. In this paper, we make the following contributions:  \n**indicates the corresponding author.  \n\n| Reference (cmn-eng, SwitchLingua)\u003Cbr>这种科学发现真是太fascinating了 |  |  |\n| --- | --- | --- |\n| Format-only baseline | CER 1.81 |  |\n|  This kind of scientific discovery is really fascinating |  |  |\n|  Chinese translated to English 太...fascinating...了frame code-switching lost  |  |  |\n| LoRA SFT (100% data)\u003Cbr> This scientific discovery is so fascinating. |  | CER 1.33 |\n| RLVR (10% data) CER 0 .00 这种科学发现真是太fascinating了\u003Cbr>Chinese preserved English insertion intact Code-switch boundary correct |  |  |\n\nFigure 1: Chinese-English code-switch sample in SwitchLingua  \n• We apply Reinforcement Learning with Verifiable Rewards (RLVR) [10] on code-switched ASR, using rewards based on CER and a script fidelity reward to directly encourage correct writing systems at code-switch boundaries.  \n• We introduce a training-time two-pass draft and refinement procedure that conditions a second decoding pass on the best draft to encourage “listen again and fix” behavior.  \n• We evaluate on the human-read subset of CS-FLEURS [11] across 10 language pairs, showing data-efficient gains over LoR","cbCaibc0EyQaz2jK","https://ap.wps.com/l/cbCaibc0EyQaz2jK","pdf",352885,2,1,5,"English","en",105,"# Introduction\n# Related work\n# Method\n## Problem Setup","[{\"question\":\"Why is code-switched ASR challenging compared with monolingual ASR?\",\"answer\":\"Code-switching requires modeling rapidly shifting acoustic and linguistic cues, and labeled code-switched training data is scarce. Boundary regions can trigger language confusion and script hallucination, making decoding harder.\"},{\"question\":\"What is the main idea behind RLVR in this approach?\",\"answer\":\"The method frames CS-ASR as a verifiable reward problem where automatic rewards can be computed from reference transcripts. Reinforcement learning then directly optimizes sequence-level transcription quality using reward signals.\"},{\"question\":\"How do the proposed rewards and two-pass decoding improve results?\",\"answer\":\"Rewards combine an error-rate component with a script fidelity reward that penalizes wrong writing systems at code-switch boundaries. A two-pass draft-and-refinement procedure conditions the second decoding pass on the best draft to encourage a “listen again and fix” behavior.\"}]",1784197560,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"reinforcement-learning-for-data-efficient-code-switched-asr","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/reinforcement-learning-for-data-efficient-code-switched-asr/84666/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is code-switched ASR challenging compared with monolingual ASR?","Question",{"text":75,"@type":76},"Code-switching requires modeling rapidly shifting acoustic and linguistic cues, and labeled code-switched training data is scarce. Boundary regions can trigger language confusion and script hallucination, making decoding harder.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the main idea behind RLVR in this approach?",{"text":80,"@type":76},"The method frames CS-ASR as a verifiable reward problem where automatic rewards can be computed from reference transcripts. Reinforcement learning then directly optimizes sequence-level transcription quality using reward signals.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the proposed rewards and two-pass decoding improve results?",{"text":84,"@type":76},"Rewards combine an error-rate component with a script fidelity reward that penalizes wrong writing systems at code-switch boundaries. A two-pass draft-and-refinement procedure conditions the second decoding pass on the best draft to encourage a “listen again and fix” behavior.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":22,"slug":137},19,"General","general"]