[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82516-en":3,"doc-seo-82516-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82516,687197100911,"Himbo","https://ap-avatar.wpscdn.com/avatar/a000239b6f1da00475?x-image-process=image/resize,m_fixed,w_180,h_180&k=1782698725881665579",8,"Research & Report","Efficient Multilingual Reasoning Transfer via Progressive Code-Switching","Large reasoning models perform strongly in English but degrade when required to reason in other languages. A transfer strategy can reuse English reasoning skills, yet prior methods often need distilled target-language traces from stronger models or costly online supervision from judge models. This work introduces PCS (Progressive Code-Switching), an efficient framework using only lightweight translation. PCS builds code-switched traces by converting a subset of English reasoning steps, then applies supervised fine-tuning and reinforcement learning with step-level language consistency curriculum to progressively reach fully target-language reasoning, reducing performance gaps while preserving accuracy.","Efficient Multilingual Reasoning Transfer via Progressive Code-Switching  \nZhijun Wang 1 ∗ , Junxiao Liu 1 , Hao Zhou 1 ,  \nHao-Ran Wei2 , Baosong Yang2 , Shujian Huang 1†  \n1National Key Laboratory for Novel Software Technology, Nanjing University, China  \n2Tongyi Lab, Alibaba Group  \n{wangzj,junxiaoliu,[zhouh}@smail.nju.edu.cn](zhouh}@smail.nju.edu.cn), {funan.whr,[yangbaosong.ybs}@alibaba-inc.com](yangbaosong.ybs}@alibaba-inc.com), [huangsj@nju.edu.cn](huangsj@nju.edu.cn)  \narXiv :2607 .00485v 1 [ cs .CL] 1 Jul 2026  \nAbstract  \nLarge reasoning models (LRMs) have achieved strong reasoning capabilities in English, yet their performance degrades significantly when required to reason in other languages. A natural solution is to transfer the model’s English reasoning ability to target languages. However, existing transfer approaches typically rely on distilled target-language reasoning traces from stronger LRMs or online supervision from external judge models, which are costly and difficult to scale. In this paper, we propose PCS (Progressive Code-Switching), a more efficient transfer framework that requires only lightweight translation  \n—without any stronger model for distillation or judging. PCS first constructs code-switched reasoning traces by translating a subset of English reasoning steps into the target language, and uses them to initialize the model’s code-switching ability via supervised fine-tuning. It then applies reinforcement learning with a step-level language consistency curriculum, progressively raising the target-language ratio until the model reasons entirely in the target language. This progressive design provides a smooth transfer path that avoids the instability and performance degradation commonly observed when directly enforcing target-language reasoning. Experiments on multiple benchmarks and five typologically diverse languages show that PCS substantially narrows the performance gap between target-language and English reasoning, yielding more language-consistent reasoning while maintaining competitive accuracy.  \nIntroduction  \nLarge Reasoning Models (LRMs) such as DeepSeekR1 (Guo et al. 2025), Qwen3 (Yang et al. 2025), and OpenAIo1 (OpenAI et al. 2024) have achieved strong performance on math and code reasoning. Reinforcement Learning with Verified Rewards (RLVR) improves such abilities through scaling, and the long chain-of-thought reasoning steps strengthen problem-solving and enhance interpretability.  \nDespite these advances, multilingual reasoning—defined as the capability to solve mathematical problems in languages other than English—remains a persistent challenge. When prompted with non-English questions, LRMs often still reason in English (Wang et al. 2025a) . This mismatch between the user’s language and the model’s reasoning process can undermine comprehensibility and trustworthiness.  \n∗Work done during internship at Tongyi Lab.†Corresponding author.  \nExisting approaches to enforce input-output language matching include Prompt Control (Tam et al. 2025), Prefix Control (Qi et al. 2025a), Supervised Fine-Tuning (SFT) on multilingual traces (Luo et al. 2025), and Reinforcement Learning (RL) with language-consistency rewards (MistralAI et al. 2025) . However, these methods exhibit distinct practical limitations. Prompt Control is often unreliable, while prefix-based constraints, though effective for certain languages, frequently compromise accuracy relative to English reasoning. Supervised Fine-Tuning is constrained by the scarcity of multilingual long-reasoning data and tends to introduce repetition artifacts. Furthermore, while RL with language-consistency rewards encourages the model to reason in the question language, performance typically remains inferior to English-based reasoning.  \nTo address these issues, we focus on transferring strong English reasoning ability to target languages. Rather than forcing an abrupt switch to fully target-language reasoning, we propose PCS (Progressive Co","cbCaioWhYhNadgCA","https://ap.wps.com/l/cbCaioWhYhNadgCA","pdf",1133354,3,1,12,"English","en",105,"# Abstract\n# Introduction\n# Methodology\n## Code-Switched Reasoning","[{\"question\":\"Why do large reasoning models perform worse on non-English reasoning tasks?\",\"answer\":\"When prompted with non-English questions, large reasoning models often continue to reason in English, creating a mismatch between the user’s language and the model’s reasoning process and reducing performance and trustworthiness.\"},{\"question\":\"What is PCS (Progressive Code-Switching) and how does it enable multilingual transfer?\",\"answer\":\"PCS treats code-switched reasoning as an intermediate transfer stage. It translates only a subset of English reasoning steps to build mixed English–target-language trajectories, then trains with supervised fine-tuning and reinforcement learning using a step-level language consistency curriculum that gradually increases target-language usage.\"},{\"question\":\"What makes PCS more efficient than prior transfer approaches?\",\"answer\":\"PCS avoids reliance on distilled target-language reasoning traces from stronger teacher models and does not require online supervision from external judge models. It only needs lightweight translation and then uses training procedures to reach target-language reasoning progressively.\"}]",1784181094,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"efficient-multilingual-reasoning-transfer-via-progressive-code-switching","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/efficient-multilingual-reasoning-transfer-via-progressive-code-switching/82516/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do large reasoning models perform worse on non-English reasoning tasks?","Question",{"text":75,"@type":76},"When prompted with non-English questions, large reasoning models often continue to reason in English, creating a mismatch between the user’s language and the model’s reasoning process and reducing performance and trustworthiness.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is PCS (Progressive Code-Switching) and how does it enable multilingual transfer?",{"text":80,"@type":76},"PCS treats code-switched reasoning as an intermediate transfer stage. It translates only a subset of English reasoning steps to build mixed English–target-language trajectories, then trains with supervised fine-tuning and reinforcement learning using a step-level language consistency curriculum that gradually increases target-language usage.",{"name":82,"@type":73,"acceptedAnswer":83},"What makes PCS more efficient than prior transfer approaches?",{"text":84,"@type":76},"PCS avoids reliance on distilled target-language reasoning traces from stronger teacher models and does not require online supervision from external judge models. It only needs lightweight translation and then uses training procedures to reach target-language reasoning progressively.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]