[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-421530-105":59,"doc-detail-421530-en":130},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":123,"head_meta":125,"extra_data":127,"updated_unix":129},105,"en","effective-vocabulary-expansion-of-multilingual-language-models-for-extremely-low-resource-languages","Effective Vocabulary Expansion of Multilingual Language Models for Extremely Low-Resource Languages","","Multilingual pre-trained language models (mPLMs) provide broad coverage for many low-resource languages, yet many target languages remain unsupported because existing approaches rarely extend mPLMs with informed vocabulary representations. This work expands model vocabulary using a target-language corpus, filters biased source-oriented vocabulary, initializes new token representations with bilingual dictionaries, and continues pre-training on the target corpus. Experiments show improvements over randomly initialized baselines in POS tagging (+0.54%) and NER (+2.60%), with robust corpus selection and no source-language degradation.",{"@graph":69,"@context":122},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/effective-vocabulary-expansion-of-multilingual-language-models-for-extremely-low-resource-languages/421530/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/effective-vocabulary-expansion-of-multilingual-language-models-for-extremely-low-resource-languages/421530.png","ImageObject",300,407,{"name":92,"@type":93},"Pentious","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-09-29","2026-09-28",true,{"@type":102,"interactionType":103,"userInteractionCount":14},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118],{"name":109,"@type":110,"acceptedAnswer":111},"Why is vocabulary expansion important for extremely low-resource languages in mPLMs?","Question",{"text":112,"@type":113},"Vocabulary layer parameters strongly affect convergence and final performance during continued pre-training. Expanding vocabulary with better-initialized representations improves learning for previously unsupported target languages.","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"How does the proposed method initialize the expanded vocabulary representations?",{"text":117,"@type":113},"It uses a target-language corpus to select and expand vocabulary, filters a subset from the original vocabulary that is biased toward the source language, and initializes representations of new tokens using bilingual dictionaries.",{"name":119,"@type":110,"acceptedAnswer":120},"What improvements does the method achieve compared with random initialization?",{"text":121,"@type":113},"Compared with a baseline using randomly initialized expanded vocabulary, the method improves POS tagging by 0.54% and NER by 2.60%, while remaining robust in training-corpus selection and not degrading the source language performance.","https://schema.org",{"og:url":83,"og:type":124,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":126,"canonical":83},"index,follow",{"doc_id":128,"site_id":62},421530,1790704537,{"code":4,"msg":5,"data":131},{"doc_id":128,"user_id":132,"nickname":92,"user_avatar":133,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":14,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":139,"language":140,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":67,"update_tm":144,"read_time":145},1374404730887,"https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8","Effective vocabulary expansion of multilingual language models for extremely low-resource languages  \nJianyu Zhenga,b,∗  \na School of Foreign Languages, University of Electronic Science and Technology of China, Chengdu, Sichuan Province, 611731 CN, China b School of Computer Science and Engineering, University of Electronic Science and Technology of China, Chengdu, Sichuan Province, 611731 CN, China  \nARTICLE INFO  \nKeywords:  \nAB STRACT  \nMultilingual pre-trained language models(mPLMs) offer significant benefits for many low-  \nmultilingual pre-trained language model continued pre-training  \nexpanded vocabulary  \nbilingual dictionary  \nresource languages. To further expand the range of languages these models can support, many works focus on continued pre-training of these models. However, few works address how to extend mPLMs to low-resource languages that were previously unsupported. To tackle this issue, we expand the model’s vocabulary using a target language corpus. We then screen out a subset from the model’s original vocabulary, which is biased towards representing the source language(e.g. English), and utilize bilingual dictionaries to initialize the representations of the expanded vocabulary. Subsequently, we continue to pre-train the mPLMs using the target language corpus, based on the representations of these expanded vocabulary. Experimental results show that our proposed method outperforms the baseline, which uses randomly initialized expanded vocabulary for continued pre-training, in POS tagging and NER tasks, achieving improvements by 0.54% and 2.60%, respectively. Furthermore, our method demonstrates high robustness in selecting the training corpora, and the models’ performance on the source language does not degrade after continued pre-training.  \n1. Introduction  \nThe emergence of multilingual pre-trained language models (mPLMs) brings significant benefits to a wider range of languages(Doddapaneni, Ramesh, Khapra, Kunchukuttan and Kumar, 2025) . These models are trained on largescale multilingual corpora through tasks such as Masked Language Modeling (MLM)(Kenton, Toutanova et al., 2019; Taylor, 1953) and Translation Language Modeling (TLM)(Conneau and Lample, 2019) within a unified neural network architecture, which allows them to not only support multiple languages but also transfer knowledge across languages, thereby alleviating disparities in language processing.  \nCurrently, the number of languages supported by mPLMs remains relatively limited. Figure 1 illustrates the number of languages supported by commonly used mPLMs(Kenton et al., 2019; Conneau, Khandelwal, Goyal, Chaudhary, Wenzek, Guzmán, Grave, Ott, Zettlemoyer and Stoyanov, 2020; Kale, Xue, Constant, Roberts, Al-Rfou, Siddhant and Barua, 2020; Qin; Eisenschlos, Ruder, Czapla, Kadras, Gugger and Howard, 2019; Huang, Liang, Duan, Gong, Shou, Jiang and Zhou; Liu, Gu, Goyal, Li, Edunov, Ghazvininejad, Lewis and Zettlemoyer, 2020; Shliazhko, Fenogenova, Tikhonova, Kozlova, Mikhailov and Shavrina, 2024; Kondratyuk and Straka, 2019; Ouyang, Wang, Pang, Sun, Tian, Wu and Wang, 2021; Fan, Bhosale, Schwenk, Ma, El-Kishky, Goyal, Baines, Celebi, Wenzek, Chaudhary et al., 2021) . It is evident that most of these models support only about one hundred languages, which is a very small proportion (roughly 1.5%) compared to the nearly 7,000 languages(Campbell and Grondona, 2008) in the world. Furthermore, those models perform poorly on low-resource languages, particularly the languages with small corpus or significant grammatical and orthographic differences from high-resource languages (e.g. English)(Ebrahimi and von der Wense, 2021) . Therefore, many works focus on the continued pre-training of pre-trained language models (PLMs) to better support low-resource languages(Wang, Ruder and Neubig, 2022; Ebrahimi and von der Wense, 2021; Dobler and De Melo, 2023) . For example, Wang et al.(Wang, Karthikeyan, Mayhew and Roth, 2020) expand the vocabulary and continue ","cbCaiafjDriQaW8R","https://ap.wps.com/l/cbCaiafjDriQaW8R","pdf",577202,13,"English","# Introduction\n## Vocabulary expansion for unsupported low-resource languages\n## Motivation: limited language coverage and weak low-resource performance\n## Prior work and initialization challenges\n## Contribution overview and experimental motivation","[{\"question\":\"Why is vocabulary expansion important for extremely low-resource languages in mPLMs?\",\"answer\":\"Vocabulary layer parameters strongly affect convergence and final performance during continued pre-training. Expanding vocabulary with better-initialized representations improves learning for previously unsupported target languages.\"},{\"question\":\"How does the proposed method initialize the expanded vocabulary representations?\",\"answer\":\"It uses a target-language corpus to select and expand vocabulary, filters a subset from the original vocabulary that is biased toward the source language, and initializes representations of new tokens using bilingual dictionaries.\"},{\"question\":\"What improvements does the method achieve compared with random initialization?\",\"answer\":\"Compared with a baseline using randomly initialized expanded vocabulary, the method improves POS tagging by 0.54% and NER by 2.60%, while remaining robust in training-corpus selection and not degrading the source language performance.\"}]","Effective Vocabulary Expansion of Multilingual Language Models for Extremely Low-Resource Languages | PDF",1790625076,33]