[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85076-en":3,"doc-seo-85076-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85076,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Echoes Across Vietnam’s Highlands, Delta, and Coast Multilingual Corpus for Cham, Khmer, and Tay-Nung","Vietnam’s ethnic minority languages remain largely absent from Natural Language Processing, and the barrier goes beyond data scarcity. Cham, Khmer, and Tay-Nung differ in script, contact with Vietnamese, and standardization, causing multilingual adaptation to learn incorrect signals. CKTN is presented as the first corpus and benchmark for these languages, with 44,367 documents and 24M subword tokens, enabling continued pretraining, category classification, and summary-document retrieval.","Echoes Across Vietnam’s Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung  \nAnh Trac Duc Dinh1*†, Khang Nhat Hoang Vo2*†, Vinh Cong Doan1 ,  \nTai Tien Ta 1 , Khoa Duc Anh Lam 1  \n1Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT), VNU-HCM, Ho Chi Minh City, Vietnam  \n2Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE  \n*Corresponding author: [anh.dinhtracduc@hcmut.edu.vn](anh.dinhtracduc@hcmut.edu.vn); [Khang.Vo@mbzuai.ac.ae](Khang.Vo@mbzuai.ac.ae)  \narXiv :2607 .08362v 1 [ cs .CL] 9 Jul 2026  \nAbstract  \nVietnam’s ethnic minority languages are almost absent from the field of Natural Lan  \nguage Processing (NLP), and the challenge goes beyond data scarcity: Cham, Khmer, and Tay-Nung differ sharply in script, Vietnamese contact, and standardization, conditions under which standard multilingual adaptation can learn the wrong signals. We introduce CKTN, the first corpus and benchmark for these languages (44,367 documents, 24M subword tokens), spanning continued pretraining, category classification, and summary-document retrieval. We show that existing multilingual encoders severely fragment these languages, and that common adaptation metrics can mislead: models may lower language-modeling loss or excel at lexical-overlap retrieval while still failing at semantic generalization across documents. We address this with a script-aware adaptation recipe-vocabulary augmentation combined with calibrated replaced-token pretraining-that prevents the discriminator from exploiting trivial script mismatches. The result is an encoder with substantially less fragmentation and the strongest classification performance among evaluated models, exposing the limits of lexical-overlap retrieval as an evaluation signal.  \n1 Introduction  \nLow-resource NLP is often framed as a problem of missing data. For many minority languages, however, scarcity is only one part of the difficulty. Languages spoken under sustained contact with a dominant national language may diverge from their higher-resource genealogical relatives, develop unstable or mixed writing practices, and appear in digital text through scripts that are poorly represented by multilingual tokenizers. In such settings, adaptation can fail even when some text is available: models may learn local subword regularities,  \n†These authors contributed equally to this work, and are project leaders. Names are ordered alphabetically.  \nlexical overlap, or script cues rather than transferable language structure.  \nVietnam’s ethnic minority languages offer a compelling testbed for this problem. Vietnam recognizes 54 ethnic groups (Van, 2002 ; Tappe, 2015); beyond the Kinh majority, over 14 million people speak minority languages spanning the Austroasiatic, Tai-Kadai, Austronesian, and Sino-Tibetan families (Edmondson and Gregerson, 2007 ; Liu et al., 2020) . We focus on three languages actively published in digital media yet almost absent from NLP: Cham, Khmer, and Tay-Nung. Together they form a natural contrast set: Cham (Austronesian, south-central coast) is traditionally written ina Brahmic-derived abugida but often appears in Latin script online; Khmer (Austroasiatic, Mekong Delta) uses a script without explicit word boundaries; and Tay-Nung (Tai-Kadai varieties of the northern highlands) uses Latin-based orthography but has limited standardized digital text. These languages thus span different regions, scripts, language families, and degrees of visible contact with Vietnamese.  \nThis diversity is not merely sociolinguistic background; it creates concrete modeling failures. Existing multilingual encoders such as mBERT, XLM-R, and RemBERT (Pires et al., 2019 ; Conneau et al., 2020 ; Chung et al., 2021) allocate little vocabulary capacity to these languages, fragmenting words into continuation pieces and weakening lexical representations. Moreover, script heterogeneity creates a specific hazard for ELECTRAsty","cbCaiq79hYX3PI1w","https://ap.wps.com/l/cbCaiq79hYX3PI1w","pdf",3904851,2,1,18,"English","en",105,"# Introduction\n# Related Work\n## Low-Resource Corpus for Vietnamese Ethnic Languages","[{\"question\":\"What problem does the paper address in NLP for Vietnam’s ethnic minority languages?\",\"answer\":\"It addresses the gap where Cham, Khmer, and Tay-Nung are largely absent from NLP, and explains that adaptation can fail due to script and standardization differences, not just data scarcity.\"},{\"question\":\"What is CKTN and what tasks does it support?\",\"answer\":\"CKTN is the first multilingual corpus and benchmark for Cham, Khmer, and Tay-Nung. It includes 44,367 documents and 24M subword tokens, supporting continued pretraining, 28-way category classification, and summary-document retrieval.\"},{\"question\":\"Why can common adaptation metrics give misleading results?\",\"answer\":\"The paper shows that models may reduce language-modeling loss or perform well on lexical-overlap retrieval without achieving semantic generalization across documents, so evaluation based on lexical overlap can overestimate real representation quality.\"}]",1784200894,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"echoes-across-vietnams-highlands-delta-and-coast-multilingual-corpus-for-cham-khmer-and-tay-nung","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/echoes-across-vietnams-highlands-delta-and-coast-multilingual-corpus-for-cham-khmer-and-tay-nung/85076/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in NLP for Vietnam’s ethnic minority languages?","Question",{"text":75,"@type":76},"It addresses the gap where Cham, Khmer, and Tay-Nung are largely absent from NLP, and explains that adaptation can fail due to script and standardization differences, not just data scarcity.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is CKTN and what tasks does it support?",{"text":80,"@type":76},"CKTN is the first multilingual corpus and benchmark for Cham, Khmer, and Tay-Nung. It includes 44,367 documents and 24M subword tokens, supporting continued pretraining, 28-way category classification, and summary-document retrieval.",{"name":82,"@type":73,"acceptedAnswer":83},"Why can common adaptation metrics give misleading results?",{"text":84,"@type":76},"The paper shows that models may reduce language-modeling loss or perform well on lexical-overlap retrieval without achieving semantic generalization across documents, so evaluation based on lexical overlap can overestimate real representation quality.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]