[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84018-en":3,"doc-seo-84018-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84018,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","CoPiT: Cognitive Pivot Translation for Digraphic Low-Resource Mongolian in the Traditional Script","Low-resource languages pose major challenges for machine translation, and Mongolian is a representative case due to its digraphic writing system. Mongolian appears in both Cyrillic and Traditional scripts, yet data availability is highly imbalanced: Cyrillic is comparatively well-resourced, while Traditional is extremely scarce and orthographically ambiguous, degrading direct translation quality. CoPiT introduces a cognitively motivated pivot-based pipeline that resolves Traditional-script ambiguity before translation by routing through Cyrillic. Across multiple backbone models and target languages, CoPiT yields robust absolute BLEU gains and consistent 1.5–1.6× COMET improvements, enabling strong open-source models to match or exceed GPT-4.1 under similar settings.","CoPiT: Cognitive Pivot Translation for Digraphic Low-Resource Mongolian  \nin the Traditional Script  \nBurte Bayarsaikhan* 1 Serynn Kim*2 Buru Chang†1  \n1 Korea University 2Hankuk University of Foreign Studies  \n{burtebay,[buru_chang}@korea.ac.kr](buru_chang}@korea.ac.kr)  \n[serynn@hufs.ac.kr](serynn@hufs.ac.kr)  \narXiv :2607 .05849v 1 [ cs .CL] 7 Jul 2026  \nAbstract  \nLow-resource languages remain challenging for machine translation, and Mongolian is a representative case. As a digraphic language, Mongolian is written in both Cyrillic and Traditional scripts, which exhibit a severe imbalance in data availability. While the Cyrillic script is relatively well-resourced, the Traditional script remains extremely data-scarce and orthographically ambiguous, leading to substantial performance degradation in direct translation. We propose CoPiT, a cognitively motivated pivotbased translation pipeline that exploits this internal resource hierarchy by routing translation through the Cyrillic script. The pipeline explicitly resolves script-induced ambiguity in the Traditional script before translation, enabling more stable and accurate meaning transfer. Across multiple backbone models and target languages, CoPiT consistently outperforms direct translation, achieving substantial absolute BLEU improvements together with consistent 1.5–1.6× COMET gains. These gains allow strong open-source models to match or outperform GPT-4.1 under comparable evaluation settings. Beyond inference-time improvements, CoPiT enables the construction of synthetic parallel data directly from Traditionalscript text, mitigating data scarcity in realistic low-resource scenarios. We release a new multi-script parallel dataset covering Mongolian in both scripts alongside English, Korean, and Russian. All datasets and code are publicly available at [https://anonymous.4open](https://anonymous.4open). science/r/anonymous_project-76C7 .  \n1 Introduction  \nLarge language models (LLMs) have evolved from primarily monolingual systems to multilingual systems through large-scale multilingual pre-training. Despite this progress, LLMs continue to perform poorly on low-resource languages (LRLs), largely  \n*Equal contribution.†Corresponding author.  \nFigure 1: Mongolian digraphia. The same content can be written in Traditional script (left), with multiple surface forms and limited resources, and in betterresourced Cyrillic script (right) .  \ndue to limited training data and sparse linguistic resources. This limitation is particularly evident in machine translation (MT), which relies heavily on large-scale parallel corpora that are often unavailable for LRLs (Raja and Vats, 2025) .  \nThese challenges are particularly pronounced in Mongolian, a low-resource language with a unique digraphic characteristic, as it is written in both Cyrillic and Traditional scripts (see Figure 1) . The Cyrillic script, which serves as the dominant modern form, is largely phonemic and exhibits a high degree of consistency in the sound representation. In contrast, the Traditional script, an archaic writing system, is written vertically, and its letter forms vary by positional context. Many phonological and morphological distinctions are not explicitly encoded in the script, such that a single written form may correspond to multiple plausible interpretations. This intrinsic ambiguity distinguishes the Traditional script structurally from its Cyrillic counterpart. For example, the Gemini-3-Pro-Preview API flags content written in the Traditional Mongolian script as harmful, as shown in Appendix A.1 . Beyond these structural differences, the two Mongolian scripts are associated with markedly different levels of linguistic and computational resources. The Cyrillic script has long served as the dominant writing system in modern Mongolian so-  \nciety following its institutional adoption through state language policies. As a result, the vast majority of textual data and language technologies are developed for ","cbCaiiiKiBq7S2Ks","https://ap.wps.com/l/cbCaiiiKiBq7S2Ks","pdf",3392747,3,1,23,"English","en",105,"# Abstract\n# Introduction","[{\"question\":\"Why does direct machine translation for Mongolian in the Traditional script perform poorly?\",\"answer\":\"The Traditional script has extreme data scarcity and orthographic ambiguity, so a single written form can correspond to multiple plausible interpretations, which causes substantial performance degradation in direct translation.\"},{\"question\":\"What is CoPiT and how does it improve translation?\",\"answer\":\"CoPiT (Cognitive Pivot Translation) uses a two-stage pipeline: first converting Traditional-script Mongolian into Cyrillic, then translating from Cyrillic to the target language. This isolates script-specific ambiguity and leverages the relatively better-resourced Cyrillic representation.\"},{\"question\":\"What impact do CoPiT results show across models and target languages?\",\"answer\":\"CoPiT consistently outperforms direct translation, producing substantial absolute BLEU improvements and consistent 1.5–1.6× COMET gains, allowing strong open-source models to match or outperform GPT-4.1 under comparable evaluation settings.\"}]",1784192040,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"copit-cognitive-pivot-translation-for-digraphic-low-resource-mongolian-in-the-traditional-script","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/copit-cognitive-pivot-translation-for-digraphic-low-resource-mongolian-in-the-traditional-script/84018/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why does direct machine translation for Mongolian in the Traditional script perform poorly?","Question",{"text":75,"@type":76},"The Traditional script has extreme data scarcity and orthographic ambiguity, so a single written form can correspond to multiple plausible interpretations, which causes substantial performance degradation in direct translation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is CoPiT and how does it improve translation?",{"text":80,"@type":76},"CoPiT (Cognitive Pivot Translation) uses a two-stage pipeline: first converting Traditional-script Mongolian into Cyrillic, then translating from Cyrillic to the target language. This isolates script-specific ambiguity and leverages the relatively better-resourced Cyrillic representation.",{"name":82,"@type":73,"acceptedAnswer":83},"What impact do CoPiT results show across models and target languages?",{"text":84,"@type":76},"CoPiT consistently outperforms direct translation, producing substantial absolute BLEU improvements and consistent 1.5–1.6× COMET gains, allowing strong open-source models to match or outperform GPT-4.1 under comparable evaluation settings.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]