[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83552-en":3,"doc-seo-83552-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83552,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","MultiSynt/MT Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages","Open web-scale pre-training corpora remain heavily concentrated in English, slowing the development of multilingual LLMs. MultiSynt/MT introduces an open synthetic parallel corpus with about 4.8T target-language tokens across 36 languages, generated by translating 100B high-quality Nemotron-CC tokens using TOWER+ and OPUS-MT/HPLTMT systems. For many European low- and medium-resource languages, it is among the largest openly available resources. On a multilingual benchmark suite, models trained on MultiSynt/MT match the final score of a native-data baseline using ~72% fewer tokens and exceed it by ~15% under a matched 100B budget. Analyses reveal evaluation blind spots where standard multiple-choice benchmarks miss translation-quality differences, and Norwegian idiomatic, culturally grounded tasks still benefit from native data. The released corpus includes row-aligned translations from multiple systems under a permissive open license, enabling controlled multilingual pretraining research and evaluation.","MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages  \narXiv :2607 .00890v 1 [ cs .CL] 1 Jul 2026  \nMaximilian Idahl1,2 , Jörg Tiedemann3 , Sampo Pyysalo4 , David Salinas5,6 , Tomasz Galica4 ,  \nShenbin Qian7 , Tudor Nicolae Mateiu8 , Zihao Li3 , Anna Lokrantz9 , Fedor Vitiugin4 , André Martins10,11,12 , Jenna Kanerva4 , Filip Ginter4 , Matthias Lindemann10 , Tim Isbister9 , Birger Moëll9 , Jonas Lindh9 , Jan Haji13 , Jenia Jitsev14,15,16,17 , Andrey Kutuzov7 , Stephan Oepen7 , Gema Ramírez-Sánchez8  \n1 ellamind 2Leibniz University Hannover 3University of Helsinki 4University of Turku  \n5ELLIS Institute Tübingen 6Prior Labs 7University of Oslo 8Prompsit Language Engineering  \n9AI Sweden 10Instituto de Telecomunicações 11Instituto Superior Técnico 12TransPerfect  \n13 Charles University 14 Ontocord 15LAION 16 Open-Ψ (Open-Sci) Collective  \n17Juelich Supercomputing Center (JSC), Research Center Juelich (FZJ)  \nAbstract  \nOpen web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approximately 4.8 trillion target-language tokens across 36 languages, produced by translating 100 billion high-quality Nemotron-CC tokens with TOWER+ and OPUS-MT/HPLTMT systems. For many medium-and lowerresource European languages, this is the largest openly available pre-training resource. On abroad multilingual benchmark suite, reference LLMs trained on MultiSynt/MT reach the final score of HPLT 2.0, a native-data baseline, using roughly 72% fewer pre-training tokens, and outperform it by approximately 15% relative at a matched 100B-token training budget. Our analyses also identify evaluation blind spots: standard multiple-choice benchmarks miss translation-quality differences that a fluency-sensitive LLM-as-judge evaluation cleanly recovers on the trained LLMs (with no fluency deficit in MultiSynt itself), and Norwegian idiomatic and culturally grounded tasks, for example, remain better served by native data. We release the corpus, including rowaligned translations from multiple systems, to support controlled research on multilingual pretraining data and evaluation.  \n1 Introduction  \nOpenly available pre-training data at scale exist primarily for English (Penedo et al., 2024 ; Su et al., 2025), with most other languages lacking volume or quality (see Section 2 for details) . This shortfall  \nhas motivated growing interest in machine translation as a means of producing multilingual pretraining data at scale, but the practice raises legitimate concerns: translated text can carry stylistic artifacts known as “translationese” and inherits the cultural reference frame of the source language, with names, places, idioms and culturally grounded knowledge in the English source remaining English-anchored after translation (Gellerstam, 1986 ; Riley et al., 2020) .  \nWe address this resource limitation by introducing MultiSynt/MT, an open multilingual synthetic parallel corpus of approximately 4.8 trillion target-language tokens covering 36 languages, produced by translating a 100B-token sample of high-quality web data from Nemotron-CC with open translation models, including TOWER+ 9Band 72B (Rei et al., 2026) and OPUS-MT (Tiedemann et al., 2024) . MultiSynt/MT is the largest openly available pre-training corpus to date for many European low- and medium-resource languages, exceeding the largest comparable native resource (HPLT 3.0; Oepen et al., 2025) by more than an order of magnitude in the lowest-resource ones (Figure 1) . For a subset of languages, we release row-aligned parallel translations from multiple systems, enabling controlled comparison under matched source data, and the corpus is released under a permissive open license.  \nAlongside the corpus, we present a balanced empirical characterization reporting both where reference LLMs trained on MultiSynt/MT outperform native multilingual baselines and w","cbCaiqvSSdgFkgaA","https://ap.wps.com/l/cbCaiqvSSdgFkgaA","pdf",721184,7,1,22,"English","en",105,"# Abstract\n# Introduction\n## Resource gap in multilingual pretraining\n## MultiSynt/MT corpus construction\n## Empirical results and evaluation observations","[{\"question\":\"What problem does MultiSynt/MT address in multilingual LLM development?\",\"answer\":\"Openly available large-scale pre-training data is concentrated in English, leaving many other languages with insufficient volume or quality. MultiSynt/MT targets this gap by providing a large synthetic parallel corpus across 36 languages.\"},{\"question\":\"How is MultiSynt/MT created and what is its scale?\",\"answer\":\"A 100B-token sample of high-quality Nemotron-CC web data is translated into target languages using open translation models including TOWER+ and OPUS-MT/HPLTMT. The corpus contains roughly 4.8 trillion target-language tokens across 36 languages.\"},{\"question\":\"What do results on multilingual benchmarks show about models trained on MultiSynt/MT?\",\"answer\":\"Reference LLMs trained on MultiSynt/MT reach the final score of a native-data baseline (HPLT 2.0) using about 72% fewer tokens, and they outperform it by around 15% when training at a matched 100B-token budget.\"}]",1784188767,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"multisyntmt-trillion-token-multi-parallel-pre-training-data-translated-across-36-languages","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/multisyntmt-trillion-token-multi-parallel-pre-training-data-translated-across-36-languages/83552/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does MultiSynt/MT address in multilingual LLM development?","Question",{"text":76,"@type":77},"Openly available large-scale pre-training data is concentrated in English, leaving many other languages with insufficient volume or quality. MultiSynt/MT targets this gap by providing a large synthetic parallel corpus across 36 languages.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is MultiSynt/MT created and what is its scale?",{"text":81,"@type":77},"A 100B-token sample of high-quality Nemotron-CC web data is translated into target languages using open translation models including TOWER+ and OPUS-MT/HPLTMT. The corpus contains roughly 4.8 trillion target-language tokens across 36 languages.",{"name":83,"@type":74,"acceptedAnswer":84},"What do results on multilingual benchmarks show about models trained on MultiSynt/MT?",{"text":85,"@type":77},"Reference LLMs trained on MultiSynt/MT reach the final score of a native-data baseline (HPLT 2.0) using about 72% fewer tokens, and they outperform it by around 15% when training at a matched 100B-token budget.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]