[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86486-en":3,"doc-seo-86486-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":20,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86486,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","GigaAM Multilingual Foundation Model for Underrepresented Languages","Multilingual ASR quality remains uneven because long-tail languages face extreme data scarcity. The paper presents GigaAM Multilingual, a Conformer encoder pretrained on 2M hours of audio with a HuBERT-style objective. A cluster-level data balancing strategy improves representation learning by reducing head-language dominance. During fine-tuning, a domain-aware sampling recipe further mitigates imbalance. Controlled comparisons show consistent gains on Kyrgyz, Kazakh, and Uzbek, including stronger performance on spontaneous speech while keeping computational efficiency.","GigaAM Multilingual: Foundation Model for Underrepresented Languages  \nAndrei Kuzmenko 1 ,∗ , Alexandr Maximenko 1 ,∗ , Aleksandr Kutsakov 1 , Georgii Gospodinov 1 ,∗∗,  \nDmitrii Bolotov 1, Oleg Kutuzov 1, Pavel Bogomolov 1, Fyodor Minkin 1  \n1 SaluteDevices, Russia  \n{andrey .kuzmenko2907, ae .maximenko, askutsakov, georgygospodinov, bolotovdm, olegkutuzov01, bobrosoft98, [minkin.fyodor](minkin.fyodor}@gmail.com)[}](minkin.fyodor}@gmail.com)[@gmail.com](minkin.fyodor}@gmail.com)  \narXiv :2607 . 10371v1 [ ee ss .AS] 11 Jul 2026  \nAbstract  \nDespite recent scaling successes, multilingual ASR performance remains highly uneven, with long-tail languages suffering from severe data scarcity. This work addresses the challenge of building robust foundation models for underrepresented Central Asian languages (Kazakh, Kyrgyz, Uzbek) . We present GigaAM Multilingual, a Conformer encoder pretrained on 2M hours of audio using a HuBERT-style objective. Crucially, we introduce a cluster-level data balancing strategy during pre-training and a domain-aware sampling method during fine-tuning to mitigate head-language dominance. In controlled comparisons, our approach outperforms strong open pretrained encoders (Whisper Large v3, Omnilingual-1B) on target languages, achieving significant gains on spontaneous speech while maintaining efficiency. We release the foundation encoder and ASR model, offering a proven recipe for effective multilingual adaptation under realistic data imbalance.  \nIndex Terms: speech recognition, self-supervised learning, multilingual, low-resource languages  \n1. Introduction  \nRecent years have seen rapid progress in multilingual automatic speech recognition (ASR), driven by scaling data, model capacity, and weakly-supervised and self-supervised training objectives [1, 2, 3, 4] . Despite these advances, recognition quality remains highly uneven across languages: performance on highresource languages is strong, while many low-resource and long-tail languages still exhibit error rates that are prohibitive for downstream applications [5] .  \nA major contributor to this disparity is data imbalance [6]: empirical data distribution favors head languages, whereas drastic equalization via naive upsampling can lead to overfitting or degradation on high-resource languages. Principled strategies for mixing and weighting data in pre-training and adaptation are thus essential.  \nIn this work, we extend GigaAM [7] – an efficient selfsupervised learner (SSL) for Russian ASR – to the multilingual setting. We pre-train the encoder on a 2M-hour corpus with highly skewed language-group proportions (Fig. 1) . Although the target Central Asian languages (Kyrgyz, Kazakh, Uzbek) are present in pre-training, their share is small, and strong general-purpose ASR models often struggle in this lowresource regime (Table 1) .  \nWe propose a two-level approach that explicitly accounts for long-tail structure. First, we incorporate group-aware reweighting during SSL pre-training, increasing the contribution of low-share language groups to the pre-training objective  \n*These authors contributed equally.  \n**indicates the corresponding author.  \nTable 1: WER (%) of GigaAM Multilingual against public multilingual ASR systems on Common Voice (CV), FLEURS, and our internal in-the-wild test sets (Sec. 4.2), best in bold.  \n\n| Lang | Split | GigaAM\u003Cbr>Multilingual | Omnilingual\u003Cbr>1B LLM ASR | Seamless M4T\u003Cbr>large v2 | Whisper\u003Cbr>large v3 |\n| --- | --- | --- | --- | --- | --- |\n| English | CV | 21.5 | 24.7 | 16.2 | 20.0 |\n|  | FLEURS | 9.4 | 7.1 | 5.8 | 3.9 |\n| Russian | CV | 5.1 | 13.6 | 9.2 | 9.1 |\n|  | FLEURS | 3.0 | 6.4 | 4.6 | 3.1 |\n|  | Internal | 6.0 | 14.6 | 16.1 | 10.1 |\n| Kazakh | CV | 13.8 | 23.7 | 23.8 | 57.8 |\n|  | FLEURS | 4.4 | 6.6 | 6.8 | 32.4 |\n|  | Internal | 15.8 | 32.2 | 62.9 | 65.2 |\n| Kyrgyz | CV | 10.2 | 21.6 | 14.3 | 95.2 |\n|  | FLEURS | 5.5 | 8.1 | 9.5 | 86.3 |\n|  | Internal | 9.8 | 25.0 | 78.3 | 102.2 |\n| Uzbek | CV | 9.2 | 32.8 | 2","cbCaiiG0Nb9VC9ZW","https://ap.wps.com/l/cbCaiiG0Nb9VC9ZW","pdf",340339,6,1,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n## Multilingual speech representation learning","[{\"question\":\"How does fine-tuning address low-resource imbalance?\",\"answer\":\"Fine-tuning uses a mixture of multiple target languages together with a domain-aware sampling data recipe, combining open corpora, synthetic speech, weakly supervised data, and internal annotations to improve adaptation and convergence on minimally covered languages.\"}]",1784212083,15,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"gigaam-multilingual-foundation-model-for-underrepresented-languages","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/gigaam-multilingual-foundation-model-for-underrepresented-languages/86486/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How does fine-tuning address low-resource imbalance?","Question",{"text":75,"@type":76},"Fine-tuning uses a mixture of multiple target languages together with a domain-aware sampling data recipe, combining open corpora, synthetic speech, weakly supervised data, and internal annotations to improve adaptation and convergence on minimally covered languages.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,106,111,114,119,122,126],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":45,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":103,"show_sort_weight":104,"slug":105},"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":45,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":45,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":98,"slug":129},19,"General","general"]