[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82866-en":3,"doc-seo-82866-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82866,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition","Extending automatic speech recognition (ASR) to low-resource African languages is limited by the large-scale effort required to collect training data. A common proposal is to exploit linguistic relatedness by sequentially adapting a model on a related auxiliary language before moving to the low-resource target. Prior gains were observed in small ASR models, but large ASR effectiveness was unclear. Controlled experiments across six factors, two Africa-centric corpora, and four large ASR models show that related auxiliary pre-adaptation brings no practically meaningful transfer improvements with limited target-language data.","Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition  \nAndrei Florian 1 Cynthia Jayne Amol2 Hope Kerubo Ombaba2 Xiaoyu Cui 1 Boniface Mwau2  \nBiatus Maina Kamau2 Lilian Diana Awuor Wanzare2 Christiane Fellbaum 1 Happy Buzaaba 1  \n1Princeton University 2Maseno University  \narXiv :2607 .048 14v 1 [ cs .CL] 6 Jul 2026  \nAbstract  \nExtending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection at scale. A promising direction is to leverage linguistic relatedness to enhance cross-lingual transfer from a related auxiliary language to the low-resource target by sequentially adapting on both. Although this strategy has shown meaningful improvements in small ASR models, its effectiveness in large ASR remains unclear. We extend this frame  \nwork to large multilingual ASR through a systematic controlled experimental design spanning six factors, two Africa-centric corpora, and four large ASR models, isolating whether linguistic relatedness reliably predicts crosslingual transfer gains in this setting. Across all conditions, pre-adaptation on related auxiliary languages yields no practically meaningful transfer improvements given minimal targetlanguage data, suggesting that linguistic relatedness alone may not reliably predict crosslingual transfer gains in large multilingual ASR, or constitute an effective strategy for extending such models to low-resource languages.  \n1 Introduction  \nAutomatic Speech Recognition (ASR) holds particular promise for extending language technologies to communities whose languages are primarily oral, as is the case for the majority of Africa’s approximately 2,000 languages (Orife et al., 2020) . Yet African languages remain profoundly underrepresented in contemporary NLP systems. Joshi et al. (2020) estimate that 88% of the world’s languages have no meaningful presence in the datasets underpinning modern language models, a category encompassing most languages spoken across the African continent, and although recent large-scale models such as OpenAI’s Whisper (Radford et al.,  \n* Correspondence to Andrei Florian and Happy Buzaaba {andrei.florian, [happy.buzaaba}@princeton.edu](happy.buzaaba}@princeton.edu).  \nFigure 1: Experimental design for the first factor using languages from the AfriVoices KE corpus (Wanzareet al., 2026) . Instances of the base Whisper Small (Radford et al., 2022) model are fine-tuned on each of four auxiliary languages, two linguistically related and two unrelated, at 1, 10, and 70 hours of labelled speech. The resulting auxiliary models, alongside the original baseline, are subsequently fine-tuned on 1, 10, and 70 hours of the target language Kalenjin, and evaluated using the Word Error Rate (WER) metric on the Kalenjin test set. Colours denote language family membership: blue (Nilotic), green (Bantu), orange (Cushitic) .  \n2022) and Meta’s Omnilingual ASR (Keren et al., 2025) have extended ASR capabilities across increasingly larger language sets, the vast majority of African languages remain systematically absent from their training corpora, leading to poor downstream performance.  \nThe root of this deficit lies in the nature of the data collection problem itself. Many African languages are primarily oral, with limited written materials and an even smaller fraction of those digitally available (Mbogho et al., 2025) . As such, the web-scraping paradigms that enable large-scale corpus construction for high-resource languages do not transfer to this setting. Instead, building corpora for low-resource African languages typically  \nrequires manual elicitation from local speakers, followed by careful transcription and linguistic annotation (Nekoto et al., 2020) . At the scale demanded by modern ASR training, this process is neither financially nor logistically viable for extending ASR models to the hundreds of African languages lacking digital ","cbCaipPfrBW7h4Mz","https://ap.wps.com/l/cbCaipPfrBW7h4Mz","pdf",2508154,4,1,16,"English","en",105,"# Abstract\n# 1 Introduction\n## Motivation: data scarcity for African languages\n## Cross-lingual transfer and shared representations\n## Prior work on relatedness and adaptation in ASR","[{\"question\":\"Why is adapting ASR to low-resource African languages difficult?\",\"answer\":\"Many African languages are primarily oral, with limited written materials and few digitally available resources, making large-scale corpus construction costly and logistically infeasible.\"},{\"question\":\"What hypothesis does the document test regarding linguistic relatedness?\",\"answer\":\"It tests whether linguistic relatedness reliably predicts cross-lingual transfer gains when sequentially adapting on related auxiliary languages before adapting to the low-resource target.\"},{\"question\":\"What do the experiments conclude for large multilingual ASR models?\",\"answer\":\"Across conditions, pre-adaptation on related auxiliary languages yields no practically meaningful transfer improvements when target-language data is minimal.\"}]",1784183543,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"evaluating-the-effect-of-linguistic-relatedness-on-cross-lingual-transfer-in-large-multilingual-automatic-speech-recognition","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/evaluating-the-effect-of-linguistic-relatedness-on-cross-lingual-transfer-in-large-multilingual-automatic-speech-recognition/82866/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is adapting ASR to low-resource African languages difficult?","Question",{"text":75,"@type":76},"Many African languages are primarily oral, with limited written materials and few digitally available resources, making large-scale corpus construction costly and logistically infeasible.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What hypothesis does the document test regarding linguistic relatedness?",{"text":80,"@type":76},"It tests whether linguistic relatedness reliably predicts cross-lingual transfer gains when sequentially adapting on related auxiliary languages before adapting to the low-resource target.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the experiments conclude for large multilingual ASR models?",{"text":84,"@type":76},"Across conditions, pre-adaptation on related auxiliary languages yields no practically meaningful transfer improvements when target-language data is minimal.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]