[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83815-en":3,"doc-seo-83815-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83815,5909877438554,"Maeve","https://ap-avatar.wpscdn.com/avatar/5600025385ad2bf12a7?_k=1778553567797529272",8,"Research & Report","Towards Digital Preservation of Efik TTS for a Low-Resource African Language","Efik, a tonal language spoken in Southeastern Nigeria by about 1.5 million native speakers and 3 million second-language speakers, is largely missing from speech synthesis research. This work introduces the first documented end-to-end text-to-speech study for Efik, building a curated single-speaker corpus of 2,632 utterances (three hours) and evaluating four neural TTS models under low-resource conditions. Native speakers rate outputs using MOS, Nat-MOS, and A-MOS; MMS-TTS attains the best MOS (3.80 ± 0.63) and yields more stable long-form speech, while tonal errors remain. Results provide a reproducible baseline and motivate tone-aware modeling.","Towards Digital Preservation of Efik: TTS for a Low-Resource African  \nLanguage  \nOffiong Bassey Edet 1 ,3 ,∗∗, Emmanuel Oyo-Ita 1 , Archibong OkonArchibong 2 , David Effanga Bassey 2,  \nMbuotidem Sunday Awak 3  \n1 University of Cross River State, Nigeria  \n2 University of Calabar, Nigeria  \n3 ML Collective  \n[offiongbassey99@gmail.com](offiongbassey99@gmail.com) , [emmanueloyoita@unicross.edu.ng](emmanueloyoita@unicross.edu.ng) ,  \n[archibongarchibong54@gmail.com](archibongarchibong54@gmail.com) , [daviba231@gmail.com](daviba231@gmail.com) , [mbuotidemawak@gmail.com](mbuotidemawak@gmail.com)  \narXiv :2607 .045 15v 1 [ cs .CL] 5 Jul 2026  \nAbstract  \nEfik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresented in speech synthesis research. We present the first documented end-to-end text-tospeech study for Efik, introducing a curated single speaker corpus of 2,632 utterances totaling three hours and a comparative evaluation of four neural models (VITS, MMS-TTS, SpeechT5, and Orpheus-TTS) under low resource conditions. Native speakers evaluated the systems using MOS, Nat-MOS, and A-MOS. MMS-TTS achieved the highest MOS of 3.80 ± 0.63 and produced more stable long form speech, though tonal errors persisted. Other models showed greater tonal and prosodic inconsistencies. These results provide a reproducible baseline and highlight the need for larger corpora and tone aware modeling for tonal African languages.  \nIndex Terms: text-to-speech, low-resource languages, Efik language, tonal languages, speech synthesis, African languages  \n1. Introduction  \nText-to-Speech (TTS) technology has made remarkable progress in producing speech that approaches human naturalness in high-resource languages [1, 2] . Modern end-to-end architectures, often combining attention mechanisms with neural vocoders, achieve high-quality synthesis when trained on largescale paired text–audio corpora. However, these successes rely on hundreds of hours of curated speech data, a requirement that remains a major obstacle for low-resource languages, where such datasets are scarce or entirely unavailable [3, 4, 5] .  \nEfik, a Lower Cross language spoken in Southeastern Nigeria, exemplifies this challenge. Despite an estimated 3 million second speakers and 1.5 million native speakers [6], Efik lacks publicly available speech datasets suitable for supervised TTStraining and remains largely absent from contemporary speech technology pipelines, reflecting broader patterns of digital language inequality [7] . While text-based NLP tasks such as machine translation for Efik language have been receiving attention, speech systems remain critically underdeveloped [8, 9] .  \nDeveloping TTS for Efik presents several technical challenges. First, severe data scarcity constrains the training of dataintensive neural architectures [10] . Second, Efik is a tonal language in which lexical and grammatical contrasts are encoded through pitch variation [11] . Accurate synthesis therefore requires careful modeling of tonal contours in addition to segmental phonetic structure. Inadequate tone realization can alter  \n**indicates the corresponding author.  \nlexical meaning, reducing intelligibility and perceived naturalness.  \nTo investigate the feasibility of neural speech synthesis under such constraints, we curated a high-quality Efik speech corpus consisting of approximately three hours of single-speaker recordings captured with a wireless microphone in controlled conditions. The dataset comprises 2,632 utterances designed to provide phonetic and tonal coverage for supervised training. Operating within this limited-data regime, we finetune and evaluate four neural TTS models to assess naturalness and intelligibility in synthesized Efik speech.  \nThis research addresses these critical gaps by developing TTS systems specifically tailored for Efik, contributing to the broader effort of digital lan","cbCaivhBiDPEsNa5","https://ap.wps.com/l/cbCaivhBiDPEsNa5","pdf",210126,5,1,6,"English","en",105,"# Abstract\n# Introduction\n# Efik Language Documentation and Computational Efforts","[{\"question\":\"What is the main goal of the study on Efik text-to-speech?\",\"answer\":\"To develop and evaluate neural end-to-end TTS for Efik under low-resource conditions, starting from a curated single-speaker corpus and testing multiple neural TTS models.\"},{\"question\":\"What dataset is introduced for training and evaluation?\",\"answer\":\"A curated single-speaker Efik corpus containing 2,632 utterances totaling about three hours, designed to cover phonetic and tonal variation for supervised training.\"},{\"question\":\"Which neural TTS model performed best and what limitation remains?\",\"answer\":\"MMS-TTS achieved the highest MOS (3.80 ± 0.63) and produced more stable long-form speech; however, tonal errors persisted and affected tonal correctness and naturalness.\"}]",1784190598,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"towards-digital-preservation-of-efik-tts-for-a-low-resource-african-language","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/towards-digital-preservation-of-efik-tts-for-a-low-resource-african-language/83815/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the main goal of the study on Efik text-to-speech?","Question",{"text":76,"@type":77},"To develop and evaluate neural end-to-end TTS for Efik under low-resource conditions, starting from a curated single-speaker corpus and testing multiple neural TTS models.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What dataset is introduced for training and evaluation?",{"text":81,"@type":77},"A curated single-speaker Efik corpus containing 2,632 utterances totaling about three hours, designed to cover phonetic and tonal variation for supervised training.",{"name":83,"@type":74,"acceptedAnswer":84},"Which neural TTS model performed best and what limitation remains?",{"text":85,"@type":77},"MMS-TTS achieved the highest MOS (3.80 ± 0.63) and produced more stable long-form speech; however, tonal errors persisted and affected tonal correctness and naturalness.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]