[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125044-en":3,"doc-seo-125044-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125044,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Systematic Literature Review of Speaker Diarization Techniques - Toward Bridging Gaps in Low-resourced Languages using Machine Learning","Speaker diarization segments audio into speaker-specific regions to identify “who spoke when,” supporting speech technologies such as ASR, speaker verification, and conversational systems. Research progress benefits heavily from annotated resources, yet low-resourced languages remain underexplored, limiting both diarization performance and subsequent ASR advances. This study targets Sarawak Malay using crowd-sourced conversational data, where missing speaker turns and transcripts hinder acoustic modeling, and applies a systematic review of ML methods.","Systematic Literature Review of Speaker Diarization Techniques: Toward Bridging Gaps in Low-resourced Languages using Machine Learning  \nMohd Zulhafiz Rahim 1, Sarah Samson Juan 1,2* and Syahrul Nizam Junaini 1  \n1Faculty of Computer Science and Information Technology, Universiti Malaysia Sarawak, 94300 Kota Samarahan, Sarawak,  \nMalaysia  \n2Data Science Centre, Universiti Malaysia Sarawak, 94300 Kota Samarahan, Sarawak, Malaysia  \n*Corresponding author: [sjsflora@unimas.my](sjsflora@unimas.my)  \nSubmitted 21 September 2024, Revised 29 October 2024, Accepted 13 November 2024, Available online 02 January 2025. Copyright © 2025 The Authors.  \nAbstract: Speaker diarization, the process of segmenting audio into speaker-specific regions, plays a critical role in various speech technologies by determining \"who spoke when\" in a conversation. This technique is particularly valuable for enhancing automatic speech recognition (ASR) and conversational artificial intelligent systems. However, its application to lowresourced languages remains underexplored, limiting not only the performance of speaker diarization among low-resourced languages, but also stagnating the advancements of ASR to low-resourced languages. This is due to the fact that speaker diarization enables speaker adaptation in ASR, crucial for maximizing the performance of ASR itself. This lack of digital resources of speaker diarization to low-resourced languages, as well as the scarcity of its implementation presents a gap between low-resourced languages and popular languages in terms of the advancements of speech technologies involving the particular languages. This paper focuses on Sarawak Malay, a low-resourced language, and presents conversational data collected through a crowd-sourced approach, which needs speaker turns and transcripts. These missing annotations create challenges for building accurate acoustic models. To address this, we conducted a systematic review of recent speaker diarization research and related machine learning techniques. Using the PRISMA methodology, we reviewed 42 articles published between 2018 and 2023. Our findings identify key machine learning models, such as i-vectors and x-vectors, and open-source tools like Pyannote, which offer promising advancements in diarization performance. Besides that, these tools have shown potential to be implemented in developing speaker diarization models for low-resourced language. By highlighting the gaps in current research for low-resourced languages, we provide a pathway for improving speaker diarization models in these underrepresented languages through machine learning techniques.  \nKeywords: Deep neural network; Low-resourced; Machine learning; Speaker diarization; x-vectors.  \n1. INTRODUCTION  \nIn the era of machine learning and artificial intelligence, speech technologies such as automatic speech recognition (ASR) ([1],[2],[3]), speaker verification ([4–7]), and conversational systems [8] have seen significant advancements. These technologies rely heavily on accurately identifying speakers within an audio stream, a task referred to as speaker diarization. Speaker diarization involves segmenting audio into speaker-specific regions to determine \"who spoke when.\" This task is essential for applications requiring speaker-specific insights, such as in meeting transcription services, voice assistants, or forensic audio analysis [9] .  \nSpeaker diarization emerged in the 1990s, initially used to identify speakers in air traffic control communications [10] or broadcast news recordings [11,12]. The technology has since evolved, incorporating advanced statistical and machine learning techniques to improve accuracy. Traditional methods such as Gaussian Mixture Models (GMMs) and Hidden Markov Models (HMMs) laid the groundwork for speaker segmentation. Still, these techniques often needed to be improved with variability in speakers' voices and environmental noise. Over time, the introduction of i-vectors [1,13]","cbCaijCDhQSGq2Rf","https://ap.wps.com/l/cbCaijCDhQSGq2Rf","pdf",714554,1,15,"English","en",105,"# Introduction\n## Speaker diarization background\n## Motivation for low-resourced languages\n# Systematic review approach\n## PRISMA-based article selection (2018–2023)\n## Key models and tools","[{\"question\":\"What problem does the paper address for low-resourced languages in speaker diarization?\",\"answer\":\"It addresses the lack of diarization digital resources and limited implementation for low-resourced languages, which restricts model performance and slows related ASR progress.\"},{\"question\":\"Why does missing annotation like speaker turns and transcripts matter?\",\"answer\":\"Missing annotations create challenges for building accurate acoustic models because speaker turn boundaries and transcription alignment are essential for training and evaluation.\"},{\"question\":\"How was the literature reviewed and what time range was covered?\",\"answer\":\"The authors used PRISMA methodology and reviewed 42 articles published between 2018 and 2023.\"}]","Systematic Literature Review of Speaker Diarization Techniques - Toward Bridging Gaps in Low-resourced Languages using Machine Learning | PDF",1785896325,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"systematic-literature-review-of-speaker-diarization-techniques-toward-bridging-gaps-in-low-resourced-languages-using-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/systematic-literature-review-of-speaker-diarization-techniques-toward-bridging-gaps-in-low-resourced-languages-using-machine-learning/125044/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address for low-resourced languages in speaker diarization?","Question",{"text":75,"@type":76},"It addresses the lack of diarization digital resources and limited implementation for low-resourced languages, which restricts model performance and slows related ASR progress.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why does missing annotation like speaker turns and transcripts matter?",{"text":80,"@type":76},"Missing annotations create challenges for building accurate acoustic models because speaker turn boundaries and transcription alignment are essential for training and evaluation.",{"name":82,"@type":73,"acceptedAnswer":83},"How was the literature reviewed and what time range was covered?",{"text":84,"@type":76},"The authors used PRISMA methodology and reviewed 42 articles published between 2018 and 2023.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]