[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-139514-105":59,"doc-detail-139514-en":131},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":124,"head_meta":126,"extra_data":128,"updated_unix":130},105,"en","supervised-machine-learning-for-speaker-diarization-by-pncc-with-lpcc-audio-coefficients","Supervised Machine Learning for Speaker Diarization by PNCC with LPCC Audio Coefficients","","Speaker diarization segregates a multi-speaker input into individual speaker signals while limiting leakage from other speakers. The study extracts audio features as a machine-learning training stage, then applies a classification stage to group those features for diarization. Linear Prediction Cepstral Coefficients (LPCC) and Power-Normalized Cepstral Coefficients (PNCC) are used independently and then recombined into a mixed feature representation. An improved Euclidean-distance strategy supports nearest-label identification. Using TIMIT audio for clustering two speakers, the method recovers female and male speech with low diarization error, achieving 1.8% (females), 2.9% (males), and 2.5% overall. Relative improvements over prior work reach 6.5%, 10%, and 8.8% respectively.",{"@graph":69,"@context":123},[70,84,106],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/supervised-machine-learning-for-speaker-diarization-by-pncc-with-lpcc-audio-coefficients/139514/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/supervised-machine-learning-for-speaker-diarization-by-pncc-with-lpcc-audio-coefficients/139514.png","ImageObject",300,407,{"name":92,"@type":93},"Chloe Bennett","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-10-04","2026-08-24",true,{"@type":102,"interactionType":103,"userInteractionCount":105},"InteractionCounter",{"@type":104},"ViewAction",11,{"@type":107,"mainEntity":108},"FAQPage",[109,115,119],{"name":110,"@type":111,"acceptedAnswer":112},"What problem does the paper address in speaker diarization?","Question",{"text":113,"@type":114},"It targets separating a mixed multi-speaker audio observation into individual speaker utterances while handling dialog and overlapped-speech cases.","Answer",{"name":116,"@type":111,"acceptedAnswer":117},"How are LPCC and PNCC used in the proposed method?",{"text":118,"@type":114},"The method computes features using LPCC and PNCC independently, then recombines them into a new mixed feature representation for diarization.",{"name":120,"@type":111,"acceptedAnswer":121},"What evaluation results are reported for clustering two speakers?",{"text":122,"@type":114},"Average diarization error rates are reported as 1.8% for females, 2.9% for males, and 2.5% overall, with relative improvements over other standard approaches up to 10% for males.","https://schema.org",{"og:url":83,"og:type":125,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":127,"canonical":83},"index,follow",{"doc_id":129,"site_id":62},139514,1787551591,{"code":4,"msg":5,"data":132},{"doc_id":129,"user_id":133,"nickname":92,"user_avatar":134,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":135,"file_id":136,"file_url":137,"file_type":138,"file_size":139,"view_count":105,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":44,"language":140,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":67,"update_tm":130,"read_time":144},962084925782,"https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc","[https://doi.org/10.33261/jaaru.2024.31.3.005](https://doi.org/10.33261/jaaru.2024.31.3.005) Association of Arab Universities Journal of Engineering Sciences (2024)31(3):37–45  \nSupervised Machine Learning for Speaker Diarization by PNCC with LPCC Audio Coefficients  \nHasan M. Kadhim 1, *, Alaa H. Ahmed 2, Alaa K. Hassan 3, and Saad T. Y. Alfalahi 4  \n1 Department of Electrical Engineering, College of Engineering, Mustansiriyah University, Baghdad, Iraq, [hasanalmgotir@uomustansiriyah.edu.iq](hasanalmgotir@uomustansiriyah.edu.iq)[ ](hasanalmgotir@uomustansiriyah.edu.iq)2 Department of Electrical Engineering, College of Engineering, Mustansiriyah University, Baghdad, Iraq, [alaa75hs@uomustansiriyah.edu.iq](alaa75hs@uomustansiriyah.edu.iq)  \n3 Department of Electrical Engineering, College of Engineering, Mustansiriyah University, Baghdad, Iraq, [alaak_eng@uomustansiriyah.edu.iq](alaak_eng@uomustansiriyah.edu.iq)  \n4 Department of Computer Engineering, MadenatAlelem University College, Baghdad, Iraq, [saad.t.yasin@mauc.edu.iq](saad.t.yasin@mauc.edu.iq)  \n* Corresponding author: Hasan M. Kadhim, and email: [hasanalmgotir@uomustansiriyah.edu.iq](hasanalmgotir@uomustansiriyah.edu.iq)[ ](hasanalmgotir@uomustansiriyah.edu.iq)[Published online: 30 September 2024](Published online: 30 September 2024)  \nAbstract— Speaker Diarization is a speech digital signal processing technique that segregates one input observation of n multi-speaker signal into an individual speech ofthose n persons. Each segregated signal belongs to one of them plus a bit of error, which is speech that belongs to other speakers. The format of that speech is a dialog because they speak non-simultaneously. By the use of speaker diarization in this research, audio features are extracted from speech. The extraction is the training stage of machine learning. The second classification stage can then decide how to divide these features into those n groups. Linear Prediction Cepstral Coefficients (LPCC) and Power-Normalized Cepstral Coefficients (PNCC) are used independently to generate their features. In this paper, the researchers re-combined these LPCC and PNCC features to form a new mixture of features. Improved Euclidian distance facilitates the job of measuring distances to identify who is the nearest label. Because PNCC is a non-inversible transformation, a small frame at the center of a large windowed frame has been regarded (because it has a reasonable weight) to obtain original speech signals. The procedure was efficient for clustering a mixture of two speaker signals, female and male from the TIMIT standard audio library, i.e., successfully recovered each person's individual speech. The average Diarization Error Rate (SDR) objective tests of the recovered speech were 1.8% for the females, 2.9% for the males, and 2.5% for the overall females and males. Compared with other standard research, the improvements were 6.5% for the females, 10% for the males, and 8. 8% for all females and males.  \nKeywords—Speaker Diarization, LPCC, PNCC, Clustering, Diarization Error Rate.  \n1. Introduction  \nSuppose there is the following spontaneous speech chat between Girl (G = White color rectangles in Figure 1) and Boy (B = Black color rectangles in Figure 1); first row in Figure 1: At first the Boy (B, the girl is silent) is speaking alone, then the Girl is speaking alone (G, the boy is silent), then both of them the girl with the boy are not speaking (there is a silent period (S) which is without rectangle in Figure 1), then both of them the girl with the boy are speaking simultaneously (Gray color (Gr) in the following figures), then there is a silent period, then the Girl is speaking (G), then both of them the girl with the boy are speaking somnolently (Gr overlapped speech between them), then the Boy (B) is speaking alone, and then both of them are speaking somnolently (Gray) . The sequence of  \nthe speech/ speaker is: B, G, S, Gr, S, G, Gr, B, and then Gr. When both of the two","cbCaifynvXKG2S0V","https://ap.wps.com/l/cbCaifynvXKG2S0V","pdf",2015625,"English","# Abstract\n# Keywords\n# Introduction\n## Problem framing: mixture, dialog, and overlapped speech\n## Overlapped-speech detection\n## Speech separation","[{\"question\":\"What problem does the paper address in speaker diarization?\",\"answer\":\"It targets separating a mixed multi-speaker audio observation into individual speaker utterances while handling dialog and overlapped-speech cases.\"},{\"question\":\"How are LPCC and PNCC used in the proposed method?\",\"answer\":\"The method computes features using LPCC and PNCC independently, then recombines them into a new mixed feature representation for diarization.\"},{\"question\":\"What evaluation results are reported for clustering two speakers?\",\"answer\":\"Average diarization error rates are reported as 1.8% for females, 2.9% for males, and 2.5% overall, with relative improvements over other standard approaches up to 10% for males.\"}]","Supervised Machine Learning for Speaker Diarization by PNCC with LPCC Audio Coefficients | PDF",23]