[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123495-en":3,"doc-seo-123495-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123495,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Hybrid Model for Speech Emotion Recognition using Mel-Frequency Cepstral Coefficients and Machine Learning Algorithms - Research Summary","Speech Emotion Recognition (SER) is a core topic in affective computing, aiming to infer human emotions from voice signals for more natural, empathetic interaction with users. Background noise, overlapping emotional cues, and speaker variability often degrade classification quality and limit practical deployment. This study proposes a lightweight hybrid SER model that uses Mel-Frequency Cepstral Coefficients (MFCC) for feature extraction and evaluates SVM, Decision Tree (DT), and K-Nearest Neighbors (KNN) classifiers. Experiments on the RAVDESS dataset with an 80/20 split show strong results, with KNN reaching 78.26% accuracy and 77.06% F1-score.","Hybrid Model for Speech Emotion Recognition using Mel-Frequency Cepstral Coefficients and Machine Learning Algorithms  \nOdi Nurdiawan*1, Dian Ade Kurnia2, Dadang Sudrajat3, Irfan Pratama4  \n1,2Informatics Management, STMIK IKMI Cirebon, Indonesia 3Informatics Engineering, STMIK IKMI Cirebon, Indonesia 4Information System, Mercu Buana University Yogyakarta, Indonesia  \n[Email:](Email:1odinurdiawan2020@gmail.com)[1](Email:1odinurdiawan2020@gmail.com)[odinurdiawan2020@gmail.com](Email:1odinurdiawan2020@gmail.com)  \nReceived : Jul 19, 2025; Revised : Sep 8, 2025; Accepted : Sep 23, 2025; Published : Oct 16, 2025  \nAbstract  \n\n| Speech Emotion Recognition (SER) is a subfield of affective computing that focuses on identifying human emotions through voice signals. Accurate emotion classification is essential for developing intelligent systems capable of interacting naturally with users. However, challenges such as background noise, overlapping emotional features, and speaker variability often reduce model performance. This study aims to develop a lightweight hybrid SER model by combining Mel-Frequency Cepstral Coefficients (MFCC) as feature representations with three machine learning algorithms: Support Vector Machine (SVM), Decision Tree (DT), and K-Nearest Neighbors (KNN) . The methodology involves audio data preprocessing, MFCC-based feature extraction, and classification using the selected algorithms. The RAVDESS dataset, consisting of 1,440 English-language audio samples across four emotions (happy, angry, sad, neutral), was used with an 80/20 train-test split to ensure class balance.. Experimental results show that the KNN model achieved the highest performance, with an accuracy of 78.26%, precision of 85.09%, recall of 78.26%, and F1-score of 77.06% . The Decision Tree model produced comparable results, while the SVM model performed poorly across all metrics. These findings demonstrate that the proposed hybrid approach is effective for recognizing emotions in speech and offers a computationally efficient alternative to deep learning models. The integration of MFCC features with multiple machine learning classifiers provides a robust framework for real-time emotion recognition applications, especially in environments with limited computing resources.\u003Cbr>Keywords : Affective Computing, Audio Classification, Decision Tree, K-Nearest Neighbors, MFCC, Speech Emotion Recognition. |\n| --- |\n| This work is an open access article and licensed under a Creative Commons Attribution-Non Commercial\u003Cbr>4.0 International License\u003Cbr> |\n\n1. INTRODUCTION  \nSpeech Emotion Recognition (SER) is a branch of emotional computing that focuses on the identification and classification of human emotions through voice signals. Emotions play a significant role in interpersonal communication, and the ability of computational systems to recognize these emotions enables more intelligent, empathetic, and human-like interactions between humans and machines [1], [2], [3] . SER technology typically employs feature extraction techniques such as MelFrequency Cepstral Coefficients (MFCC), pitch, and chroma to capture the acoustic and emotional patterns present in speech. MFCC is among the most commonly used features due to its ability to efficiently represent the spectral characteristics of the human voice. The implementation of SER technology spans multiple domains, including voice-based customer service systems, adaptive virtual assistants, learning systems that respond to students’ emotional states, and early detection of psychological conditions in the field of mental health[4], [5], [6], [7] .  \nDespite the rapid development of Speech Emotion Recognition (SER) technology, there are still several technical challenges that remain key concerns in its advancement. One of the main challenges is the presence of noise or acoustic disturbances from the surrounding environment, which can affect the  \nquality of voice signals and reduce classification accuracy. In additio","cbCaioVrlcIPJthc","https://ap.wps.com/l/cbCaioVrlcIPJthc","pdf",698059,1,16,"English","en",105,"# Introduction\n## Speech Emotion Recognition Background\n## Technical Challenges\n## Related Work and Motivation","[{\"question\":\"What problem does the study address in speech emotion recognition?\",\"answer\":\"It tackles performance loss caused by background noise, overlapping emotional features, and speaker variability that reduce accurate emotion classification from voice signals.\"},{\"question\":\"Which features and classifiers are used in the proposed hybrid model?\",\"answer\":\"The model extracts features using Mel-Frequency Cepstral Coefficients (MFCC) and classifies emotions using SVM, Decision Tree (DT), and K-Nearest Neighbors (KNN).\"},{\"question\":\"How were the experiments conducted and what dataset was used?\",\"answer\":\"Experiments used the RAVDESS dataset with 1,440 English audio samples across four emotions, using an 80/20 train-test split to maintain class balance.\"}]","Hybrid Model for Speech Emotion Recognition using Mel-Frequency Cepstral Coefficients and Machine Learning Algorithms - Research Summary | PDF",1785816851,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"hybrid-model-for-speech-emotion-recognition-using-mel-frequency-cepstral-coefficients-and-machine-learning-algorithms-research-summary","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/hybrid-model-for-speech-emotion-recognition-using-mel-frequency-cepstral-coefficients-and-machine-learning-algorithms-research-summary/123495/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-04",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the study address in speech emotion recognition?","Question",{"text":76,"@type":77},"It tackles performance loss caused by background noise, overlapping emotional features, and speaker variability that reduce accurate emotion classification from voice signals.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which features and classifiers are used in the proposed hybrid model?",{"text":81,"@type":77},"The model extracts features using Mel-Frequency Cepstral Coefficients (MFCC) and classifies emotions using SVM, Decision Tree (DT), and K-Nearest Neighbors (KNN).",{"name":83,"@type":74,"acceptedAnswer":84},"How were the experiments conducted and what dataset was used?",{"text":85,"@type":77},"Experiments used the RAVDESS dataset with 1,440 English audio samples across four emotions, using an 80/20 train-test split to maintain class balance.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":29,"slug":119},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]