[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124437-en":3,"doc-seo-124437-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124437,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Voice-based prediction of prediabetes using classical machine learning models - Original research","Voice analysis offers a non-invasive and accessible path for screening, yet its transferability across populations remains unclear. This study evaluates whether voice-based machine learning models can identify individuals with prediabetes and how well the learned patterns generalize beyond the training context. Participants recorded a standard spoken phrase via a mobile app, while glycemic status was measured using HbA1c. Acoustic features were extracted and models were trained with feature selection and imbalance handling, then validated with cross-validation and independent testing.","TYPE Original Research PUBLISHED 27 November 2025 DOI 10.3389/fcdhc.2025.1697769  \nOPEN ACCESS  \nEDITED BY  \nValeria Grancini,  \nFondazione IRCCS Ca' Granda Ospedale Maggiore Policlinico, Italy  \nREVIEWED BY  \nAbir Elbeji,  \nLuxembourg Institute of Health, Luxembourg Ichwanul Muslim Karo Karo,  \nState University of Medan, Indonesia  \n*CORRESPONDENCE  \nJessica Oreskovic  \n[joreskovic@klick.com](joreskovic@klick.com)  \nRECEIVED 02 September 2025  \nREVISED 05 November 2025  \nACCEPTED 14 November 2025  \nPUBLISHED 27 November 2025  \nCITATION  \nOreskovic J, Fazli G, Varma V, Malik K, Kaufman J and Fossat Y (2025) Voice-based prediction of prediabetes using classical machine learning models.  \nFront. Clin. Diabetes Healthc. 6:1697769 .  \ndoi: 10.3389/fcdhc.2025.1697769  \nCOPYRIGHT  \n© 2025 Oreskovic, Fazli, Varma, Malik, Kaufman and Fossat. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY) . The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.  \nVoice-based prediction of prediabetes using classical machine learning models  \nJessica Oreskovic 1*, Ghazal Fazli 2, Vanita Varma 3, Kinza Malik 4, Jaycee Kaufman 1 and Yan Fossat 1  \n1 Klick Applied Sciences, Klick, Inc., Toronto, ON, Canada, 2 Department of Geography, Geomatics and the Environment, University of Toronto Mississauga, Mississauga, ON, Canada, 3Center for Innovation in Health and Wellness, Humber Polytechnic, Toronto, ON, Canada, 4 Faculty of Health &Life Sciences, Humber Polytechnic, Toronto, ON, Canada  \nIntroduction: Prediabetes is a highly prevalent metabolic condition that signiﬁcantly increases the risk of developing type 2 diabetes and cardiovascular disease. Despite its clinical importance, over 80% of individuals with prediabetes remain undiagnosed. Voice analysis has emerged as a non-invasive, accessible method for disease screening, with prior work showing promising results in detecting hypertension and type 2 diabetes from acoustic features. This study investigates whether voice-based machine learning models can identify individuals with prediabetes and evaluates the generalizability of these models across populations.  \nMethods: Participants were recruited from clinical sites in India and a community college in Canada. All participants recorded the same spoken phrase multiple times daily via a mobile app, and glycemic status was assessed using HbA1c levels. Voice recordings were preprocessed to remove silence and trimmed to exclude potentially uninformative sections. A total of 167 acoustic features were extracted from each sample using Librosa, scipy, and parselmouth. Features were averaged per participant. Sex-speciﬁc models were developed under six experimental conﬁgurations varying by dataset balance (age/BMI-matched vs. unbalanced) and BMI inclusion. Feature selection was conducted using L1-regularized logistic regression (LASSO), and SMOTE was applied during training to address class imbalance. Twelve machine learning classiﬁers were evaluated using leave-one-subject-out cross-validation (LOSO-CV) on the India dataset. Final models were tested on a holdout India subset and the independent Canada dataset.  \nResults: In cross-validation, the best female model (XGBoost, balanced, no BMI) achieved a balanced accuracy of 0.78, and the best male model (Random Forest, balanced, no BMI) achieved 0 .68. However, holdout set testing identiﬁed different optimal conﬁgurations for generalization: the male XGBoost model trained on an unbalanced dataset outperformed the cross-validated model. In the Canada dataset, models failed to generalize effectively, with several conﬁgurations unable to correctly identify prediabetic participant","cbCaipL5UxHhHniM","https://ap.wps.com/l/cbCaipL5UxHhHniM","pdf",485944,1,11,"English","en",105,"# Introduction\n## Prediabetes burden and need for screening\n# Methods\n## Participants and data collection\n## Feature extraction and preprocessing\n## Model building and evaluation\n# Results\n## Cross-validation performance\n## Holdout and independent dataset generalization\n# Discussion\n## Implications and limitations","[{\"question\":\"What is the main goal of this study on prediabetes prediction?\",\"answer\":\"To determine whether voice-based machine learning models can identify people with prediabetes and to assess whether those models generalize across populations.\"},{\"question\":\"How were voice data and prediabetes status obtained?\",\"answer\":\"Participants repeatedly recorded the same spoken phrase using a mobile app, and glycemic status was determined using HbA1c levels.\"},{\"question\":\"What factors affected model performance and generalizability?\",\"answer\":\"The study varied dataset balance and BMI inclusion, used feature selection with LASSO and SMOTE for class imbalance, and found that models often performed worse when tested across geographic or demographic boundaries, including failure to generalize effectively to the Canada dataset.\"}]","Voice-based prediction of prediabetes using classical machine learning models - Original research | PDF",1785822293,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"voice-based-prediction-of-prediabetes-using-classical-machine-learning-models-original-research","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/voice-based-prediction-of-prediabetes-using-classical-machine-learning-models-original-research/124437/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of this study on prediabetes prediction?","Question",{"text":75,"@type":76},"To determine whether voice-based machine learning models can identify people with prediabetes and to assess whether those models generalize across populations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How were voice data and prediabetes status obtained?",{"text":80,"@type":76},"Participants repeatedly recorded the same spoken phrase using a mobile app, and glycemic status was determined using HbA1c levels.",{"name":82,"@type":73,"acceptedAnswer":83},"What factors affected model performance and generalizability?",{"text":84,"@type":76},"The study varied dataset balance and BMI inclusion, used feature selection with LASSO and SMOTE for class imbalance, and found that models often performed worse when tested across geographic or demographic boundaries, including failure to generalize effectively to the Canada dataset.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]