[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117910-en":3,"doc-seo-117910-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117910,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Comparisons of the prediction models for undiagnosed diabetes between machine learning versus traditional statistical methods","A performance comparison evaluates machine learning–based and traditional statistics–based prediction models for undiagnosed type 2 diabetes in Korean adults. The study uses 2014–2020 KNHANES data (N=32,827): 2014–2018 for training and internal validation, and 2019–2020 for external validation. Model discrimination is assessed with ROC AUC. Feature sets using sex, age, resting heart rate, and waist circumference improve AUC relative to traditional methods, with higher gains under expanded anthropometric and lifestyle variables. Results indicate ML models using accessible non-invasive measurements may better identify undiagnosed diabetes earlier.","[www. nature.com/scientificreports](www. nature.com/scientificreports)  \nOPEN  \nComparisons ofthe prediction models for undiagnosed diabetes between machine learning versus traditional statistical methods  \nSeong Gyu Choi1,8, Minsuk Oh1,2,8, Dong–Hyuk Park1, Byeongchan Lee3, Yong‑ho Lee4, Sun Ha Jee5 & Justin Y. Jeon1,2,6,7*  \nWe compared the prediction performance of machine learning‑based undiagnosed diabetes prediction models with that of traditional statistics‑based prediction models. We used the 2014–2020 Korean National Health and Nutrition Examination Survey (KNHANES) (N = 32,827). The KNHANES 2014– 2018 data were used as training and internal validation sets and the 2019–2020 data as external validation sets. The receiver operating characteristic curve area under the curve (AUC) was used to compare the prediction performance of the machine learning‑based and the traditional statistics‑ based prediction models. Using sex, age, resting heart rate, and waist circumference as features, the machine learning‑based model showed a higherAUC (0.788 vs. 0.740) than that of the traditional statistical‑based prediction model. Using sex, age, waist circumference, family history of diabetes, hypertension, alcohol consumption, and smoking status as features, the machine learning‑based prediction model showed a higherAUC (0.802 vs. 0.759) than the traditional statistical‑based prediction model. The machine learning‑based prediction model using features for maximum prediction performance showed a higherAUC (0.819 vs. 0.765) than the traditional statistical‑based prediction model. Machine learning‑based prediction models using anthropometric and lifestyle measurements may outperform the traditional statistics‑based prediction models in predicting undiagnosed diabetes.  \nAbbreviations  \nWC  \nWHtR  \nRHR  \nDRSKNHANES ROC  \nAUC  \nML  \nTS  \nPPV  \nNPV  \nWaist circumference Waist to height ratio Resting heart rate Diabetes risk score  \nKorean National Health and Nutrition Examination Survey Receiver operating characteristic  \nArea under the ROC curve Machine learning Traditional statistics Positive predictive value Negative predictive value  \n1Department of Sports Industry Studies, Yonsei University, Seoul, Republic of Korea. 2Frontier Research Institute of Convergence Sports Science, Yonsei University, Seoul, Republic of Korea. 3Gauss Labs, Seoul, Republic of Korea. 4Department of Internal Medicine, Yonsei University College of Medicine, Seoul, Republic of Korea. 5Institute for Health Promotion, Graduate School of Public Health, Yonsei University, Seoul, Republic of Korea. 6Exercise Medicine Center for Diabetes and Cancer Patients, ICONS, Seoul, Republic of Korea. 7Cancer Prevention Center Shinchon Severance, Yonsei University College of Medicine, Shinchon-Dong, Seodaemun-Gu, Seoul 120-749, Republic of Korea. 8These authors contributed equally: Seong Gyu Choi and Minsuk Oh. *email: [jjeon@yonsei.ac.kr](jjeon@yonsei.ac.kr)  \n[www. nature.com/scientificreports/](www. nature.com/scientificreports/)  \nPLR  \nNLR  \nSHAP LightGBMXGBoost AdaBoost Bagging  \nPositive likelihood ratio Negative likelihood ratio  \nShapely additive explanation Light gradient boosting machine Extreme gradient boosting machine Adaptive boosting Bootstrapping and aggregating  \nThe Diabetes Fact Sheet in Korea 2020 from the Korean Diabetes Association reported that the prevalence of type 2 diabetes (hereafter “diabetes”) in Korean adults aged ≥ 30 years in 2018 was 13.8%(approximately 4.9 million)1. However, detecting diabetes is challenging, given the asymptomatic state at an early stage of diabetes. Consequently, many cases of diabetes are not diagnosed until after one’s diabetes complications have deteriorated2, and the optimal timing of diabetes treatment is often delayed3,4.  \nTherefore, it is imperative to identify an easy and accessible diabetes prediction at an early stage to effectively treat and manage diabetes and prevent its complications. Growing evidence has sugg","cbCaidRvWM5idP2p","https://ap.wps.com/l/cbCaidRvWM5idP2p","pdf",1499128,1,11,"English","en",105,"# Background\n# Methods\n## Data source and validation strategy\n## Features and model comparison\n# Results\n# Implications","[{\"question\":\"What data and validation approach were used to compare the models?\",\"answer\":\"The analysis used KNHANES 2014–2020 data (N=32,827). Training and internal validation used 2014–2018, while external validation used 2019–2020.\"},{\"question\":\"How was prediction performance evaluated?\",\"answer\":\"Receiver operating characteristic (ROC) AUC was used to compare discrimination between machine learning–based and traditional statistics–based models.\"},{\"question\":\"Which variables improved the machine learning model’s AUC compared with traditional statistics?\",\"answer\":\"Using sex, age, resting heart rate, and waist circumference increased AUC. Further gains were reported when adding family history of diabetes, hypertension, alcohol consumption, and smoking status, and when using features targeting maximum performance.\"}]","Comparisons of the prediction models for undiagnosed diabetes between machine learning versus traditional statistical methods | PDF",1785680336,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"comparisons-of-the-prediction-models-for-undiagnosed-diabetes-between-machine-learning-versus-traditional-statistical-methods","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/comparisons-of-the-prediction-models-for-undiagnosed-diabetes-between-machine-learning-versus-traditional-statistical-methods/117910/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What data and validation approach were used to compare the models?","Question",{"text":76,"@type":77},"The analysis used KNHANES 2014–2020 data (N=32,827). Training and internal validation used 2014–2018, while external validation used 2019–2020.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How was prediction performance evaluated?",{"text":81,"@type":77},"Receiver operating characteristic (ROC) AUC was used to compare discrimination between machine learning–based and traditional statistics–based models.",{"name":83,"@type":74,"acceptedAnswer":84},"Which variables improved the machine learning model’s AUC compared with traditional statistics?",{"text":85,"@type":77},"Using sex, age, resting heart rate, and waist circumference increased AUC. Further gains were reported when adding family history of diabetes, hypertension, alcohol consumption, and smoking status, and when using features targeting maximum performance.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]