[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128733-en":3,"doc-seo-128733-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128733,1099523882182,"Eliana","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","Comparative Analysis of Machine Learning Models for Chronic Disease Prediction - A Multimodel Study On Diabetes, Hypertension, and Stroke","This thesis applies supervised machine learning to predict three common chronic diseases—diabetes, hypertension, and stroke—using real-world clinical and health data. A complete modeling pipeline is built, including data preprocessing, feature engineering, transformation, collinearity reduction, and class balancing. Logistic Regression, KNN, Random Forest, and XGBoost are trained with cross-validation and holdout test sets. XGBoost achieves the best overall performance, notably recall and AUC, while Logistic Regression offers a strong interpretable baseline.","UCLA  \nUCLA Electronic Theses and Dissertations  \nTitle  \nComparative Analysis of Machine Learning Models for Chronic Disease Prediction: AMultimodel Study On Diabetes, Hypertension, and Stroke  \nPermalink  \n[https://escholarship.org/uc/item/0jq522zw](https://escholarship.org/uc/item/0jq522zw)  \nAuthor  \nWang, Jiaheng  \nPublication Date  \n2025  \nPeer reviewed|Thesis/dissertation  \n[eScholarship.org](eScholarship.org) Powered by the California Digital Library  \nUniversity of California  \nUNIVERSITY OF CALIFORNIA Los Angeles  \nComparative Analysis of Machine Learning Models for Chronic Disease Prediction:  \nA Multimodel Study On Diabetes, Hypertension, and Stroke  \nA thesis submitted in partial satisfaction of the requirements for the degree Master of Applied Statistics and Data Science  \nby  \nJiaheng Wang  \n2025  \n© Copyright by Jiaheng Wang 2025  \nABSTRACT OF THE THESIS  \nComparative Analysis of Machine Learning Models  \nfor Chronic Disease Prediction:  \nA Multimodel Study On Diabetes, Hypertension, and Stroke  \nby  \nJiaheng Wang  \nMaster of Applied Statistics and Data Science  \nUniversity of California, Los Angeles, 2025  \nProfessor Guang Cheng, Chair  \nThis paper exemplifies how supervised machine learning can be used to predict three common chronic diseases (diabetes, hypertension, and stroke) utilizing clinical and health data from the real world. A comprehensive modeling pipeline was developed, encompassing data preprocessing, feature engineering, transformation, collinearity reduction, and class balancing. Logistic Regression, KNN, Random Forest, and XGBoost are considered as modeling techniques and used both cross-validation and holdout test sets. XGBoost demonstrated superior performance compared to other models, especially recall and AUC, while Logistic Regression served as a strong, interpretable baseline. Hypertension models achieved near-perfect results—likely due to clear class separability confirmed by PCA—and tree-based models improved stroke prediction by capturing complex nonlinear relationships. Within the predictor variables, many variables appeared within the models (regardless of the disease) and demonstrated commonality, supporting the development of possible integrated screening and pragmatic public health intervention with evidence to show how the data drove decisions.  \nThe thesis of Jiaheng Wang is approved.  \nHongquan Xu Yingnian Wu Guang Cheng, Committee Chair  \nUniversity of California, Los Angeles 2025  \nTABLE OF CONTENTS  \n1 Introduction ...................................... 1  \n1.1 Background and Motivations ........................... 1  \n1.2 Research Objectives ................................ 2  \n2 Data Preprocessing and Exploratory Data Analysis ............. 4  \n2.1 Data Source .................................... 4  \n2.2 Variables Explanation .............................. 4  \n2.3 Feature Selection ................................. 5  \n2.3.1 Correlation Heatmap ........................... 6  \n2.3.2 Feature Importance ............................ 7  \n2.3.3 Full Feature Ranking ........................... 8  \n2.3.4 Feature Dropped ............................. 8  \n2.4 Skewness Check and Transformation ...................... 9  \n2.4.1 Log1p Transformation .......................... 10  \n2.4.2 Yeo-Johnson Transformation ....................... 10  \n2.4.3 Transformation Comparison ....................... 11  \n2.5 Variance Inflation Factor ............................. 12  \n3 Modeling and Evaluation .............................. 14  \n3.1 Logistic Regression ................................ 14  \n3.2 K-Nearest Neighbors ............................... 15  \n3.3 Random Forest .................................. 15  \n3.4 XGBoost ...................................... 16  \n3.5 Model Comparison ................................ 17  \n3.5.1 Model Performance Metrics ....................... 17  \n3.5.2 Problem facing and Model Selection ................... 18  \n3.5.3 Model Selection .......................","cbCairsbiIO7YSZe","https://ap.wps.com/l/cbCairsbiIO7YSZe","pdf",1231365,2,1,39,"English","en",105,"# Introduction\n## Background and Motivations\n## Research Objectives\n# Data Preprocessing and Exploratory Data Analysis\n## Data Source\n## Variables Explanation\n## Feature Selection\n## Skewness Check and Transformation\n## Variance Inflation Factor\n# Modeling and Evaluation\n## Logistic Regression\n## K-Nearest Neighbors\n## Random Forest\n## XGBoost\n## Model Comparison\n## Feature Importance\n# Conclusion\n## Summary\n## Future Work\n# References","[{\"question\":\"What diseases are modeled in this thesis?\",\"answer\":\"The thesis predicts diabetes, hypertension, and stroke using clinical and health data collected from real-world settings.\"},{\"question\":\"Which machine learning models are compared?\",\"answer\":\"Logistic Regression, KNN, Random Forest, and XGBoost are evaluated using both cross-validation and holdout test sets.\"},{\"question\":\"Why does XGBoost perform best?\",\"answer\":\"XGBoost shows superior recall and AUC, while the Hypertension results are near-perfect, attributed to clear class separability confirmed by PCA and tree-based models capturing complex nonlinear patterns for stroke prediction.\"}]","Comparative Analysis of Machine Learning Models for Chronic Disease Prediction - A Multimodel Study On Diabetes, Hypertension, and Stroke | PDF",1786002937,98,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"comparative-analysis-of-machine-learning-models-for-chronic-disease-prediction-a-multimodel-study-on-diabetes-hypertension-and-stroke","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,48,51],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":20},"https://docshare.wps.com/document/","Document",{"item":49,"name":12,"@type":44,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":44,"position":53},"https://docshare.wps.com/document/comparative-analysis-of-machine-learning-models-for-chronic-disease-prediction-a-multimodel-study-on-diabetes-hypertension-and-stroke/128733/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-24","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What diseases are modeled in this thesis?","Question",{"text":76,"@type":77},"The thesis predicts diabetes, hypertension, and stroke using clinical and health data collected from real-world settings.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which machine learning models are compared?",{"text":81,"@type":77},"Logistic Regression, KNN, Random Forest, and XGBoost are evaluated using both cross-validation and holdout test sets.",{"name":83,"@type":74,"acceptedAnswer":84},"Why does XGBoost perform best?",{"text":85,"@type":77},"XGBoost shows superior recall and AUC, while the Hypertension results are near-perfect, attributed to clear class separability confirmed by PCA and tree-based models capturing complex nonlinear patterns for stroke prediction.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]