[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126157-en":3,"doc-seo-126157-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},126157,3985741905716,"Rowan","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",7,"Healthcare","Advanced Machine Learning Techniques for Cardiovascular Disease Risk Prediction","Cardiovascular diseases remain a leading global cause of mortality, requiring accurate predictive approaches for early detection and prevention. This study applies machine learning methods—Logistic Regression, K-Nearest Neighbors, Random Forest, and XGBoost—to estimate CVD risk using 69,997 observations with demographic, clinical, and lifestyle factors. Preprocessing uses one-hot encoding and feature scaling. Performance is assessed with accuracy, precision, recall, F1-score, and AUC-ROC. XGBoost achieves the best accuracy at 74%, while Random Forest reaches 73% and KNN performs at 66%, highlighting the value of ensemble models for clinical and public health decision-making.","University of Central Florida  \nSTARS  \nData Science and Data Mining  \nJanuary 2025  \nAdvanced Machine Learning Techniques for Cardiovascular Disease Risk Prediction  \nGodfred Ahenkroa Kesse  \nUniversity of Central Florida, [go262554@ucf.edu](go262554@ucf.edu)  \n Part of the Data Science Commons  \nFind similar works at: [https://stars.library.ucf.edu/data-science-mining](https://stars.library.ucf.edu/data-science-mining)  \nUniversity of Central Florida Libraries [http://library.ucf.edu](http://library.ucf.edu)  \nThis Article is brought to you for free and open access by STARS. It has been accepted for inclusion in Data Science and Data Mining by an authorized administrator of STARS. For more information, [please contact](please contact STARS@ucf.edu)[ STARS@ucf.edu](please contact STARS@ucf.edu).  \nSTARS Citation  \nKesse, Godfred Ahenkroa, \"Advanced Machine Learning Techniques for Cardiovascular Disease Risk Prediction\"(2025) . Data Science and Data Mining. 31.  \n[https://stars.library.ucf.edu/data-science-mining/31](https://stars.library.ucf.edu/data-science-mining/31)  \nAdvanced Machine Learning Techniques for Cardiovascular Disease Risk Prediction  \nGodfred Ahenkroa Kesse  \nDepartment of Statistics and Data Science  \nUniversity of Central Florida  \nAbstract—Cardiovascular diseases (CVDs) remain a leading global cause of mortality, necessitating advanced predictive models to aid early detection and prevention. This study explores the application of machine learning techniques, including Logistic Regression, K-Nearest Neighbors (KNN), Random Forest, and XGBoost, to predict CVD risk using a dataset of 69,997 observations encompassing demographic, clinical, and lifestyle factors. Data preprocessing involved one-hot encoding of categorical variables and scaling to ensure compatibility with all models. Model performance was evaluated using metrics such as accuracy, precision, recall, F1-score, and AUC-ROC. Among the models, XGBoost demonstrated the highest accuracy at 74%, leveraging its gradient-boosting framework to effectively handle feature interactions and imbalanced data. Random Forest, with an accuracy of 73%, provided insights into feature importance, highlighting systolic blood pressure and age as critical predictors. In contrast, KNN exhibited lower performance at 66%, attributed to its sensitivity to scaling and high-dimensional data. These fndings underscore the potential of ensemble methods like XGBoost and Random Forest in clinical decision-making and public health strategies for mitigating CVD risks.  \nIndex Terms—Cardiovascular Diseases (CVD), Machine Learning, Random Forest, XGBoost, AUC-ROC, Predictive Modeling, Feature Importance, Linear Classifers, Feature Independence, Precision, Accuracy, Recall, Gradient Boosting, Generalization.  \nI. INTRODUCTION CARDIOVASCULAR diseases (CVDs), according to  \nWorld Health Organization [1] are the leading cause of death globally, accounting for millions of deaths annually. These conditions, which include heart attacks, strokes, and heart failure, are caused by a combination of genetic, behavioral, and environmental factors. Key risk factors for CVDs include high blood pressure, high cholesterol levels, smoking, obesity, diabetes, and a sedentary lifestyle. Additionally, age and family history signifcantly contribute to susceptibility, with older individuals and those with a family history of CVDs at higher risk (Giovanni et al. 2020)[2] . Understanding these risk factors is critical for identifying individuals at risk of developing CVDs and implementing early intervention strategies.  \nWith the serious complications of CVDs such as heart failure, stroke, and peripheral artery disease, quality of life can be signifcantly deteriorated and increase healthcare costs. Identifying high-risk individuals is essential for preventing these adverse outcomes and reducing the global burden of CVDs. Researchers now increasingly rely on machine learning (ML) methods to classify whether an","cbCaijUsHK8cYfvC","https://ap.wps.com/l/cbCaijUsHK8cYfvC","pdf",3839101,5,1,11,"English","en",105,"# Abstract\n# Introduction\n## Cardiovascular disease burden and risk factors\n## Machine learning models for CVD classification\n## Logistic regression as a baseline\n## Random forest and feature importance\n## XGBoost and handling imbalanced data","[{\"question\":\"Which machine learning models are used to predict cardiovascular disease risk?\",\"answer\":\"The study evaluates Logistic Regression, K-Nearest Neighbors (KNN), Random Forest, and XGBoost.\"},{\"question\":\"What dataset size and features support the prediction task?\",\"answer\":\"The work uses 69,997 observations containing demographic, clinical, and lifestyle factors.\"},{\"question\":\"Which model performs best and how is performance evaluated?\",\"answer\":\"XGBoost achieves the highest accuracy (74%). Models are compared using accuracy, precision, recall, F1-score, and AUC-ROC.\"}]","Advanced Machine Learning Techniques for Cardiovascular Disease Risk Prediction | PDF",1785903455,28,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"advanced-machine-learning-techniques-for-cardiovascular-disease-risk-prediction","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/healthcare/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/advanced-machine-learning-techniques-for-cardiovascular-disease-risk-prediction/126157/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-24","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"Which machine learning models are used to predict cardiovascular disease risk?","Question",{"text":77,"@type":78},"The study evaluates Logistic Regression, K-Nearest Neighbors (KNN), Random Forest, and XGBoost.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"What dataset size and features support the prediction task?",{"text":82,"@type":78},"The work uses 69,997 observations containing demographic, clinical, and lifestyle factors.",{"name":84,"@type":75,"acceptedAnswer":85},"Which model performs best and how is performance evaluated?",{"text":86,"@type":78},"XGBoost achieves the highest accuracy (74%). Models are compared using accuracy, precision, recall, F1-score, and AUC-ROC.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,119,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":117,"slug":118},40,"healthcare",{"id":120,"doc_module":4,"doc_module_name":47,"category_name":121,"show_sort_weight":122,"slug":123},8,"Research & Report",30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":20,"slug":139},19,"General","general"]