[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123057-en":3,"doc-seo-123057-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123057,8796095461610,"Oliver","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Advanced Machine Learning Techniques for Predicting Heart Disease - A Comparative Analysis Using the Cleveland Heart Disease Dataset","Heart disease prediction supports timely diagnosis and intervention, helping reduce adverse outcomes. This study evaluates multiple machine learning models on the Cleveland Heart Disease dataset, covering LSTM networks, Random Forest, Gradient Boosting, XGBoost, and Logistic Regression. Data preprocessing addresses missing values, categorical transformations, and target binarization. Model performance is measured with AUC-ROC, F1-score, recall, accuracy, and precision, while SHAP explains feature importance. Results indicate XGBoost achieves the highest accuracy (90%) and AUC-ROC (0.94).","Advanced Machine Learning Techniques for Predicting Heart Disease: A Comparative Analysis Using the Cleveland Heart Disease Dataset  \nDhadkan SHRESTHA  \nDepartment of Computer Science, Texas State University, 601 University Dr, San Marcos, TX 78666, United States  \n[Emails: gsu7@txstate.edu](Emails: gsu7@txstate.edu); [shresthadhadkan10@gmail.com](shresthadhadkan10@gmail.com)  \n* Author to whom correspondence should be addressed;  \nReceived: 11 July 2024/Accepted: 26 September 2024/ Published online: 29 September 2024  \nAbstract  \nThe ability to predict heart illness was essential for prompt diagnosis and treatment. Using the Cleveland Heart Disease dataset, this study tested a number of machine learning models, including LSTM networks, Random Forest, Gradient Boosting, XGBoost, and Logistic Regression. In order to handle missing values, transform categorical variables, and binarize the target variable, the dataset underwent pre-processing. AUC-ROC, F1-score, recall, accuracy, and precision were used to assess each model. SHAP values shed light on the significance of each characteristic. The results showed that XGBoost was the most accurate model, exceeding the other models with an accuracy of 90% and an AUC-ROC of 0.94. This study highlighted the potential of advanced machine learning techniques for improving heart disease prediction and contributed to the development of better diagnostic tools for patient care.  \nKeywords: Heart Disease Prediction; Machine Learning; XGBoost; Gradient Boosting; Long Short-Term Memory (LSTM); SHapley Additive exPlanations (SHAP)  \nIntroduction  \nHeart diseases have become the leading cause of death worldwide, taking hundreds of thousands of lives annually. Its early prediction will immensely reduce its prevalence and result in better outcomes by allowing early interventions [1] . Of late, with the development of machine learning and artificial intelligence, medicine-related diagnostics have opened up newer avenues for predictive analytics in healthcare [2] .  \nKnown benchmarks, one of which is the Cleveland Heart Disease dataset, provide a ground for testing machine-learning models concerning heart disease prediction [3] . The Cleveland Heart Disease dataset focuses on patients who undergo cardiac catheterization at the Cleveland Clinic Foundation. It includes both male and female patients, aged from 29 to 77 years old, showing different grades of heart disease risk factors. The data were collected between 1981 and 1984 by Robert Detrano, M.D., Ph.D., at the V.A. Medical Center, Long Beach, and Cleveland Clinic Foundation. This is a publicly available dataset from the UCI Machine Learning Repository [https://archive.ics.uci.edu/ml/datasets/heart](https://archive.ics.uci.edu/ml/datasets/heart) +disease. It includes a comprehensive set of features like age, sex, chest pain type, resting blood pressure, serum cholesterol, fasting blood sugar, resting electrocardiographic results, maximum heart rate achieved, exercise-induced angina, ST depression induced by exercise relative to rest, the slope of the peak exercise ST segment, the number of major vessels colored by fluoroscopy, and thalassemia. Such abroad feature space makes this dataset very suitable for training machine learning models, providing insight into their predictive capability [4, 5] .  \nIn this paper, I will implement several machine learning models and then compare their performance concerning heart disease prediction. These include traditional methods on the one hand, such as logistic regression, and advanced ones on the other hand, such as random forest, gradient boosting, and XGBoost, together with  \nLSTM networks. All these models have unique benefits: Logistic Regression confers simplicity and interpretability, while methods such as Random Forest and Gradient Boosting are based on ensemble methods with very complex interactions among features. It is because of their high performance and robustness that boosting techniques, spec","cbCaihhcLao79HAB","https://ap.wps.com/l/cbCaihhcLao79HAB","pdf",501444,1,12,"English","en",105,"# Abstract\n# Introduction\n## Cleveland Heart Disease dataset\n## Machine learning models and comparison focus","[{\"question\":\"Which machine learning models are compared for heart disease prediction?\",\"answer\":\"The study compares LSTM networks, Random Forest, Gradient Boosting, XGBoost, and Logistic Regression on the Cleveland Heart Disease dataset.\"},{\"question\":\"How is the dataset prepared before model training?\",\"answer\":\"The dataset undergoes preprocessing to handle missing values, transform categorical variables, and binarize the target variable.\"},{\"question\":\"How are model performances evaluated and interpreted?\",\"answer\":\"Performance is assessed using AUC-ROC, F1-score, recall, accuracy, and precision. SHAP values are used to interpret the contribution of each feature.\"}]","Advanced Machine Learning Techniques for Predicting Heart Disease - A Comparative Analysis Using the Cleveland Heart Disease Dataset | PDF",1785814435,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"advanced-machine-learning-techniques-for-predicting-heart-disease-a-comparative-analysis-using-the-cleveland-heart-disease-dataset","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/advanced-machine-learning-techniques-for-predicting-heart-disease-a-comparative-analysis-using-the-cleveland-heart-disease-dataset/123057/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which machine learning models are compared for heart disease prediction?","Question",{"text":75,"@type":76},"The study compares LSTM networks, Random Forest, Gradient Boosting, XGBoost, and Logistic Regression on the Cleveland Heart Disease dataset.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is the dataset prepared before model training?",{"text":80,"@type":76},"The dataset undergoes preprocessing to handle missing values, transform categorical variables, and binarize the target variable.",{"name":82,"@type":73,"acceptedAnswer":83},"How are model performances evaluated and interpreted?",{"text":84,"@type":76},"Performance is assessed using AUC-ROC, F1-score, recall, accuracy, and precision. SHAP values are used to interpret the contribution of each feature.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]