[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124228-en":3,"doc-seo-124228-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124228,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Utilising Exploratory Data Analysis and Machine Learning Algorithms for Heart Disease Analysis and Prediction","Heart disease is a major and potentially fatal condition that requires early detection to support timely treatment and reduce risk. This research uses exploratory data analysis (EDA) and machine learning to study factors linked with heart disease and support risk mitigation. Key predictors identified include number of major arteries stained by fluoroscopy, chest pain forms, maximum heart rate, exercise-induced angina, peak exercise ST slope, and activity-related ST depression. Using the UCI heart disease dataset, six classifiers are compared, with Random Forest achieving the highest accuracy of 85.25%.","Utilising Exploratory Data Analysis and Machine Learning Algorithms for Heart Disease Analysis and Prediction  \nHumra Khan 1, P . Singh2  \n1,2Amity School of Engineering and Technology, Amity University (AUUP), Lucknow, (Uttar Pradesh), India.  \n[1](1 khanhumra024@gmail.com)[ khanhumra024@gmail.com](1 khanhumra024@gmail.com), [2](2 pawansingh51279@gmail.com)[ pawansingh51279@gmail.com](2 pawansingh51279@gmail.com)  \n\n| How to cite this paper: H. Khan and P. Singh, “Utilising Exploratory Data Analysis and Machine Learning Algorithms for Heart Disease Analysis and Prediction,” Journal of Management and Service Science (JMSS), Vol. 04, Iss. 01, S. No. 056, pp. 1-9, 2024.\u003Cbr>[https://doi.org/10.54060/a2zjourna](https://doi.org/10.54060/a2zjourna) |\n| --- |\n| ls.jmss.56\u003Cbr>Received: 09/06/2023\u003Cbr>Accepted: 10/03/2024\u003Cbr>Online First: 25/04/2024\u003Cbr>Published: 25/04/2024\u003Cbr>Copyright © 2024 The Author(s) . This work is licensed under the Creative Commons Attribution International License (CC BY 4.0) . [http://creativecommons.org/licens](http://creativecommons.org/licens) |\n| es/by/4 .0/\u003Cbr>  Open Access  |\n\nAbstract  \nAs one of the most common and potentially fatal diseases in the world, heart disease must be detected early for proper treatment. With exploratory data analysis (EDA) and machine learning algorithms for predictive analysis, this research project seeks to thoroughly investigate the different aspects that contribute to heart disease. This will enable prompt diagnosis and risk mitigation. Numerous crucial features affecting the diagnosis of heart disease have been found through in-depth exploratory analysis of data. Among these features, the number of major arteries stained by fluoroscopy, the various forms of chest pain, the maximum heart rate reached, exercise-induced angina, the slope of the peak exercise ST segment, and the ST depression brought on by activity relative to rest stand out as most significant factors. Clinicians can learn a great deal about a patient's risk of developing heart disease by carefully examining these characteristics. In order to put this research's predictive component into practice, machine learning classifiers are built using the UCI heart disease dataset, which contains important variables pertaining to cardiac health. For comparison analysis, six different methods are used: Random Forest (RF), Gradient Boost (GB), K-Nearest Neighbour (KNN), Decision Tree (DT), Support Vector Machine (SVM), and Logistic Regression (LR). After conducting a comprehensive analysis, it has been determined that the Random Forest classifier has the best accuracy rate, attaining a remarkable 85.25%.  \nKeywords  \nExploratory Data Analysis (EDA), Machine Learning, Heart Disease Analysis, Heart Disease Prediction  \n1. Introduction  \nEarly detection of heart disease is detrimental to lowering heart-related issues and safeguarding it from catastrophic risks. Heart disease symptoms can include physical weakness, breathing problems, chest pain, etc. An expert's symptom analysis report, physical laboratory results, and the patient's medical history are used in conjunction with invasive tests, to diagnose cardiac issues [1, 2] . Exploratory Data Analysis (EDA) aids in recognising patterns and highlights the most important characteristics of the data. EDA is ultimately used to check a theory or validate a claim [3] . Machine Learning (ML) along with exploratory data analysis aim to increase the effectiveness and accuracy of medical professionals' work [4] . ML approaches can help with clinical management in various medical applications, including tumour or cancer cell identification [5], natural language processing, and medical picture analysis [6, 7] . The percentage of mortality from heart disease has decreased because of these machine learning-based expert medical decision-making systems [8] .  \nThe objective of this paper is to analyse and visualise the heart disease dataset in order to better understand it and identif","cbCaiqLSjYcP5YAK","https://ap.wps.com/l/cbCaiqLSjYcP5YAK","pdf",570907,1,9,"English","en",105,"# Introduction\n## Objectives and problem framing\n## Dataset and machine learning approach\n# Literature Review\n## Prior work on health prediction and heart disease\n## Importance of preprocessing and data quality\n## Reported model accuracies and comparisons","[{\"question\":\"What is the main goal of this paper on heart disease analysis?\",\"answer\":\"To analyze and visualize the heart disease dataset, identify hidden trends and correlations, and build a prediction model that estimates a person’s risk using feature variables.\"},{\"question\":\"Which exploratory factors are highlighted as most significant for diagnosis?\",\"answer\":\"The study emphasizes features such as number of major arteries stained by fluoroscopy, chest pain types, maximum heart rate, exercise-induced angina, peak exercise ST slope, and ST depression caused by activity relative to rest.\"},{\"question\":\"How are prediction models evaluated and which method performs best?\",\"answer\":\"Performance is assessed using confusion matrix metrics including Accuracy, Precision, Recall, and F1-Score. Random Forest achieves the best accuracy rate at 85.25%.\"}]","Utilising Exploratory Data Analysis and Machine Learning Algorithms for Heart Disease Analysis and Prediction | PDF",1785821119,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"utilising-exploratory-data-analysis-and-machine-learning-algorithms-for-heart-disease-analysis-and-prediction","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/utilising-exploratory-data-analysis-and-machine-learning-algorithms-for-heart-disease-analysis-and-prediction/124228/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of this paper on heart disease analysis?","Question",{"text":75,"@type":76},"To analyze and visualize the heart disease dataset, identify hidden trends and correlations, and build a prediction model that estimates a person’s risk using feature variables.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which exploratory factors are highlighted as most significant for diagnosis?",{"text":80,"@type":76},"The study emphasizes features such as number of major arteries stained by fluoroscopy, chest pain types, maximum heart rate, exercise-induced angina, peak exercise ST slope, and ST depression caused by activity relative to rest.",{"name":82,"@type":73,"acceptedAnswer":83},"How are prediction models evaluated and which method performs best?",{"text":84,"@type":76},"Performance is assessed using confusion matrix metrics including Accuracy, Precision, Recall, and F1-Score. Random Forest achieves the best accuracy rate at 85.25%.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]