[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122079-en":3,"doc-seo-122079-105":29,"detail-sidebar-cat-0-en-105":89},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":20,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},122079,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",7,"Healthcare","Predicting the Risk of Getting a Stroke Using a Machine Learning Tool","Predicting stroke risk by leveraging Python-based machine learning to analyze a structured stroke dataset. The workflow applies data preprocessing steps such as cleaning missing values, standardizing integer variables, and categorizing continuous features, followed by exploratory data analysis using correlation matrices. Age, hypertension, and heart disease emerge as the strongest correlated features with stroke incidence. The study then trains and tests selected algorithms to output interpretable predictions supported by graphical representations.","| PREDICTING THE RISK OF\u003Cbr>GETTING A\u003Cbr>STROKE USING A MACHINE\u003Cbr>LEARNING TOOl\u003Cbr>MADE BY MARK AGYAPONG |  |  |  | DATA PREPROCESSING\u003Cbr>The process of converting data to a common format to enable users to process and analyze it.\u003Cbr> Data cleaning-Removing all rows with missing values\u003Cbr> Data standardization-Converting integer values to categorical or binanary\u003Cbr> Data catergorisation-Converting continuous into categorical data\u003Cbr>Image 1. Raw data of stroke dataset\u003Cbr>\u003Cbr>\u003Cbr>Image 2. After data preprocessing\u003Cbr> |  |  |\n| --- | --- | --- | --- | --- | --- | --- |\n|  | INTRODUCTION\u003Cbr>Over the years, there have been many ways used in predicting different health outcomes. The introduction of machine learning in healthcare provides a more accurate means of measuring these outcomes using artificial intelligence. Stroke is a disease that has constantly caused a lot of deaths and chronic illness causing governments billions of dollars. For this reason, there have been numerous studies done in predicting stroke. The main goal of t he research is to explore and predict stroke occurrences in the dataset using machine learning tool Python. Python, which is a collection of machine learning algorithms is a very efficient tool that is used for data mining activities (Coursera, 2022) . It is important to note that the dataset being used has previously been analyzed using machine learning by several researchers. |  |  |  |  |  |\n|  | OBJECTIVE\u003Cbr>The main goal of t he research is to predict the risk of a person getting a stroke using a machine learning tool. Using different portions and attributes of the dataset understudy, this will train them to determine which of the attributes are strongly associated with an increased risk of getting a stroke. These will be done by classifying and clustering different aspects of the dataset sets and also most importantly using graphical representations to have a clear overview of both increased and decreased risks. Therefore, the project objectives formulated are:\u003Cbr> To determine the risk factors associated with stroke using Python algorithms.\u003Cbr> Compare the accuracy of risks level with or without feature selection using the attributes.\u003Cbr>METHODOLOGY\u003Cbr>The chosen programming language for this thesis is Python. The goal is to utilize machine learning algorithms to predict the likelihood of stroke occurrence. The predicted outcomes will be visually presented in graphs to facilitate clear interpretation. The machine learning that will be used in thesis will be:\u003Cbr> Decision tree\u003Cbr> Naive bayes\u003Cbr> Linear regression\u003Cbr> Support Vector Machine\u003Cbr>Dataset on stroke information was taken from Kaggle, a specialized online platform that is built for data scientists and machine learning lovers. The dataset used contains 5110 rows with 12 attributes.\u003Cbr> Patient ID\u003Cbr> Gender  Age |  |  | RESEARCH / FINDINGS\u003Cbr>This section presents the results of the exploratory data analysis ( EDA) . The algorithms will be used after the EDA.\u003Cbr>CORRELATION MATRIX\u003Cbr>A correlation matrix created to show the correrlation coeffiectuens between attributes. This will idenify any strong correlated attributes. .\u003Cbr>AGE\u003Cbr>The correlation matrix shows a strong 23% correlation between age and stroke, suggesting that older people are more likely to experience a stroke.\u003Cbr>HYPERTENSION\u003Cbr>The correlation matrix shows an 14% correlation between hypertension and stroke, indicating that hypertension plays a role in increasing the risk of stroke.\u003Cbr>HEART DISEASE\u003Cbr>The correlation matrix indicates an 14% correlation between heart disease and stroke, suggesting that heart disease is a factor that may increase the risk of stroke.\u003Cbr>CONCLUSION\u003Cbr>The main goal of t his research is to predict the risk of stroke using Python algorithms. The dataset contains 12 attributes that underwent preprocessing steps, including data cleaning, standardization, and categorization, before conducting exploratory data analysis ( EDA) . The EDA revealed ","cbCaifMNiWCrPPuI","https://ap.wps.com/l/cbCaifMNiWCrPPuI","pdf",407088,1,"English","en",105,"# Introduction\n# Objective\n# Methodology\n# Research / Findings\n## Correlation matrix\n# Conclusion","[{\"question\":\"What is the main goal of the stroke prediction study?\",\"answer\":\"To predict an individual’s risk of getting a stroke using a machine learning tool in Python, and to identify attributes strongly associated with higher risk.\"},{\"question\":\"Which machine learning algorithms are included in the methodology?\",\"answer\":\"Decision tree, Naive Bayes, Linear regression, and Support Vector Machine are used to model stroke likelihood.\"},{\"question\":\"What preprocessing and analysis steps are applied before training?\",\"answer\":\"The dataset undergoes cleaning, standardization, and categorization, then exploratory data analysis is performed using a correlation matrix to find strongly related features.\"}]","Predicting the Risk of Getting a Stroke Using a Machine Learning Tool | PDF",1785808720,3,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":84,"head_meta":86,"extra_data":88,"updated_unix":27},"predicting-the-risk-of-getting-a-stroke-using-a-machine-learning-tool","",{"@graph":35,"@context":83},[36,52,66],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":28},"https://docshare.wps.com/document/healthcare/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/predicting-the-risk-of-getting-a-stroke-using-a-machine-learning-tool/122079/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":60,"encodingFormat":59,"isAccessibleForFree":61,"interactionStatistic":62},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":63,"interactionType":64,"userInteractionCount":4},"InteractionCounter",{"@type":65},"ViewAction",{"@type":67,"mainEntity":68},"FAQPage",[69,75,79],{"name":70,"@type":71,"acceptedAnswer":72},"What is the main goal of the stroke prediction study?","Question",{"text":73,"@type":74},"To predict an individual’s risk of getting a stroke using a machine learning tool in Python, and to identify attributes strongly associated with higher risk.","Answer",{"name":76,"@type":71,"acceptedAnswer":77},"Which machine learning algorithms are included in the methodology?",{"text":78,"@type":74},"Decision tree, Naive Bayes, Linear regression, and Support Vector Machine are used to model stroke likelihood.",{"name":80,"@type":71,"acceptedAnswer":81},"What preprocessing and analysis steps are applied before training?",{"text":82,"@type":74},"The dataset undergoes cleaning, standardization, and categorization, then exploratory data analysis is performed using a correlation matrix to find strongly related features.","https://schema.org",{"og:url":50,"og:type":85,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":87,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":90},[91,95,99,103,108,113,116,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":92,"show_sort_weight":93,"slug":94},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":96,"show_sort_weight":97,"slug":98},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":100,"show_sort_weight":101,"slug":102},"Exam",70,"exam",{"id":104,"doc_module":4,"doc_module_name":45,"category_name":105,"show_sort_weight":106,"slug":107},5,"Comic",60,"comic",{"id":109,"doc_module":4,"doc_module_name":45,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":114,"slug":115},40,"healthcare",{"id":117,"doc_module":4,"doc_module_name":45,"category_name":118,"show_sort_weight":119,"slug":120},8,"Research & Report",30,"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":104,"slug":136},19,"General","general"]