[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119327-en":3,"doc-seo-119327-105":29,"detail-sidebar-cat-0-en-105":89},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":20,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},119327,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",7,"Healthcare","Machine Learning and SHAP Interpretability for Chronic Disease Understanding - Diabetes Prediction - Results and Deployment","Non-communicable diseases, especially diabetes, remain a major global health challenge due to the complexity of medical data and the limitations of traditional approaches for prediction and management. This study applies machine learning to diabetes risk estimation and uses hyperparameter tuning for model development. SHAP is incorporated to interpret feature contributions and improve clinical insight. The workflow includes SMOTE for class imbalance, model evaluation with accuracy/precision/recall/F1, and a Streamlit web interface enabling real-time, explainable predictions for healthcare providers.","| Machine Learning and SHAP Interpretability for Chronic Disease Understanding\u003Cbr>Nnaemeka Charles Igwe, Khandaker Mamun Ahmed\u003Cbr>DAKOTA STATE UNIVERSITY |  |\n| --- | --- |\n\nIntroduction  \n• Non-communicable diseases (NCDs), such as diabetes, are major global health concerns influenced by various health parameters and lifestyle choices.  \n• Traditional methods struggle to efficiently predict and manage these conditions due to the complexity and diversity of medical data.  \n• There is a need to leverage machine learning algorithms and modern computational tools to accurately predict diabetes, improve diagnosis, and provide actionable insights for better healthcare outcomes.  \n• There are various health parameters and lifestyle choices responsible for diabetes.  \n• Key health parameters include age, insulin level, body mass index (BMI), and family history.  \nProblem Statement  \n• Our study predicts non-communicable diseases (NCDs), such as diabetes, by leveraging machine learning algorithms.  \n• We leverage hyperparameter tuning techniques for model development and SHapley Additive exPlanation (SHAP) for results interpretations.  \nProject Goals  \n• Train machine learning models for diabetes prediction  \n• Utilize SHAP for feature importance and model interpretability.  \n• Handle class imbalance using SMOTE.  \n• Identify key factors affecting diabetes risk.  \n• Develop a user-friendly web interface for healthcare providers using STREAMLIT.  \nSpecifications  \n• Data: We use the Pima Indian dataset that includes raw (diabetes.csv) and processed datasets (updated_diabetes_dataset.csv) and SMOTE to handle class imbalance.  \n• Key Tools: Python for scripting, Streamlit for building a web app interface, and SHAP for explaining black box AI model predictions.  \nMethodology  \n• Data Preparation: We use the PIMA Indian diabetes dataset, preprocess data and select feature using a correlation matrix.  \n• Model Training and Evaluation: We train different machine learning models (SVM, KNN, Logistic Regression(LR)), fine-tune the hyperparameters, perform cross-validation, and evaluate the models’ performance using metrics like accuracy, precision, recall, and F1-score.  \n• Deployment and Interpretability: We deploy a Streamlit-based web app for real-time predictions and utilize SHAP to explain feature importance and model decisions.  \n•  \n•  \n•  \n•  \nFig 1: The overall architecture of our proposed method  \nResults  \nFig 3. F-1 score, Recall, Accuracy, Precision of SVM, KNN, and LR models  \nFig 5: Individual feature’s contribution to the prediction.  \nFig 4: SHAP interpretation of features importance in predicting diabetes using the SVM model  \nFig 6: Compares the performance of LinearSVC, KNN, and LR after applying SMOTE  \n•  \n•  \n•  \n•  \n•  \n•  \n Discussion  \nFigure 2 and 3 show the comparison of SVM, KNN, and LR models shows that SVM and KNN consistently outperform LR in key metrics like accuracy, precision, recall, and F1-score.  \nSVM displayed the most robust performance overall, while KNN was competitive in precision and recall. LR, although less effective, demonstrated acceptable results in simpler scenarios.  \nGlucose is the most influential feature, with higher values strongly contributing to a positive diabetes prediction. Age and BMI are also significant predictors, where higher values generally indicate an increased risk. The plot visually distinguishes high (pink) and low (blue) feature values and their corresponding SHAP values, showing how individual features influence the model's predictions.  \nLinearSVC performs the best, with the highest Accuracy (0 .76) and F1-Score (0 .63) .  \nConclusion  \nThis study highlights the importance of explainability and user-centric tools in deploying AI for healthcare cyber-physical systems.  \nML models (SVM, KNN, and LR) were applied to predict diabetes using the Pima Indian Diabetes dataset. SVM performed best, achieving 76% accuracy and the highest F1-score, while KNN followed with 75% accurac","cbCaiurwFAhNXpAE","https://ap.wps.com/l/cbCaiurwFAhNXpAE","pdf",326858,1,"English","en",105,"# Introduction\n# Problem Statement\n# Project Goals\n# Specifications\n# Methodology\n## Data Preparation\n## Model Training and Evaluation\n## Deployment and Interpretability\n# Results\n# Discussion\n# Conclusion\n# Reference","[{\"question\":\"How does the study predict diabetes risk?\",\"answer\":\"It trains multiple machine learning models (SVM, KNN, and Logistic Regression) using the Pima Indian diabetes dataset and evaluates them with standard classification metrics.\"},{\"question\":\"Why is SHAP used in this project?\",\"answer\":\"SHAP explains model decisions by identifying individual feature contributions to diabetes predictions, improving interpretability of the otherwise black-box outputs.\"},{\"question\":\"How is class imbalance handled and how is the system deployed?\",\"answer\":\"SMOTE is applied during dataset preparation to balance classes. A Streamlit-based web app is deployed to provide real-time predictions with SHAP-based explanations.\"}]","Machine Learning and SHAP Interpretability for Chronic Disease Understanding - Diabetes Prediction - Results and Deployment | PDF",1785723726,3,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":84,"head_meta":86,"extra_data":88,"updated_unix":27},"machine-learning-and-shap-interpretability-for-chronic-disease-understanding-diabetes-prediction-results-and-deployment","",{"@graph":35,"@context":83},[36,52,66],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":28},"https://docshare.wps.com/document/healthcare/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/machine-learning-and-shap-interpretability-for-chronic-disease-understanding-diabetes-prediction-results-and-deployment/119327/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":60,"encodingFormat":59,"isAccessibleForFree":61,"interactionStatistic":62},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":63,"interactionType":64,"userInteractionCount":20},"InteractionCounter",{"@type":65},"ViewAction",{"@type":67,"mainEntity":68},"FAQPage",[69,75,79],{"name":70,"@type":71,"acceptedAnswer":72},"How does the study predict diabetes risk?","Question",{"text":73,"@type":74},"It trains multiple machine learning models (SVM, KNN, and Logistic Regression) using the Pima Indian diabetes dataset and evaluates them with standard classification metrics.","Answer",{"name":76,"@type":71,"acceptedAnswer":77},"Why is SHAP used in this project?",{"text":78,"@type":74},"SHAP explains model decisions by identifying individual feature contributions to diabetes predictions, improving interpretability of the otherwise black-box outputs.",{"name":80,"@type":71,"acceptedAnswer":81},"How is class imbalance handled and how is the system deployed?",{"text":82,"@type":74},"SMOTE is applied during dataset preparation to balance classes. A Streamlit-based web app is deployed to provide real-time predictions with SHAP-based explanations.","https://schema.org",{"og:url":50,"og:type":85,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":87,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":90},[91,95,99,103,108,113,116,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":92,"show_sort_weight":93,"slug":94},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":96,"show_sort_weight":97,"slug":98},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":100,"show_sort_weight":101,"slug":102},"Exam",70,"exam",{"id":104,"doc_module":4,"doc_module_name":45,"category_name":105,"show_sort_weight":106,"slug":107},5,"Comic",60,"comic",{"id":109,"doc_module":4,"doc_module_name":45,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":114,"slug":115},40,"healthcare",{"id":117,"doc_module":4,"doc_module_name":45,"category_name":118,"show_sort_weight":119,"slug":120},8,"Research & Report",30,"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":104,"slug":136},19,"General","general"]