[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120637-en":3,"doc-seo-120637-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120637,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Predicting Diabetes Mellitus with Machine Learning Techniques - Research Study","This study tackles the accurate identification of diabetes mellitus using online-accessible and real-world diagnostic data. Machine learning models, including Support Vector Machine, Random Forest, Naïve Bayes, eXtreme Gradient Boosting, and Deep Neural Network, are trained and evaluated on the PIMA Indian Diabetes and NHANES 1999-2016 datasets. The workflow includes rigorous preprocessing: null value handling, outlier management, class-imbalance treatment, and normalization. Results show RF reaches 79% binary accuracy on PIMA with BORUTA-selected features, while XGBoost achieves 92% binary and 91% multiclass accuracy on NHANES 1999-2016.","Journal of Engineering Technology and Applied Physics  \nPredicting Diabetes Mellitus with Machine  \nLearning Techniques  \nTong Hau Lee*, Ng Hu and Harannesh Arul Ananthan  \nFaculty of Computing and Informatics, Multimedia University, 63100 Cyberjaya, Selangor, Malaysia.  \n*[Corresponding author:](Corresponding author: hltong@mmu.edu.my)[ hltong@mmu.edu.my](Corresponding author: hltong@mmu.edu.my), ORCiD: 0000-0002-3128-585X  \n[https://doi.org/10.33093/jetap.2024.6.1.12](https://doi.org/10.33093/jetap.2024.6.1.12)  \nManuscript Received: 6 October 2023, Accepted: 20 December 2023, Published: 15 March 2024  \nAbstract—This study addresses the challenge of accurately identifying diabetes mellitus in individuals. Utilizing accessible online and real-world diagnostic data, we employ machine learning models, including Support Vector Machine, Random Forest, Naïve Bayes, eXtreme Gradient Boosting, and Deep Neural Network, on the PIMA Indian Diabetes and NHANES 1999-2016 datasets. Rigorous data pre-processing steps were conducted, handling null values, outliers, and imbalanced data together with data normalization. Our results reveal that the RF model achieves a 79% accuracy for binary classification on the PIMA Indian Diabetes dataset, using a 60:40 train-test split with BORUTA selected features. Meanwhile, the XGBoost model excels on the NHANES 1999-2016 dataset, achieving 92% accuracy for binary and 91% for multiclass classification respectively.  \nKeywords—Diabetes Mellitus, Machine Learning, Accuracy  \nI. INTRODUCTION  \nDiabetes Mellitus (DM) poses a significant global health challenge, affecting millions and exhibiting a rising prevalence, leading to severe health consequences [1] . DM encompasses Type 1, Type 2, and Gestational Diabetes, each with distinct characteristics and impacts, demanding early detection and effective management [2] .  \nUtilizing Machine Learning (ML) techniques, the purpose of this research is to contribute to DM prediction based on patient medical data, ultimately advancing early intervention and patient care. The study addresses crucial questions, including feature relevance, model selection, and appropriate evaluation metrics. Expected outcomes encompass the identification of critical predictive features, a comparative assessment of various ML methods, validation of model performance using real-world datasets like PIMA Indian Diabetes and National Health and Nutrition Examination Survey  \n(NHANES) 1999-2016 datasets, and insights into factors impacting DM prediction.  \nThe project ’s scope involves an in-depth analysis of ML techniques and models, with a focus on feature relevance, model selection, and performance evaluation. This research will utilize real-world datasets, including PIMA Indian Diabetes and NHANES 1999-2016, to validate andrefine the predictive models. This research is motivated by the pressing need to enhance DM prevention, management, and patient outcomes. This research aims to develop accurate prediction models for the benefit of individuals and healthcare professionals alike.  \nThis research intended to find the specific attributes in the patient's medical data hold greater significance in predicting DM. Besides, among the diverse array of ML models presently accessible, pinpoint the suitable models for predicting DM compared to alternative existing approaches.  \nII. LITERATURE REVIEW  \nA comprehensive literature review was conducted to augment the project’s foundational knowledge. This review encompassed two critical aspects: first, understanding the global background and trends of DM, as discussed in the previous chapter; and second, examining prior research by various scholars involving DM prediction through ML techniques. This chapter provides a detailed account of the insights and findings derived from this extensive review.  \nA. Datasets Used  \nPrevious research on predicting DM using ML has leveraged various datasets. The widely adopted dataset in this domain is the PIMA India","cbCaipzkyf99XIkX","https://ap.wps.com/l/cbCaipzkyf99XIkX","pdf",519091,1,9,"English","en",105,"# Introduction\n## Diabetes Mellitus challenge and model motivation\n## Research questions and expected outcomes\n# Literature Review\n## Datasets Used\n## Data Preprocessing Methods Used","[{\"question\":\"Which datasets are used to predict diabetes mellitus in this study?\",\"answer\":\"The study uses the PIMA Indian Diabetes dataset and the NHANES 1999-2016 dataset, with additional references to other real-world datasets mentioned in the literature review.\"},{\"question\":\"What machine learning models are evaluated?\",\"answer\":\"Support Vector Machine, Random Forest, Naïve Bayes, XGBoost, and a Deep Neural Network are employed for prediction and comparison.\"},{\"question\":\"What preprocessing steps are applied before model training?\",\"answer\":\"The workflow includes handling null values, managing outliers, addressing imbalanced data, and applying data normalization.\"}]","Predicting Diabetes Mellitus with Machine Learning Techniques - Research Study | PDF",1785731032,23,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"predicting-diabetes-mellitus-with-machine-learning-techniques-research-study","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/predicting-diabetes-mellitus-with-machine-learning-techniques-research-study/120637/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which datasets are used to predict diabetes mellitus in this study?","Question",{"text":75,"@type":76},"The study uses the PIMA Indian Diabetes dataset and the NHANES 1999-2016 dataset, with additional references to other real-world datasets mentioned in the literature review.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What machine learning models are evaluated?",{"text":80,"@type":76},"Support Vector Machine, Random Forest, Naïve Bayes, XGBoost, and a Deep Neural Network are employed for prediction and comparison.",{"name":82,"@type":73,"acceptedAnswer":83},"What preprocessing steps are applied before model training?",{"text":84,"@type":76},"The workflow includes handling null values, managing outliers, addressing imbalanced data, and applying data normalization.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]