[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121250-en":3,"doc-seo-121250-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121250,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Application of machine learning in diabetes prediction based on electronic health record data analysis","Electronic health records enable predictive medicine through data-driven learning methods, and this study proposes an improved machine learning approach for diabetes risk prediction. The work refines an integrated model and validates its effectiveness using experimental results on a test dataset. Prediction accuracy reaches 77.7%, indicating strong generalization capability. Results demonstrate solid performance for diabetes prediction while highlighting opportunities to further enhance accuracy and reliability through future research.","Application of machine learning in diabetes prediction based on electronic health record data analysis  \nZihan Yang  \nSchool of Electronics and Computer Science, University of Southampton, SO17 1BJ Southampton, United Kingdom  \nAbstract. With the application of electronic health records (EHRs) in the medical field, the use of machine learning to predict disease has become oneof the important research hotspots in the healthcare industry. This study introduces an improved machine learning model specifically designed to predict diabetes risk, with the aim of improving the accuracy of predictions.  \nThe purpose of the study is not only to refine the model, but also to evaluate the performance of the model according to the experimental results. The integrated model was used in this experiment, and the prediction accuracy of diabetes reached 77.7%, showing strong generalization ability on the test data set. These results show that the model performs well at predicting diabetes, but there is still room for further improvement. While presenting the current research results, this study also Outlines future research directions, focusing on further improving the accuracy and reliability of the model. Th is research contributes to the development of machine learning in healthcare, specifically improving disease prediction models through advanced data analysis techniques.  \n1 Introduction  \nThe adoption of electronic health record (EHR) systems is becoming more widespread, and the use of machine learning in these systems has also grown significantly. This includes predicting the patient's condition, estimating the likelihood of disease, and identifying the numerous factors that have the greatest impact on the patient's health.  \nThis model could not only help patients self-assess their risk of disease, but also reduce the burden on doctors. Because the database contains a large amount of data based on clinical history and medical images, doctors can easily draw on the experience of previous cases to apply specific medical interventions, which means that it can improve the accuracy of diagnosis [1] .  \nWith relatively little patient data recorded, the quality of these models varies. This paper aims to explore the impact of various advanced machine learning models on the construction of disease data systems, with a primary focus on accuracy. In this study , taking diabetes asan example, the logistic regression method was used to select the most significant feature  \nCorresponding author: [zhy22zachary@gmail.com](zhy22zachary@gmail.com)  \n© The Authors, published by EDP Sciences. This is an open access article distributed under the terms of the Creative Commons Attribution License 4.0 ([https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/)).  \nvectors affecting the progression of diabetes. Modeling was performed using four key feature vectors, examining the application of deep learning models, neural network models, and integrated models in predicting the accuracy of diabetes outcomes, and ultimately selecting the best model.  \n2 Methods  \nThe data used in this experiment comes from the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK) dataset on Kaggle. All patients in the dataset are women aged 21 and above [2] .  \nIn this experiment, NLP (Natural Language Processing) techniques were first used to clean the data, ensuring that irrelevant data and noise in the raw text do not affect model performance.  \nThe data was structured by organizing the information into fixed fields (columns) such as age, BMI, and family history of diabetes. This ensured consistency and completeness in the data used by the model.  \nLogistic regression is used to select four different feature vectors from the data set. Logistic regression is usually easy to implement and intuitively understands the importance of features by estimating the size of the coefficients [3]. The higher the coefficient, the greater ","cbCaibk67HLlc8um","https://ap.wps.com/l/cbCaibk67HLlc8um","pdf",306790,1,6,"English","en",105,"# Introduction\n# Methods\n## Data source\n## Feature selection\n## Model evaluation and visualization\n# Experimental results","[{\"question\":\"What is the main goal of the study?\",\"answer\":\"The study aims to introduce and improve a machine learning model for predicting diabetes risk and to evaluate its performance through experiments.\"},{\"question\":\"What data source and preprocessing steps are used?\",\"answer\":\"The experiment uses the NIDDK dataset from Kaggle, applies NLP techniques to clean raw text data, and structures the data into fixed columns such as age, BMI, and family history.\"},{\"question\":\"How well does the model perform?\",\"answer\":\"The integrated model achieves 77.7% prediction accuracy on the test dataset, showing strong generalization, with scope for further improvement in accuracy and reliability.\"}]","Application of machine learning in diabetes prediction based on electronic health record data analysis | PDF",1785734620,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"application-of-machine-learning-in-diabetes-prediction-based-on-electronic-health-record-data-analysis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/application-of-machine-learning-in-diabetes-prediction-based-on-electronic-health-record-data-analysis/121250/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of the study?","Question",{"text":75,"@type":76},"The study aims to introduce and improve a machine learning model for predicting diabetes risk and to evaluate its performance through experiments.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What data source and preprocessing steps are used?",{"text":80,"@type":76},"The experiment uses the NIDDK dataset from Kaggle, applies NLP techniques to clean raw text data, and structures the data into fixed columns such as age, BMI, and family history.",{"name":82,"@type":73,"acceptedAnswer":83},"How well does the model perform?",{"text":84,"@type":76},"The integrated model achieves 77.7% prediction accuracy on the test dataset, showing strong generalization, with scope for further improvement in accuracy and reliability.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]