[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128648-en":3,"doc-seo-128648-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},128648,962084925782,"Ava Thompson","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Cross Validation Machine Learning Model Predicts More Accurate - A Comparative Study of Heart Disease - Using Linear Regression, Support Vector Machine, K Neighbors and Random Forest Models","This primary research paper investigates how cross-validation improves machine learning model performance for heart disease prediction. The study contrasts train-test split evaluation with cross-validation, transforming datasets and applying LR, SVM, KNN, and Random Forest models. Accuracy and related metrics are computed to estimate average predictive performance and guide model selection. Results show strong cross-validated gains (about 5–13%) and identify logistic regression and KNN as top performers, with reported 81% accuracy, while linear regression yields the lowest accuracy among the four models. Random forest reaches an F1 score of 95%.","Cross Validation Machine Learning Model Predicts More Accurate: A Comparative Study of Heart Disease Using Linear Regression, Support Vector Machine, K Neighbors and  \nRandom Forest Models  \nYagyanath Rimal1 , Siddhartha Paudel 2 , Abeer Alsadoon3,4,5 , Madhav Prasad Koirala 1 ,  \nSumeet Gill6  \n1 Pokhara University, Nepal([rimal.yagya@gmail.com](rimal.yagya@gmail.com))  \n2 IOE, Pulchowk Campus, Patan, Nepal([paudelsiddhartha36@gmail.com](paudelsiddhartha36@gmail.com))  \n3 School of Computing Mathematics and Engineering, Charles Sturt University (CSU), School of Computer Data and Mathematical Sciences, Australia  \n4Western Sydney University (WSU), Sydney, Australia  \n5Asia Pacific International College (APIC), Sydney, Australia([alsadoon.abeer@gmail.com](alsadoon.abeer@gmail.com))  \n1 Pokhara University, Nepal ([mploirala@pu.edu.np](mploirala@pu.edu.np))  \n6 Maharshi Dayanand University Rohtak([drsumeetgill@mdrohtak.ac.in](drsumeetgill@mdrohtak.ac.in))  \nCorrespondence Author: Yagyanath Rimal, Pokhara University, Nepal,  \n[rimal.yagya@gmail.com](rimal.yagya@gmail.com)  \nAbstract: This primary research paper focuses on using cross-validation, where each iteration of test data is uniquely structured to ensure optimal model performance by combining weak learners for improved model final accuracy. In the machine learning process, data is commonly split into two sets: a training set comprising 70% of the data anda test set comprising 30% . Cross-validation is then utilized for training and evaluation, often involving reusing previous data sets. This research study transforms the original datasets and cross-validating comparative analysis using LR, SVM, KNN, and RF methodologies to predict heart disease. The  \nobjective is to easily identify the average accuracy of model predictions and subsequently make recommendations for model selection based on both cross-validated increased (5 to 13%) and non-cross-validated approaches. From comparing each model’s accuracy scores, it is found that the logistic regression and k-nearest neighbour models achieved the highest accuracy of 81% among the four models.  \nSimilarly, the random forest model attained an F1 score of 95%, indicating the highest accuracy score from the enhanced heart disease sample. These findings can be further corroborated using learning curve validation.  \nConversely, the linear regression model exhibited the lowest accuracy of 84% among the four machine learning models.  \nKeywords- Machine Learning, Crossvalidation, Accuracy-precision, Learning Curve, Health informatics, Bio-signal processing  \n1. Introduction  \nMachine learning involves crafting models based on training datasets, which are evaluated using testing datasets of unseen samples. While the train-test split is a common practice for dividing research datasets into training and testing sets, it is often less preferable for model prediction. Another option is splitting available data (training/testing) sets before with some ratio 70:30 splits where the programmer builds the model using training data and then whose value is further tested with test unseen datasets. This approach achieves greater accuracy than the initial option, but it might not be suitable if a student asks questions beyond the chapters taught to attain the highest grade. Cross-validation is a method of training a model by storing some portion of the sample data set of each split and the rest of the data set to train the model (Maldonado et al. ) . Similarly, the authors(Mahesh et al. )  \nexamined stratified cross-validation employed to split the data, ensuring a similar distribution of target outputs among prediction samples, thereby yielding the best average score. The holdout method functions by reserving a portion of the training dataset for model validation. In contrast, stratified nfold cross-validation effectively handles imbalanced datasets, ensuring each fold contains a proportional representation of each output class. Leave-p-out cross-v","cbCaijIU3FavnUq8","https://ap.wps.com/l/cbCaijIU3FavnUq8","pdf",671257,4,1,18,"English","en",105,"# Introduction\n## Data splitting strategies (train-test, holdout, stratified n-fold)\n## Cross-validation variants (leave-p-out, bias-variance balance)\n## Related methods and evaluation concepts\n# Abstract and objectives\n## Model comparison (LR, SVM, KNN, RF)\n## Accuracy results and selection recommendations","[{\"question\":\"How does cross-validation improve model performance in this study?\",\"answer\":\"Cross-validation iterates testing on uniquely structured data splits to obtain optimal model performance. It also enables more reliable evaluation than a single train-test split by reusing dataset portions across folds.\"},{\"question\":\"Which machine learning models are compared for heart disease prediction?\",\"answer\":\"The study compares Linear Regression (LR), Support Vector Machine (SVM), K Neighbors (KNN), and Random Forest (RF) using cross-validated comparative analysis.\"},{\"question\":\"What accuracy and performance results are reported for the four models?\",\"answer\":\"Logistic regression and KNN achieve the highest accuracy at 81% among the four models. Random forest reaches an F1 score of 95%, while linear regression shows the lowest accuracy at 84% in the reported comparison.\"}]","Cross Validation Machine Learning Model Predicts More Accurate - A Comparative Study of Heart Disease - Using Linear Regression, Support Vector Machine, K Neighbors and Random Forest Models | PDF",1786002292,45,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"cross-validation-machine-learning-model-predicts-more-accurate-a-comparative-study-of-heart-disease-using-linear-regression-support-vector-machine-k-neighbors-and-random-forest-models","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/cross-validation-machine-learning-model-predicts-more-accurate-a-comparative-study-of-heart-disease-using-linear-regression-support-vector-machine-k-neighbors-and-random-forest-models/128648/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-25","2026-08-06",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"How does cross-validation improve model performance in this study?","Question",{"text":76,"@type":77},"Cross-validation iterates testing on uniquely structured data splits to obtain optimal model performance. It also enables more reliable evaluation than a single train-test split by reusing dataset portions across folds.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which machine learning models are compared for heart disease prediction?",{"text":81,"@type":77},"The study compares Linear Regression (LR), Support Vector Machine (SVM), K Neighbors (KNN), and Random Forest (RF) using cross-validated comparative analysis.",{"name":83,"@type":74,"acceptedAnswer":84},"What accuracy and performance results are reported for the four models?",{"text":85,"@type":77},"Logistic regression and KNN achieve the highest accuracy at 81% among the four models. Random forest reaches an F1 score of 95%, while linear regression shows the lowest accuracy at 84% in the reported comparison.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]