[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121068-en":3,"doc-seo-121068-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121068,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Comparative analysis of machine learning models for breast cancer prediction and diagnosis - A dual-dataset approach","Breast cancer causes substantial mortality worldwide, and its complex nature creates major needs for rapid diagnosis and reliable prognosis. This study compares seven machine learning classifiers—logistic regression, support vector machine, k-nearest neighbor, decision tree, random forest, Naïve Bayes, and artificial neural network—using the Wisconsin breast cancer dataset and a broader breast cancer dataset. Results show distinct best-performing models per dataset and motivate robust preprocessing to address class imbalance and missing values. The dual-dataset design supports more dependable predictive assessment for clinical decision support.","Comparative analysis of machine learning models for breast cancer prediction and diagnosis: A dual-dataset approach  \nMuhammad Zeerak Awan1, Muhammad Shoaib Arif2,3, Mirza Zain Ul Abideen4,  \nKamaleldin Abodayeh2  \n1Centre for AI and Big Data, Namal University Mianwali, Mianwali, Pakistan 2Department of Mathematics and Sciences, College of Humanities and Sciences, Prince Sultan University, Riyadh, Saudi Arabia 3Department of Mathematics, Air University, PAF Complex E-9, Islamabad, Pakistan  \n4Department of Biotechnology, Quaid-i-Azam University, Islamabad, Pakistan  \nArticle history:  \nReceived Jan 22, 2024 Revised Feb 16, 2024 Accepted Mar 10, 2024  \nKeywords:  \nBreast cancer Data mining  \nDatasets Machine learning Model evaluation  \nCorresponding Author:  \nBreast cancer is ranked as a significant cause of mortality among females globally. Its complex nature poses principal challenges for physicians and researchers for rapid diagnosis and prognosis. Hence, machine learning algorithms are employed to forecast and identify diseases. This study discusses the comparative analysis of seven machine learning models, e.g., logistic regression (LR), support vector machine (SVM), k-nearest neighbor classifier (KNN), decision tree classifier (DT), random forest classifier (RF), Naïve Bayes (NB), and artificial neural network (ANN) to predict breast cancer using Wisconsin breast cancer and breast cancer datasets. In the Wisconsin breast cancer dataset, KNN depicted 99% accuracy, followed by RF (98%), SVM (96%), NB (96%), LR (96%), ANN (93%), and DT (92%) . On the contrary, in the breast cancer (BC) dataset, the highest accuracy was achieved by LR at 83%, and the lowest was achieved by DT (65%), which depicted that the numeric dataset WBC has better accuracy than the breast cancer dataset.  \nThis is an open access article under the CC BY-SA license.  \nMuhammad Shoaib Arif  \nDepartment of Mathematics and Sciences, College of Humanities and Sciences, Prince Sultan University Riyadh, 11586, Saudi Arabia  \n[Email: marif@psu.edu.sa](Email: marif@psu.edu.sa)  \nArticle Info ABSTRACT  \n1. INTRODUCTION  \nBreast cancer continues to be a significant global health issue and a leading cause of death among women globally. The cause of occurrence involves genetic and environmental factors [1] . 25% of hereditary cases are due to mutations affecting high penetrant genes, e.g., HER2, BRCA1, BRCA2, TP53, PTEN, CDH1, and STK11, and moderate penetrant genes, e.g., CHEK2, BRIP1, ATM, and PALB2 [2] . GLOBOCAN 2020 data depicts the estimated number of new cases in women as 2.3 million, with 6.9% mortality over the five years respectively [3] . According to the LLR (log-linear Regression) model, women above the age of 75 are more likely to get Breast Cancer, followed by women between the ages of 55 and 64 [4]-[6] .  \nThe complex phenomenon of a varied nature requires accurate and timely diagnostic and prognostic approaches. Thus, using machine learning (ML) approaches in the medical sector is essential to help forecast by analyzing and configuring data. Because of their strong classification results, many researchers utilize these algorithms to address complex problems [7] . Data mining using machine learning is being utilized in the clinical domains to arrange and comprehend extensive data more readily using a computer-assisted detection (CAD) system that employs machine learning techniques to give reliable Breast Cancer diagnosis [8].  \nLuckily, most Breast Cancer data is open source and available on data repositories for the medical research community, e.g., the Wisconsin Breast Cancer Dataset on Kaggle and the Breast Cancer Dataset on the UCI Machine Learning Repository. So, researchers are using them for analysis and prediction by applying Machine Learning algorithms [9] . So, researchers and physicians are using machine learning algorithms to develop effective predictive breast cancer detection and prognosis models. The Wisconsin (diagnostic) dataset was u","cbCaid622YpyXRTd","https://ap.wps.com/l/cbCaid622YpyXRTd","pdf",832287,1,13,"English","en",105,"# Article Info\n## Abstract\n## 1. Introduction\n## 2. Method","[{\"question\":\"Which machine learning models are compared in the study for breast cancer prediction and diagnosis?\",\"answer\":\"The study compares logistic regression, support vector machine, k-nearest neighbor, decision tree, random forest, Naïve Bayes, and artificial neural network.\"},{\"question\":\"How do the models perform differently on the Wisconsin dataset versus the breast cancer dataset?\",\"answer\":\"On the Wisconsin breast cancer dataset, KNN achieves the highest accuracy (about 99%), while on the breast cancer dataset the highest accuracy is reported for logistic regression (about 83%) and the lowest for decision tree (about 65%).\"},{\"question\":\"Why does the study emphasize preprocessing before model training?\",\"answer\":\"Preprocessing is used to handle dataset issues such as uneven class distributions and missing values, aiming to make model training robust and reduce bias.\"}]","Comparative analysis of machine learning models for breast cancer prediction and diagnosis - A dual-dataset approach | PDF",1785733562,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"comparative-analysis-of-machine-learning-models-for-breast-cancer-prediction-and-diagnosis-a-dual-dataset-approach","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/comparative-analysis-of-machine-learning-models-for-breast-cancer-prediction-and-diagnosis-a-dual-dataset-approach/121068/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which machine learning models are compared in the study for breast cancer prediction and diagnosis?","Question",{"text":75,"@type":76},"The study compares logistic regression, support vector machine, k-nearest neighbor, decision tree, random forest, Naïve Bayes, and artificial neural network.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the models perform differently on the Wisconsin dataset versus the breast cancer dataset?",{"text":80,"@type":76},"On the Wisconsin breast cancer dataset, KNN achieves the highest accuracy (about 99%), while on the breast cancer dataset the highest accuracy is reported for logistic regression (about 83%) and the lowest for decision tree (about 65%).",{"name":82,"@type":73,"acceptedAnswer":83},"Why does the study emphasize preprocessing before model training?",{"text":84,"@type":76},"Preprocessing is used to handle dataset issues such as uneven class distributions and missing values, aiming to make model training robust and reduce bias.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]