[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122216-en":3,"doc-seo-122216-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122216,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","The Effect of Feature Selection Methods on Machine Learning Model Performance - A Comparative Study for Breast Cancer Prediction","Breast cancer remains a major health challenge in many developing countries, so early detection is critical for effective treatment. Machine learning models can estimate risk using regular diagnostic data, yet not all features contribute meaningfully to prediction. This study compares the impact of feature selection methods on model performance for breast cancer prediction. Seven algorithms (KNN, Naive Bayes, Decision Trees, SVM, Logistic Regression, Neural Network, Random Forest) are evaluated with F-test, Mutual Information, and Spearman correlation on the WDBC dataset from UCI, showing that feature selection improves accuracy, especially for Logistic Regression and Neural Network.","THE EFFECT OF FEATURE SELECTION METHODS ON MACHINE LEARNING MODEL PERFORMANCE: A COMPARATIVE STUDY FOR BREAST CANCER PREDICTION  \nDiman Siddiq Hassan  \nComputer Science Department, College of Science, University of Zakho, Zakho, Kurdistan Region, IraqCorresponding author [email:diman.hassan@uoz.edu.krd](email:diman.hassan@uoz.edu.krd)  \nReceived: 12 Nov 2024 / Accepted:11 Jan., 2025/ Published:13 feb., 2025. [https://doi.org/10.25271/sjuoz.2025.13.1.1429](https://doi.org/10.25271/sjuoz.2025.13.1.1429)  \nABSTRACT:  \nDeveloping countries often face a high incidence of breast cancer, making early detection vital for effective treatment. The risk of developing breast cancer can be evaluated using machine learning methods and regular diagnostic data. In cancer datasets, there is a wealth of patient information, but not all of it is valuable for predicting cancer. This highlights the significance of feature selection methods in uncovering the relevant data. In this field, many studies have attempted to predict the different types of breast tumours, since it is important to diagnose breast cancer medication accurately. This paper aims to perform a comparison such that to show the effect of different feature selection methods on the accuracy of various existing machine learning algorithms. The study focuses on seven machine learning algorithms: K-Nearest Neighbors (KNN), Naive Bayes (NB), Decision Trees (DT), Support Vector Machines (SVM), Logistic Regression (LR), Neural Network (NN), and Random Forest (RF) . The feature selection techniques examined include F-test Feature Selection, Mutual Information (MI), and Spearman Correlation Coefficient. The dataset used for the experiments is the Wisconsin Diagnostic Breast Cancer (WDBC) dataset, which is publicly available from the UCI Repository. The findings reveal that when feature selection is implemented, the LR and NN algorithms demonstrate superior accuracy and perform exceptionally well across other metrics compared to the other models.  \nKEYWORDS: Breast Cancer; Machine Learning; Feature Selection; Breast Cancer Diagnostic Dataset.  \n1. INTRODUCTION  \nCancer is one of the deadliest diseases in the world. The latest statistics about this disease were reported in 2023 (Zhou et al., 2024), listing ten types of cancers, including breast cancer diagnosed in women. Breast cancer has been and still is the most common type of cancer that has affected a high percentage of women around the world at approximately 31% . It is considered the first type of cancer that causes deaths in women and is ranked fifth in terms of all cancer deaths around the world. It was the reason for 685,000 deaths in 2020, and that number increased to around 963,000 deaths in 2021, exceeding lung cancer with approximately 2.3 million new cases of this disease, according to the World Health Organization (WHO) (Bray et al., 2024) . The percentage of these cancer cases was 25%, and the death cases among women were 17% around the world (Zhou et al., 2024) . The abnormal growth of the breast cell is called a tumour, which is divided into two types: malignant and benign. The former is cancerous, while the latter is non-cancerous. Despite the incomprehension of the causes of breast cancer in women, several factors and attributes were contributed as the reasons for this disease, such as family history, problems in the inside uterine environment, adolescent exposures, pregnancy problems, gene mutation, alcohol and tobacco consumption, and childbearing at advanced maternal ages, specifically in developing countries (Uddin et al., 2023) .  \nConsequently, to reduce the rate of breast cancer cases and to prevent mortality in women, it is important to make regular visits to health professionals for screening, treatment, and accurate examination in clinical health. However, misdiagnosis may occur, which reduces the opportunity for early recovery, or it may as well be that there is a shortage in the number of health experts. Also, ","cbCaioAUVzVQNbdl","https://ap.wps.com/l/cbCaioAUVzVQNbdl","pdf",712446,1,12,"English","en",105,"# Introduction\n## Breast cancer burden and need for early detection\n## Machine learning approaches for diagnosis\n## Study objective and methodology overview\n# Abstract\n## Dataset and feature selection techniques\n## Compared learning algorithms\n## Key findings","[{\"question\":\"Why is early breast cancer detection important in this study?\",\"answer\":\"Early detection enables more effective treatment and recovery, which is emphasized as vital because misdiagnosis and limited expert availability can delay outcomes.\"},{\"question\":\"Which dataset and feature selection methods are used?\",\"answer\":\"Experiments use the Wisconsin Diagnostic Breast Cancer (WDBC) dataset from the UCI Repository. Feature selection methods include F-test Feature Selection, Mutual Information, and Spearman Correlation Coefficient.\"},{\"question\":\"What machine learning algorithms are compared, and which performed best with feature selection?\",\"answer\":\"The study evaluates K-Nearest Neighbors, Naive Bayes, Decision Trees, Support Vector Machines, Logistic Regression, Neural Network, and Random Forest. With feature selection, Logistic Regression and Neural Network show superior accuracy and strong results across other metrics.\"}]","The Effect of Feature Selection Methods on Machine Learning Model Performance - A Comparative Study for Breast Cancer Prediction | PDF",1785809411,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-effect-of-feature-selection-methods-on-machine-learning-model-performance-a-comparative-study-for-breast-cancer-prediction","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-effect-of-feature-selection-methods-on-machine-learning-model-performance-a-comparative-study-for-breast-cancer-prediction/122216/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is early breast cancer detection important in this study?","Question",{"text":75,"@type":76},"Early detection enables more effective treatment and recovery, which is emphasized as vital because misdiagnosis and limited expert availability can delay outcomes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which dataset and feature selection methods are used?",{"text":80,"@type":76},"Experiments use the Wisconsin Diagnostic Breast Cancer (WDBC) dataset from the UCI Repository. Feature selection methods include F-test Feature Selection, Mutual Information, and Spearman Correlation Coefficient.",{"name":82,"@type":73,"acceptedAnswer":83},"What machine learning algorithms are compared, and which performed best with feature selection?",{"text":84,"@type":76},"The study evaluates K-Nearest Neighbors, Naive Bayes, Decision Trees, Support Vector Machines, Logistic Regression, Neural Network, and Random Forest. With feature selection, Logistic Regression and Neural Network show superior accuracy and strong results across other metrics.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]