[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119265-en":3,"doc-seo-119265-105":29,"detail-sidebar-cat-0-en-105":89},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":20,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},119265,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Comparative Analysis of Feature Selection and Machine Learning Models for Breast Cancer Risk Prediction","Comparative analysis develops accurate and efficient breast cancer risk prediction models by examining feature selection strategies alongside multiple machine learning classifiers. The study uses a breast cancer dataset from the UCI Machine Learning Repository, applies data cleaning and one-hot encoding, and then evaluates correlation-based feature selection, PCA dimensionality reduction, and recursive feature elimination. Model performance is measured with accuracy, precision, recall, and F1-score, showing that feature selection meaningfully improves predictive quality. Results indicate that Random Forest and SVM perform strongest and that mean perimeter, mean texture, and mean smoothness are highly influential.","|  | Comparative Analysis of Feature Selection and Machine Learning Models for Breast\u003Cbr>Cancer Risk Prediction\u003Cbr>\u003Cbr>Tonmoy Roy\u003Cbr>Faculty Advisor: Dr. Zsolt Ugray Data Analytics & Information Systems Department\u003Cbr>Utah State University |  |  |  |\n| --- | --- | --- | --- | --- |\n| ABSTRACT\u003Cbr>\u003Cbr>The comparative analysis of featureselection and machine learning models for breast cancer risk prediction aims to develop accurate and efficient models for diagnosing breast cancer.\u003Cbr>In this analysis, we explore different feature selection techniques and machine learning models to identify the most effective feature combination for breast cancer risk prediction.\u003Cbr>Our study provides valuable insights into the importance of feature selection and model selection in developing accurate breast cancer risk prediction models. Using the three features provide high accuracy to detect breast cancer.\u003Cbr>Overall, this study highlights the importance of combining feature selection techniques with machine learning algorithms to develop accurate and efficient models for breast cancer risk prediction.\u003Cbr>The findings ofthis study are a very simple method to diagnose to improve breast cancer, ultimately leading to better outcomes for patients.\u003Cbr>Key Factors\u003Cbr>\u003Cbr>• Data Preprocessing\u003Cbr>• Feature Selection\u003Cbr>• Dimensionality Reduction\u003Cbr>• Machine Learning Models\u003Cbr>• Evaluation Metrics\u003Cbr>Limitations\u003Cbr>\u003Cbr>• The dataset is not representative of all breast cancer cases.\u003Cbr>• The dataset did not include important variables, such as genetic markers, that could potentially improve the accuracy of breast cancer risk prediction models. | INTRODUCTION\u003Cbr>\u003Cbr>Breast cancer is the most commonly diagnosed cancer in women worldwide and one of the leading causes of death among women. Early detection and accurate prediction of breast cancer risk are critical for successful treatment and management of the disease. Machine learning models have shown great potential in predicting breast cancer risk based on various features and risk factors. However, selecting the most relevant features for model training can significantly impact the model's accuracy and efficiency.\u003Cbr>\u003Cbr>Methodology\u003Cbr>• Data Collection\u003Cbr>The dataset used in this study was obtained from the UCI Machine Learning Repository. The dataset contains 31 attributes that are used to predict the diagnosis of breast cancer. All the attributes described the geometric features of the tumor. Name some of the variables :\u003Cbr>❑\u003Cbr>❑\u003Cbr>❑\u003Cbr>❑\u003Cbr>mean radius\u003Cbr>mean texture\u003Cbr>mean_perimeter\u003Cbr>mean  area\u003Cbr>❑ mean  smoothness\u003Cbr>• Data Preprocessing\u003Cbr>Data preprocessing involves cleaning the data, removing missing values, and encoding categorical variables. We used one-hot encoding to encode categorical variables.\u003Cbr>• Feature Selection\u003Cbr>Feature selection is a crucial step in machine learning models. It helps to identify the most important features that contribute to the prediction. We used three different featureselection techniques: correlation-based feature selection, principal component analysis, and recursive feature elimination.\u003Cbr>• Model Selection\u003Cbr>We evaluated the performance of eight different machine learning models: Logistic Regression, Decision Tree, Random Forest, K-Nearest Neighbor, Ada Boost Classifier, Gradient Boosting, Classifier XGB Classifier, and Support Vector Machine\u003Cbr>• Evaluation Metrics\u003Cbr>We used the following evaluation metrics to assess the performance of the machine learning models:\u003Cbr>❑ Accuracy\u003Cbr>❑ Precision\u003Cbr>❑ Recall\u003Cbr>❑ F1 score | RESULTS\u003Cbr>\u003Cbr>The performance of different feature selection and machine learning models was evaluated in terms of accuracy, precision, recall, and F1-score for breast cancer risk prediction. The following results were obtained:\u003Cbr>• Machine Learning Models: The performance of different machine learning models was evaluated with respect to the selected features. The Random Forest (RF) and SVM models equally outperformed other mode","cbCaihn4d8KWWqKg","https://ap.wps.com/l/cbCaihn4d8KWWqKg","pdf",654171,1,"English","en",105,"# Abstract\n# Introduction\n# Methodology\n## Data Collection\n## Data Preprocessing\n## Feature Selection\n## Model Selection\n## Evaluation Metrics\n# Results\n# Discussion\n# Conclusions","[{\"question\":\"What is the main goal of this study on breast cancer risk prediction?\",\"answer\":\"To compare feature selection methods and machine learning models, aiming to produce accurate and efficient breast cancer risk prediction. It focuses on identifying the most effective feature combinations for diagnosis.\"},{\"question\":\"Which feature selection techniques are used?\",\"answer\":\"Correlation-based feature selection, principal component analysis (PCA), and recursive feature elimination (RFE) are applied to select or reduce features for model training.\"},{\"question\":\"Which models and features show the best results?\",\"answer\":\"Random Forest (RF) and SVM outperform other models in accuracy and F1-score. The most important features are mean perimeter, mean texture, and mean smoothness.\"}]","Comparative Analysis of Feature Selection and Machine Learning Models for Breast Cancer Risk Prediction | PDF",1785723396,3,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":84,"head_meta":86,"extra_data":88,"updated_unix":27},"comparative-analysis-of-feature-selection-and-machine-learning-models-for-breast-cancer-risk-prediction","",{"@graph":35,"@context":83},[36,52,66],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":28},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/comparative-analysis-of-feature-selection-and-machine-learning-models-for-breast-cancer-risk-prediction/119265/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":22,"description":14,"dateModified":60,"datePublished":60,"encodingFormat":59,"isAccessibleForFree":61,"interactionStatistic":62},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":63,"interactionType":64,"userInteractionCount":4},"InteractionCounter",{"@type":65},"ViewAction",{"@type":67,"mainEntity":68},"FAQPage",[69,75,79],{"name":70,"@type":71,"acceptedAnswer":72},"What is the main goal of this study on breast cancer risk prediction?","Question",{"text":73,"@type":74},"To compare feature selection methods and machine learning models, aiming to produce accurate and efficient breast cancer risk prediction. It focuses on identifying the most effective feature combinations for diagnosis.","Answer",{"name":76,"@type":71,"acceptedAnswer":77},"Which feature selection techniques are used?",{"text":78,"@type":74},"Correlation-based feature selection, principal component analysis (PCA), and recursive feature elimination (RFE) are applied to select or reduce features for model training.",{"name":80,"@type":71,"acceptedAnswer":81},"Which models and features show the best results?",{"text":82,"@type":74},"Random Forest (RF) and SVM outperform other models in accuracy and F1-score. The most important features are mean perimeter, mean texture, and mean smoothness.","https://schema.org",{"og:url":50,"og:type":85,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":87,"canonical":50},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":90},[91,95,99,103,108,113,118,121,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":92,"show_sort_weight":93,"slug":94},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":96,"show_sort_weight":97,"slug":98},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":100,"show_sort_weight":101,"slug":102},"Exam",70,"exam",{"id":104,"doc_module":4,"doc_module_name":45,"category_name":105,"show_sort_weight":106,"slug":107},5,"Comic",60,"comic",{"id":109,"doc_module":4,"doc_module_name":45,"category_name":110,"show_sort_weight":111,"slug":112},6,"Technology",50,"technology",{"id":114,"doc_module":4,"doc_module_name":45,"category_name":115,"show_sort_weight":116,"slug":117},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":119,"slug":120},30,"research-report",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":124,"slug":125},9,"Religion & Spirituality",20,"religion-spirituality",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":124,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":104,"slug":136},19,"General","general"]