[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126308-en":3,"doc-seo-126308-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},126308,2336475104957,"Seraphina","https://ap-avatar.wpscdn.com/avatar/22000c4c6bd8a5076e1?x-image-process=image/resize,m_fixed,w_180,h_180&k=1787554080175789136",8,"Research & Report","The Effect of Feature Selection on Machine Learning Classification - Research Report","High-dimensional datasets often cause overfitting and make machine learning model construction computationally expensive. This study applies feature selection as a dimensionality reduction strategy to improve classification performance. Five feature selection methods—Chi-Square (CS), Information Gain (IG), Genetic Algorithm (GA), Particle Swarm Optimization (PSO), and LASSO—are evaluated with three classifiers: Naïve Bayes, Extreme Gradient Boosting (XGB), and Random Forest (RF). Experiments use the Heart Attack Analysis & Prediction Dataset across three best-feature scenarios and assess accuracy, precision, recall, F1-score, AUC, and training time. Results show feature selection effectively improves predictive performance, with CS and IG in the filter category performing best with XGB, while training time decreases by 23.5% and technique-specific selection outperforms intersection-based selection.","INTERNATIONAL JOURNAL ON INFORMATICS VISUALIZATION  \n[journal homepage :](journal homepage : www.joiv.org/index.php/joiv)[ www.joiv.org/index.php/joiv](journal homepage : www.joiv.org/index.php/joiv)  \nThe Effect of Feature Selection on Machine Learning Classification  \nJasman Pardede a,*, Rio Dwianto a  \na Department of Informatics, Institut Teknologi Nasional (Itenas) Bandung, Bandung, Indonesia Corresponding author:*[jasman@itenas.ac.id](jasman@itenas.ac.id)  \nAbstract—High-dimensional datasets can lead to overfitting and computationally expensive model building on machine learning. This study uses a dimensionality reduction technique, namely feature selection techniques, to overcome these problems. Five feature selection methods were used, i.e., Chi-Square (CS), Information Gain (IG), Genetic Algorithm (GA), Particle Swarm Optimization (PSO), and Least Absolute Shrinkage and Selection Operator (LASSO), and three classifier methods viz. Naïve Bayes, Extreme Gradient Boosting (XGB), and RF Classifier. The dataset used is the Heart Attack Analysis & Prediction Dataset. In this study, three scenarios ofthe best feature selection were carried out, namely: 1. selection of the best feature using a specific feature selection, 2. the intersection of selection of the best feature from the same category, 3. the intersection of selection of the best feature from the five proposed feature selection methods. The performance model is measured using accuracy, precision, recall, f1-score, AUC, and training time. This study reveals that feature selection is very effective in improving the performance of prediction models. Based on the experiment results, the best feature selection is CS and IG in the Filter Category with the XGB model. The best feature selected improved the performance of accuracy, precision, recall, f1-score, and AUC, i.e., 1.7%, 1%, 2.3%, 1.6%, and 0.2%, respectively. Meanwhile, training time requirements decreased by 23.5%. Feature selection with specific techniques performs better than feature selection by selecting the best features from the same category feature selection technique or various other feature selection methods.  \nKeywords—Machine learning; feature selection; Chi-Square; information gain; classification; extreme gradient boosting.  \nManuscript received 21 Jul. 2024; revised 19 Nov. 2024; accepted 10 Dec. 2024. Date of publication 31 Jul. 2025.  \nInternational Journal on Informatics Visualization is licensed under a Creative Commons Attribution-Share Alike 4.0 International License.  \nI. INTRODUCTION  \nResearchers often have to deal with complex datasets with many features and high-dimensional data. High-dimensional data causes several problems in the learning model, such asthe learning model having difficulty but optimal performance because the more features used [1] and the more complex a machine learning model must make [2], [3]. Highdimensional data can also cause overfitting because there are many feature configurations even though we only have limited data [4]. Data with large dimensions are challenging to process computationally (computationally expensive) concerning memory and time [5]. To overcome these problems, researchers use the dimensionality reduction technique, namely the feature selection technique [1]-[4], [6] .  \nFeature selection is an essential pre-processing technique to select influential features in a dataset [1], [7]-[9]. Featureselection is used to select influential features, remove irrelevant features in the dataset attributes, fast computation time, and improve the performance of classification methods [10] . Feature selection algorithms can be divided into three groups, namely Filter, Wrapper, and Embedded Selector [1],  \n[7], [9], [11]. This study uses Filter-based feature selection techniques, namely Chi-Square (CS) [6], [12]-[14] and Information Gain (IG) [15]-[17], Wrapper namely Genetic Algorithm (GA) [18], [19] and Particle Swarm Optimization (PSO) [20], and Embedded Sel","cbCaiob4LELUSUUB","https://ap.wps.com/l/cbCaiob4LELUSUUB","pdf",3706449,6,1,11,"English","en",105,"# Introduction\n## Feature selection for high-dimensional data\n## Feature selection methods: Filter, Wrapper, Embedded\n# Related Work\n## CS and CART results\n## IG with Naïve Bayes\n## GA with Naïve Bayes\n## PSO with Naïve Bayes\n## LASSO with Random Forest\n# Method and Experimental Scenarios\n## Feature selection methods and classifiers\n## Best-feature selection scenarios\n# Evaluation Metrics\n## Accuracy, precision, recall, F1-score, AUC\n## Training time comparison\n# Results and Findings\n## Best feature selection: CS and IG with XGB\n## Performance improvements and reduced training time","[{\"question\":\"Why is feature selection important for machine learning classification in high-dimensional datasets?\",\"answer\":\"High-dimensional data increases complexity, can cause overfitting, and requires more memory and time. Feature selection reduces dimensions by keeping influential features and removing irrelevant ones, improving classification performance and computation efficiency.\"},{\"question\":\"Which feature selection methods and classifiers are compared in the study?\",\"answer\":\"The study compares CS, IG, GA, PSO, and LASSO as feature selection methods, and evaluates them with Naïve Bayes, Extreme Gradient Boosting (XGB), and Random Forest (RF) classifiers.\"},{\"question\":\"What are the best-performing feature selection results reported?\",\"answer\":\"The best feature selection is CS and IG in the filter category using the XGB model. The selected features improve accuracy, precision, recall, F1-score, and AUC, and reduce training time by 23.5%.\"}]","The Effect of Feature Selection on Machine Learning Classification - Research Report | PDF",1785904371,28,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"the-effect-of-feature-selection-on-machine-learning-classification-research-report","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/the-effect-of-feature-selection-on-machine-learning-classification-research-report/126308/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-22","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"Why is feature selection important for machine learning classification in high-dimensional datasets?","Question",{"text":77,"@type":78},"High-dimensional data increases complexity, can cause overfitting, and requires more memory and time. Feature selection reduces dimensions by keeping influential features and removing irrelevant ones, improving classification performance and computation efficiency.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"Which feature selection methods and classifiers are compared in the study?",{"text":82,"@type":78},"The study compares CS, IG, GA, PSO, and LASSO as feature selection methods, and evaluates them with Naïve Bayes, Extreme Gradient Boosting (XGB), and Random Forest (RF) classifiers.",{"name":84,"@type":75,"acceptedAnswer":85},"What are the best-performing feature selection results reported?",{"text":86,"@type":78},"The best feature selection is CS and IG in the filter category using the XGB model. The selected features improve accuracy, precision, recall, F1-score, and AUC, and reduce training time by 23.5%.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,112,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":108,"doc_module":4,"doc_module_name":47,"category_name":109,"show_sort_weight":110,"slug":111},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":108,"slug":139},19,"General","general"]