[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118870-en":3,"doc-seo-118870-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},118870,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Impact of Data Balancing and Feature Selection on Machine Learning-based Network Intrusion Detection - research abstract","Unbalanced datasets in supervised machine learning bias models toward majority classes, reducing recognition of minority classes. Network intrusion detection datasets such as NSL-KDD and UNSW-NB15 show attack composition imbalance, commonly addressed through oversampling and synthetic data generation. This work applies SMOTE and ADASYN to construct balanced minority-class samples and uses recursive feature elimination (RFE) to select informative features. Results indicate SMOTE outperforms ADASYN on highly imbalanced UNSW-NB15, while RFE slightly reduces accuracy but improves training speed. A decision tree classifier achieves higher recognition rates than random forest and KNN.","INTERNATIONAL JOURNAL ON INFORMATICS VISUALIZATION  \n[journal homepage : www.joiv.org/index.php/joiv](journal homepage : www.joiv.org/index.php/joiv)  \nImpact of Data Balancing and Feature Selection on Machine Learning  \nbased Network Intrusion Detection  \nAzhari Shouni Barkaha,b,*, Siti Rahayu Selamatb, Zaheera Zainal Abidin b, Rizki Wahyudia  \na Department of Informatics, Universitas Amikom Purwokerto, Purwokerto Utara, Banyumas, 55127, Indonesia bFakulti Teknologi Maklumat dan Komunikasi (FTMK), Universiti Teknikal Malaysia Melaka, Melaka, Malaysia Corresponding author:*[azhari@amikompurwokerto.ac.id](azhari@amikompurwokerto.ac.id)  \nAbstract—Unbalanced datasets are a common problem in supervised machine learning. It leads to a deeper understanding of the majority of classes in machine learning. Therefore, the machine learning model is more effective at recognizing the majority classes than the minority classes. Naturally, imbalanced data, such as disease data and data networking, has emerged in real life. DDOS is oneof the network intrusions found to happen more often than R2L. There is an imbalance in the composition of network attacks in Intrusion Detection System (IDS) public datasets such as NSL-KDD and UNSW-NB15. Besides, researchers propose many techniques to transform it into balanced data by duplicating the minority class and producing synthetic data. Synthetic Minority Oversampling Technique (SMOTE) and Adaptive Synthetic (ADASYN) algorithms duplicate the data and construct synthetic data for the minority classes. Meanwhile, machine learning algorithms can capture the labeled data's pattern by considering the input features. Unfortunately, not all the input features have an equal impact on the output (predicted class or value). Some features are interrelated and misleading. Therefore, the important features should be selected to produce a good model. In this research, we implement therecursive feature elimination (RFE) technique to select important features from the available dataset. According to the experiment, SMOTE provides a better synthetic dataset than ADASYN for the UNSW-B15 dataset with a high level of imbalance. RFE featureselection slightly reduces the model's accuracy but improves the training speed. Then, the Decision Tree classifier consistently achievesa better recognition rate than Random Forest and KNN.  \nKeywords—Intrusion detection; feature selection; imbalance; SMOTE; ADASYN.  \nManuscript received 23 Jul. 2022; revised 26 Dec. 2022; accepted 14 Jan. 2023. Date of publication 31 Mar. 2023.  \nInternational Journal on Informatics Visualization is licensed under a Creative Commons Attribution-Share Alike 4.0 International License.  \nI. INTRODUCTION  \nWith the current high level of internet usage, network attacks pose a serious threat. The attacks are evolving in line with the advance of computing capacity. To ensure the safety of data communication, defensive action must be taken. Therefore, researchers in network defense are working hard all the time to encounter new types of attacks.  \nThe important task in network security is to recognize the type of attack. The attack dataset is evolving due to the introduction of new attack techniques. Even though the indicator variables (features) are similar, the type of network intrusion is evolving. Network security researchers provide datasets allowing the machine to recognize attack classes automatically. Many researchers provide KDD99 and NSLKDD [1], [2], [3], [4], while other researchers provide UNSW-NB15 [2], [3], [5], [6], [7], [8] and Liu provide CICIDS2017 [3], and others provide CICDDS001 [9], [10].  \nBased on publicly available datasets, many researchers develop methods and tools to recognize network intrusions, such as random forest, decision tree, logistic regression, KNN, and ANN. The common problems identified in many academic papers are that certain classes of attacks have rarely happened. Therefore, the available data is limited, while othe","cbCaian8PhehJiUN","https://ap.wps.com/l/cbCaian8PhehJiUN","pdf",3678118,1,"English","en",105,"# Introduction\n## Network security and intrusion recognition\n## Public IDS datasets and class imbalance\n## Data balancing via resampling and synthetic oversampling\n## Feature reduction and recursive feature elimination","[{\"question\":\"Why do unbalanced datasets affect machine learning intrusion detection performance?\",\"answer\":\"They cause models to learn majority classes more deeply, leading to weaker recognition of minority attack classes and overall lower classification performance.\"},{\"question\":\"How do SMOTE and ADASYN address class imbalance?\",\"answer\":\"They generate synthetic samples for minority classes by duplicating and constructing new data points, improving representation compared with simple duplication.\"},{\"question\":\"What is the role of recursive feature elimination (RFE) in this research?\",\"answer\":\"RFE selects more informative features from the dataset; experiments show it slightly reduces accuracy but improves training speed by removing less useful or misleading features.\"}]","Impact of Data Balancing and Feature Selection on Machine Learning-based Network Intrusion Detection - research abstract | PDF",1785720708,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"impact-of-data-balancing-and-feature-selection-on-machine-learning-based-network-intrusion-detection-research-abstract","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/impact-of-data-balancing-and-feature-selection-on-machine-learning-based-network-intrusion-detection-research-abstract/118870/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":22,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do unbalanced datasets affect machine learning intrusion detection performance?","Question",{"text":75,"@type":76},"They cause models to learn majority classes more deeply, leading to weaker recognition of minority attack classes and overall lower classification performance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do SMOTE and ADASYN address class imbalance?",{"text":80,"@type":76},"They generate synthetic samples for minority classes by duplicating and constructing new data points, improving representation compared with simple duplication.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the role of recursive feature elimination (RFE) in this research?",{"text":84,"@type":76},"RFE selects more informative features from the dataset; experiments show it slightly reduces accuracy but improves training speed by removing less useful or misleading features.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":28,"slug":126},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":28,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]