[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121144-en":3,"doc-seo-121144-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121144,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Machine Learning in Network Intrusion Detection - A Cross-Dataset Generalization Study","Network Intrusion Detection Systems (NIDS) are essential for cybersecurity, and their ability to generalize across different networks determines real-world effectiveness. This study analyzes machine-learning-based NIDS generalization via cross-dataset experimentation using four classifiers and four datasets: CIC-IDS-2017, CSE-CIC-IDS2018, LycoS-IDS2017, and LycoS-Unicas-IDS2018. Training and testing on the same dataset yields near-perfect performance, while cross-dataset accuracy drops to near random for most attack-dataset pairs. Data visualization highlights anomalies that obstruct transfer of learned patterns to new scenarios, emphasizing the need to address data heterogeneity.","Received 21 August 2024, accepted 24 September 2024, date of publication 3 October 2024, date of current version 14 October 2024. Digital Object Identifier 10.1109/ACCESS.2024.3472907  \nMachine Learning in Network Intrusion Detection: A Cross-Dataset Generalization Study  \nMARCO CANTONE, CLAUDIO MARROCCO,(Member, IEEE), AND ALESSANDRO BRIA Department of Electrical and Information Engineering, University of Cassino and Southern Latium, 03043 Cassino, Italy  \nCorresponding author: Marco Cantone ([marco.cantone@unicas.it](marco.cantone@unicas.it))  \nThis work was supported in part by Italian Ministry of University, Ministry of Education, University and Research (MIUR) Program‘‘Department of Excellence’’ Law under Grant 232/216; and in part by the ‘‘Innovative Ph.D. Programs for Public Administration under Grant D.M.351/2022 .’’  \nABSTRACT Network Intrusion Detection Systems (NIDS) are a fundamental tool in cybersecurity. Their ability to generalize across diverse networks is a critical factor in their effectiveness and a prerequisite for real-world applications. In this study, we conduct a comprehensive analysis on the generalization of machinelearning-based NIDS through an extensive experimentation in a cross-dataset framework. We employ four machine learning classifiers and utilize four datasets acquired from different networks: CIC-IDS- 2017, CSE-CIC-IDS2018, LycoS-IDS2017, and LycoS-Unicas-IDS2018 . Notably, the last dataset is a novel contribution, where we apply corrections based on LycoS-IDS2017 to the well-known CSE-CIC-IDS2018 dataset. The results show nearly perfect classification performance when the models are trained and tested on the same dataset. However, when training and testing the models in a cross-dataset fashion, the classification accuracy is largely commensurate with random chance except for a few combinations of attacks and datasets. We employ data visualization techniques in order to provide valuable insights on the patterns in the data. Our analysis unveils the presence of anomalies in the data that directly hinder the classifiers capability to generalize the learned knowledge to new scenarios. This study enhances our comprehension of the generalization capabilities of machine-learning-based NIDS, highlighting the significance of acknowledging data heterogeneity.  \nINDEX TERMS CIC-IDS2017, cross-dataset, CSE-CIC-IDS2018, generalization, intrusion detection system, machine learning.  \nI. INTRODUCTION  \nThe rapid expansion of network interconnections has led to a corresponding growth in the cyber threat landscape, attracting the interest of an increasing number of cyber attackers. This has resulted in the disruption of essential services with significant economic consequences. Globally, cybercrime is estimated to have an impact of around 1 trillion dollars in 2020, with an increase of more than 50% compared to 2018 [1] . It is therefore necessary to use systems and strategies to counter this phenomenon [2],[3],[4] .  \nThe associate editor coordinating the review of this manuscript and approving it for publication was Xueqin Jiang .  \nNetwork Intrusion Detection Systems (NIDS) are specifically designed to identify intrusions analyzing network traffic, enabling targeted entities to take timely actions against potential threats [5] . They can be grouped into three categories: statistics-based, knowledge-based, and MachineLearning-based (ML-based) [6], [7] . The statistics-based paradigm concerns the meticulous examination of each record within a dataset, with the goal of building a statistical model that encompasses the established norms of user behavior within a network [8],[9]. In contrast, the knowledge-based approach is based on attempting to identify the requested actions by leveraging existing system data, including protocol specifications and network traffic instances. Knowledgebased NIDS, guided by human-defined rules, deduces  \n􀀊 2024 The Authors. This work is licensed under a Creative Commons Attrib","cbCais15O23KGdDB","https://ap.wps.com/l/cbCais15O23KGdDB","pdf",3278101,1,20,"English","en",105,"# Introduction\n## Network Intrusion Detection Systems (NIDS)\n## ML-based NIDS: supervised vs. unsupervised learning","[{\"question\":\"What problem does the study address about network intrusion detection?\",\"answer\":\"It evaluates how well machine-learning-based NIDS generalize when trained and tested on different datasets from different networks, which is required for practical deployment.\"},{\"question\":\"Which datasets are used in the cross-dataset experiments?\",\"answer\":\"The study uses CIC-IDS-2017, CSE-CIC-IDS2018, LycoS-IDS2017, and LycoS-Unicas-IDS2018, with the last dataset created via corrections based on LycoS-IDS2017 applied to CSE-CIC-IDS2018.\"},{\"question\":\"What do the results show when models are trained and tested on the same dataset versus across datasets?\",\"answer\":\"Models achieve nearly perfect classification when trained and tested on the same dataset. In contrast, cross-dataset classification accuracy is largely comparable to random guessing except for a few specific attack-dataset combinations.\"}]","Machine Learning in Network Intrusion Detection - A Cross-Dataset Generalization Study | PDF",1785734066,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-in-network-intrusion-detection-a-cross-dataset-generalization-study","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-in-network-intrusion-detection-a-cross-dataset-generalization-study/121144/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the study address about network intrusion detection?","Question",{"text":75,"@type":76},"It evaluates how well machine-learning-based NIDS generalize when trained and tested on different datasets from different networks, which is required for practical deployment.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which datasets are used in the cross-dataset experiments?",{"text":80,"@type":76},"The study uses CIC-IDS-2017, CSE-CIC-IDS2018, LycoS-IDS2017, and LycoS-Unicas-IDS2018, with the last dataset created via corrections based on LycoS-IDS2017 applied to CSE-CIC-IDS2018.",{"name":82,"@type":73,"acceptedAnswer":83},"What do the results show when models are trained and tested on the same dataset versus across datasets?",{"text":84,"@type":76},"Models achieve nearly perfect classification when trained and tested on the same dataset. In contrast, cross-dataset classification accuracy is largely comparable to random guessing except for a few specific attack-dataset combinations.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]