[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119172-en":3,"doc-seo-119172-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119172,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Practical Performance of a Distributed Processing Framework for Machine-Learning-based NIDS","Network Intrusion Detection Systems (NIDSs) identify intrusion attacks within network traffic and alert administrators, especially to support detection of unknown threats. This work addresses a gap in prior research by implementing a distributed processing framework for machine-learning-based NIDSs and evaluating end-to-end performance when common ML classifiers are deployed. Five representative classifiers—Decision Tree, Random Forest, Naive Bayes, SVM, and kNN—are built on the framework and benchmarked using throughput and latency measurements. Results reveal classifier-dependent processing behavior and pinpoint subsystem bottlenecks that limit pipeline speed.","Practical Performance of a Distributed Processing Framework for Machine-Learning-based NIDS  \nMaho Kajiura  \nDepartment of Computer Science and Engineering Toyohashi University of Technology Toyohashi, Japan [kmaho@dsl.cs.tut.ac.jp](kmaho@dsl.cs.tut.ac.jp)  \nJunya Nakamura  \nInformation and Media Center Toyohashi University of Technology Toyohashi, Japan [junya@imc.tut.ac.jp](junya@imc.tut.ac.jp)  \narXiv :2405 . 13066v1 [ cs .CR] 20 May 2024  \nAbstract—Network Intrusion Detection Systems (NIDSs) detect intrusion attacks in network traffic. In particular, machinelearning-based NIDSs have attracted attention because of their high detection rates of unknown attacks. A distributed processing framework for machine-learning-based NIDSs employing a scalable distributed stream processing system has been proposed in the literature. However, its performance, when machine-learningbased classifiers are implemented has not been comprehensively evaluated. In this study, we implement five representative classifiers (Decision Tree, Random Forest, Naive Bayes, SVM, and kNN) based on this framework and evaluate their throughput and latency. By conducting the experimental measurements, we investigate the difference in the processing performance among these classifiers and the bottlenecks in the processing performance of the framework.  \nIndex Terms—machine-learning, network intrusion detection system, distributed processing, network security  \nI. INTRODUCTION  \nA network-based intrusion detection system (NIDS) detects intrusion attacks in network traffic and notifies the network administrator. NIDS is considered an effective defense mechanism against cyber attacks. Traditional NIDSs detect abnormal traffic that matches intrusion attack patterns, called signatures, stored in a system database. However, this approach cannot detect unknown attacks because their patterns do not match the signatures. To overcome this limitation, machine-learningbased NIDSs (MLNIDSs) have been proposed in recent years [1]–[3] . MLNIDSs detect both known and unknown attacks by building machine-learning models that include learned attack patterns based on known attacks.  \nSeveral frameworks that assist in the implementation of MLNIDSs have been proposed [4], [5] . These frameworks implemented all the functions necessary for MLNIDS on scalable distributed processing systems to process network traffic efficiently. When the network traffic increases, the MLNIDS performance can be flexibly adapted by adding nodes. They demonstrated the effectiveness of the framework by building MLNIDSs and evaluating their performance in terms of throughput and latency.  \nHowever, both authors in [4], [5] did not evaluate the effectiveness of the framework when a machine-learning classifier  \nThis work was supported by the Hibi Science Foundation, the Naito Science & Engineering Foundation, and Tokai Foundation for Technology.  \nis implemented. Furthermore, existing MLNIDSs also often focus on classifier performance, with only a few focusing on processing speed. Consequently, the volume of traffic MLNIDS can process was not investigated. As a result, system sizing, which is necessary if the framework is actually used for MLNIDSs, becomes difficult.  \nIn this paper, we construct an MLNIDS by implementing five representative classifiers based on the framework proposed in [4] and evaluate their throughput and latency. Based on this evaluation, we identify the differences in the processing performance among the classifiers and the bottlenecks in the processing performance in the framework.  \nThe experimental results show that the processing speed and classifier performance are highly dependent on the typeof classifier. Using appropriate machine-learning algorithms, the load on the MLNIDS can be reduced while maintaining high classifier performance. We found that Zeek [6], which constructs sessions from the network traffic, Logstash [7], which performs the classification process using machinelea","cbCaigpyaxR64Mjl","https://ap.wps.com/l/cbCaigpyaxR64Mjl","pdf",336382,1,7,"English","en",105,"# Introduction\n## Background on NIDS and ML-based NIDS\n## Motivation and contributions\n# Related Work\n## Datasets for ML-based intrusion detection\n## Existing MLNIDS frameworks","[{\"question\":\"What problem does the study address in distributed ML-based NIDS frameworks?\",\"answer\":\"Prior work proposed distributed processing frameworks but did not comprehensively evaluate performance when specific machine-learning classifiers are implemented. The study fills this gap by measuring throughput and latency across multiple representative classifiers.\"},{\"question\":\"Which classifiers are implemented and evaluated?\",\"answer\":\"The framework is used to implement Decision Tree, Random Forest, Naive Bayes, SVM, and kNN classifiers. Their impact on processing throughput and latency is then benchmarked experimentally.\"},{\"question\":\"Where do the main processing bottlenecks occur?\",\"answer\":\"The evaluation identifies bottlenecks in framework subsystems associated with Zeek session construction, Logstash classification, and Elasticsearch result storage.\"}]","Practical Performance of a Distributed Processing Framework for Machine-Learning-based NIDS | PDF",1785722904,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"practical-performance-of-a-distributed-processing-framework-for-machine-learning-based-nids","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/practical-performance-of-a-distributed-processing-framework-for-machine-learning-based-nids/119172/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the study address in distributed ML-based NIDS frameworks?","Question",{"text":75,"@type":76},"Prior work proposed distributed processing frameworks but did not comprehensively evaluate performance when specific machine-learning classifiers are implemented. The study fills this gap by measuring throughput and latency across multiple representative classifiers.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which classifiers are implemented and evaluated?",{"text":80,"@type":76},"The framework is used to implement Decision Tree, Random Forest, Naive Bayes, SVM, and kNN classifiers. Their impact on processing throughput and latency is then benchmarked experimentally.",{"name":82,"@type":73,"acceptedAnswer":83},"Where do the main processing bottlenecks occur?",{"text":84,"@type":76},"The evaluation identifies bottlenecks in framework subsystems associated with Zeek session construction, Logstash classification, and Elasticsearch result storage.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]