[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128162-en":3,"doc-seo-128162-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128162,549768702563,"Sage","https://ap-avatar.wpscdn.com/avatar/8000c4aa63b76e948b?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786536092046926083",8,"Research & Report","Detailed Analyses and Efficient Identification of Malware Evidence in CLaMP Dataset based on Machine Learning Approaches","Malware constitutes a major threat in computer networks and cyberspace, and while signature-based detection has notable limitations, many machine learning (ML) approaches still suffer from inadequate detection rates due to dataset pattern size and characteristics. This study trains and tests malware classification models using the CLaMP dataset, starting with comprehensive exploratory analysis to understand distributions and guide preprocessing. Two scenarios are evaluated: models built without cleaning or feature selection, and models built after cleaning plus Recursive Feature Elimination. Naive Bayes and Logistic Regression are tuned and compared using accuracy, precision, recall, F1-score, and AUC.","Detailed Analyses and Efficient Identification of Malware Evidence in CLaMP Dataset based on Machine Learning Approaches  \n1M. O. Ayinla  \nDepartment of Computer Science Kwara State College of Education, Ilorin, Nigeria  \nEmail: mo.ayinla [AT] [kwcoeilorin.edu.ng](kwcoeilorin.edu.ng)  \n2A. M. Oyelakin  \nDepartment of Computer Science Crescent University, Abeokuta, Nigeria  \nEmail: moruff.oyelakin [AT] [cuab.edu.ng](cuab.edu.ng)  \n3U. A. Adeniyi  \nDepartment of Cyber Security Airforce Institute of Technology, Kaduna, Nigeria Email: adedayousman [AT] [afit.edu.ng](afit.edu.ng)  \n4K. O. Tajudeen  \nDepartment of Computer Science Al-Hikmah University, Ilorin, Nigeria Email: kotajudeen [AT] [alhikmah.edu.ng](alhikmah.edu.ng)  \n5O. J. Olaleye  \nDepartment of Computer Science & Information Technology Bells University of Technology, Ota, Nigeria Email: ojolaleye [AT] [bellsuniversity.edu.ng](bellsuniversity.edu.ng)  \nAbstract—Malware is a malicious software that is used to launch attacks of different types in computer networks and cyber space. Several signature and machine learning-based approaches have been used for the identification of malware types in the past. However,signature-based detection approaches have been reported to have serious limitations which gave room for machine learning-based malware identification techniques to be more popular. Despite the promises of the ML methods in the identification of malware evidence, some of the ML approaches in literature have poor detection rates which can be as a result of the size and nature of the patterns in the datasets used. This study used a dataset named CLaMP for the training and testing of the malware classification models. Firstly, comprehensive exploratory analyses of the dataset were carried out with a view to understanding the data distributions in it better and make informative decisions on how to pre-process and apply it for malware identification. During the experimentations, two scenarios were established before feeding the data into the learning algorithms. Scenario 1 involves building malware identification model without data cleaning and feature selection while scenario 2 involves the cleaning of the data and selection of promising features for building the models.In scenario 2, Recursive Feature Elimination (RFE) technique was used for selecting the promising attributes which were used to build the two malware classification models. Naive Bayes (NB) and Logistic Regression (LR) algorithms were used for building the models. The hyper parameters of the two selected algorithms were varied and the models tested and validated severally before optimal performances were arrived at. The results of the models were compared based on the selected metrics, namely: accuracy,  \nprecision, recall, f1-score and Area Under the Curve (AUC). Experimental results showed that in the scenario 1, where the dataset was not pre-processed and all the attributes were used for the model building, poor results were obtained by both models in all metrics except in recall. However, NB-based malware identification model slightly performed better than LR in all the metrics. It was also discovered that both NB and LR-based malware identification models performed well in scenario 2 when the dataset was pre-processed and promising features were selected using RFE. This study concluded that the detailed exploratory analyses, data cleaning and feature subset selection methods helped in achieving promising results from the malware identification models  \nKeywords- Malware Identification; Machine Learning Algorithms; Feature Selection; Windows PE headers  \nI. INTRODUCTION  \nMalware of different types are used to launch various attacks in computer networks [1] . Generally, malware is a term for all types of malicious program. It is the type of software that is used with the aim of attempting to breach the security policy of computer system or network with respect to Confidentiality, Integrity or Availabil","cbCaitwROVk8KuVS","https://ap.wps.com/l/cbCaitwROVk8KuVS","pdf",518768,1,7,"English","en",105,"# Abstract\n# Keywords\n# I. Introduction","[{\"question\":\"What was the purpose of the exploratory analysis in this study?\",\"answer\":\"The study performed exploratory analyses to understand dataset distributions and to make informed decisions on preprocessing steps for malware identification.\"},{\"question\":\"How do Scenario 1 and Scenario 2 differ in model building?\",\"answer\":\"Scenario 1 builds models without data cleaning and feature selection, while Scenario 2 includes cleaning and selecting promising features using Recursive Feature Elimination (RFE).\"},{\"question\":\"Which algorithms and evaluation metrics were used to compare malware identification performance?\",\"answer\":\"Naive Bayes and Logistic Regression were used, and results were compared using accuracy, precision, recall, F1-score, and AUC.\"}]","Detailed Analyses and Efficient Identification of Malware Evidence in CLaMP Dataset based on Machine Learning Approaches | PDF",1785945209,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"detailed-analyses-and-efficient-identification-of-malware-evidence-in-clamp-dataset-based-on-machine-learning-approaches","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/detailed-analyses-and-efficient-identification-of-malware-evidence-in-clamp-dataset-based-on-machine-learning-approaches/128162/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-24","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What was the purpose of the exploratory analysis in this study?","Question",{"text":76,"@type":77},"The study performed exploratory analyses to understand dataset distributions and to make informed decisions on preprocessing steps for malware identification.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How do Scenario 1 and Scenario 2 differ in model building?",{"text":81,"@type":77},"Scenario 1 builds models without data cleaning and feature selection, while Scenario 2 includes cleaning and selecting promising features using Recursive Feature Elimination (RFE).",{"name":83,"@type":74,"acceptedAnswer":84},"Which algorithms and evaluation metrics were used to compare malware identification performance?",{"text":85,"@type":77},"Naive Bayes and Logistic Regression were used, and results were compared using accuracy, precision, recall, F1-score, and AUC.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]