[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121644-en":3,"doc-seo-121644-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121644,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Are Machine Learning Models for Malware Detection Ready for Prime Time? - Research-to-Deployment Reproducibility and Robustness","Machine learning models for malware detection achieve near-perfect results in academic settings but often fail to reproduce in real-world deployments. The paper examines why research evaluation conditions differ from practical deployment and how to design experiments that mimic operational scenarios. It also outlines ways to assess a model’s robustness over time and highlights pitfalls that hinder translating performance metrics into trustworthy, mainline defense tools.","VU Research Portal  \nAre Machine Learning Models for Malware Detection Ready for Prime Time?  \nCavallaro, Lorenzo; Kinder, Johannes; Pendlebury, Feargus; Pierazzi, Fabio; Massacci, Fabio; Bodden, Eric; Sabetta, Antonino  \npublished in  \nIEEE Security and Privacy 2023  \nDOI (link to publisher)  \n10.1109/MSEC.2023.3236543  \ndocument version  \nPublisher's PDF, also known as Version of record  \ndocument license  \nArticle 25fa Dutch Copyright Act  \nLink to publication in VU Research Portal  \ncitation for published version (APA)  \nCavallaro, L. , Kinder, J. , Pendlebury, F. , Pierazzi, F. , Massacci, F. , Bodden, E. , & Sabetta, A. (2023) . Are Machine Learning Models for Malware Detection Ready for Prime Time? IEEE Security and Privacy, 21(2), 53- 56. [https://doi.org/10.1109/MSEC.2023.3236543](https://doi.org/10.1109/MSEC.2023.3236543)  \nGeneral rights  \nCopyright and moral rights for the publications made accessible in the public portal are retained by the authors and/or other copyright owners and it is a condition of accessing publications that users recognise and abide by the legal requirements associated with these rights.  \n• Users may download and print one copy of any publication from the public portal for the purpose of private study or research.  \n• You may not further distribute the material or use it for any profit-making activity or commercial gain  \n• You may freely distribute the URL identifying the publication in the public portal  \nTake down policy  \nIf you believe that this document breaches copyright please contact us providing details, and we will remove access to the work immediately and investigate your claim.  \nE-mail address:  \n[vuresearchportal.ub@vu.nl](vuresearchportal.ub@vu.nl)  \n[Download date: 03](Download date: 03) . Aug. 2026  \n\n| \u003Cbr>BUILDING SECURITY IN |\n| --- |\n|  Editors: Fabio Massacci, [fabio.massacci@ieee.org | Eric](fabio.massacci@ieee.org | Eric) Bodden, eric.bodden@uni-paderborn.de | Antonino Sabetta, [as@sabetta.com](as@sabetta.com) |\n\nAre Machine Learning Models for Malware Detection Ready for Prime Time?  \nLorenzo Cavallaro | University College London Johannes Kinder | Bundeswehr University Feargus Pendlebury | University College London Fabio Pierazzi | King’s College London  \nWe investigate why the performance of machine learning models for malware detection observed in alab setting often cannot be reproduced in practice. We discuss how to set up experiments mimicking a practical deployment and how to measure the robustness of a model over time.  \nI n academic research on malware  \ndetection, machine learning-based techniques have become the de facto standard. Over the years, many papers havebeenpublishedonthetopic,using everything in the machine learning toolbox, from traditional models1, 2 to more recent neural network architectures.3, 4, 5 With published accuracies and other performance numbers regularly close to the 100% mark, you could not be blamed for thinking that the problem is practically solved. Why would anyone be interested in continuing to use malware signaturesand simple patterns?  \nCannot Reproduce?  \nYet, practitioners find themselves disillusioned when trying to put machine learning models from research into practice. In real-world deployments, machine learning-based malware classifiers are known to often be  \nDigital Object Identifier 10.1109/MSEC.2023.3236543 Date of current version: 15 March 2023  \nunreliable and are not trusted to act asthe mainline of defense.  \nWhere does this discrepancy come from? It would be too easy toshrug this off as expected technical challenges in technology transfer; we must get to the bottom of the issue or risk exacerbating the reproducibility crisis decried in machine learning-based science.6 So, what’sthe deal? If the same model performs so differently in production, then the lab settings that yield near-perfect performance must not be representative of how such models are deployed.  \nThe typical research project on machine learning","cbCaigjGyFbOYNwu","https://ap.wps.com/l/cbCaigjGyFbOYNwu","pdf",288863,1,5,"English","en",105,"# Introduction\n## Reproducibility gap in practice\n# Research setup and common pitfalls\n## Base rate fallacy and evaluation metrics","[{\"question\":\"Why do malware-detection ML models perform differently in production than in lab settings?\",\"answer\":\"The gap arises because lab evaluation conditions are often not representative of how models are deployed. Differences in experimental setup and assumptions can make reported performance non-transferable.\"},{\"question\":\"How should experiments be designed to better reflect practical deployment?\",\"answer\":\"Experiments should mimic deployment conditions and evaluation procedures, so that measured performance aligns with operational realities. This includes using setups that better represent real-world conditions.\"},{\"question\":\"What is the base rate fallacy and how does it affect malware detection metrics?\",\"answer\":\"Metrics can be misleading when the class ratio in the evaluation dataset differs from the real world. For example, a high precision or F1 can occur even when the classifier is effectively useless if malware dominates the test set.\"}]","Are Machine Learning Models for Malware Detection Ready for Prime Time? - Research-to-Deployment Reproducibility and Robustness | PDF",1785805901,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"are-machine-learning-models-for-malware-detection-ready-for-prime-time-research-to-deployment-reproducibility-and-robustness","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/are-machine-learning-models-for-malware-detection-ready-for-prime-time-research-to-deployment-reproducibility-and-robustness/121644/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do malware-detection ML models perform differently in production than in lab settings?","Question",{"text":75,"@type":76},"The gap arises because lab evaluation conditions are often not representative of how models are deployed. Differences in experimental setup and assumptions can make reported performance non-transferable.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How should experiments be designed to better reflect practical deployment?",{"text":80,"@type":76},"Experiments should mimic deployment conditions and evaluation procedures, so that measured performance aligns with operational realities. This includes using setups that better represent real-world conditions.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the base rate fallacy and how does it affect malware detection metrics?",{"text":84,"@type":76},"Metrics can be misleading when the class ratio in the evaluation dataset differs from the real world. For example, a high precision or F1 can occur even when the classifier is effectively useless if malware dominates the test set.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":21,"slug":137},19,"General","general"]