[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121736-en":3,"doc-seo-121736-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121736,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Are Machine Learning Models for Malware Detection Ready for Prime Time?","Machine learning has become the default approach in academic malware detection, supported by papers reporting near-perfect accuracies. Yet real deployments often show low reliability and limited trust, creating a gap between lab performance and operational effectiveness. The document analyzes key reasons reproducibility fails: misleading metrics under class imbalance (Base Rate Fallacy), dataset collection from mismatched time periods causing models to learn spurious non-causal differences, and random train/test splits that leak future information due to shifting sample distributions over time.","Are Machine Learning Models for Malware Detection Ready for Prime Time?  \nLorenzo Cavallaro ∣ University College London, United Kingdom Johannes Kinder ∣ Bundeswehr University Munich, Germany Feargus Pendlebury ∣ University College London, United Kingdom Fabio Pierazzi ∣ King’s College London, United Kingdom  \nIn academic research on malware detection, machine learning-based techniques have become the de-facto standard. Over the years, many papers have been published on the topic, using everything in the machine learning toolbox from traditional models [1, 2] to more recent neural network architectures [3, 4, 5] . With published accuracies and other performance numbers regularly close to the 100% mark, you could not be blamed for thinking that the problem is practically solved. Why would anyone be interested in continuing to use malware signatures and simple patterns?  \nCannot Reproduce?  \nYet, practitioners find themselves disillusioned when trying to put machine learning models from research into practice. In real-world deployments, machine learning-based malware classifiers are known to often be unreliable and are not trusted to act as the main line of defense.  \nWhere does this discrepancy come from? It would be too easy to shrug this off as expected techncal challenges in technology transfer; we must get to the bottom of the issue or risk exacerbating the reproducibility crisis decried in machine learning-based science [6] . So, what’s the deal? If the same model performs so differently in production, then the lab settings that yield near-perfect performance must not  \nbe representative of how such models are deployed.  \nThe typical research project on machine learning-based malware detection goes as follows: (1) obtain a malware dataset, from an academic project or public sources; (2) collect a dataset of benign software, e.g., through scraping repositories or app stores; (3) engineer or learn features to represent the malicious and benign apps in your datasets; and (4) train a classifier while following common good-practice guidelines, for instance to prevent overfitting [7] . Unfortunately, this seemingly straightforward approach comes with several pitfalls that may prevent your results from translating to practice [8, 9] .  \nThe first and maybe the most obvious issue is that performance metrics are influenced by the relative sizes of the classes, i.e., the ratio of malware to benign software in the evaluation dataset. To illustrate, imagine a silly malware detection model that classifies everything as malware. If your dataset consists of 95% malware and 5% benign software, then the model’s performance will actually be quite decent! Widely used performance metrics are precision, recall, and 􀁆1—the harmonic mean of the two. With precision being the fraction of true malware in everything the classifier detected as malware, we would get to a respectable 95%, because only 5% of  \nthe detections would be false. Recall is the percentage of actual malware detected, which would even be a perfect 100% . The 􀁆1 then comes out at an impressive 97%, although the model is completely useless in practice. Now, this may be an extreme example, but the effect is gradual and gets worse the more the class ratio at testing time differs from the class ratio in the real world. This effect of misrepresenting the ratio between classes is also known as the Base Rate Fallacy [10] . In most practical settings, there is vastly more benign software than malware, making real datasets heavily imbalanced toward the benign class. A classifier that is good at detecting malware but misclassifies much benign software as malicious therefore has poor performance in practice. When tested on equal amounts of malware and benign software, or even on a majority of malware, its performance will be inflated.  \nA second issue is that, as a result of the dataset collection procedure, malicious and benign software datasets often end up being from separate time period","cbCaifL5J5w3fiUu","https://ap.wps.com/l/cbCaifL5J5w3fiUu","pdf",1193119,1,5,"English","en",105,"# Cannot Reproduce?\n## Why ML Research Performance Fails in Practice\n## Typical ML Malware Detection Pipeline\n## Pitfall 1: Metrics Affected by Class Ratio (Base Rate Fallacy)\n## Pitfall 2: Datasets from Different Time Periods\n## Pitfall 3: Random Train/Test Splits and Future Knowledge Leakage","[{\"question\":\"Why can random train/test splitting be misleading for malware detection?\",\"answer\":\"Random splits allow the training set to include samples from the future relative to the test set’s time context, so the classifier benefits from knowledge unavailable in realistic scenarios.\"}]","Are Machine Learning Models for Malware Detection Ready for Prime Time? | PDF",1785806566,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"are-machine-learning-models-for-malware-detection-ready-for-prime-time","",{"@graph":36,"@context":77},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/are-machine-learning-models-for-malware-detection-ready-for-prime-time/121736/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"Why can random train/test splitting be misleading for malware detection?","Question",{"text":75,"@type":76},"Random splits allow the training set to include samples from the future relative to the test set’s time context, so the classifier benefits from knowledge unavailable in realistic scenarios.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,101,106,111,114,119,122,126],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":21,"slug":129},19,"General","general"]