[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83157-en":3,"doc-seo-83157-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83157,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","SA-DRL Security-Aware Deep Reinforcement Learning for Ransomware Detection with Asymmetric Reward Design","Ransomware encrypts victim data in seconds, yet many machine learning detectors use symmetric losses that penalize false negatives (missed detections) and false positives (benign alarms) equally. In operations, a false negative can cause irreversible data loss and major recovery cost, while a false positive is typically reversible. To align optimization with this asymmetric risk, the study proposes SA-DRL, embedding FN–FP cost asymmetry into reinforcement learning rewards and using security-optimal model selection plus adaptive sample weighting. Experiments across DRL agents and validation protocols show markedly lower false-negative rates.","SA-DRL: Security-Aware Deep Reinforcement Learning for Ransomware Detection with Asymmetric Reward Design Jannatul Ferdousa , Rafiqul Islamb and Md Zahidul Islamc,d  \na School of Computing, Mathematics and Engineering, Charles Sturt University, Wagga Wagga, NSW, 2650, Australia b School of Computing, Mathematics and Engineering, Charles Sturt University, Albury, NSW, 2640, Australia c School of Computing, Mathematics and Engineering, Charles Sturt University, Bathurst, NSW, 2795, Australia dAI and Cyber Futures Center, Charles Sturt University, Panorama Avenue, Bathurst, NSW, 2795, Australia  \narXiv :2607 .06880v 1 [ cs .CR] 8 Jul 2026  \nARTICLE INFO  \nKeywords:  \nRansomware detection  \nDeep reinforcement learning Asymmetric reward design False-negative minimization Behavioral ransomware analysis DDQN  \nAB STRACT  \nRansomware encrypts victim data in seconds; however, current machine learning detectors use symmetric loss functions that equally penalize missed detections, known as false negatives (FN), and benign false alarms, known as false positives (FP) . This assumption is misaligned with operational reality, where an FN causes irreversible data loss and high recovery costs, while an FP is reversible. Detection is further complicated by ransomware variants with diverse behaviors, reducing the effectiveness of fixed supervised weighting strategies. To address these challenges, this study proposesa Security-Aware Deep Reinforcement Learning (SA-DRL) framework that embeds FN–FP costasymmetry into the reinforcement learning reward signal, optimizing detection policies to minimize missed detections. The framework also introduces a Security-Optimal Model Selection (SOMS) criterion and an adaptive sample-weighting mechanism through episode-level random permutation. Four DRL agents, DQN, DDQN, PPO, and A2C, were trained using a symmetric baseline reward (􀁒1) and a security-aware asymmetric reward (􀁒2) that penalizes FNs more heavily. Each configuration was evaluated with four discount factors, five-fold cross-validation, and three random seeds, resulting in 480 training runs on a balanced dataset. The SOMS criterion prioritizes minimizing the false-negative rate (FNR), maximizing the F1-score, and minimizing training time. Results show that asymmetric reward shaping improves detection performance. The SOMS-selected configuration, DDQN with 􀁒2 and 􀀍 = 0 . 1, achieved an FNR of 0.0080, an F1-score of 0.9915, and an AUC of 0.998, reducing missed detections by 67.6% compared with the best baseline model, MLP with FNR = 0.0247. 􀁒2 reduced the mean FNR by 43% relative to 􀁒1 across all configurations. These findings highlight the importance of reward-function design in security-sensitive detection systems. SA-DRL is the first framework to combine runtime features, asymmetric reward design, and a security-first model-selection criterion with statistical validation for ransomware detection.  \n1. Introduction  \nRansomware has emerged as one of the most operationally destructive and economically consequential categories of malware in the modern threat landscape. Unlike conventional malware that silently exfiltrates data, ransomware weaponizes cryptographic extortion by encrypting victim files and demanding payment, typically in cryptocurrency, to restore access. The impact extends well beyond ransom payment alone. In 2025, the global average cost of ransomware recovery reached $1.53 million, including downtime, investigation, regulatory exposure, and reputational damage [4] . Ransomware was involved in 44% of all data breaches analyzed in the Verizon 2025 Data Breach Investigations Report, representing a 12-percentage-point increase over the previous year [47] . Furthermore, global ransomware damage costs are projected to exceed $57 billion annually in 2025 and surpass $275 billion by 2031, with attacks occurring approximately every two seconds [32] . Modern ransomware campaigns are increasingly operated through ransomware-as-a-service (RaaS) ","cbCaivm7I6ruD48I","https://ap.wps.com/l/cbCaivm7I6ruD48I","pdf",3602675,3,1,22,"English","en",105,"# Introduction\n## Security-asymmetric costs in ransomware detection\n## Proposed SA-DRL approach","[{\"question\":\"How were the proposed methods evaluated in the study?\",\"answer\":\"Four DRL agents (DQN, DDQN, PPO, A2C) were trained using both a symmetric baseline reward and a security-aware asymmetric reward, then evaluated with multiple discount factors, five-fold cross-validation, three random seeds, and 480 training runs on a balanced dataset.\"}]",1784185660,55,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"sa-drl-security-aware-deep-reinforcement-learning-for-ransomware-detection-with-asymmetric-reward-design","",{"@graph":36,"@context":77},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/sa-drl-security-aware-deep-reinforcement-learning-for-ransomware-detection-with-asymmetric-reward-design/83157/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"How were the proposed methods evaluated in the study?","Question",{"text":75,"@type":76},"Four DRL agents (DQN, DDQN, PPO, A2C) were trained using both a symmetric baseline reward and a security-aware asymmetric reward, then evaluated with multiple discount factors, five-fold cross-validation, three random seeds, and 480 training runs on a balanced dataset.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]