[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83735-en":3,"doc-seo-83735-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83735,549758252649,"Ivy","https://ap-avatar.wpscdn.com/avatar/8000253669c5317157?_k=1778319167496531819",8,"Research & Report","ThreatVisionAI A Hybrid CNN-ViT Framework for Image-Based Malware Classification","Traditional malware detection struggles to generalize against obfuscated or previously unseen threats. ThreatVisionAI presents a hybrid malware family classification framework that combines a raw-image CNN, a wavelet-based CNN, and a Vision Transformer (ViT) to capture complementary spatial, frequency-domain, and global relational cues from malware images. Wavelet features provide multi-scale frequency discrimination for closely related families, while ViT models long-range dependencies. On the Malimg dataset, it reaches 98.01% accuracy and 0.9742 weighted F1, with Grad-CAM interpretability.","ThreatVisionAI: A Hybrid CNN-ViT Framework for Image-Based Malware Classification  \nAllyson Taylor  \nDepartment of Mathematics & Computer Science University of North Carolina Pembroke  \nPrashanth BusiReddyGari  \nDepartment of Mathematics & Computer Science University of North Carolina Pembroke  \narXiv :2607 .03653v 1 [ cs .CR] 4 Jul 2026  \nAbstract—Traditional malware detection methods struggle to generalize to obfuscated or previously unseen threats. This paper introduces ThreatVisionAI, a hybrid malware family classification framework that integrates a raw-image CNN, a waveletbased CNN, and a Vision Transformer (ViT) to capture complementary spatial, frequency-domain, and global relational features in malware images. The wavelet-based CNN captures multi-scale frequency information that helps distinguish closely related families, while the ViT branch models long-range dependencies across the image. Evaluated on the Malimg dataset, ThreatVisionAI achieves 98.01% accuracy and a weighted F1 score of 0.9742, with wavelet-domain features providing measurable gains on minority and visually similar families. These results confirm that frequency-aware and transformer-based representations improve image-based malware family classification.  \nIndex Terms—Malware Detection, Deep Learning, Convolutional Neural Networks, Vision Transformer, Explainable AI  \nI. INTRODUCTION  \nThe cybersecurity threat landscape faces constant pressure from malware that uses obfuscation, polymorphic transformations, and zero-day vulnerabilities to evade traditional defenses [1], [2] . Signature-based systems depend on known patterns and fail against unseen threats [3], [4] . Heuristic approaches offer broader coverage but produce high false-positive rates that undermine operational trust [3],[4] . These limitations motivate solutions that can adapt to modern malware ecosystems while maintaining high accuracy [5], [6] .  \nImage-based malware classification offers a promising alternative. By transforming binary files into visual representations, these methods capture structural relationships useful for distinguishing obfuscated and polymorphic variants atthe family level [7] . Accurate family classification enables incident responders to reuse knowledge about behavior and countermeasures for related samples, supporting faster triage and more targeted defense.  \nHowever, image-based classification faces several challenges. The volume and diversity of malware require scalable solutions that generalize across families and platforms [5],[6] . Attackers continuously adapt through obfuscation and camouflage, making reliable family assignment difficult [1],[2] . Many AI-driven frameworks also lack interpretability, leaving analysts without clear explanations for model decisions [8], [9] .  \nMany existing frameworks also fail to exploit frequency information useful for distinguishing closely related families.  \nCNNs capture local spatial patterns and ViTs model global structure, but both typically discard high-frequency details through conventional pooling or strided convolutions [10],[11] . Recent hybrid CNN–Transformer models such as LeViTMC [12] and ConvNeXt–Swin [13] improve over singlebranch architectures, but every branch in these frameworks operates exclusively in pixel space. No existing hybrid integrates a dedicated frequency-domain branch alongside both a spatial CNN and a global-context Transformer, leaving directional textural differences between visually similar families unexploited.  \nTo address these gaps, we propose ThreatVisionAI, a hybrid framework combining a raw-image CNN, a waveletbased CNN, and a ViT. The wavelet CNN captures spatial and frequency-domain characteristics, the ViT models global dependencies, and weighted soft voting fuses their predictions without additional trainable parameters.  \nThis research makes the following contributions:  \n1) Proposes ThreatVisionAI, a hybrid malware family classification framework combining a raw-image CN","cbCainjt5GvIrhRW","https://ap.wps.com/l/cbCainjt5GvIrhRW","pdf",784749,3,1,6,"English","en",105,"# Introduction\n# Literature Review","[{\"question\":\"What problem does ThreatVisionAI address in malware family classification?\",\"answer\":\"Traditional defenses often fail to generalize to obfuscated, polymorphic, or previously unseen malware. ThreatVisionAI targets more reliable family assignment by improving feature representation for image-based malware.\"},{\"question\":\"How does the framework combine CNN, wavelet-based CNN, and ViT?\",\"answer\":\"It uses a raw-image CNN for spatial patterns, a wavelet-based CNN to capture multi-scale frequency information, and a ViT branch to model global dependencies. Predictions are fused using weighted soft voting without additional trainable parameters.\"},{\"question\":\"What results does ThreatVisionAI achieve on the Malimg dataset and what explains the gains?\",\"answer\":\"It achieves 98.01% accuracy and a weighted F1 score of 0.9742. Wavelet-domain features deliver measurable improvements, especially for minority and visually similar malware families, and Grad-CAM supports the interpretation of model behavior.\"}]",1784190093,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"threatvisionai-a-hybrid-cnn-vit-framework-for-image-based-malware-classification","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/threatvisionai-a-hybrid-cnn-vit-framework-for-image-based-malware-classification/83735/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does ThreatVisionAI address in malware family classification?","Question",{"text":75,"@type":76},"Traditional defenses often fail to generalize to obfuscated, polymorphic, or previously unseen malware. ThreatVisionAI targets more reliable family assignment by improving feature representation for image-based malware.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the framework combine CNN, wavelet-based CNN, and ViT?",{"text":80,"@type":76},"It uses a raw-image CNN for spatial patterns, a wavelet-based CNN to capture multi-scale frequency information, and a ViT branch to model global dependencies. Predictions are fused using weighted soft voting without additional trainable parameters.",{"name":82,"@type":73,"acceptedAnswer":83},"What results does ThreatVisionAI achieve on the Malimg dataset and what explains the gains?",{"text":84,"@type":76},"It achieves 98.01% accuracy and a weighted F1 score of 0.9742. Wavelet-domain features deliver measurable improvements, especially for minority and visually similar malware families, and Grad-CAM supports the interpretation of model behavior.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]