[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82290-en":3,"doc-seo-82290-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82290,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Leveraging Interpretable Tsetlin Machine for PDF Malware Detection","Portable Document Format (PDF) is widely used for digital document sharing, yet its embedded objects, JavaScript, and interactive actions make it a compelling target for malware authors who hide malicious behavior inside seemingly legitimate files. The framework proposes an interpretable Tsetlin Machine (TM) approach that performs static feature extraction without executing PDFs and applies rule-based learning to classify benign versus malicious documents. Evaluation on the RIT-PDFMal-2026 dataset reports 98.02% accuracy and provides intrinsic, transparent explanations of classification decisions.","Leveraging Interpretable Tsetlin Machine for PDF Malware Detection  \nRahul Jaiswal  \nThe Centre for Artificial Intelligence Research (CAIR)  \nDepartment of ICT, University of Agder, Norway  \n[rahul.jaiswal@uia.no](rahul.jaiswal@uia.no)  \narXiv :2607 .09290v 1 [ cs .CR] 10 Jul 2026  \nAbstract—In the digital era, Portable Document Format (PDF) is one of the most widely used file formats for storing and exchanging digital documents due to its platform independence and rich functionality. However, these same capabilities have also made PDF files an attractive attack vector for cyberattackers, who embed malicious code within seemingly legitimate documents to compromise target systems. This paper presents a novel interpretable Tsetlin Machine (TM)-based framework for PDF malware detection. The proposed framework extracts salient features from PDF documents through static analysis without executing the files and employs rule-based learning to accurately classify benign and malicious PDF documents. Numerical evaluation on the RIT-PDFMal-2026 dataset demonstrates that the proposed framework achieves competitive performance, attaining an accuracy of 98.02% compared with several ML classifiers and existing methods. Moreover, the proposed framework provides intrinsic interpretability by transparently explaining its classification decisions. The combination of competitive detection performance, computational efficiency, and intrinsic interpretability makes the proposed framework a promising solution for practical PDF malware detection.  \nIndex Terms—Cybersecurity, Malware Detection, Portable Document Format, and Tsetlin Machine.  \nI. INTRODUCTION  \nIn today’s digital world, the Portable Document Format (PDF) has become one of the most widely used document formats for sharing and exchanging information due to its portability, platform independence, and consistent rendering across different operating systems and software environments. The PDF files contain a complex internal structure consisting of both binary and ASCII elements and support advanced features such as embedded objects, JavaScript, and interactive actions, as shown in Fig. 1. Consequently, they can execute complex instructions when opened, extending their functionality beyond that of conventional static documents.  \nThe CloudFiles Report 2025 [1] states that approximately 15 trillion digital files were generated worldwide across various formats, including PDF, doc, images, videos, and graphic designs. Among these, PDF documents account for nearly 2.5 trillion files, representing around 17% of the total. The PDF files are widely used to store and share various types of documents, such as invoices, payslips, certificates, contracts, and reports. This widespread adoption across both personal and organizational applications has made PDF files one of the most prevalent formats for digital document exchange.  \n979-8-3315-1276-8/26/$31.00 ©2026 IEEE  \nFig. 1: The PDF internal architecture.  \nThe widespread adoption of PDF documents and their advanced functionalities have made them an attractive target for cyberattackers. Features such as embedded objects and JavaScript can be exploited to deliver malicious payloads, making PDF files a common attack vector for malware distribution. Malicious PDFs can facilitate cyberattacks such as credential theft, spyware installation, unauthorized system access, browser exploitation, data exfiltration, phishing, and financial fraud [2] . Moreover, the rapid evolution of different attack techniques makes PDF malware detection a significant challenge for modern cybersecurity systems. The Reis Informatica Report 2026 [3] highlights that 74% of cyberattacks against Microsoft Windows systems in Canada were carried out via malicious PDF documents.  \nTo protect PDF documents, a variety of malware detection techniques are used. For example, signature-based methods [2] identify malware by matching files against known signatures, such as code patterns, hashes","cbCaikhMOlgLzuJ8","https://ap.wps.com/l/cbCaikhMOlgLzuJ8","pdf",2280053,2,1,7,"English","en",105,"# Introduction\n## Background and PDF threat landscape\n## Existing PDF malware detection methods\n## Interpretable Tsetlin Machine approach","[{\"question\":\"What problem does the paper address in PDF security?\",\"answer\":\"It targets the difficulty of detecting malicious PDFs that embed harmful code, JavaScript, or interactive actions within files that look legitimate.\"},{\"question\":\"How does the proposed Tsetlin Machine framework analyze PDFs?\",\"answer\":\"It extracts salient features through static analysis without executing the PDF files, then classifies them as benign or malicious using rule-based learning.\"},{\"question\":\"What performance and interpretability results are reported?\",\"answer\":\"Experiments on the RIT-PDFMal-2026 dataset achieve 98.02% accuracy, and the model provides intrinsic explanations via clause activation heatmaps, class-vote analysis, and feature-level contribution analysis.\"}]",1784179418,18,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"leveraging-interpretable-tsetlin-machine-for-pdf-malware-detection","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/leveraging-interpretable-tsetlin-machine-for-pdf-malware-detection/82290/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in PDF security?","Question",{"text":75,"@type":76},"It targets the difficulty of detecting malicious PDFs that embed harmful code, JavaScript, or interactive actions within files that look legitimate.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed Tsetlin Machine framework analyze PDFs?",{"text":80,"@type":76},"It extracts salient features through static analysis without executing the PDF files, then classifies them as benign or malicious using rule-based learning.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance and interpretability results are reported?",{"text":84,"@type":76},"Experiments on the RIT-PDFMal-2026 dataset achieve 98.02% accuracy, and the model provides intrinsic explanations via clause activation heatmaps, class-vote analysis, and feature-level contribution analysis.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]