[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84595-en":3,"doc-seo-84595-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84595,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Forensic-Oriented Intrusion Detection Using Synthetic Network Traffic Data and Explainable Artificial Intelligence","Digital forensic investigations of network intrusions require analytical outputs that are traceable, reproducible, and court-defensible, yet conventional machine-learning pipelines use original evidence for training and produce opaque classifications without instance-level justification. This paper proposes a forensic-oriented intrusion detection framework that unifies synthetic data generation, binary classification, and explainable AI in a single pipeline governed by ISO/IEC 27037, 27041, 27042 and NIST SP 800-86, ensuring evidentiary integrity while enabling interpretable forensic reporting.","Forensic-Oriented Intrusion Detection Using Synthetic Network Traffic Data and Explainable Artificial Intelligence  \nJosé Luis Vela¹ *, Carmen Pellicer¹  \n¹CESTE, University Center, Zaragoza, Spain  \n*Corresponding author: José Luis Vela Alonso ([j](jlvela@ceste.com)[lvela@ceste.com](jlvela@ceste.com))  \nHighlights  \n• Forensic pipeline integrates synthetic data and XAI under ISO/IEC standards.  \n• Framework aligns with ISO 27037, 27042 and NIST SP 800-86 requirements.  \n• Synthetic-trained XGBoost achieves F1-macro=0.96 on real network traffic.  \n• KS testing confirms synthetic data utility with privacy preservation.  \n• SHAP maps flow features to forensic indicators across three attack types.  \nAbstract  \nDigital forensic investigations of network intrusions require analytical outputs that are traceable, reproducible, and court-defensible—requirements that existing machine learning pipelines do not satisfy because they treat original evidence as training data and produce opaque classifications without instance-level justification. This paper presents a forensic-oriented intrusion detection framework that resolves both problems simultaneously, constituting the first integration of synthetic data generation, binary classification, and explainability within a single pipeline explicitly governed by ISO/IEC 27037, ISO/IEC 27041, ISO/IEC 27042, and NIST SP 800-86.  \nThe framework operationalises the ISO/IEC 27037 requirement for strict separation between original digital evidence and derived analytical artefacts — a requirement rarely implemented in machine learning workflows. Original datasets are treated as immutable, hash-verified artefacts; all training operates on parameterized synthetic derivatives generated via SDV + CTGAN. XGBoost binary classification provides high-performance detection on tabular network flow data, and SHAP TreeExplainer produces instance-level feature attributions that map statistical predictions to observable network behaviour for forensic reporting.  \nTrain-on-Synthetic, Test-on-Real (TSTR) evaluation on CICIDS2017 achieves F1-macro = 0.96, within cross-validation variance of the real-data baseline (0.97) . Kolmogorov–Smirnov testing confirms synthetic privacy preservation (mean |KS| = 0.38) alongside operational utility. Cross-dataset validation on UNSW-NB15 and Kitsune identifies feature space dimensionality as the primary determinant of synthetic training effectiveness, establishing a practical deployment boundary of approximately 30 numeric flow-level features. SHAP attributions for Brute Force, Port Scan, and DoS attacks are consistent across real and synthetic instances, confirming that synthetic training preserves the forensically relevant attack fingerprints required for expert witness testimony.  \nKeywords: digital forensic investigation; network traffic analysis; synthetic data; explainable artificial intelligence; SHAP; XGBoost  \n1. Introduction  \nNo prior framework integrates synthetic data generation, machine learning-based intrusion detection, and forensic explainability within a single pipeline governed by established digital forensic standards— this paper presents that framework. Digital forensic investigation of network traffic occupies a position of increasing institutional and legal significance: its outputs are scrutinised by internal audit boards, regulatory bodies, and courts, and must meet evidentiary standards that predictive performance metrics alone cannot satisfy (Casey, 2011; Pollitt, 2010) . A model that classifies network flows as malicious with high accuracy provides an investigator with little practical value if the basis of each classification cannot be explained, independently reproduced, or defended under cross-examination. This constraint distinguishes forensic applications from general intrusion detection: accuracy is necessary but not sufficient.  \nTwo structural barriers prevent the forensic adoption of machine learning-based intrusion detection. The first is a d","cbCairIsIvrGqDOH","https://ap.wps.com/l/cbCairIsIvrGqDOH","pdf",879677,3,1,23,"English","en",105,"# Introduction\n## Forensic need and limitations of existing intrusion detection\n## Forensic-oriented framework design principles\n## Synthetic data generation and explainable classification integration","[{\"question\":\"Why do traditional machine-learning intrusion detection pipelines fall short for digital forensics?\",\"answer\":\"They often treat original evidence as training data and output opaque classifications that cannot be independently explained, reproduced, or defended in evidentiary settings.\"},{\"question\":\"How does the proposed framework maintain evidentiary integrity?\",\"answer\":\"It enforces strict separation between immutable original evidence and derived analytical artefacts, using integrity-hash verification for originals and parameterized synthetic derivatives for training and analysis with logged transformations.\"},{\"question\":\"What evaluation results demonstrate the effectiveness and forensic relevance of synthetic training?\",\"answer\":\"Train-on-synthetic/test-on-real evaluation on CICIDS2017 reports F1-macro=0.96 near the real-data baseline, KS testing indicates privacy preservation, and SHAP attributions remain consistent across real and synthetic instances for multiple attack types.\"}]",1784196994,58,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"forensic-oriented-intrusion-detection-using-synthetic-network-traffic-data-and-explainable-artificial-intelligence","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/forensic-oriented-intrusion-detection-using-synthetic-network-traffic-data-and-explainable-artificial-intelligence/84595/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-21","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do traditional machine-learning intrusion detection pipelines fall short for digital forensics?","Question",{"text":75,"@type":76},"They often treat original evidence as training data and output opaque classifications that cannot be independently explained, reproduced, or defended in evidentiary settings.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed framework maintain evidentiary integrity?",{"text":80,"@type":76},"It enforces strict separation between immutable original evidence and derived analytical artefacts, using integrity-hash verification for originals and parameterized synthetic derivatives for training and analysis with logged transformations.",{"name":82,"@type":73,"acceptedAnswer":83},"What evaluation results demonstrate the effectiveness and forensic relevance of synthetic training?",{"text":84,"@type":76},"Train-on-synthetic/test-on-real evaluation on CICIDS2017 reports F1-macro=0.96 near the real-data baseline, KS testing indicates privacy preservation, and SHAP attributions remain consistent across real and synthetic instances for multiple attack types.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]