[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85337-en":3,"doc-seo-85337-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85337,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Graph-Based Structural Evaluation of LLM-Translated Adversary Emulation Procedures","Adversary emulation plans describe attacker procedures using MITRE ATT&CK techniques, privilege requirements, and expected observable telemetry. Cross-platform translation is needed for defender evaluation, but LLMs may preserve technique labels while breaking structural and evidence-level equivalence, yielding unusable coverage. Graph-Based Structural Evaluation (GBSE) models each procedure as a directed attributed graph and computes normalized Graph Edit Distance across progressively stricter layers: technique, tactic, telemetry class, and Sigma logsource. Experiments on a 29-step ALPHV/BlackCat Windows!Linux plan show technique and tactic fidelity preserved (GED=0), telemetry fidelity reduced, while Sigma-layer matching fully recovers.","arXiv :2607 . 1 15 17v 1 [ cs .CR] 13 Jul 2026  \nFUJITSU RESEARCH OF EUROPE · Security Science Research Group  \nTechnical White Paper | In collaboration with MITRE Research | 10 June 2026  \n\n| Graph-Based Structural Evaluation of LLM-Translated Adversary Emulation Procedures\u003Cbr>This technical contribution supports the MITRE white paper titled: Evaluating LLMs for Impact-Faithful Translation of Adversary Behavior Across Operating\u003Cbr>Systems |\n| --- |\n| Ahmed M. Elmisery\u003Cbr>Security Science Research Group\u003Cbr>Fujitsu Research of Europe Limited\u003Cbr>Abstract\u003Cbr>Adversary emulation plans specify multi-step attacker procedures at the level of MITREATT&CK techniques, privilege requirements, and observable telemetry. Translating such plans across operating systems is necessary for cross-platform defender evaluation, and large language models (LLMs) make automated translation tractable. They also introduce a qualityassurance problem: a translation that renames tools while retaining source-platform constructs yields no usable coverage for defenders of the target platform. Binary question-based scoring tends to overestimate how faithful these translations are, because it measures countable properties rather than structural, observable, or rule-level equivalence.\u003Cbr>Graph-Based Structural Evaluation (GBSE) addresses this gap. Each procedure is modelled as a directed attributed graph, and fidelity is computed as a normalized Graph Edit Distance (GED) across four progressively stricter node-matching layers: technique (L0), tactic (L1), telemetry class (L2), and Sigma logsource (L3) . Applied to the full 29-step ALPHV/BlackCat Windows!Linux plan, with a genuine native-Windows control reconstructed step-for-step from the conversion record and a Linux variant taken unmodified from the LLM output, the framework finds: technique and tactic structure is preserved across the OS boundary (GED=0, Simstruct =1 .000); telemetry fidelity drops to Simstruct =0 .897 (GED=3), driven by three steps that emit an unmapped observable class or drift in telemetry; and independent Sigma-layer matching recovers to Simstruct =1 .000 across all 29 steps. Every state classifies as Medium Fidelity (best composite S=0 .674), and the deployment gate (S􀀕0 .80, requiring technical realism 􀀕 0.990 against the measured 0.43) is unreachable at current evaluation quality. The layer scores are numerically identical to those obtained under an earlier same-procedure design, which confirms that the layers decompose cleanly along the OS-abstraction axis.\u003Cbr>The framework includes a bipartite-GED implementation, a telemetry-intent parser that derives structured observable classes from free-text annotations, and a validated library of\u003Cbr>49 Sigma detection rules (19 Linux, 30 Windows) that gives complete ATT&CK technique coverage of the procedure and passes the Sigma specification validator with zero findings. A supplementary analysis recovers a genuine technique-level divergence (for example RDP-based external access reassigned to unencrypted exfiltration, and credential-store access reassigned to remote-system discovery) that a procedure-aligned view necessarily suppresses. All numerical results in this paper were reproduced from the reference implementation and asserted against the recorded pipeline outputs.\u003Cbr>Keywords: adversary emulation, MITRE ATT&CK, graph edit distance, large language models, cross-platform translation, Sigma, detection engineering, CALDERA. |\n\nNotation  \n\n| Symbol | Definition |\n| --- | --- |\n| G = (V, E, φV , φE ) | Directed attributed procedure graph: step nodes, dependency edges, node/edge attribute maps. |\n| Gc , Gv | Control graph (human-validated source) and variant graph (LLM translation) . |\n| GED(Gc , Gv) | Minimum-cost edit path transforming Gc into Gv . |\n| Simstruct (Gc , Gv) | 1 􀀀 GED/ max(jVc j , jVvj) . Trial normaliser max(29 , 29) = 29 . |\n| szss(A, B) | jA \\ Bj/ max(jAj , j Bj), the Szymkiewicz–Simpson overlap coeﬀicient. |\n| θtac = θsig =","cbCaismECWR9DDaT","https://ap.wps.com/l/cbCaismECWR9DDaT","pdf",219891,2,1,19,"English","en",105,"# Abstract\n# Introduction and Motivation\n# Notation\n# Graph-Based Structural Evaluation (GBSE)\n## Node-matching layers: technique, tactic, telemetry class, Sigma logsource\n## Case study and fidelity results\n## Implementation and Sigma rule validation","[{\"question\":\"Why is translating adversary emulation procedures across operating systems necessary?\",\"answer\":\"Defender evaluation in mixed environments requires the same multi-step procedure expressed on both platforms. Hand-authoring the second version is slow, so cross-platform translation enables scalable evaluation.\"},{\"question\":\"What problem does graph-based evaluation solve compared with question-based scoring?\",\"answer\":\"Question-based scoring can overestimate fidelity because it measures countable properties, not structural, observable, or rule-level equivalence. GBSE instead evaluates equivalence using graph structure and progressively stricter matching layers.\"},{\"question\":\"How is fidelity computed in GBSE?\",\"answer\":\"Each procedure is represented as a directed attributed graph, and fidelity is derived from a normalized Graph Edit Distance. Matching is computed through four layers: technique (L0), tactic (L1), telemetry class (L2), and Sigma logsource (L3).\"}]",1784202593,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"graph-based-structural-evaluation-of-llm-translated-adversary-emulation-procedures","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/graph-based-structural-evaluation-of-llm-translated-adversary-emulation-procedures/85337/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is translating adversary emulation procedures across operating systems necessary?","Question",{"text":75,"@type":76},"Defender evaluation in mixed environments requires the same multi-step procedure expressed on both platforms. Hand-authoring the second version is slow, so cross-platform translation enables scalable evaluation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem does graph-based evaluation solve compared with question-based scoring?",{"text":80,"@type":76},"Question-based scoring can overestimate fidelity because it measures countable properties, not structural, observable, or rule-level equivalence. GBSE instead evaluates equivalence using graph structure and progressively stricter matching layers.",{"name":82,"@type":73,"acceptedAnswer":83},"How is fidelity computed in GBSE?",{"text":84,"@type":76},"Each procedure is represented as a directed attributed graph, and fidelity is derived from a normalized Graph Edit Distance. Matching is computed through four layers: technique (L0), tactic (L1), telemetry class (L2), and Sigma logsource (L3).","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]