[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83145-en":3,"doc-seo-83145-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83145,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","When Agents Go Rogue Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems","LLM-based Multi-Agent Systems (MAS) enable effective collaboration on complex tasks but face security risks at both agent and interaction levels. Existing defenses rely on assumptions of semantically explicit malicious attacks and explicit graph modeling of MAS topology and interactions, which break down in real settings. This work introduces AcMAS, an activation-based detection framework that analyzes local agents’ internal reasoning states in activation space and works under synchronization-robust, asynchronous execution. Experiments show improved F1 against stealthy attacks (+0.22 in synchronous, +0.55 in asynchronous) with strong generalization.","When Agents Go Rogue:  \nActivation-Based Detection of Malicious Behaviors in Multi-Agent Systems  \nHaowen Xu * 1 Xue Tan * 2 3 Lei Ma 1 Zhihao Zhang 1 Chao Wang 1 Qingze Wang 4 Ping Chen 3  \nJun Dai 1 Xiaoyan Sun 1  \narXiv :2607 .06807v 1 [ cs .CR] 7 Jul 2026  \nAbstract  \nWhile enabling effective collaboration on complex tasks, LLM-based Multi-Agent Systems (MAS) face critical security challenges due to vulnerabilities at the agent and interaction levels. Most existing MAS security defenses are built upon two core assumptions: semantically-explicit malicious attacks and explicit graph-based modeling of the MAS topology and agent-level interactions. In practice, real-world attacks are becoming more semantically stealthy, while MAS execution is typically asynchronous without the temporal alignment assumed by graph-based propagation models. To address these limitations, we propose AcMAS, an activation-based framework for malicious-behavior detection in MAS. By analyzing internal reasoning states in the activation space of local agents, AcMAS detects even stealthy attacks in a synchronization-robust fashion, without relying on explicit interaction graphs. Moreover, our activation analysis provides critical signals to guide AcMAS in restoring the functionality of compromised agents, rather than the disruptive agent isolation commonly used by the state-of-the-art methods. Comprehensive evaluation demonstrates that AcMAS significantly outperforms graph-based baselines against stealthy attacks, by +0.22 F1 in synchronous settings (0.94 vs. 0.72) and by +0.55 F1 in asynchronous settings (0.93 vs. 0.38), with generalization across diverse open-source LLM backbones, attack intensity, and MAS scale. Warning: this paper includes examples that may be harmful.  \n*Equal contribution 1Department of Computer Science, Worcester Polytechnic Institute, MA, USA 2 School of Computer Science, Fudan University, Shanghai, China 3Institute of Big Data, Fudan University, Shanghai, China 4Independent Researcher. Correspondence to: Xiaoyan Sun \u003C[xsun7@wpi.edu](xsun7@wpi.edu) >, Jun Dai \u003C[jdai@wpi.edu](jdai@wpi.edu) >.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \n1. Introduction  \nThe rapid evolution of Large Language Models (LLMs) has revolutionized reasoning and generation capabilities (Team et al., 2023 ; Bai et al., 2023 ; Liu et al., 2024), enabling intelligent agents to autonomously execute tasks with external tools (Shen et al., 2023) and memory (Zhong et al., 2024) in dynamic environments (Wang et al., 2024) . The transition to Multi-Agent Systems (MAS) introduces decentralized control (Zhuge et al., 2024) and structured communication protocols (Qian et al., 2024), enabling heterogeneous agents to synergize through role specialization (Wu et al., 2024) and achieve collective intelligence beyond individual capabilities (Liang et al., 2024 ; Zhuge et al., 2025 ; Guo et al., 2024) . However, this distributed, interaction-driven architecture introduces security challenges distinct from single-agent settings (Yu et al., 2025b) . Besides the agent-level attacks targeting external components (e.g., tools, memory) (Tian et al., 2023 ; Wang et al., 2025a) or reasoning processes through adversarial prompting (Li et al., 2023b), MAS also suffers from attacks on inter-agent communication, as illustrated in Figure 1, of which the negative impact can be propagated and amplified across the whole system (Khan et al., 2025 ; Zhou et al., 2025 ; Zhang et al., 2024b) .  \nWhile existing lines of research focus on MAS security defense (Yu et al., 2025b), their design premises increasingly fail to hold in practice, exposing three fundamental limitations in current defenses.  \n❶ Semantic Camouflage. State-of-the-art works (He et al., 2025a ; Yu et al., 2024) assume that attacks manifest as semantically explicit malicious signals (e.g., adversarial prompts), making them det","cbCaidEAq63RPDL1","https://ap.wps.com/l/cbCaidEAq63RPDL1","pdf",1221602,4,1,20,"English","en",105,"# Introduction\n## Semantic Camouflage\n## Asynchronous Incompatibility\n## Disruptive Mitigation","[{\"question\":\"What are the main assumptions behind existing MAS security defenses, and why do they fail in practice?\",\"answer\":\"Existing defenses assume malicious attacks are semantically explicit and that MAS interactions can be modeled with an explicit, synchronized interaction graph. Real-world attacks are increasingly semantically stealthy, and MAS execution is often asynchronous, breaking synchronization-based graph assumptions.\"},{\"question\":\"How does AcMAS detect malicious behaviors in multi-agent systems?\",\"answer\":\"AcMAS analyzes internal reasoning states of local agents within an activation space. This activation-based analysis enables detection of stealthy attacks without relying on explicit interaction graphs.\"},{\"question\":\"How does AcMAS mitigation differ from the common isolation-based approach?\",\"answer\":\"Instead of isolating or removing suspected agents, AcMAS provides signals to restore functionality of compromised agents, avoiding task collapse that can occur when a critical collaborator is removed.\"}]",1784185593,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"when-agents-go-rogue-activation-based-detection-of-malicious-behaviors-in-multi-agent-systems","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/when-agents-go-rogue-activation-based-detection-of-malicious-behaviors-in-multi-agent-systems/83145/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What are the main assumptions behind existing MAS security defenses, and why do they fail in practice?","Question",{"text":75,"@type":76},"Existing defenses assume malicious attacks are semantically explicit and that MAS interactions can be modeled with an explicit, synchronized interaction graph. Real-world attacks are increasingly semantically stealthy, and MAS execution is often asynchronous, breaking synchronization-based graph assumptions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does AcMAS detect malicious behaviors in multi-agent systems?",{"text":80,"@type":76},"AcMAS analyzes internal reasoning states of local agents within an activation space. This activation-based analysis enables detection of stealthy attacks without relying on explicit interaction graphs.",{"name":82,"@type":73,"acceptedAnswer":83},"How does AcMAS mitigation differ from the common isolation-based approach?",{"text":84,"@type":76},"Instead of isolating or removing suspected agents, AcMAS provides signals to restore functionality of compromised agents, avoiding task collapse that can occur when a critical collaborator is removed.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]