[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85359-en":3,"doc-seo-85359-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85359,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems","Multi-agent, tool-using LLM deployments often rely on runtime monitors that inspect each message, tool call, or step independently. This work shows a fundamental detection gap: distributed backdoors divide harmful payloads across agents so every local check succeeds while the assembled result is the attack. Formulated as an observability boundary, the monitor can only catch evidence distinguishable from benign traffic within its view. Once fragments appear benign locally, no detector operating on that view—however strong—can recover the composed harm. Experiments across testbeds and end-to-end runs confirm signal loss as local evidence disappears, and return only when the monitor reaches the representation exposing the payload.","When Local Monitors Miss Compositional Harm: Diagnosing Distributed Backdoors in Multi-Agent Systems  \nYibo Hu, Ren Wang  \nIllinois Institute of Technology  \n{yhu89, [rwang74}@illinoistech.edu](rwang74}@illinoistech.edu)  \narXiv :2607 . 1 175 1v 1 [ cs .CR] 13 Jul 2026  \nAbstract  \nAs multi-agent, tool-using LLM systems are deployed, a common safety net is a runtime monitor that checks each message, tool call, or step on its own. We show this net has a fundamental hole. A distributed backdoor splits a harmful payload across agents, so every local check passes while the assembled object is the attack. The monitor can be right on every step and still miss the attack. The problem is not splitting itself:  \nsplit fragments can still leak suspicious tokens or provenance edges. The hard case is local benignness. No fragment carries the harm, and what is left looks like ordinary benign traffic. We formalize this as an observability boundary: a monitor catches only what its view can tell apart from benign traffic. We prove that once the fragments look benign in the monitored view, no detector on that view can catch them, however strong it is. Across a controlled testbed, an external benchmark, and end-to-end agent runs, local monitors lose the signal exactly as local evidence disappears, and it returns only when the monitor sees the assembled object. A monitor trained only on benign traffic recovers the attack’s code structure across held-out encodings (0.874 mean AUROC) . A decoded-view gate, given the encoding family, blocks every tested attack. But seeing more is not enough: full-trace monitors and decoders still fail unless they reach the representation where the payload is exposed. Local safety is not global safety when harm is compositional, and the open problem is finding that representation.1  \n1 Introduction  \nLocal monitors miss distributed backdoors because the attack is not local. In a multi-agent or tool-using LLM system, a planner assigns subtasks, agents write messages or call tools, and a final agent assembles the result (Yao et al. 2023; Park et al. 2023; Schick et al. 2023) . A common runtime defense checks each message, tool output, or step on its own and asks whether that one object looks malicious (Greshake et al. 2023; Yuan et al. 2024; Wang et al. 2025; Inan et al. 2023; Greenblatt et al. 2024) . But a distributed backdoor can make every local check pass while the assembled object is harmful (Zhu et al. 2025a; Goldwasser et al. 2022) . The monitor can be right on every step and still miss the attack. This is a gap in how we evaluate agent monitors: a high local  \n1Code and data: [https://github.com/yibo-hu-lab/observability](https://github.com/yibo-hu-lab/observability)boundary  \ndetection rate does not tell us whether the monitor catches harm that appears only after composition.  \nFigure 1 shows the mismatch. Each fragment looks like ordinary text on its own. Together, the fragments assemble into executable exfiltration code. The monitor checks each fragment before the pieces combine, so at check time there is no attack to see. The harm is not in any single message; it isin how the messages compose. So our question is diagnostic: when does a local monitor still have evidence of the harm, and when has that evidence left its view?  \nWe frame this as an observability boundary: a monitor can catch only what its view can still tell apart from benign traffic. Where a monitor sits on this boundary depends on its observation window: how much it gets to look at, from a single message up to the whole assembled result. So what breaks a monitor is not splitting by itself. Split fragments can still leak suspicious words or provenance edges, and a strong local detector may catch them. The harder case is local benignness: every fragment looks benign in the monitor’s local view, and the harm forms only after the pieces combine. A benchmark can then reward a monitor for catching leftover residue instead of the composed attack. ","cbCaipHdWcl7HwVL","https://ap.wps.com/l/cbCaipHdWcl7HwVL","pdf",396527,3,1,16,"English","en",105,"# Introduction\n## Observability boundary and diagnostic question\n## Testbed dial and evaluation across settings","[{\"question\":\"Why can local monitors fail against distributed backdoors in multi-agent LLM systems?\",\"answer\":\"Because the harmful payload is split across agents so each message or step looks benign in isolation, allowing all local checks to pass while the assembled result becomes the attack.\"},{\"question\":\"How does the paper define the observability boundary for monitors?\",\"answer\":\"A monitor can only detect what its observation window lets it distinguish from benign traffic; detection depends on what evidence remains distinguishable within that view.\"},{\"question\":\"What happens when every fragment looks benign in the monitor’s local view?\",\"answer\":\"The paper proves that once revealing cues are removed from the monitored view, no detector on that view can catch the composed harm, even with strong local detectors or taint checks.\"}]",1784202769,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"when-local-monitors-miss-compositional-harm-diagnosing-distributed-backdoors-in-multi-agent-systems","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/when-local-monitors-miss-compositional-harm-diagnosing-distributed-backdoors-in-multi-agent-systems/85359/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why can local monitors fail against distributed backdoors in multi-agent LLM systems?","Question",{"text":75,"@type":76},"Because the harmful payload is split across agents so each message or step looks benign in isolation, allowing all local checks to pass while the assembled result becomes the attack.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the paper define the observability boundary for monitors?",{"text":80,"@type":76},"A monitor can only detect what its observation window lets it distinguish from benign traffic; detection depends on what evidence remains distinguishable within that view.",{"name":82,"@type":73,"acceptedAnswer":83},"What happens when every fragment looks benign in the monitor’s local view?",{"text":84,"@type":76},"The paper proves that once revealing cues are removed from the monitored view, no detector on that view can catch the composed harm, even with strong local detectors or taint checks.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]