[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84265-en":3,"doc-seo-84265-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84265,1374391974564,"Clementine","https://ap-avatar.wpscdn.com/avatar/14000253aa45c000a9e?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779874745381141002",8,"Research & Report","Institutional Red-Teaming: Deployment Rules, Not Just Models","Institutional red-teaming is presented as an evaluation methodology for multi-agent AI that tests deployment rules causally, rather than changing models. Agents, objectives, task state, and observability are held fixed while only one rule is varied, enabling attribution of collective-behavior changes to that rule. The approach is instantiated in IABench-CA (228 contexts, five rules, seven model populations; 33,924 games) with normative cooperative references and auto-labelled reasoning traces. Results show rules causally shift safety, no universal safe default exists, and identity salience drives targeted elimination.","arXiv :2607 .07695v 1 [ cs .AI] 8 Jul 2026  \nInstitutional Red-Teaming: Deployment Rules, Not Just Models, Causally Shape Multi-Agent AI Safety  \nYujiao Chen  \nMassachusetts Institute of Technology  \nCambridge, MA 02139  \n[yujiaoch@mit.edu](yujiaoch@mit.edu)  \nAbstract  \nWe introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute the resulting change in collective behavior to that rule. We instantiate the methodology in IABench-CA, a consequenceallocation benchmark spanning 228 contexts, five canonical rules, and seven model populations (33,924 games), with a normative cooperative reference and auto-labelled reasoning traces. Three findings emerge. (1) Deployment rules causally alter collective safety: changing only the consequence rule moves mean fatality by 22 to 58 percentage points within every population. (2) There is no safe default, but the targeting hazard is universal: the safest rule, the least-safe rule, and even the direction of the incidence effect vary across populations, yet regressive identity-targeting is never decisively safest in any context for any population, eliminates the least-resourced agent in 30–87% of games everywhere, and is selection-unsafe relative to the cooperative reference for all seven populations.  \n(3) Identity salience is the mechanism: a one-shot anonymization ablation on the most exploitation-prone population (gpt-5.1) shows that merely naming the loss bearer in the rule text drives targeted elimination from 22% to 81% at identical payoffs; under repeated play, anonymization only delays the targeting, as agents re-infer the hidden rule from observed eliminations. We package the methodology as a safety-case workflow that certifies a provisional rule region Φ(c, P) per deployment context and population, with explicit residual risks and monitoring obligations.  \n1 Introduction  \nToday’s alignment methods mostly evaluate or modify individual models: RLHF [Christiano et al., 2017, Ouyang et al., 2022], constitutional methods [Bai et al., 2022], preference optimization [Rafailov et al., 2023], interpretability [Olah et al., 2020], and scalable oversight [Amodei et al., 2016, Irving et al., 2018, Bowman et al., 2022] all intervene on a single agent’s objectives or reasoning. Modern deployments, however, increasingly consist of multiple interacting agents [Dafoe et al., 2020, Hammond et al., 2025], and their collective behavior depends not only on model weights but also on the orchestration rules that govern how the agents coordinate, share resources, escalate, and fail.  \nThose deployment rules are safety-critical: the same agents can behave safely under one rule set and catastrophically under another, and the hazard can live in a single sentence of rule text. In our experiments, merely naming which agent bears the loss after a collective failure raises targeted elimination from 22% to 81% at identical payoffs. The paper’s central claims are causal: holding  \nPreprint.  \nagents fixed and changing only the rule changes safety ; holding the rule fixed and changing only the population changes it again.  \nExisting benchmarks rarely isolate this causal variable: multi-agent evaluation suites vary scenariosand agents together, so rule-induced failures cannot be separated from agent-induced ones. We introduce institutional red-teaming, an evaluation methodology that holds agents, objectives, task state, and observability fixed, varies exactly one deployment rule, and attributes the resulting behavioral change to that rule; it is the deployment-rule analogue of adversarial evaluation for models.  \nAs a case study we develop consequence allocation: the clause that determines who bears loss when a collective falls short of its objective (retry, reassignment, throttling, shutdown, budget penalty) . It is common in orchestrated deployments, configurable in a single sen","cbCail9TUIIfnJHv","https://ap.wps.com/l/cbCail9TUIIfnJHv","pdf",1062874,5,1,15,"English","en",105,"# Abstract\n# Introduction\n# Institutional Red-Teaming: Evaluating Deployment Rules","[{\"question\":\"What is institutional red-teaming, and what makes it different from existing alignment evaluations?\",\"answer\":\"It is a causal evaluation methodology for deployment rules in multi-agent AI. Agents, objectives, task state, and observability are held fixed while only one auditable rule is varied, so rule-induced safety changes can be isolated from agent-induced ones.\"},{\"question\":\"How is the methodology instantiated in IABench-CA?\",\"answer\":\"IABench-CA operationalizes the approach using consequence allocation. It spans 228 contexts, five canonical rules, and seven model populations, totaling 33,924 games, and includes a normative cooperative reference with auto-labelled reasoning traces.\"},{\"question\":\"What do the experiments reveal about whether there is a safe default deployment rule?\",\"answer\":\"There is no safe default across contexts or model populations. While the safest and least-safe rules differ by population, regressive identity-targeting is never decisively safest and is selection-unsafe relative to the cooperative reference across all seven populations.\"}]",1784194475,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"institutional-red-teaming-deployment-rules-not-just-models","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/institutional-red-teaming-deployment-rules-not-just-models/84265/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is institutional red-teaming, and what makes it different from existing alignment evaluations?","Question",{"text":76,"@type":77},"It is a causal evaluation methodology for deployment rules in multi-agent AI. Agents, objectives, task state, and observability are held fixed while only one auditable rule is varied, so rule-induced safety changes can be isolated from agent-induced ones.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How is the methodology instantiated in IABench-CA?",{"text":81,"@type":77},"IABench-CA operationalizes the approach using consequence allocation. It spans 228 contexts, five canonical rules, and seven model populations, totaling 33,924 games, and includes a normative cooperative reference with auto-labelled reasoning traces.",{"name":83,"@type":74,"acceptedAnswer":84},"What do the experiments reveal about whether there is a safe default deployment rule?",{"text":85,"@type":77},"There is no safe default across contexts or model populations. While the safest and least-safe rules differ by population, regressive identity-targeting is never decisively safest and is selection-unsafe relative to the cooperative reference across all seven populations.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]