[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81671-en":3,"doc-seo-81671-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81671,1649267921044,"Ava Thompson","https://us-avatar.wpscdn.com/avatar/1800007509477c92dfb?_k=1782875107921204101",8,"Research & Report","Contract Based Compositional Shielding for Safe Multi Agent Reinforcement Learning","Safe coordination problems arise in multi-agent reinforcement learning when no single agent can unilaterally guarantee global safety: an action’s admissibility depends on other agents’ evolving behaviors. Decentralised shields enforce safety at runtime, yet purely factorised permissions can eliminate optimal team behavior that is safe only through coordination. This work provides deterministic guarantees for decentralised execution by using a shared global LTLsafe specification ϕ and contract tuples of local obligations certified jointly for projection into action masks.","arXiv :2606 . 14 130v2 [ cs .LG] 9 Jul 2026  \nContract-Based Compositional Shielding for Safe Multi-Agent Reinforcement Learning  \nOmar Adalat 1 , Edwin Hamel-De le Court 1 ,2 , and Francesco Belardinelli 1  \n1 Imperial College London, London SW7 2AZ, UK  \n2 University of Manchester, Manchester M13 9PL, UK  \n{o.adalat24,[e.hamel.de-le-court](e.hamel.de-le-court) , [francesco.belardinelli}@imperial.ac.uk](francesco.belardinelli}@imperial.ac.uk)  \nAbstract. Safe coordination problems surface in multi-agent reinforcement learning when global safety cannot be enforced by any agent unilaterally: the admissibility of one agent’s action may depend on the dynamics of other agents. Decentralised shields can enforce safety at runtime, but purely factorised permissions often exclude optimal team behaviour that is safe only through coordination. We study deterministic safety guarantees for agents trained and deployed under decentralised execution, recovering team-optimal safe behaviour without centralised runtime control. Agents have a shared global specification ϕ in the safety fragment of Linear Temporal Logic (LTLsafe ), and select among tuples of local LTLsafe obligations whose conjunction implies the global specification ϕ . Each agent may rely on the other agents’ local obligations as assumptions because the whole contract tuple is certified simultaneously and allows projection into local action masks. At learning time, a non-stationary multi-armed bandit chooses among a library of local LTLsafe obligations to select the tuple that optimises team reward, all without forgoing end-to-end safety. We evaluate the approach across 6 environments and 15 algorithmic variants.  \nKeywords: Safe Multi-Agent Reinforcement Learning · Decentralised Learning · Compositional Verification · Shielding  \n1 Introduction  \nAssuring the safety of learning cooperative agents requires reconciling two demands: agents should optimise a shared task objective, whilst satisfying safety constraints during training and deployment. In safe coordination problems [11, 23], the admissibility of one agent’s action depends on the non-stationary policies of the other agents in the shared environment, so reasoning that treats teammate choices as arbitrary must discard behaviour that is safe only under coordinated actions. Multi-Agent Reinforcement Learning (MARL) offers a powerful framework for sequential decision-making under uncertainty in a shared environment,  \nby leveraging sampling-based methods to iteratively refine policies in stochastic Paper accepted to the 23rd European Conference on Multi-Agent Systems.  \n2 O. Adalat et al.  \n✓ centrally safe × centrally unsafe  enabled safe action  unsafe  \n⋆ locally rejected optimum  safe but not enabled here  \n(a) Central shield (b) Unilateral masks (c) Refined obligation  \np 1  \n1  \n0  \nSafe (x) {(0 , 0) ,(1 , 0) ,(0 , 1)} .  \np 1  \n1  \n0  \n0 1  \np2  \nSafe1 (x) × Safe2 (x) ={(0 , 0)} .  \np 1  \n1  \n0  \n0 1  \np2  \nφ2 : Always (p2 = 0)  \npermits 1 ∈ SC1 (x) .  \nFig. 1. A 2 × 2 safe-coordination instance in which unilateral local masks lose the safe optimal action, while a refined obligation recovers it through teammate commitment.  \ngames [18] . Stochastic games are commonly studied under cooperative, competitive, and mixed strategic dynamics. Many multi-agent systems naturally involve cooperative tasks, such as rescue drones [9] and autonomous warehouses [28], which makes the cooperative setting a natural target for safe MARL. However, standard reward penalties are insufficient to verify that behaviour is safe once a policy is deployed [14] . From the formal methods standpoint, shielding [16] is a popular technique for enforcing safety during both training and deployment by pre-emptively masking out unsafe actions that may lead to a specification violation or post-posedly replacing unsafe actions [3] .  \nExample 1 . Figure 1 gives a two-agent safe-coordination instance. At state x, each agent chooses an action pi ∈ {0, 1}, ","cbCaicDrcWvnWr8X","https://ap.wps.com/l/cbCaicDrcWvnWr8X","pdf",3427675,3,1,34,"English","en",105,"# Introduction\n## Contributions\n# Contract and Decomposition Approach\n## Example Safe Coordination Instance","[{\"question\":\"Why do safe coordination problems make unilateral action masking insufficient in multi-agent reinforcement learning?\",\"answer\":\"Because the admissibility of one agent’s action can depend on the other agents’ non-stationary policies, unilateral reasoning may discard behaviors that are only safe under coordinated teammate actions.\"},{\"question\":\"How does the method recover team-optimal safe behavior under decentralised execution?\",\"answer\":\"It uses a shared global safety specification ϕ in LTLsafe and selects tuples of local LTLsafe obligations whose conjunction implies ϕ. The full contract tuple is certified simultaneously, enabling sound projection into decentralised runtime action masks.\"},{\"question\":\"What role does the non-stationary multi-armed bandit play during learning?\",\"answer\":\"At learning time, a non-stationary multi-armed bandit chooses among a library of local LTLsafe obligations to form the tuple that optimizes team reward while preserving end-to-end safety.\"}]",1784175324,86,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"contract-based-compositional-shielding-for-safe-multi-agent-reinforcement-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/contract-based-compositional-shielding-for-safe-multi-agent-reinforcement-learning/81671/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do safe coordination problems make unilateral action masking insufficient in multi-agent reinforcement learning?","Question",{"text":75,"@type":76},"Because the admissibility of one agent’s action can depend on the other agents’ non-stationary policies, unilateral reasoning may discard behaviors that are only safe under coordinated teammate actions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the method recover team-optimal safe behavior under decentralised execution?",{"text":80,"@type":76},"It uses a shared global safety specification ϕ in LTLsafe and selects tuples of local LTLsafe obligations whose conjunction implies ϕ. The full contract tuple is certified simultaneously, enabling sound projection into decentralised runtime action masks.",{"name":82,"@type":73,"acceptedAnswer":83},"What role does the non-stationary multi-armed bandit play during learning?",{"text":84,"@type":76},"At learning time, a non-stationary multi-armed bandit chooses among a library of local LTLsafe obligations to form the tuple that optimizes team reward while preserving end-to-end safety.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]