[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81657-en":3,"doc-seo-81657-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81657,16904993612988,"Olivia Brown","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Heterogeneous Information-Bottleneck Coordination Graphs for Multi-Agent Reinforcement Learning","Coordination graphs guide which agents exchange information in cooperative multi-agent reinforcement learning, but existing sparse-graph methods use heuristic edge-uniform topology criteria and apply structurally blind bottlenecks to message bandwidth. Heterogeneous Information-Bottleneck Coordination Graphs (HIBCG) jointly learns group-aware connectivity and allocates differentiated communication capacity. It couples a group-aligned block-diagonal prior with per-agent message compression on the learned graph. Results on SMACv1, SMACv2, and MAgent Battle (up to 100 agents) show strong gains on heterogeneous multi-role tasks, improved convergence, and theory-backed diagnostics without a separate mutual-information estimator.","arXiv :2605 . 17393v2 [ cs .AI] 10 Jul 2026  \nHeterogeneous Information-Bottleneck Coordination Graphs for Multi-Agent Reinforcement Learning  \nWei Duan, Junyu Xuan, En Yu, Xiaoyu Yang, and Jie Lu  \nAustralian Artificial Intelligence Institute (AAII), University of Technology Sydney {wei. duan, junyu. xuan, en. yu-1, [jie. lu](jie. lu}@uts. edu. au)[}](jie. lu}@uts. edu. au)[@uts. edu. au](jie. lu}@uts. edu. au), [xiaoyu. yang-3@student. uts. edu. au](xiaoyu. yang-3@student. uts. edu. au)  \nAbstract  \nCoordination graphs specify which agents exchange information in cooperative multi-agent reinforce  \nment learning (MARL) . Existing sparse-graph methods, however, rely on heuristic, edge-uniform criteria  \nfor topology and structurally blind bottlenecks for message bandwidth, and lack a principled way to  \njointly learn heterogeneous connectivity and allocate differentiated communication capacity.  \nWe propose Heterogeneous Information-Bottleneck Coordination Graphs (HIBCG), which  \ncouples a group-aligned block-diagonal prior—assigning heterogeneous edge density per group block—  \nwith per-agent message compression on the learned graph. Theoretically, we show that the group-aligned  \nprior is never worse than a flat isotropic prior, that the structural penalty decomposes additively across  \ngroup blocks, and that minimising the standard TD loss lower-bounds IB relevance without a separate  \nmutual-information estimator. Across nine scenarios in SMACv1, SMACv2, and MAgent Battle (up  \nto 100 agents), HIBCG attains the strongest results on heterogeneous multi-role maps, scales where  \nseveral baselines fail to converge, and is validated by ablations and theory-aligned diagnostics of the  \ngroup-aligned prior and dual-path design.  \nKeywords. Multi-agent reinforcement learning, coordination graph, graph learning, information bottleneck, graph neural network.  \n1 Introduction  \nCooperative multi-agent reinforcement learning (MARL) relies on agents exchanging task-relevant information so that local decisions can be coordinated toward a shared team objective [1–4] . A coordination graph makes this communication selective by specifying which agent pairs exchange information at each step [5–7] . By keeping only informative links, sparse coordination graphs reduce noise aggregation, lower computation, and generalise better to unseen states [8–11] . The core open problem is therefore not whether to use a graph, but how to learn a graph whose topology is faithful to the underlying coordination structure of the task. Existing graph learners answer this question with homogeneous criteria—fixed thresholds [10], top-k selection [9], variance payoffs [8], attention scores [12,13], or dynamic factor-graph generation policies [14]—that treat every link with the same rule.  \nTwo limitations remain. (1) Topology learning is edge-uniform and heuristic: these methods lack a principled mechanism to learn heterogeneous connectivity—e.g., how dense intra-group edges should be relative to inter-group links when agents form functional sub-teams. (2) Message bandwidth is structurally blind: even after a topology is learned, surviving edges share the same representational capacity or pass through a single global bottleneck [15–18] . In heterogeneous deployments, agents collaborating within a functional sub-team (e.g., co-located mobile robots synchronising a joint manoeuvre) require dense, high-bandwidth messaging to maintain coherence, whereas cross-team links primarily carry sparse status orhandoff signals. Yet, no existing method allocates differentiated communication capacity to these structurally distinct agent relationships.  \nTo address both limitations, we propose Heterogeneous Information-Bottleneck Coordination Graphs (HIBCG), which jointly learns heterogeneous connectivity and allocates differentiated message  \nFigure 1: Edge-uniform vs. heterogeneous coordination graphs. Left: existing methods apply edgeuniform sparsification to all links (unif","cbCaietzls0vj1jN","https://ap.wps.com/l/cbCaietzls0vj1jN","pdf",2475702,3,1,44,"English","en",105,"# Introduction\n## Problem of edge-uniform topology learning\n## Problem of structurally blind message bandwidth\n## Proposed method: HIBCG\n# Contributions\n## Heterogeneous coordination-graph learning\n## Theoretical guarantees via GIB","[{\"question\":\"What problem does HIBCG address in multi-agent reinforcement learning coordination graphs?\",\"answer\":\"HIBCG addresses two gaps in prior work: edge-uniform, heuristic topology learning and a structurally blind bottleneck that does not differentiate communication capacity across structurally different agent relationships.\"},{\"question\":\"How does HIBCG learn heterogeneous connectivity between agents?\",\"answer\":\"HIBCG uses a group-aligned block-diagonal prior for topology learning, encouraging denser intra-group edges and requiring cross-group edges to justify their information cost.\"},{\"question\":\"How does HIBCG control message bandwidth once the graph topology is learned?\",\"answer\":\"HIBCG applies per-agent message compression on the learned graph, allocating per-agent feature bandwidth so surviving channels carry task-relevant information instead of passing through a single global bottleneck.\"}]",1784175210,111,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"heterogeneous-information-bottleneck-coordination-graphs-for-multi-agent-reinforcement-learning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/heterogeneous-information-bottleneck-coordination-graphs-for-multi-agent-reinforcement-learning/81657/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does HIBCG address in multi-agent reinforcement learning coordination graphs?","Question",{"text":75,"@type":76},"HIBCG addresses two gaps in prior work: edge-uniform, heuristic topology learning and a structurally blind bottleneck that does not differentiate communication capacity across structurally different agent relationships.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does HIBCG learn heterogeneous connectivity between agents?",{"text":80,"@type":76},"HIBCG uses a group-aligned block-diagonal prior for topology learning, encouraging denser intra-group edges and requiring cross-group edges to justify their information cost.",{"name":82,"@type":73,"acceptedAnswer":83},"How does HIBCG control message bandwidth once the graph topology is learned?",{"text":84,"@type":76},"HIBCG applies per-agent message compression on the learned graph, allocating per-agent feature bandwidth so surviving channels carry task-relevant information instead of passing through a single global bottleneck.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]