[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117325-en":3,"doc-seo-117325-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117325,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","MANSA - Learning Fast and Slow in Multi-Agent Systems","Multi-agent reinforcement learning (MARL) shows strong scalability with independent learning (IL), yet IL can be inefficient and may fail when agents must coordinate. Centralised learning (CL) enables quick coordination but is often too expensive for real-world use, and value-based CL may require strict representational constraints that degrade performance. This paper proposes MANSA, a plug-and-play IL framework that activates CL only in states requiring coordination. The added switching agent learns which states to use CL, preserves cooperative convergence, reduces computation, and optimizes a fixed CL-call budget.","MANSA: Learning Fast and Slow in Multi-Agent Systems  \nDavid Mguni 1 Haojun Chen 2 Taher Jafferjee 1 Jianhong Wang 3 Longfei Yue 2 Xidong Feng 1 4 Stephen McAleer 5 Feifei Tong 1 Jun Wang 1 4 Yaodong Yangy 2  \nAbstract  \nIn multi-agent reinforcement learning (MARL), independent learning (IL) often shows remarkable performance and easily scales with the number of agents. Yet, using IL can be inefﬁcient and runs the risk of failing to successfully train, particularly in scenarios that require agents to coordinate their actions. Using centralised learning (CL) enables MARL agents to quickly learn how to coordinate their behaviour but employing CL everywhere is often prohibitively expensive in real-world applications. Besides, using CL in value-based methods often needs strong representational constraints (e.g. individual-global-max condition) that can lead to poor performance if violated. In this paper, we introduce a novel plug & play IL framework named Multi-Agent Network Selection Algorithm (MANSA) which selectively employs CL only at states that require coordination. At its core, MANSA has an additional agent that uses switching controls to quickly learn the best states to activate CL during training, using CL only where necessary and vastly reducing the computational burden of CL. Our theory proves MANSA preserves cooperative MARL convergence properties, boosts IL performance and can optimally make use of a ﬁxed budget on the number CL calls. We show empirically in Level-based Foraging (LBF) and StarCraft Multi-agent Challenge (SMAC) that MANSA achieves fast, superior and more reliable performance while making 40% fewer CL calls in SMAC and using CL at only 1% CL calls in LBF.  \n1Huawei R&D 2Institute for AI, Peking University 3University of Manchester 4University College, London 5Independent Researcher. Correspondence to: \u003C[davidmguni@hotmail.com](davidmguni@hotmail.com) >,  \n\u003C[j.wang@ucl.ac.uk](j.wang@ucl.ac.uk)>, \u003C[yaodong.yang@pku.edu.cn](yaodong.yang@pku.edu.cn) >.  \nProceedings of the 40 th International Conference on Machine Learning, Honolulu, Hawaii, USA. PMLR 202, 2023 . Copyright 2023 by the author(s) .  \n1. Introduction  \nMulti-agent reinforcement learning (MARL) has emerged as a powerful framework that enables autonomous agents to complete various tasks in areas such as autonomous driving (Zhou et al., 2021), swarm robotics (Mguni et al., 2018; 2019) and smart grids (Wang et al., 2021; Qiu et al., 2021; 2022) . Among MARL methods are a class of algorithms known as independent learners (IL) e.g. independent Q learning (Tan, 1993) . IL decomposes a MARL problem with N agents into N decentralised single-agent problems. In this way, each agent treats other agents as part of the environment which provides a straightforward way of training agents in a decentralised manner. Since the agents ignore other agents, IL can be trained quickly as each agent's learning process is contingent on only its local observations and own actions. This is efﬁcient in scenarios that require only weak interactions between agents (Kok & Vlassis, 2004) .  \nDespite these apparent beneﬁts, training MARL using IL has several formidable drawbacks: with no ability to observe the actions of other agents, random occurrences of successful coordination among IL agents are improbable, causing IL methods to sometimes struggle in tasks that require coordination (Hernandez-Leal et al., 2017) . Also, ignoring other agents' inﬂuence on the system means from the agent's perspective, the environment can appear non-stationary which precludes convergence guarantees (Yang & Wang, 2020) .  \nOn the other hand, MARL learners can be trained in simulated environments in which agents can be provided with other agents' observations and other state information. Centralised training and decentralised execution (CT-DE) (Kraemer & Banerjee, 2016; Foerster et al., 2018; McAleer et al., 2022) is a framework that uses a centralised critic that exploits global information du","cbCairTCgTzdeTqH","https://ap.wps.com/l/cbCairTCgTzdeTqH","pdf",1866567,1,28,"English","en",105,"# Introduction\n## Independent learning (IL) in MARL\n## Centralised training and decentralised execution (CT-DE)\n## Complexity growth in CT-DE and value decomposition limits","[{\"question\":\"What problem does MANSA address in multi-agent reinforcement learning?\",\"answer\":\"MANSA addresses the inefficiency and potential failure of independent learning when coordination is required, while also avoiding the high cost of using centralised learning everywhere.\"},{\"question\":\"How does MANSA decide when to use centralised learning?\",\"answer\":\"MANSA uses an additional switching agent with switching controls that learns which states require coordination, activating CL only for those states during training.\"},{\"question\":\"What does the paper report in experiments on benchmark environments?\",\"answer\":\"On Level-based Foraging (LBF) and StarCraft Multi-agent Challenge (SMAC), MANSA achieves fast, superior, and more reliable performance while using fewer CL calls—about 40% fewer in SMAC and CL at only 1% of calls in LBF.\"}]","MANSA - Learning Fast and Slow in Multi-Agent Systems | PDF",1785675194,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"mansa-learning-fast-and-slow-in-multi-agent-systems","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/mansa-learning-fast-and-slow-in-multi-agent-systems/117325/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does MANSA address in multi-agent reinforcement learning?","Question",{"text":75,"@type":76},"MANSA addresses the inefficiency and potential failure of independent learning when coordination is required, while also avoiding the high cost of using centralised learning everywhere.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does MANSA decide when to use centralised learning?",{"text":80,"@type":76},"MANSA uses an additional switching agent with switching controls that learns which states require coordination, activating CL only for those states during training.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the paper report in experiments on benchmark environments?",{"text":84,"@type":76},"On Level-based Foraging (LBF) and StarCraft Multi-agent Challenge (SMAC), MANSA achieves fast, superior, and more reliable performance while using fewer CL calls—about 40% fewer in SMAC and CL at only 1% of calls in LBF.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]