[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86185-en":3,"doc-seo-86185-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86185,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Multi-Agent LLMs Fail to Explore Each Other","Exploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language model agents can explore effectively when interacting with one another. The study shows modern LLM agents fail, producing myopic and polarized interaction patterns that reduce coordination quality and increase regret. A Multi-Agent Exploration problem is formalized as a partially observable stochastic game, requiring agents to probe peers, infer capabilities, and learn interaction strategies. MultiAgent Contextual Exploration (MACE) improves exploration and downstream performance across diversity settings and analysis links exploration value to agent diversity.","arXiv :2607 . 11250v1 [ cs .MA] 13 Jul 2026  \nMulti-Agent LLMs Fail to Explore Each Other  \nHyeong Kyu Choi1 , Jiatong Li1 , Wendi Li1 , Xin Eric Wang2 and Sharon Li1  \n1University of Wisconsin–Madison, 2University of California, Santa Barbara  \nAbstract  \nExploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language model (LLM) agents can explore effectively when interacting with one another. We show that modern LLM agents fail to do so, often exhibiting myopic and polarized interaction patterns that lead to suboptimal coordination and increased regret. We formalize this challenge as the Multi-Agent Exploration problem, modeling it as a partially observable stochastic game (POSG) problem in which agents must probe peers to infer their capabilities and identify effective interaction strategies. To address this, we introduce MultiAgent Contextual Exploration (MACE), a lightweight framework that explicitly promotes exploration through structured peer selection. Across both contextual and parametric diversity settings, MACE substantially improves exploration behavior and downstream task performance. We further show theoretically that the value of exploration increases with agent diversity. Overall, our results highlight a fundamental limitation of current LLM agents and underscore the importance of explicitly guided exploration for reliable multi-agent autonomy. Code will be released in [https://github.com/deeplearning-wisc/mace](https://github.com/deeplearning-wisc/mace).  \nContact: {froilanchoi, [sharonli](sharonli}@cs.wisc.edu)[}](sharonli}@cs.wisc.edu)[@cs.wisc.edu](sharonli}@cs.wisc.edu)  \n[1](1). Introduction  \nLarge language models are increasingly deployed not only as isolated assistants, but as autonomous agents embedded in multi-agent systems: communicating, delegating, and making decisions in a decentralized fashion to accomplish complex tasks (Du et al., 2024 ; Gao et al., 2024 ; Hong et al., 2023 ; Li et al., 2023 ; Qian et al., 2024 ; Wu et al., 2024) . This paradigm is rapidly expanding: open-world platforms now instantiate heterogeneous agent populations that interact autonomously to discuss problems, divide labor, and collectively reason toward shared goals (Jiang et al., 2026 ; Park et al., 2023 ; Piao et al., 2025 ; Zhang et al., 2026) . Such systems implicitly rely on the assumption that agents are capable of autonomous decision-making, and that meaningful collective behavior can emerge from agent-to-agent interactions.  \nBut what is required for agents to act autonomously in a reliable manner? A large body of literature in intrinsic motivation and reinforcement learning has demonstrated that exploration is foundational to autonomous behavior (Burda et al., 2019 ; Pathak et al., 2017) . Exploration is not merely a mechanism for improving task performance, but the core process by which agents self-generate goals, discover novel states, and acquire reusable strategies without external guidance (Baker et al., 2019 ; Chentanez et al., 2004 ; Forestier et al., 2022) .  \nThe importance of exploration is further amplified in multi-agent settings, where the environment is shaped not only by environmental dynamics, but also by the behaviors and capabilities of other agents, which are often heterogeneous and can only be revealed through direct interaction. Thus, in such settings, agents must proactively explore peers to identify effective collaborators, uncover complementary information, and adapt as the interaction landscape evolves.  \nYet despite the centrality of exploration to reliable autonomy, we find that current LLM agents fail to explore effectively, even in the simplest possible settings. We begin with a simple autonomous multi-agent setting—a controlled two-armed bandit experiment, where an LLM agent must repeatedly choose between two peers with unknown success rates and must infer the better one through exploration. Rather than accumulating evidence and ","cbCaig5FVsHYrY2u","https://ap.wps.com/l/cbCaig5FVsHYrY2u","pdf",2277740,1,52,"English","en",105,"# Abstract\n# Introduction\n## Exploration in Multi-Agent Settings\n## Observed Failure Mode in LLM Agents\n## Formalizing the Multi-Agent Exploration Problem\n## Introducing MACE","[{\"question\":\"What problem does the paper investigate about LLM multi-agent systems?\",\"answer\":\"It investigates whether LLM agents can explore effectively when interacting with other agents, and how exploration affects reliable autonomy.\"},{\"question\":\"How do current LLM agents behave in the simplified two-armed bandit experiment?\",\"answer\":\"They commit prematurely to a peer within the first few rounds and persist with that choice regardless of whether it is correct, instead of converging to the better peer.\"},{\"question\":\"What is MACE and how does it improve multi-agent exploration?\",\"answer\":\"MACE is a lightweight framework that guides exploration via structured peer selection, improving exploration behavior and downstream task performance across contextual and parametric diversity settings.\"}]",1784209209,131,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"multi-agent-llms-fail-to-explore-each-other","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/multi-agent-llms-fail-to-explore-each-other/86185/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":11},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper investigate about LLM multi-agent systems?","Question",{"text":75,"@type":76},"It investigates whether LLM agents can explore effectively when interacting with other agents, and how exploration affects reliable autonomy.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do current LLM agents behave in the simplified two-armed bandit experiment?",{"text":80,"@type":76},"They commit prematurely to a peer within the first few rounds and persist with that choice regardless of whether it is correct, instead of converging to the better peer.",{"name":82,"@type":73,"acceptedAnswer":83},"What is MACE and how does it improve multi-agent exploration?",{"text":84,"@type":76},"MACE is a lightweight framework that guides exploration via structured peer selection, improving exploration behavior and downstream task performance across contextual and parametric diversity settings.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]