[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82081-en":3,"doc-seo-82081-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82081,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Offline Nash Solvers Meet Online Tree Search in Multi-Agent Games on Graphs","Computing Nash equilibrium policies in multi-agent Pursuit–Evasion games is hindered by exponential growth in joint state and action spaces as the number of agents increases. Existing methods either depend on offline approximations that adapt poorly during execution or use online planning that struggles with large branching factors. This work introduces PrimitiveGuided Tree Search (PGTS), a hybrid method combining offline exact Nash computation on tractable sub-games with online tree search at each time step to guide expansion and estimate leaf values.","arXiv :2607 .08892v 1 [ cs .GT] 9 Jul 2026  \nOffline Nash Solvers Meet Online Tree Search in Multi-Agent Games on Graphs  \nMukesh Kumar⋆ , Yue Guan, and Panagiotis Tsiotras  \nSchool of Aerospace Engineering, Georgia Institute of Technology, Atlanta, GA, USA {mukeshkumar361,yguan44,[tsiotras}@gatech.edu](tsiotras}@gatech.edu)  \nAbstract. Computing Nash equilibrium policies in multi-agent PursuitEvasion games (PEG) is challenging due to the exponential growth of the joint state and action spaces with the number of agents. Existing approaches either rely on offline equilibrium approximations, which may lack adaptability during execution, or online planning methods, which suffer from large branching factors. In this work, we propose PrimitiveGuided Tree Search (PGTS), a hybrid framework that integrates offline exact Nash equilibrium computation with online tree search: PGTS first solves a collection of smaller, tractable sub-games offline; at deployment, PGTS performs online tree search at each time step, using the optimal sub-game policies and value functions to guide tree expansion and estimate leaf-node values. Extensive experiments on varied graph topologies, including real-world networks, demonstrate that PGTS significantly outperforms state-of-the-art learning and heuristic baselines, while maintaining robust performance against adversaries.  \nKeywords: Multi-Agent Systems · Pursuit–Evasion Games · Monte Carlo Tree Search.  \n1 Introduction  \nMulti-agent decision-making under adversarial conditions arises in a wide range of applications, including monitoring and surveillance, urban security, border patrol, and autonomous pursuit tasks [3,11] . A canonical model for studying such interactions in discrete settings is the graph-based Pursuit–Evasion game (PEG)[8,15] . The standard solution concept for such games is the Nash equilibrium. While the existence of equilibrium policies is well established [5], computing equilibria exactly in realistic multi-agent settings remains a significant challenge. As the number of agents increases, both the state space and the joint action space grow rapidly. This combinatorial growth renders exact equilibrium computation impractical even for moderately sized environments, motivating the development of scalable approximation and planning techniques.  \nExisting approaches to approximating solutions in multi-agent settings can be broadly classified into three categories: offline, online, and hybrid. Offline approaches, such as Policy Space Response Oracles (PSRO) [2] and its variants,  \n⋆ Corresponding author.  \n2 M. Kumar et al.  \nprovide a general framework for computing approximate equilibria by iteratively training best-response policies and maintaining a growing meta-game over the policy space. In the Pursuit-Evasion context, MT-PSRO [13] combines PSRO with multi-task pretraining, and uses MAPPO [23] as the best-response oracle. While offline methods enable fast decision-making during execution, their training complexity can grow quickly with the number of agents. Moreover, the agents’ performance can degrade significantly when they encounter adversarial behaviors that differ from those encountered during training.  \nOnline planning methods, such as Monte Carlo Tree Search (MCTS) [19], adapt to changing situations and reason about strategic interactions at execution time by simulating a large number of future trajectories. Although originally introduced to solve sequential in-turn games, extensions of MCTS to simultaneous-move games have been explored through Simultaneous-Move MCTS (SM-MCTS) [20] . Unlike offline methods, online search methods can reason directly about the current game state, but their decision quality depends critically on the ability of the search process to identify and explore promising future trajectories. In multi-agent games, this exploration becomes increasingly difficult because the number of joint actions, and hence the branching factor, grows exponentially with th","cbCaigU4IKGxNkAZ","https://ap.wps.com/l/cbCaigU4IKGxNkAZ","pdf",1152487,1,20,"English","en",105,"# Introduction\n## Offline Approaches\n## Online Planning Methods\n## Hybrid Approaches\n# Proposed Method (PGTS)","[{\"question\":\"Why is computing Nash equilibrium policies difficult in multi-agent Pursuit–Evasion games?\",\"answer\":\"Exact equilibrium computation becomes impractical because both the joint state space and the joint action space grow exponentially with the number of agents.\"},{\"question\":\"What limitations do offline and online approaches have in this setting?\",\"answer\":\"Offline methods can train efficiently at execution time but may degrade under adversarial behaviors not seen during training. Online planning adapts at execution time but can suffer from very large branching factors in multi-agent games.\"},{\"question\":\"How does PGTS combine offline Nash solving with online tree search?\",\"answer\":\"PGTS decomposes the original game into smaller primitive sub-team games solved exactly offline, then uses the resulting equilibrium policies and value functions to guide online tree search and rollout expansion at each time step.\"}]",1784178122,50,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"offline-nash-solvers-meet-online-tree-search-in-multi-agent-games-on-graphs","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/offline-nash-solvers-meet-online-tree-search-in-multi-agent-games-on-graphs/82081/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is computing Nash equilibrium policies difficult in multi-agent Pursuit–Evasion games?","Question",{"text":75,"@type":76},"Exact equilibrium computation becomes impractical because both the joint state space and the joint action space grow exponentially with the number of agents.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What limitations do offline and online approaches have in this setting?",{"text":80,"@type":76},"Offline methods can train efficiently at execution time but may degrade under adversarial behaviors not seen during training. Online planning adapts at execution time but can suffer from very large branching factors in multi-agent games.",{"name":82,"@type":73,"acceptedAnswer":83},"How does PGTS combine offline Nash solving with online tree search?",{"text":84,"@type":76},"PGTS decomposes the original game into smaller primitive sub-team games solved exactly offline, then uses the resulting equilibrium policies and value functions to guide online tree search and rollout expansion at each time step.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":28,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]