[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-140735-105":59,"doc-detail-140735-en":125},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":118,"head_meta":120,"extra_data":122,"updated_unix":124},105,"en","a-competition-winning-deep-reinforcement-learning-agent-in-microrts-abstract-and-training-methods","A Competition Winning Deep Reinforcement Learning Agent in microRTS - abstract and training methods","","Deep reinforcement learning has advanced in real-time strategy games, yet adoption in the academic IEEE microRTS competitions remained limited because training demands are high and agent development is complex. RAISocketAI is presented as the first DRL agent to win the IEEE microRTS competition, repeatedly defeating prior winners in a benchmark setting. The approach relies on iterative policy fine-tuning, transfer learning to specific maps, and efficient bootstrapping via imitation learning with behavior cloning plus DRL refinement.",{"@graph":69,"@context":117},[70,84,100],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/a-competition-winning-deep-reinforcement-learning-agent-in-microrts-abstract-and-training-methods/140735/",{"url":83,"name":65,"@type":85,"author":86,"headline":65,"publisher":89,"fileFormat":92,"inLanguage":63,"description":67,"dateModified":93,"datePublished":94,"encodingFormat":92,"isAccessibleForFree":95,"interactionStatistic":96},"DigitalDocument",{"name":87,"@type":88},"Oliver","Person",{"url":74,"name":90,"@type":91},"DocShare","Organization","application/pdf","2026-09-11","2026-08-24",true,{"@type":97,"interactionType":98,"userInteractionCount":19},"InteractionCounter",{"@type":99},"ViewAction",{"@type":101,"mainEntity":102},"FAQPage",[103,109,113],{"name":104,"@type":105,"acceptedAnswer":106},"Why has deep reinforcement learning been limited in microRTS competitions so far?","Question",{"text":107,"@type":108},"Training requires substantial resources, and creating and debugging DRL agents is complex. The paper also highlights practical constraints like action computation within tight time windows and sparse delayed win/loss rewards.","Answer",{"name":110,"@type":105,"acceptedAnswer":111},"What makes RAISocketAI effective in winning the IEEE microRTS competition?",{"text":112,"@type":108},"RAISocketAI uses iterative fine-tuning of a base policy and transfer learning to specific maps. It also selects among multiple policy networks based on the map and compute capabilities.",{"name":114,"@type":105,"acceptedAnswer":115},"How does the work propose training future DRL agents more economically?",{"text":116,"@type":108},"It suggests transfer learning to specific maps and further bootstrapping using imitation learning. Behavior cloning followed by fine-tuning with DRL is presented as an efficient way to start from demonstrated competitive behaviors.","https://schema.org",{"og:url":83,"og:type":119,"og:title":65,"og:site_name":90,"og:description":67},"article",{"robots":121,"canonical":83},"index,follow",{"doc_id":123,"site_id":62},140735,1787607766,{"code":4,"msg":5,"data":126},{"doc_id":123,"user_id":127,"nickname":87,"user_avatar":128,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":129,"file_id":130,"file_url":131,"file_type":132,"file_size":133,"view_count":19,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":56,"language":134,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":135,"faqs":136,"seo_title":137,"seo_description":67,"update_tm":124,"read_time":138},8796095461610,"https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c","A Competition Winning Deep Reinforcement Learning Agent in microRTS  \nScott Goodfriend  \nBerkeley, CA  \n[goodfriend.scott@gmail.com](goodfriend.scott@gmail.com)  \narXiv :2402 .08 1 12v2 [ cs .LG] 2 Jan 2025  \nAbstract—Scripted agents have predominantly won the five previous iterations of the IEEE microRTS (µRTS) competitions hosted at CIG and CoG. Despite Deep Reinforcement Learning (DRL) algorithms making significant strides in real-time strategy (RTS) games, their adoption in this primarily academic competition has been limited due to the considerable training resources required and the complexity inherent in creating and debugging such agents. RAISocketAI is the first DRL agent to win the IEEE microRTS competition. In a benchmark without performance constraints, RAISocketAI regularly defeated the two prior competition winners. This first competition-winning DRL submission can be a benchmark for future microRTS competitions and a starting point for future DRL research. Iteratively fine-tuning the base policy and transfer learning to specific maps were critical to RAISocketAI’s winning performance. These strategies can be used to economically train future DRL agents. Further work in Imitation Learning using Behavior Cloning and fine-tuning these models with DRL has proven promising as an efficient way to bootstrap models with demonstrated, competitive behaviors.  \nIndex Terms—Machine learning, Games, Artificial Intelligence  \nI. INTRODUCTION  \nDeep reinforcement learning (DRL) has proven to be powerful at solving complex problems requiring several steps to achieve a goal, such as Atari games [1], continuous control tasks [2], and even real-time strategy (RTS) games like StarCraft II [3] . The StarCraft II grandmaster agent AlphaStar was trained with thousands of CPUs and GPUs/TPUs for several weeks. RTS games are particularly challenging for DRL for several reasons: (1) the observation and action spaces are large and varied with different terrain and unit types; (2) each unit type can have different actions and abilities; (3) each action can control several units at once; (4) rewards are sparse (win, loss, or tie) and delayed by possibly several thousand timesteps; (5) winning requires combining tactical (micro) and strategic (macro) decisions; (6) actions must be computed within a reasonable time window; (7) the agent might not have full visibility of the game state (i.e., fog of war); and (8) events in the game might be non-deterministic.  \nmicroRTS (stylized as µRTS) is a minimalist, open-source, two-player, zero-sum RTS game testbed designed for research purposes [4] . It includes many aspects of RTS games, simplified: different unit types, unit-specific actions, terrain, resource collection and utilization to build units, and unit-to-unit combat where units have different strengths and weaknesses.  \n979-8-3503-5067-8/24/$31.00 ©2024 IEEE  \nmicroRTS also supports fog of war and non-determinism; however, these were disabled for the IEEE-CoG 2023 microRTS competition.  \nThe IEEE microRTS competitions have been hosted at the Conference on Games (CoG) nearly every year since 2019 and at the Conference on Computational Intelligence and Games (CIG) before that since 2017 [5] . Competitors submit an agent that plays against other submissions and baselinesin a round-robin tournament on 12 different maps: 8 Open (known beforehand, Fig. 1) and 4 Hidden (unknown until after the competition results are released) . Agents are supposed to submit actions every step within 100 ms. Without GPU acceleration, this is a significant constraint for deep neural network agents.  \nThis paper describes how the RAISocketAI agent 1 was trained and became the first DRL agent to win the microRTS competition by winning at CoG in 2023 . The agent chooses between 7 policy networks based on the map and compute capabilities. The significant training time (70 GPU-days) combined with the general difficulty in debugging and fine-tuning a DRL implementation co","cbCaiqcONUykXtYd","https://ap.wps.com/l/cbCaiqcONUykXtYd","pdf",1169584,"English","# Introduction\n## MicroRTS competition setup and constraints\n## Paper contributions and RAISocketAI overview\n# Related Work\n## MicroRTS-Py","[{\"question\":\"Why has deep reinforcement learning been limited in microRTS competitions so far?\",\"answer\":\"Training requires substantial resources, and creating and debugging DRL agents is complex. The paper also highlights practical constraints like action computation within tight time windows and sparse delayed win/loss rewards.\"},{\"question\":\"What makes RAISocketAI effective in winning the IEEE microRTS competition?\",\"answer\":\"RAISocketAI uses iterative fine-tuning of a base policy and transfer learning to specific maps. It also selects among multiple policy networks based on the map and compute capabilities.\"},{\"question\":\"How does the work propose training future DRL agents more economically?\",\"answer\":\"It suggests transfer learning to specific maps and further bootstrapping using imitation learning. Behavior cloning followed by fine-tuning with DRL is presented as an efficient way to start from demonstrated competitive behaviors.\"}]","A Competition Winning Deep Reinforcement Learning Agent in microRTS - abstract and training methods | PDF",48]