[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84036-en":3,"doc-seo-84036-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84036,13056703019404,"Miles","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","Intercepting an Agile Target with Net-Carrying Drones using Competitive Multi-Agent Reinforcement Learning","Intercepting agile drone targets with a team of net-carrying interceptors is formulated as a competitive Multi-Agent Reinforcement Learning (MARL) problem. The study addresses nonstationarity and catastrophic forgetting by training pursuers and an evader with MAPPO and Prioritized Fictitious Self Play (PFSP). Policies are learned in a high-fidelity simulator using low-level CTBR commands for agile flight. Evaluations compare catch rate, time-to-catch, and crash rates against heuristic baselines, with ablations confirming PFSP robustness and the importance of low-level controls, plus emerging cooperative tactics.","Intercepting an Agile Target with Net-Carrying Drones using Competitive Multi-Agent Reinforcement Learning  \nTimothe Gavin∗†‡, and Murat Bronz†  \n∗ IAS, Thales LAS, Rungis, France  \n†Dynamic Systems, OPTIM, Fe´de´ration ENAC ISAE-SUPAERO ONERA, Universit e´ de Toulouse, Toulouse, France  \n‡RIS, LAAS-CNRS Toulouse, France  \n[timothee.gavin@thalesgroup.fr](timothee.gavin@thalesgroup.fr) , [murat.bronz@enac.fr](murat.bronz@enac.fr)  \narXiv :2607 .05939v 1 [ cs .RO] 7 Jul 2026  \nAbstract—This article presents a solution to intercept an agile drone by a team of agile drone carrying catching nets. We formulate the problem as a competitive Multi-Agent Reinforcement Learning (MARL) task. To address the problem of nonstationarity and catastrophic forgetting of agents overfitting to the current opponent strategy, we train the pursuers and the evader using Multi-Agent Proximal Policy Optimization (MAPPO) with Prioritized Fictitious Self Play (PFSP). We train the agents in a high-fidelity simulator using low-level control commands, collective thrust and body rates (CTBR), to achieve agile flights for both the pursuers and the evader. We compare the performance of the trained policies in terms of catch rate, time to catch and crash rates, against heuristic baselines and show that our solution outperforms them. Ablation studies show that PFSP lead to more robust policies that can adapt to different opponent strategies, and that a low-level control commands are crucial for learning performing strategies in the pursuit-evasion task. Finally, a qualitative analysis of the learned behaviours highlights the emergence of cooperative tactics among the pursuers.  \nI. INTRODUCTION  \nIntercepting agile aerial targets with autonomous drones is a key challenge in robotics and security. As Unmanned Aerial Vehicle (UAV) incursions into restricted airspace become more common, applications such as airspace protection, infrastructure security, and event safety demand solutions that can neutralize unauthorized drones with minimal collateral damage [1] . Fleets of interceptor drones with capture-nets area promising option, but their success depends on advanced control and coordination against evasive targets.  \nClassical interception methods rely on accurate models, preplanned strategies, or predictable target behaviour [2] . However, modern quad-rotor drones can perform highly dynamic manoeuvres, and will actively evade capture, making classical methods ineffective [3] .  \nRecent advances in deep reinforcement learning (RL) have shown that drones can learn agile flight behaviours directly from interaction with the environment. RL-trained policies achieved superhuman performance in drone racing, where agents navigate challenging courses with dynamic manoeuvres [4] . However, drone racing typically involves static or predictable gates, whereas interception requires agents to respond to adversarial targets that actively evade capture. Atthe same time, Multi-Agent RL (MARL) has demonstrated remarkable success in competitive environments, particularly in complex games [5], [6] .  \nFig. 1: A multi-agent competitive reinforcement learning approach to train both a team of pursuers and an evader drone for agile pursuit-evasion using catching-nets. Both pursuers and evader learn low-level control policies that enable them to perform dynamic manoeuvres in a high-fidelity simulator.  \nIn this work, we formulate agile drone interception between a team of net-carrying pursuers and an evasive target as a competitive MARL task. The main contributions of this paper with respect to the literature are:  \n• A competitive MARL framework for cooperative interception of an agile evader by multiple pursuer drones using low-level control.  \n• A Prioritized Fictitious Self-Play training scheme for both the pursuers and the evader to learn robust policies against varied opponents.  \n• Extensive simulation results demonstrating stronger performance than standard baselines.  \nThe remain","cbCaict7egXy87oJ","https://ap.wps.com/l/cbCaict7egXy87oJ","pdf",1306541,4,1,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"How is the agile drone interception task modeled in this work?\",\"answer\":\"The task is formulated as a competitive Multi-Agent Reinforcement Learning problem, with net-carrying pursuers working against an evasive target (evader).\"},{\"question\":\"Why are MAPPO and Prioritized Fictitious Self Play (PFSP) used?\",\"answer\":\"They are used to mitigate nonstationarity and catastrophic forgetting caused by overfitting to the current opponent strategy, improving robustness across varied opponents.\"},{\"question\":\"What information supports learning agile behaviors for both teams?\",\"answer\":\"Agents are trained in a high-fidelity simulator using low-level control commands based on CTBR (collective thrust and body rates), enabling dynamic manoeuvres in the pursuit-evasion setting.\"}]",1784192169,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"intercepting-an-agile-target-with-net-carrying-drones-using-competitive-multi-agent-reinforcement-learning","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":20},"https://docshare.wps.com/document/intercepting-an-agile-target-with-net-carrying-drones-using-competitive-multi-agent-reinforcement-learning/84036/",{"url":51,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"How is the agile drone interception task modeled in this work?","Question",{"text":74,"@type":75},"The task is formulated as a competitive Multi-Agent Reinforcement Learning problem, with net-carrying pursuers working against an evasive target (evader).","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"Why are MAPPO and Prioritized Fictitious Self Play (PFSP) used?",{"text":79,"@type":75},"They are used to mitigate nonstationarity and catastrophic forgetting caused by overfitting to the current opponent strategy, improving robustness across varied opponents.",{"name":81,"@type":72,"acceptedAnswer":82},"What information supports learning agile behaviors for both teams?",{"text":83,"@type":75},"Agents are trained in a high-fidelity simulator using low-level control commands based on CTBR (collective thrust and body rates), enabling dynamic manoeuvres in the pursuit-evasion setting.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]