[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86510-en":3,"doc-seo-86510-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86510,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning","Robust motion planning in dense traffic requires autonomous vehicles to handle rare, safety-critical interactions that are underrepresented in naturalistic driving logs. Existing adversarial training approaches often depend on external scenario generators, heuristic perturbations, or simulator-heavy rollouts, which limits integration with modern autoregressive planners. The work formulates adversarially robust planner learning as a constrained min-max game and introduces Adversarial World Modeling (AWM), a multi-agent self-play fine-tuning framework with a decoupled solver. Experiments on nuPlan and InterPlan show transferable adversarial interactions and competitive closed-loop performance across nominal and long-tail scenarios, supported by theoretical analysis.","arXiv :2607 . 10630v1 [ cs .RO] 12 Jul 2026  \nWorld Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning  \nTong Nie,1,2,† Yuewen Mei,2,† Junlin He,1 Yihong Tang,3,4 Jian Sun,2,B Wei Ma1,B  \n1The Hong Kong Polytechnic University, 2Tongji University,  \n3McGill University, 4Mila-Quebec AI Institute  \n[tong.nie@connect.polyu.hk](tong.nie@connect.polyu.hk) [wei.w.ma@polyu.edu.hk](wei.w.ma@polyu.edu.hk)  \n† Equal contribution. B Corresponding authors.  \nAbstract  \nRobust motion planning in dense traffic requires autonomous vehicles to interact in rare and safety-critical scenarios that are underrepresented in naturalistic driving data. Although adversarial training offers a feasible solution, existing methods often rely on external scenario generators, heuristic perturbations, or simulator-heavy rollouts, which makes them difficult to integrate with modern autoregressive planners. Here, we cast adversarially robust planner learning as a constrained min-max game and propose Adversarial World Modeling (AWM) , a theoretically grounded multi-agent self-play fine-tuning framework. Since solving the exact game is intractable, AWM introduces a principled decoupled solver. In the inner minimization, the planner’s predictive world model is converted into a role-conditioned adversary that learns sparse, scene-adaptive attack coalitions via counterfactual credit assignment. In the outer maximization, the ego planner optimizes a regret-aware robust best response against the frozen AWM, utilizing tail-risk weighting and reference-anchored trust regions to improve hard-case recovery while preserving nominal driving behavior. Experiments on the nuPlan and InterPlan benchmarks demonstrate that our method generates transferable adversarial interactions and yields a robust planner that achieves competitive closed-loop performance in both nominal and highly interactive long-tail scenarios. Theoretical analysis justifies the decoupled solver and the main optimization components.  \n1 Introduction  \nDeveloping robust motion planners for closed-loop autonomous driving remains a central open problem [44, 23] . Recently, the leading paradigm has shifted from rule-based systems toward generative motion modeling [38, 33, 46, 53, 49] . Drawing inspiration from Large Language Models (LLMs), modern discrete architectures [46, 54, 43, 52, 48] treat map elements and motion primitivesas discrete tokens, exhibiting remarkable performance when trained on large-scale human driving logs. However, these autoregressive planners are primarily optimized against nominal, naturalistic data distributions. Consequently, they remain brittle on interactive long-tail cases that are rare but safety-critical, such as aggressive cut-ins, coordinated blocking, and collision-inducing maneuvers from surrounding vehicles [11, 25] . To expose and rectify these vulnerabilities, adversarial training offers a compelling solution by actively searching hard scenarios to robustify the planner [15] .  \nDespite the promise of adversarial training, existing frameworks are not well aligned to modern autoregressive architecture. Many existing adversarial frameworks rely on external adversary agents, heuristic scenario perturbations, or simulator-intensive rollouts [51, 29, 26, 40, 30] . These attacks are often designed around low-dimensional control policies, fixed objects of interest, or handcrafted rules. As a result, they are infeasible to scale to the high-dimensional tokenized generation paradigm.  \nPreprint.  \nRather than forcing incompatible external attackers into this paradigm, we argue that an endogenous solution lies in the environment simulator itself. We observe that behavioral World Models [46, 52], which inherently learn multi-agent traffic evolution as sequential token predictions, are naturally suited to bridge this gap. Operating on the same autoregressive architecture, motion vocabulary, and scene representation as the ego planner, they can be repur","cbCaiuWtMpvgDAcO","https://ap.wps.com/l/cbCaiuWtMpvgDAcO","pdf",2998507,6,1,36,"English","en",105,"# Abstract\n# Introduction\n## From nominal training to rare safety-critical interactions\n## Limitations of existing adversarial training\n## Endogenous environment adversaries via world models\n## Optimization challenges in multi-agent min-max training\n# Method: Adversarial World Modeling (AWM)","[{\"question\":\"What problem does the paper address in motion planning?\",\"answer\":\"The paper targets brittleness in autonomous driving when planners face rare but safety-critical interactive scenarios that are underrepresented in naturalistic data.\"},{\"question\":\"Why are existing adversarial training methods difficult to use with autoregressive planners?\",\"answer\":\"They often rely on external adversary agents, heuristic perturbations, or simulator-intensive rollouts and are designed for low-dimensional control or handcrafted attack rules, making them hard to scale to tokenized autoregressive generation.\"},{\"question\":\"What is Adversarial World Modeling (AWM) and how does it work?\",\"answer\":\"AWM casts robust planner learning as a constrained min-max game and uses a decoupled solver. It converts a predictive world model into a role-conditioned adversary for inner minimization and trains the ego planner with regret-aware robust best responses in the outer maximization.\"}]",1784212285,91,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"world-models-as-adversaries-multi-agent-self-play-fine-tuning-for-robust-motion-planning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/world-models-as-adversaries-multi-agent-self-play-fine-tuning-for-robust-motion-planning/86510/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What problem does the paper address in motion planning?","Question",{"text":76,"@type":77},"The paper targets brittleness in autonomous driving when planners face rare but safety-critical interactive scenarios that are underrepresented in naturalistic data.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why are existing adversarial training methods difficult to use with autoregressive planners?",{"text":81,"@type":77},"They often rely on external adversary agents, heuristic perturbations, or simulator-intensive rollouts and are designed for low-dimensional control or handcrafted attack rules, making them hard to scale to tokenized autoregressive generation.",{"name":83,"@type":74,"acceptedAnswer":84},"What is Adversarial World Modeling (AWM) and how does it work?",{"text":85,"@type":77},"AWM casts robust planner learning as a constrained min-max game and uses a decoupled solver. It converts a predictive world model into a role-conditioned adversary for inner minimization and trains the ego planner with regret-aware robust best responses in the outer maximization.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]