[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84812-en":3,"doc-seo-84812-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84812,2336464648322,"Aria","https://ap-avatar.wpscdn.com/avatar/2200025388227c56fec?_k=1778556882303663488",8,"Research & Report","MIRA Multiplayer Interactive World Models with Representation Autoencoders","MIRA introduces a multiplayer world model for highly dynamic environments driven by complex physical interactions. Unlike single-agent world models that treat other agents as part of the environment, it conditions on multiple agents’ action streams to correctly attribute scene changes to the right player while maintaining temporal coherence. Trained on 10,000 hours of Rocket League gameplay, the 5B-parameter latent diffusion model generates real-time four-player matches at 20 FPS on a single Nvidia B200 GPU and stays stable far beyond the short-clip training horizon, enabling physically grounded evaluations and long-horizon rollouts.","arXiv :2607 .05352v2 [ cs .CV] 7 Jul 2026  \nMIRA  \nMultiplayer Interactive World Models with Representation Autoencoders  \nCore contributors  \nAnthony Hu∗,1, Václav Volhejn∗,2, Adrien Ramanana Rahary∗,2,‡, Chris Mulder∗,1, Aditya Makkar 1 , Alyx Liao2 , Amélie Royer2 , Manu Orsini2  \nContributors  \nAdam Jelley 1 , Eloi Alonso 1 , Florian Laurent 1 , Fredrik Norén 1 , James Swingos 1 , Jan Hünermann 1 , Kent Rollins 1 , Lucas Hosseini 1 , Matthieu Le Cauchois 1 , Maxim Peter 1 , Pim de Witte 1 , Tim Brown 1 , Vincent Micheli 1 , Moritz Böhle2 , Gabriel de Marmiesse2 , Viktoriia Sharmanska3 , Lucia Specia3 , Michael Black3 , Patrick Pérez2  \n∗ equal contribution  \n1 General Intuition 2 Kyutai 3 Epic Games  \nWe introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents as part of the environment, ours conditions on the action streams of multiple agents, learning to attribute changes in the scene to the correct player and to stay coherent under arbitrary combinations of their actions. We study this problem in the game of Rocket League, where players compete and cooperate under fast, tightly coupled dynamics. Trained on 10,000 hours of gameplay collected with publicly available bots, our 5-billion-parameter latent diffusion model generates four-player matches in real time, producing 20 frames per second on a single Nvidia B200 GPU. Although trained only on short clips, its rollouts stay stable far beyond the training horizon: distributional quality holds steady out to five minutes, the longest horizon we measure, and in practice we observe rollouts continuing for hours with no sign of collapse. We systematically investigate the central design choices: the video codec, the generative objective, and the multiplayer conditioning scheme. In addition, we characterize how behavior changes with model and data scale, including the capabilities that emerge and the failure modes that persist. We further develop targeted evaluations that probe the model’s physical understanding rather than visual appearance alone. To support continued research on multiplayer world models, we release our dataset, our full training and inference codebase, and a live demo.  \nLive demo: [https://mira-wm.com](https://mira-wm.com)  \nBlog: [https://mira-wm.com/blog-post](https://mira-wm.com/blog-post)  \nCode: [https://github.com/mira-wm/mira](https://github.com/mira-wm/mira)  \nDataset: [https://huggingface.co/datasets/kyutai/rocket-science](https://huggingface.co/datasets/kyutai/rocket-science)  \n1 Introduction  \nPredicting how a scene will evolve in response to actions is a core ability of any embodied agent (LeCun, 2022 ; Silver and Sutton, 2025) . A world model learns this ability directly from experience, typically by building an internal latent representation of the environment and predicting future latent states conditioned on actions. Once learned, the world model can act as a controllable simulator (Ha and Schmidhuber, 2018 ; Micheli et al. , 2023 ; Russell et al. , 2025 ; Hafner et al. , 2025b): agents can be trained inside it, evaluated against it, and used to imagine the consequences of candidate actions before committing to them in the real environment.  \nHowever, for such imagined rollouts to be useful, the world model must be faithful to the environment it represents (Alonso et al. , 2024) . As we ultimately aim to build agents that operate in real environments,  \n‡École nationale des ponts et chaussées  \nFigure 1 World model imagination. Rows are the four players’ viewpoints (Players 1–4) and columns show three timesteps in the trajectory (t0 → t0 + 6 s → t0 + 8 s) . From the initial state at t0 and the players’ action streams, the model imagines a dynamic game scene that captures the physical interplay between the players, the ball, and the scene. The rollout is temporally consistent: the in-game clock counts down with elapsed time (4","cbCaijfRE1m5HcPQ","https://ap.wps.com/l/cbCaijfRE1m5HcPQ","pdf",9836565,1,59,"English","en",105,"# Introduction\n## World model fundamentals\n## Faithfulness for planning and control\n## Multi-agent necessity\n# Method overview\n## Latent space simulation","[{\"question\":\"What makes MIRA different from single-player world models?\",\"answer\":\"MIRA conditions on the action streams of multiple agents and attributes scene changes to the correct player, instead of treating other agents as part of the environment. This improves applicability to explicit multi-participant control settings.\"},{\"question\":\"How is MIRA evaluated, and where is it applied?\",\"answer\":\"MIRA is studied in the Rocket League setting, with experiments focused on physical understanding rather than visual appearance alone. It includes targeted evaluations and examines design choices such as the video codec and generative objective.\"},{\"question\":\"What evidence is given for long-horizon stability?\",\"answer\":\"Rollouts remain stable beyond the training horizon: distributional quality holds steady out to five minutes (the longest measured horizon). The document also reports that rollouts can continue for hours without signs of collapse in practice.\"}]",1784198412,149,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"mira-multiplayer-interactive-world-models-with-representation-autoencoders","",{"@graph":35,"@context":84},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/mira-multiplayer-interactive-world-models-with-representation-autoencoders/84812/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What makes MIRA different from single-player world models?","Question",{"text":74,"@type":75},"MIRA conditions on the action streams of multiple agents and attributes scene changes to the correct player, instead of treating other agents as part of the environment. This improves applicability to explicit multi-participant control settings.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How is MIRA evaluated, and where is it applied?",{"text":79,"@type":75},"MIRA is studied in the Rocket League setting, with experiments focused on physical understanding rather than visual appearance alone. It includes targeted evaluations and examines design choices such as the video codec and generative objective.",{"name":81,"@type":72,"acceptedAnswer":82},"What evidence is given for long-horizon stability?",{"text":83,"@type":75},"Rollouts remain stable beyond the training horizon: distributional quality holds steady out to five minutes (the longest measured horizon). The document also reports that rollouts can continue for hours without signs of collapse in practice.","https://schema.org",{"og:url":51,"og:type":86,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":88,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":105,"slug":137},19,"General","general"]