[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86304-en":3,"doc-seo-86304-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86304,1099514068365,"Aurelia","https://ap-avatar.wpscdn.com/avatar/10000253d8d9f28188e?_k=1776742907772140068",8,"Research & Report","Xiaomi Robotics U0 Unified Embodied Synthesis with World Foundation Model","Recent foundation image and video generation models generalize well, yet direct use in embodied robotics is constrained by demands for multi-view consistency, geometric and physical coherence, and robot embodiment compatibility. Xiaomi-Robotics-U0 introduces a 38B-parameter multimodal autoregressive model that unifies embodied synthesis with foundation image/video generation, jointly optimizing text-to-image, image editing, embodied scene generation/transfer, and embodied video generation. It preserves world-foundation generalization while adapting to embodied settings, enabling high-quality multi-view scene generation across embodiments and structured, controllable transfer with fine-grained editing. Results achieve state-of-the-art performance on single-step and sequential tasks, including ranking first on World Arena for embodied video generation and boosting π0.5 out-of-distribution success on real-world manipulation.","arXiv :2607 . 11643v1 [ cs .RO] 13 Jul 2026  \nXiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model  \nXiaomi Robotics 1  \nAbstract  \nRecent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consistency, geometric coherence, and robot embodiment constraints. Existing methods typically adapt foundation models with limited robot data, often sacrificing visual knowledge acquired during large-scale pre-training.  \nWe present Xiaomi-Robotics-U0 , a 38-billion-parameter multimodal autoregressive model for unified embodied synthesis. It treats embodied generation as an extension of foundation image and video generation and jointly optimizes text-to-image generation, image editing, embodied scene generation, embodied transfer, and embodied video generation. This unified framework preserves the generalization of the pre-trained world foundation model while adapting it to embodied settings. Xiaomi-Robotics-U0 is the first model to support high-quality multi-view scene generation across multiple robot embodiments and to introduce structured, controllable embodied transfer for fine-grained editing while preserving multi-view consistency and interaction dynamics. It achieves state-of-the-art results on single-step and sequential generation tasks, outperforming GPT-Image- 2.0 in human evaluations of embodied scene generation and transfer, ranking first on World Arena for embodied video generation, and improving the out-of-distribution success rate of π0.5 from 36 .9% to 63 .2% on challenging real-world manipulation tasks. These results show that foundation world models can serve both as embodied world models and scalable data engines for embodied intelligence. Code and checkpoints are available at [https://robotics.xiaomi.com/](https://robotics.xiaomi.com/)[ ](https://robotics.xiaomi.com/)xiaomi-robotics-u0 .html.  \n1 Introduction  \nRecent foundation image and video generation models [4, 13 , 16 , 27 , 42 , 45 , 55 , 60 , 64] have made remarkable progress in semantic understanding, controllable generation, and visual reasoning through training with data on the Internet. Large-scale generative models are now capable of synthesizing highly realistic images and videos from various multimodal inputs, demonstrating impressive generalization far beyond the distribution of their training data. Such capabilities make foundation generative models an attractive starting point for embodied intelligence [34, 50 , 66 , 70], where robots are required to reason about complex environments and imagine future interactions before acting.  \nHowever, embodied generation [30, 32 , 36] introduces challenges that differ fundamentally from conventional image and video synthesis. Unlike natural image generation, embodied scenarios require strict multi-view  \nconsistency, accurate geometric and physical coherence across cameras, explicit robot embodiment constraints, 1 See Contributions section for full author list. Please send correspondence [to mi-robotics@xiaomi.com](to mi-robotics@xiaomi.com).  \nFigure 1 Embodied and general capabilities of Xiaomi-Robotics-U0 . The rectangle corresponds to the initial observations for the same embodiment, the pairwise transfer sample, and the keyframes within a video for the embodied capabilities. All frames are referenced and generated images.  \nand temporally consistent interaction dynamics. The generated observations must remain compatible with robot kinematics, camera calibration, and downstream manipulation policies rather than merely appearing visually realistic. Consequently, directly applying existing foundation image or video generation models to embodied scenarios often leads to inconsistent geometry, implausible robot states, and poor compatibility with robot control.  \nRecent embodied world models [1, 50 , 72] attempt to bridge this gap by continually adapting pre","cbCaihTJ35Gxxyou","https://ap.wps.com/l/cbCaihTJ35Gxxyou","pdf",38481668,3,1,39,"English","en",105,"# Introduction\n# Unified Embodied Synthesis Model","[{\"question\":\"What key limitations prevent directly applying foundation image/video generation models to embodied robotics?\",\"answer\":\"Embodied settings require strict multi-view consistency, accurate geometric and physical coherence, and explicit robot embodiment constraints. Generated observations must also remain compatible with robot kinematics, camera calibration, and downstream manipulation policies.\"},{\"question\":\"How does Xiaomi-Robotics-U0 unify foundation generation and embodied generation?\",\"answer\":\"It reformulates embodied synthesis as an extension of foundation image/video generation and uses a single unified autoregressive training objective. The model jointly optimizes text-to-image, image editing, embodied scene generation, embodied transfer, and embodied video generation.\"},{\"question\":\"What improvements does Xiaomi-Robotics-U0 demonstrate on embodied generation and transfer benchmarks?\",\"answer\":\"It achieves state-of-the-art results on single-step and sequential generation tasks and outperforms GPT-Image-2.0 in human evaluations for embodied scene generation and transfer. It ranks first on World Arena for embodied video generation and raises π0.5 success rate on challenging real-world manipulation tasks from 36.9% to 63.2%.\"}]",1784210344,98,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"xiaomi-robotics-u0-unified-embodied-synthesis-with-world-foundation-model","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/xiaomi-robotics-u0-unified-embodied-synthesis-with-world-foundation-model/86304/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What key limitations prevent directly applying foundation image/video generation models to embodied robotics?","Question",{"text":75,"@type":76},"Embodied settings require strict multi-view consistency, accurate geometric and physical coherence, and explicit robot embodiment constraints. Generated observations must also remain compatible with robot kinematics, camera calibration, and downstream manipulation policies.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Xiaomi-Robotics-U0 unify foundation generation and embodied generation?",{"text":80,"@type":76},"It reformulates embodied synthesis as an extension of foundation image/video generation and uses a single unified autoregressive training objective. The model jointly optimizes text-to-image, image editing, embodied scene generation, embodied transfer, and embodied video generation.",{"name":82,"@type":73,"acceptedAnswer":83},"What improvements does Xiaomi-Robotics-U0 demonstrate on embodied generation and transfer benchmarks?",{"text":84,"@type":76},"It achieves state-of-the-art results on single-step and sequential generation tasks and outperforms GPT-Image-2.0 in human evaluations for embodied scene generation and transfer. It ranks first on World Arena for embodied video generation and raises π0.5 success rate on challenging real-world manipulation tasks from 36.9% to 63.2%.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]