[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86207-en":3,"doc-seo-86207-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86207,13056703019662,"Evangeline","https://ap-avatar.wpscdn.com/avatar/be000253a8e92610077?_k=1778726343310543188",8,"Research & Report","Towards Predictive, Aligned, and Scalable Robot Learning","Learning extends beyond memorization by reasoning about novel problems through possibilities. This work presents Lumo-2, a latent world-action model that generates actions by reasoning over physically grounded world dynamics in latent space. The learned latent transitions capture action-inducible future possibilities and provide a unified substrate for cross-modal alignment. Action quality is tied to latent-space geometry, and reconstruction-based objectives misalign control performance. A multi-stage modality pre-alignment progressively aligns actions with latent dynamics, vision, and language, improving scaling and out-of-distribution generalization. Empirically, Lumo-2 improves over VLA and WAM baselines on long-horizon and dexterous manipulation.","arXiv :2607 . 11270v1 [ cs .RO] 13 Jul 2026  \nTowards Predictive, Aligned, and Scalable Robot Learning  \nAstribot Team  \n[research@astribot.com](research@astribot.com)  \n[Project Page: www.astribot.com/research/Lumo2](Project Page: www.astribot.com/research/Lumo2)  \nAuthor List in Contributions  \nAbstract  \nLearning, at its core, extends beyond memorization to the ability to reason and to approach novel problems by navigating a space of possibilities. In this work, we introduce Lumo-2, a latent world-action model that generates actions by reasoning over world dynamics in latent space. The learned latent world dynamics capture physically grounded visual transitions, naturally encoding action-inducible future possibilities and providing a unified substrate for cross-modal alignment. This formulation enables predictive reasoning akin to world modelling, while remaining lightweight and focused on the evolution of physical dynamics relevant to control. Central to our approach is the hypothesis that action generation quality is governed by the geometry of the latent space. We observe that standard reconstruction-based tokenization objectives for action induce representations biased toward low-level signal fidelity, prone to misalignment between reconstruction quality and downstream control performance. To address this limitation, we propose a multi-stage modality pre-alignment strategy, in which action representations are progressively aligned with latent world dynamics, vision, and language. This process enforces cross-modal consistency, promotes abstraction, and induces a semantically structured latent space conducive to predictive reasoning and enables improved scaling properties. We provide a systematic empirical study of latent world modelling and modality alignment, analyzing their roles in scaling laws and out-of-distribution generalization. Our results demonstrate that Lumo-2 achieves consistent gains over strong vision-languageaction (VLA) and world-action model (WAM) baselines, with pronounced improvements in challenging real-world tasks that require temporal reasoning, physical understanding, or high control complexity such as long-horizon and dexterous manipulation. These findings suggest that structured multi-modal alignment, coupled with predictive reasoning, is a fundamental principle for advancing generalizable embodied intelligence.  \n1. Introduction  \n“As projecting, understanding is the mode of being of Da-sein in which it is its possibilities as possibilities.\"  \n—Martin Heidegger, Being and Time  \nAs Martin Heidegger argues, human existence (Da-sein) is fundamentally defined by living in and through possibilities. Understanding is an active projection of oneself into what one can be, where these possibilities remain inherently open. In embodied settings, true robotic intelligence should similarly extend beyond memorizing experience (i.e., training data) . It requires the capability to actively project toward future physical possibilities and to ground action in a coherent, aligned representation of perception, language, and physical dynamics.  \nIn our prior work, Lumo-1 (Tang et al., 2025a) employs explicit structured textual reasoning to guide action generation through coarse-to-fine planning. While effective, this formulation suffers from limited flexibility, poor scalability, and high inference latency. In this work, we introduce Lumo-2, which replaces explicit reasoning with latent reasoning. Specifically, we construct a latent world dynamics space to serve as an implicit reasoning bridge, reducing token complexity while providing a higher-capacity representation for capturing rich spatiotemporal dynamicsand action-relevant dependencies. This latent space is trained to be physically grounded, naturally forming a structured space of action-inducible possibilities to which a learning agent, such as a robot, can align-akin to possessing an internal world model.  \nRecent advances in vision–language–action (VLA) ","cbCaiiZIPypVtLpS","https://ap.wps.com/l/cbCaiiZIPypVtLpS","pdf",51004076,5,1,49,"English","en",105,"# Introduction\n## From explicit planning to latent reasoning\n## Motivation from VLA and WAM limitations\n## Multi-stage modality pre-alignment\n# Lumo-2 overview","[{\"question\":\"What is Lumo-2 and how does it generate robot actions?\",\"answer\":\"Lumo-2 is a latent world-action model that generates actions by reasoning over physically grounded latent world dynamics. It projects future possibilities in latent space and outputs actions to drive robot control.\"},{\"question\":\"Why do reconstruction-based action tokenization objectives cause issues?\",\"answer\":\"The work argues that reconstruction-focused objectives bias representations toward low-level fidelity. This can misalign reconstruction quality with downstream control performance due to latent-space geometry.\"},{\"question\":\"What does multi-stage modality pre-alignment do?\",\"answer\":\"It progressively aligns action representations with latent world dynamics, visual observations, and language. This enforces cross-modal consistency, promotes abstraction, and induces a semantically structured latent space for predictive reasoning and better scaling.\"}]",1784209463,123,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"towards-predictive-aligned-and-scalable-robot-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/towards-predictive-aligned-and-scalable-robot-learning/86207/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is Lumo-2 and how does it generate robot actions?","Question",{"text":76,"@type":77},"Lumo-2 is a latent world-action model that generates actions by reasoning over physically grounded latent world dynamics. It projects future possibilities in latent space and outputs actions to drive robot control.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Why do reconstruction-based action tokenization objectives cause issues?",{"text":81,"@type":77},"The work argues that reconstruction-focused objectives bias representations toward low-level fidelity. This can misalign reconstruction quality with downstream control performance due to latent-space geometry.",{"name":83,"@type":74,"acceptedAnswer":84},"What does multi-stage modality pre-alignment do?",{"text":85,"@type":77},"It progressively aligns action representations with latent world dynamics, visual observations, and language. This enforces cross-modal consistency, promotes abstraction, and induces a semantically structured latent space for predictive reasoning and better scaling.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]