[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82591-en":3,"doc-seo-82591-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82591,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","Structured 4D Latent Predictive Model for Robot Planning","Video predictive models are increasingly used in robotics to support task generalization, long-horizon planning, and flexible decision-making. Existing methods often rely on 2D video sequences and therefore miss true 3D geometric understanding, which harms spatial reasoning and physical consistency. A Structured 4D Latent Predictive Model is proposed to forecast a scene’s evolving 3D structure in a structured latent space, conditioned on observations and text instructions, then translate predicted futures into executable actions via goal-conditioned inverse dynamics.","Structured 4D Latent Predictive Model for Robot Planning  \nZhiyi Li 1 Peilin Wu * 2 Xiaoshen Han * 2 Ruojin Cai 2 Yilun Du 2  \narXiv :2607 .0 1 166v 1 [ cs .RO] 1 Jul 2026  \nAbstract  \nVideo predictive models are emerging as a powerful paradigm in robotics, offering a promising path toward task generalization, long-horizon planning, and flexible decision-making. However, prevailing approaches often operate on 2D video sequences, inherently lacking the 3D geometric understanding necessary for precise spatial reasoning and physical consistency. We introduce a Structured 4D Latent Predictive Model, which predicts the evolution of a scene’s 3D structure in a structured latent space conditioned on observations and textual instructions. Our representation encodes the scene holistically and can be decoded into diverse 3D formats, enabling a more complete and 3D consistent scene understanding. This structured 4D latent predictive model serves as a planner, generating future scenes that are translated into executable actions by a goal-conditioned inverse dynamics module.  \nExperiments demonstrate that our model generates futures with strong visual quality, substantially better 3D consistency and multi-view coherence compared to state-of-the-art video-based planners. Consequently, our full planning pipeline achieves superior performance on complex manipulation tasks, exhibits robust generalization to novel visual conditions, and proves effective on real-world robotic platforms. Our website is available at [https://structured-4d-model.github.io/](https://structured-4d-model.github.io/) .  \n1. Introduction  \nLearning a general-purpose agent that can solve a wide range of real-world tasks has always been a central goal in robotics. However, progress is constrained by the scarcity of large-scale, task-diverse, and interactive robotic data required to train such agents. As a result, recent work has focused on policy-based agents that directly map observa-  \n*Equal contribution 1MIT 2Harvard University. Correspondence to: Zhiyi Li \u003C[zhiyi24@mit.edu](zhiyi24@mit.edu) >.  \nPreprint. July 2, 2026.  \ntions to actions (Lillicrap et al., 2015 ; Chi et al., 2023 ; Zhao et al., 2023 ; Xiong et al., 2021) . Although these end-to-end policies can perform well in narrow, well-instrumented settings, they commonly fail to generalize under even modest distribution shifts, such as changes in lighting, viewpoint, or the composition of unseen tasks.  \nAn alternative paradigm is to learn a dynamics model that predicts the consequences of actions, enabling planning and improving generalization. Traditionally, dynamic models have been used extensively in various robotic tasks such as locomotion and manipulation, enabling effective planners like model predictive control (MPC) to operate on top of them (Garcia et al., 1989 ; Mayne et al., 2000 ; Qi et al., 2025) . Recent work in robot learning revives this approach by learning video generative models from large scale datasets and combining them with planners or inverse dynamics modules (Du et al., 2023b ; Yang et al., 2023) . Such models predict how the environment will evolve conditioned on text or task specifications. This makes decision-making more interpretable and flexible, facilitates long-horizon planning, and improves generalization to unseen tasks and environments. However, video-based predictive models are inherently 2D and operate in pixel space, resulting in physical inconsistencies with the real 3D world and limiting accurate spatial understanding. This limitation becomes especially problematic in fine-grained manipulation tasks where accurate 3D cues are essential (Zhu et al., 2024 ; Ke et al., 2024) .  \nModeling the dynamics of a scene directly in 3D is challenging. Traditional 3D representations, such as point clouds and meshes, preserve geometry but lose rich visual detail necessary for semantic understanding. Photorealistic representations like Neural Radiance Fields (NeRFs) (Mildenhall et al., 2","cbCaifjmteMWnV3X","https://ap.wps.com/l/cbCaifjmteMWnV3X","pdf",11219805,2,1,15,"English","en",105,"# Introduction\n## Problem of 2D video-based predictive models\n## Motivation for structured 3D dynamics modeling\n## Proposed structured 4D latent predictive model","[{\"question\":\"What problem does the Structured 4D Latent Predictive Model address?\",\"answer\":\"It addresses the limitation of 2D video-based predictive models that lack inherent 3D geometric understanding, leading to physical inconsistency and weaker spatial reasoning. The model instead forecasts dynamic 3D structure in a structured latent space.\"},{\"question\":\"How does the model incorporate observations and instructions?\",\"answer\":\"The model predicts future scene evolution in a structured latent space conditioned on current observations and textual instructions. This enables goal-directed planning using the encoded 3D dynamics.\"},{\"question\":\"How are predicted futures converted into robot actions?\",\"answer\":\"Predicted future scenes are translated into executable actions by a goal-conditioned inverse dynamics module. This planning pipeline is evaluated in simulation and on real robotic platforms.\"}]",1784181693,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"structured-4d-latent-predictive-model-for-robot-planning","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/structured-4d-latent-predictive-model-for-robot-planning/82591/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the Structured 4D Latent Predictive Model address?","Question",{"text":75,"@type":76},"It addresses the limitation of 2D video-based predictive models that lack inherent 3D geometric understanding, leading to physical inconsistency and weaker spatial reasoning. The model instead forecasts dynamic 3D structure in a structured latent space.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the model incorporate observations and instructions?",{"text":80,"@type":76},"The model predicts future scene evolution in a structured latent space conditioned on current observations and textual instructions. This enables goal-directed planning using the encoded 3D dynamics.",{"name":82,"@type":73,"acceptedAnswer":83},"How are predicted futures converted into robot actions?",{"text":84,"@type":76},"Predicted future scenes are translated into executable actions by a goal-conditioned inverse dynamics module. This planning pipeline is evaluated in simulation and on real robotic platforms.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]