[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82592-en":3,"doc-seo-82592-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82592,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","World from Motion Generative Dynamic Gaussian Reconstruction from Monocular Video","World from Motion presents a method for generating freely renderable dynamic 3D Gaussian representations from monocular videos. The approach conditions a video model on dense, pixel-aligned renderings that jointly encode appearance, geometry, and 3D scene motion across input and target camera trajectories to correct artifacts and complete missing regions from an initial reconstruction. Training uses aligned multiview video pairs and dynamic 3DGS with simulated monocular artifacts. At test time, generations are distilled back into a consistent high-quality dynamic 3DGS, improving novel-view synthesis and underlying motion, and achieving new state-of-the-art 4D reconstruction with strong in-the-wild generalization.","arXiv :2607 .0 1202v 1 [ cs .CV] 1 Jul 2026  \nWorld from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video  \nLiyuan Zhu1,2 Shengyu Huang2 Amrita Mazumdar2 Tianye Li2 Zan Gojcic2  \nGordon Wetzstein1 Iro Armeni1 Shalini De Mello2 Alex Trevithick2  \n1 Stanford University 2NVIDIA  \n[https://research.nvidia.com/labs/amri/projects/world-from-motion/](https://research.nvidia.com/labs/amri/projects/world-from-motion/)  \nInput Monocular Video Initial Dynamic 3DGS Multi-view Video Samples Refined Dynamic 3DGS Initial Dynamic 3DGS  \nt = 100 t = 60 t = 20  \nProgression of World from Motion  \nFigure 1: World from Motion reconstructs a dynamic 3DGS world from the camera and scene motion in a single monocular video. From an input video and an initial reconstruction produced by MoSca [23], our video model generates novel views that are distilled into a refined reconstruction. Our method faithfully recovers observed structure and synthesizes plausible novel-view dynamics, in turn improving the underlying scene motion, as visualized on the right.  \nAbstract  \nWe present World from Motion, a method for generating freely renderable dynamic 3D Gaussian representations from monocular videos. Our approach conditions a video model on dense, pixel-aligned renderings that encode appearance, geometry, and 3D scene motion along both input and target camera trajectories to correct rendering artifacts and fill in missing regions from an initial reconstruction. To train this model, we construct a dataset of aligned multiview video pairs and dynamic 3DGS representations, with simulated artifacts characteristic of monocular reconstruction. At test time, we distill the model’s generations, including newly observed regions and motions, back into a single consistent, high-quality dynamic 3DGS, improving both novel-view synthesis and the underlying 3D motion. Our method sets a new state of the art in 4D reconstruction and seamlessly generalizes to in-the-wild videos with large viewpoint changes and dynamic motions.  \n1 Introduction  \nThe human visual system excels at inferring the 4D structure of the dynamic world from limited 2D observations. This internal \"world model\" combines parallax and stereo cues with learned priors,  \nPreprint.  \nallowing humans to reason about precise geometry and resolve inherent ambiguities in shape, scale, and motion. As a computational realization of this capability, monocular 4D reconstruction enables important applications across robotic simulation, spatial perception, and immersive AR/VR. This process aims to reconstruct a model of the dynamic world’s geometry, appearance, and motion for rendering at arbitrary viewpoints and times.  \nTo replicate these human capabilities, classical 3D vision techniques [36, 45, 14] exploit geometric constraints that are highly effective under sufficient parallax and static scene assumptions. Notably, the outputs of these methods generally improve with more observations. Building on these techniques, the state-of-the-art 4D reconstruction methods [52, 23, 53] accurately recover well-observed static regions and motion estimates with dynamic 3D Gaussian splatting (3DGS) [22, 34] . However, they struggle to infer novel views and unobserved regions of dynamic scene elements.  \nGenerative models offer the opposite tradeoff: they can resolve underconstrained regions by filling in missing appearance, geometry, and motion using learned priors. However, while state-of-the-art video generative models [2, 13] can synthesize high-fidelity frames, they often lack the multiview consistency required for precise 4D reconstruction, leading to inconsistent structure when lifted into 3D. Recent hybrid methods [37, 58] attempt to combine the strengths of both paradigms, but their results generally lack sufficient multiview consistency for precise dynamic reconstruction (see Figure 4) . Consequently, existing frameworks are forced to choose between geometric rigor and generative expressivity, fai","cbCaieCDSyqAPqvC","https://ap.wps.com/l/cbCaieCDSyqAPqvC","pdf",7273641,1,22,"English","en",105,"# Introduction\n## Motivation and central question\n## Challenges in dynamic 4D reconstruction\n## Proposed approach: World from Motion","[{\"question\":\"What problem does World from Motion address in monocular dynamic 4D reconstruction?\",\"answer\":\"It targets the difficulty of inferring novel views and unobserved regions of dynamic scene elements from a single monocular video while maintaining geometric precision and multiview consistency.\"},{\"question\":\"How does the method use dynamic 3D Gaussian representations during generation?\",\"answer\":\"It conditions a video model on dense, pixel-aligned renderings derived from an initial dynamic 3DGS, encoding appearance, geometry, and 3D motion across both input and target camera trajectories.\"},{\"question\":\"What happens at test time to produce the final reconstruction?\",\"answer\":\"The model generates novel views including newly observed regions and motions, then distills these generations back into a single consistent, high-quality dynamic 3DGS that improves both novel-view synthesis and the underlying scene motion.\"}]",1784181696,55,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"world-from-motion-generative-dynamic-gaussian-reconstruction-from-monocular-video","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/world-from-motion-generative-dynamic-gaussian-reconstruction-from-monocular-video/82592/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does World from Motion address in monocular dynamic 4D reconstruction?","Question",{"text":75,"@type":76},"It targets the difficulty of inferring novel views and unobserved regions of dynamic scene elements from a single monocular video while maintaining geometric precision and multiview consistency.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the method use dynamic 3D Gaussian representations during generation?",{"text":80,"@type":76},"It conditions a video model on dense, pixel-aligned renderings derived from an initial dynamic 3DGS, encoding appearance, geometry, and 3D motion across both input and target camera trajectories.",{"name":82,"@type":73,"acceptedAnswer":83},"What happens at test time to produce the final reconstruction?",{"text":84,"@type":76},"The model generates novel views including newly observed regions and motions, then distills these generations back into a single consistent, high-quality dynamic 3DGS that improves both novel-view synthesis and the underlying scene motion.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]