[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84002-en":3,"doc-seo-84002-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84002,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator","Embodied navigation seeks agents that understand multimodal goals, reason over 3D space, and reliably reach targets in real environments. Progress is limited by the scarcity of scalable, high-fidelity, physically grounded interactive worlds. Real scans offer realism but lack scale, while synthetic simulators scale easily yet suffer from significant sim-to-real gaps. Image2Sim builds high-quality interactive environments from posed RGB-D sequences using neural 3D feature-Gaussians and geometry-aware panoramic rendering, then scales data generation to large instruction- and action-aligned navigation datasets.","arXiv :2607 .05765v 1 [ cs .CV] 7 Jul 2026  \nImage2Sim: Scaling Embodied Navigation via Generative Neural Simulator  \nZihan Wang1 , Seungjun Lee1 , Yinghao Xu2 , Gim Hee Lee1  \n1National University of Singapore, 2HKUST  \n[zihan.wang@u.nus.edu](zihan.wang@u.nus.edu) , [gimhee.lee@nus.edu.sg](gimhee.lee@nus.edu.sg)  \nAbstract  \nEmbodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lack of scalable, high-fidelity, and physically grounded interactive environments. Although real-world scanned datasets offer visual realism, they are limited by scale. In contrast, synthetic simulators scale more easily but often exhibit large sim-to-real gaps. We introduce Image2Sim, a real-time neural simulation framework that constructs high-quality interactive environments from posed RGB-D image sequences. The central idea is to decouple 3D spatial anchoring from photorealistic observation synthesis. For scene construction, Image2Sim uses a feed-forward feature Gaussian model that lifts posed RGB-D observations into a 3D feature-Gaussian representation in a single pass. For rendering, we propose a Geometry-Aware One-Step Pixel Flow model that transforms sparse and noisy Gaussian projections into high-quality panoramic RGB-D observations. Image2Sim also serves as a fully automated embodied data engine that generates high-fidelity observations, executable actions, and diverse navigation instructions at scale. It converts large collections of videos and images into near 20K interactive scenes and synthesizes more than 10 million navigation training samples. Navigation models trained entirely in these neural environments achieve strong improvements on major benchmarks and transfer effectively to real-world zero-shot settings. These results suggest that scalable neural simulation can serve as a practical training substrate for embodied navigation at scale. Project page: [github.com/MrZihan/Image2Sim](github.com/MrZihan/Image2Sim)  \n1 Introduction  \nLarge language models [1–3], vision foundation models [4–6], and visual generative models [7–9] suggest that scaling data can lead to substantial gains in capability. However, similar scaling in embodied navigation [10–14] has not delivered comparable progress in robust real-world generalization, largely due to the lack of suitable data. A key limitation lies in the availability of scalable sources of high-fidelity, physically grounded, and interactive 3D environments.  \nExisting data sources expose a fundamental tradeoff between visual fidelity and scalability. Training data is typically drawn from two primary sources: real-world 3D scans [15–18] and synthetic procedural environments [19–21] . Real-world scans provide strong visual fidelity, but are expensive to acquire and difficult to scale. Procedural environments improve scalability, but often introduce substantial sim-to-real gaps due to unrealistic assets, layouts, and rendering statistics. Recent generative video models [22–24] offer another potential source of realistic visual data. However, they struggle to maintain the rigid-body consistency and explicit collision structure required for reliable closed-loop interaction. These limitations suggest that scalable embodied data requires both real-world visual fidelity and explicit 3D physical grounding.  \nPreprint.  \n(a) Traditional Navigation Data Pipeline  \nScanned or Synthetic 3D Meshes  \n(b) Image2Sim Neural Simulator (Ours)  \nLabor-intensive  \nHuman Annotation  \nFigure 1: Comparison of (a) traditional navigation data pipeline and (b) our Image2Sim framework.  \nThese failure modes share a common cause: existing pipelines often couple 3D geometry and visual synthesis within a single mechanism, forcing a tradeoff among scale, physical grounding, and visual fidelity. To overcome this tradeoff, we introduce Image2Sim, a real-time neural environment engine that c","cbCaicitnSyud7Ss","https://ap.wps.com/l/cbCaicitnSyud7Ss","pdf",3307959,5,1,24,"English","en",105,"# Abstract\n# Introduction\n# Related Data Challenges\n# Image2Sim Framework\n## Neural Scene Construction\n## Geometry-Aware One-Step Rendering\n## Embodied Data Engine and Instruction Generation","[{\"question\":\"What data does Image2Sim generate at scale for training?\",\"answer\":\"It generates near 20K interactive scenes and synthesizes more than 10 million navigation training samples, including executable actions and navigation instructions aligned with generated trajectories, enabling effective training and real-world zero-shot transfer.\"}]",1784191954,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"image2sim-scaling-embodied-navigation-via-generative-neural-simulator","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/image2sim-scaling-embodied-navigation-via-generative-neural-simulator/84002/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"What data does Image2Sim generate at scale for training?","Question",{"text":76,"@type":77},"It generates near 20K interactive scenes and synthesizes more than 10 million navigation training samples, including executable actions and navigation instructions aligned with generated trajectories, enabling effective training and real-world zero-shot transfer.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":85},[86,90,94,98,101,106,111,114,119,122,126],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":29,"slug":100},"Comic","comic",{"id":102,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},9,"Religion & Spirituality",20,"religion-spirituality",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":120,"show_sort_weight":117,"slug":121},"World Cup","world-cup",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":123,"slug":125},10,"Lifestyle","lifestyle",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":20,"slug":129},19,"General","general"]