[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83135-en":3,"doc-seo-83135-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83135,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","ROBOSNAP One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and Evaluation","Recovering real-world scenes as interactive simulation environments can enable generalizable robot learning and reproducible policy evaluation. However, constructing scenes that are both physically stable and visually faithful remains slow and expensive. ROBOSNAP presents a real-to-sim framework that converts a single RGB image into a simulation-ready scene using a layered design: collision-aware interaction assets for stable robot contact and a 3D Gaussian splatting visual layer to preserve faithful background appearance. Experiments on DROID scenes and real-robot tasks validate reliable trajectory replay, task-specific synthetic data generation, and meaningful sim-real correlation via the new DROID-Sim dataset.","􀂃ROBOSNAP: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and  \nEvaluation  \nShujie Zhang 1 4 * Jingkun Yi 1 3 * Weipeng Zhong 1 2 Zirui Zhou4 Yangkun Zhu 1 Hanqing Wang 1 Xudong Xu 1 † Weinan Zhang 1 2 Chunhua Shen 1 3  \n1 Shanghai AI Laboratory 2 Shanghai Jiao Tong University 3Zhejiang University  \n4Tsinghua University  \nHomepage: [https://robosnap.github.io](https://robosnap.github.io)  \narXiv :2607 .06699v 1 [ cs .RO] 7 Jul 2026  \nReal Sim Real Sim  \nFigure 1: From a single RGB image, ROBOSNAP reconstructs a reusable simulation-ready scene with interactive physical assets and visual context. The recovered scenes support trajectory replay (top-left), task-specific data generation and augmentation (top-right, bottom-left), and policy evaluation with meaningful sim-real correlation (bottom-right) .  \nAbstract: Recovering real-world scenes as interactive simulation environments can enable generalizable robot learning and reproducible policy evaluation. However, constructing scenes that are both physically stable and visually faithful remains slow and expensive. In this work, we present ROBOSNAP, a real-to-sim framework that turns a single RGB image into a simulation-ready scene. The key idea is a layered design that separates the physics-critical interaction area from the surrounding visual context: collision-aware foreground assets are refined for stable robot interaction, while a 3D Gaussian splatting visual layer preserves faithful background appearance under novel views. Experiments on DROID scenes and real-robot tasks show that ROBOSNAP achieves reliable trajectory replay in therecovered scenes, supports task-specific synthetic data generation for policy training, and yields meaningful sim-real correlation for policy evaluation. To further support real-to-sim research, we introduce DROID-Sim, a real-to-sim companion dataset constructed from 564 real-world scenes in DROID. Extensive experiments suggest that the value of real-to-sim methods lies not only in high-fidelity visual  \n*Equal contribution. †Corresponding author.  \nreconstruction, but in turning real environments into reusable infrastructure for robot learning and evaluation.  \nKeywords: Real-to-Sim-to-Real, Robot Data Generation, Vision-LanguageAction Models, Robot Manipulation  \n1 Introduction  \nRecent robot foundation models formulate manipulation as a conditional action generation task from visual, linguistic, and proprioceptive inputs [1, 2, 3, 4, 5] . As these models scale, large-scale training data and reproducible evaluation have become critical bottlenecks. Although real-world datasets and benchmarks provide substantial physical demonstrations and standardized protocols [6, 7, 8, 9], scalable data acquisition and flexible policy evaluation remain costly, labor-intensive, and hardwarebound. Simulation offers a complementary path for scalable data synthesis, scene augmentation, and repeatable policy assessment, thereby motivating the construction of interactive simulation scenes that are both physically plausible and visually faithful to real-world deployment environments.  \nExisting approaches only partially satisfy these requirements. Procedural and generative scene synthesis methods have scaled simulatable environments for robot learning [10, 11, 12, 13, 14], but they primarily focus on creating diverse simulation scenes rather than recovering reusable interactive replicas of specific in-the-wild real-world scenes. Reconstruction-based real-to-sim methods improve scene alignment but often require multi-view capture or manual refinement [15, 16, 17, 18, 19] . Recent single-image systems reduce the capture burden [20, 21, 22], but their outputs typically target narrower endpoints: retrieval-based digital cousins, task or demonstration synthesis, or partially recovered scenes with static background. As a result, they do not generally recover persistent simulation worlds that can be re-rendered, edited, and reused from new viewpoints,","cbCaigOs6GkG57FF","https://ap.wps.com/l/cbCaigOs6GkG57FF","pdf",12585778,4,1,24,"English","en",105,"# Introduction\n## Related Work\n### 3D Generation and Scene Synthesis","[{\"question\":\"What is ROBOSNAP and what problem does it address?\",\"answer\":\"ROBOSNAP is a one-shot real-to-sim framework that turns a single RGB image into a simulation-ready scene. It targets the need for physically stable and visually faithful environments that can support generalizable robot learning and reproducible evaluation.\"},{\"question\":\"How does ROBOSNAP build a simulation-ready scene from one RGB image?\",\"answer\":\"It reconstructs the interaction area as collision-aware objects and support surfaces, aligns it to a gravity-consistent frame, and refines object poses to fix floating artifacts and unstable contacts. The surrounding context is modeled as a separate visual layer using background completion, 3D Gaussian splatting, and scene lighting.\"},{\"question\":\"What does DROID-Sim contribute to real-to-sim research?\",\"answer\":\"DROID-Sim is a real-to-sim companion dataset built from 564 real-world scenes in DROID. It provides reusable simulation environments, enabling experiments that evaluate trajectory replay, task-specific synthetic data generation, and sim-real correlation for policy evaluation.\"}]",1784185526,60,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"robosnap-one-shot-real-to-sim-scene-generation-for-generalizable-robot-learning-and-evaluation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/robosnap-one-shot-real-to-sim-scene-generation-for-generalizable-robot-learning-and-evaluation/83135/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is ROBOSNAP and what problem does it address?","Question",{"text":75,"@type":76},"ROBOSNAP is a one-shot real-to-sim framework that turns a single RGB image into a simulation-ready scene. It targets the need for physically stable and visually faithful environments that can support generalizable robot learning and reproducible evaluation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does ROBOSNAP build a simulation-ready scene from one RGB image?",{"text":80,"@type":76},"It reconstructs the interaction area as collision-aware objects and support surfaces, aligns it to a gravity-consistent frame, and refines object poses to fix floating artifacts and unstable contacts. The surrounding context is modeled as a separate visual layer using background completion, 3D Gaussian splatting, and scene lighting.",{"name":82,"@type":73,"acceptedAnswer":83},"What does DROID-Sim contribute to real-to-sim research?",{"text":84,"@type":76},"DROID-Sim is a real-to-sim companion dataset built from 564 real-world scenes in DROID. It provides reusable simulation environments, enabling experiments that evaluate trajectory replay, task-specific synthetic data generation, and sim-real correlation for policy evaluation.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":29,"slug":108},5,"Comic","comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]