[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86065-en":3,"doc-seo-86065-105":29,"detail-sidebar-cat-0-en-105":87},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86065,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","3D Scene Graph Prediction Generating Hierarchical Models from Partially Observed Environments","Generating realistic 3D indoor scenes is a key challenge in computer vision and robotics, yet existing approaches mainly generate object layouts within a single room and leave room-level structure underexplored. This paper addresses partially explored environments by predicting unseen scene parts to support exploration or object search. It proposes a top-down hierarchical 3D scene graph framework with a room layer and an object layer, using mixed-domain graph diffusion for room categories, boundaries, and traversability, and diffusion-based object layout prediction conditioned on room geometry. Experiments show strong generalization to out-of-distribution partial floor plans and successful pipeline evaluation on robot-collected real scenes.","3D Scene Graph Prediction: Generating Hierarchical Models from  \nPartially Observed Environments  \nSiyi Hu 1 , Jared Strader1 ,2 , Hyungtae Lim 1 , and Luca Carlone 1  \narXiv :2607 . 10879v1 [ cs .RO] 12 Jul 2026  \nAbstract—Generating realistic 3D indoor scenes is an area of growing interest in computer vision and robotics. Existing methods, often motivated by applications such as interior design, generally focus on object layout generation within a single room. The generation of high-level scene structure, such as room-level layout and traversability, remains underexplored despite its importance for robotics applications. In this paper, we consider the case where a robot has explored part of an environment and needs to predict the unexplored parts to support downstream tasks such as exploration or object search. We propose a top-down framework for synthesizing hierarchical 3D scene graphs, including a room layer—describing the floor plan and traversability—and an object layer modeling object layouts within each room. For the room layer, we propose a novel mixed-domain graph diffusion model jointly predicting room categories, floor boundaries, and traversability between rooms. Via corruption and masking, this model supports partial constraints such as incomplete floor plans, avoiding the need for partially observed training data. For the object layer, we integrate an existing mixed discrete-continuous diffusion model for joint prediction of object categories, locations, sizes, and orientations within each room given the floor plan. We compare our method with state-of-the-art occupancy-based and LLMbased floor plan generation methods on a standard benchmark. Compared with an occupancy-based learning baseline, our method generalizes substantially better to out-of-distribution partial floor plans. We also demonstrate our integrated prediction pipeline on real-world scenes from robot-collected data, enabling prediction beyond explored areas.  \nI. INTRODUCTION  \nScene generation and prediction play a crucial role in robotics and computer vision, enabling machines to understand, anticipate, and interact with their environments. Advancements in deep learning and graph-based methods have significantly improved scene understanding, enabling autonomous systems to reason about spatial relationships and predict plausible scene configurations to support tasks such as exploration, navigation, and object search.  \n3D Scene Graphs (3DSGs), particularly hierarchical scene graphs [1],[16] that represent 3D scenes at different levels of abstraction (e.g., objects, rooms), have been widely adopted in robotics as a lightweight representation of 3D environments. The hierarchy provides organization for scene understanding and reasoning [16] . However, occlusions and partial observability are common in robotic exploration, making it  \n*This work was partially funded by the ARL DCIST program, the ONR RAPID program, and National Research Foundation of Korea (No. RS- 2024-00461409) .  \n1Laboratory for Information & Decision Systems (LIDS), Massachusetts Institute of Technology, Cambridge, MA 02139, USA. {siyi, jstrader, shapelim, [lcarlone}@mit.edu](lcarlone}@mit.edu)  \n2Department of Electrical and Computer Engineering, Oakland University, Rochester, MI 48309, [USA.](USA. jstrader@oakland.edu)[ jstrader@oakland.edu](USA. jstrader@oakland.edu)  \n(b) Completed floor plan from (c) Layered view of the com  \npartial observations. pleted apartment scene.  \nFig. 1: Real-world hierarchical 3D scene graph prediction. Gray patches denote the observed portion of the scene; colored polygons represent predicted rooms; orange dots indicate detected and green dots indicate synthesized objects.  \nchallenging to predict what lies in unexplored regions and plan efficient exploration strategies. Without the ability to reason about unobserved portions of the environment, robots cannot make informed decisions about where to explore next or anticipate the spatial structure ","cbCaimXDSkb594p4","https://ap.wps.com/l/cbCaimXDSkb594p4","pdf",2437006,1,20,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"How is the proposed approach structured for prediction?\",\"answer\":\"It uses a top-down hierarchical 3D scene graph framework with a room layer (room-level layout, traversability, and connectivity) and an object layer (object categories and their 3D layouts within each room).\"},{\"question\":\"How does the method handle missing or incomplete observations?\",\"answer\":\"For the room layer, it introduces a masking strategy within a mixed-domain graph diffusion model, allowing flexible conditioning on missing inputs without requiring partially observed training data.\"}]",1784208234,50,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":82,"head_meta":84,"extra_data":86,"updated_unix":27},"3d-scene-graph-prediction-generating-hierarchical-models-from-partially-observed-environments","",{"@graph":35,"@context":81},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/3d-scene-graph-prediction-generating-hierarchical-models-from-partially-observed-environments/86065/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77],{"name":72,"@type":73,"acceptedAnswer":74},"How is the proposed approach structured for prediction?","Question",{"text":75,"@type":76},"It uses a top-down hierarchical 3D scene graph framework with a room layer (room-level layout, traversability, and connectivity) and an object layer (object categories and their 3D layouts within each room).","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the method handle missing or incomplete observations?",{"text":80,"@type":76},"For the room layer, it introduces a masking strategy within a mixed-domain graph diffusion model, allowing flexible conditioning on missing inputs without requiring partially observed training data.","https://schema.org",{"og:url":51,"og:type":83,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":85,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":88},[89,93,97,101,106,110,115,118,122,125,129],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Exam",70,"exam",{"id":102,"doc_module":4,"doc_module_name":45,"category_name":103,"show_sort_weight":104,"slug":105},5,"Comic",60,"comic",{"id":107,"doc_module":4,"doc_module_name":45,"category_name":108,"show_sort_weight":28,"slug":109},6,"Technology","technology",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":116,"slug":117},30,"research-report",{"id":119,"doc_module":4,"doc_module_name":45,"category_name":120,"show_sort_weight":21,"slug":121},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":21,"slug":124},"World Cup","world-cup",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":126,"slug":128},10,"Lifestyle","lifestyle",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":102,"slug":132},19,"General","general"]