[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82374-en":3,"doc-seo-82374-105":29,"detail-sidebar-cat-0-en-105":82},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82374,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Hydra++ Real-Time Hierarchical 3D Scene Graph Construction With Object-Level Shape Estimation","3D scene graphs encode objects, places, and their spatial relationships, yet existing systems often approximate geometry using centroids, bounding boxes, partial point clouds, or class-level CAD priors, limiting instance-specific detail. Hydra++ integrates learning-based, category-agnostic object shape estimation into a hierarchical 3D scene graph pipeline. The framework adds a reprojection-mask consistency check (RMCC) to filter degenerate outputs, evaluates slower modular estimators to study latency and generalization tradeoffs, and supports hybrid LiDAR-camera sensing for robust outdoor reconstruction. Experiments in simulation and real campus scenes improve object- and scene-level fidelity.","arXiv :2607 .09455v1 [ cs .CV] 10 Jul 2026  \nHydra++: Real-Time Hierarchical 3D Scene Graph Construction With Object-Level Shape Estimation  \nHyungtae Lim 1 , Nathan Hughes 1 , Xihang Yu 1 , Ruihan Xu 1 ,  \nYun Chang 1 , Jingnan Shi 1 , Rajat Talak2 , and Luca Carlone 1  \nIndoor (RGB-D Camera)  \n5 m  \nOutdoor (3D LiDAR + RGB/RGB-D Camera)  \n(b)  \n10 m   \nFig. 1. 3D scene graphs with object-level shape reconstruction generated by Hydra++ . (a) Indoor environment using the uHumans2 dataset [1] . Zoom-in views show full object shapes reconstructed from partial observations. Each color represents a semantic class (e.g., couches in pink, plants in green, trash bins in purple, and chairs in red) . The object shape estimation network in our framework captures intra-class variation, as illustrated by the reconstructed trash bins and chairs with different geometries. (b) Outdoor campus scene using the hybrid LiDAR-camera configuration. Our method reconstructs objects of various sizes, from a small fire hydrant to a large car, under realistic sensing conditions.  \nAbstract—3D scene graphs provide a hierarchical abstraction of environments by encoding spatial entities (e.g., objects, places) and their relationships. However, existing scene graph systems model object geometry coarsely, relying on partial point clouds or class-level CAD templates, which limits instancespecific shape detail. This paper presents Hydra++, a systemlevel investigation into how learning-based object shape estimators can be integrated into a hierarchical 3D scene graph pipeline. Hydra++ incorporates category-agnostic shape estimation and a reprojection-mask consistency check (RMCC) to reject degenerate predictions from partial observations or imprecise segmentation. In its default CRISP-based configuration, Hydra++ performs online scene graph construction; slower estimators such as SAM3D are evaluated as modular alternatives to demonstrate generalization-latency tradeoffs. Furthermore, to address the challenges of sparse and noisy depth measurements in outdoor environments, Hydra++ supports a hybrid LiDAR-camera configuration for largescale operation, improving scene-level reconstruction quality. Experiments in both simulation and real-world outdoor campus scenarios demonstrate that Hydra++ improves object- and scene-level reconstruction quality. Project page is available at [https://hydra-plusplus.github.io/](https://hydra-plusplus.github.io/).  \n1H. Lim, N. Hughes, X. Yu, R. Xu, Y. Chang, J. Shi, and L. Carlone are with the Laboratory for Information & Decision Systems, Massachusetts Institute of Technology, Cambridge, MA, USA, {shapelim, na26933, jimmyyu, multyxu, yunchang, jnshi, [lcarlone}@mit.edu](lcarlone}@mit.edu)  \n[2](2 R. Talak is with the Department of Electrical and Computer Engineering)[ R. Talak is with the Department of Electrical and Computer Engineering](2 R. Talak is with the Department of Electrical and Computer Engineering)[ ](2 R. Talak is with the Department of Electrical and Computer Engineering)at the National University of Singapore, Singapore, {[talak@nus.edu.sg}](talak@nus.edu.sg})  \nThis work was partially funded by the National Research Foundation of Korea (NRF) grant funded by the Korean government (MSIT) (No. RS- 2024-00461409) and by the Department of the Air Force Artificial Intelligence Accelerator, accomplished under Cooperative Agreement Number FA8750-19-2-1000 .  \nI. INTRODUCTION  \nModeling 3D scenes is crucial for many robotics tasks, from navigation to manipulation [2] . With the rise of deep learning and large language models (LLMs), many studies have examined semantic mapping, which captures semantic information about the environment to support reasoning and interaction [3]–[6] . In particular, recent 3D scene graph approaches encode semantics by treating entities (e.g., objects, rooms, and agents) as nodes, and relationships as edges (e.g., adjacency, inclusion, and support), yielding a lightweight and efficient abstract","cbCaisN3Z8RgvI0T","https://ap.wps.com/l/cbCaisN3Z8RgvI0T","pdf",33670130,2,1,"English","en",105,"# Introduction\n## Motivation and Problem\n## Hydra++ Approach and System Design\n## Contributions","[{\"question\":\"What sensing setup does Hydra++ support for large-scale outdoor scenes?\",\"answer\":\"Hydra++ supports a hybrid LiDAR-camera configuration, combining metric-scale depth from a 3D LiDAR with image-based shape estimation to improve both object-level and scene-level reconstruction quality.\"}]",1784179997,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":77,"head_meta":79,"extra_data":81,"updated_unix":27},"hydra-real-time-hierarchical-3d-scene-graph-construction-with-object-level-shape-estimation","",{"@graph":35,"@context":76},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,46,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":20},"https://docshare.wps.com/document/","Document",{"item":47,"name":12,"@type":42,"position":48},"https://docshare.wps.com/document/research-report/",3,{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/hydra-real-time-hierarchical-3d-scene-graph-construction-with-object-level-shape-estimation/82374/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70],{"name":71,"@type":72,"acceptedAnswer":73},"What sensing setup does Hydra++ support for large-scale outdoor scenes?","Question",{"text":74,"@type":75},"Hydra++ supports a hybrid LiDAR-camera configuration, combining metric-scale depth from a 3D LiDAR with image-based shape estimation to improve both object-level and scene-level reconstruction quality.","Answer","https://schema.org",{"og:url":50,"og:type":78,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":80,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":83},[84,88,92,96,101,106,111,114,118,121,125],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":45,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":45,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":28,"slug":117},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":119,"show_sort_weight":28,"slug":120},"World Cup","world-cup",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":122,"slug":124},10,"Lifestyle","lifestyle",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":97,"slug":128},19,"General","general"]