[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84928-en":3,"doc-seo-84928-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},84928,687197207639,"Asher","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models","Existing 3D scene-grounded large language models concentrate on question answering in simplified single-room 3D scenes, without robust reasoning across real household settings with multiple interconnected rooms and varied object categories. CAIRN introduces a topology-aware 3D-LLM for multi-room scene understanding, aligning transformer attention with scene hierarchy to expose object relations and room connectivity. It enriches object tokens using graph neural networks, adds learned room tokens, and applies hierarchical attention with geometric bias. CAIRN is built on CAIRN-MR, a benchmark on HM3D with grounding, captioning, and four QA tasks to evaluate intra- to cross-room reasoning, achieving strong gains over prior 3D-LLMs while staying competitive on single-room benchmarks.","arXiv :2607 .06534v2 [ cs .CV] 13 Jul 2026  \nCAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models  \nHe Liang1 ∗ Chenyang Ma1 Yiming Zhang3 Sangyun Shin1 Andrew Markham 1 Niki Trigoni 1 Yuhang He2†  \n1University of Oxford 2Microsoft Research 3 Simon Fraser University  \n{he.liang, [chenyang.ma](chenyang.ma) , sangyun.shin, andrew.markham, [niki.trigoni}@cs.ox.ac.uk](niki.trigoni}@cs.ox.ac.uk)  \n[yuhanghe@microsoft.com](yuhanghe@microsoft.com) [yza440@sfu.ca](yza440@sfu.ca)  \nAbstract  \nExisting 3D scene-grounded Large Language Models (3D-LLMs) focus on answer  \ning questions grounded in simplified single-room 3D scenes, lacking the ability to  \nreason over real-world household environments containing multiple interconnected  \nrooms and diverse object categories. We introduce CAIRN, a topology-aware 3D  \nLLM for multi-room 3D scene understanding. CAIRN aligns transformer attention  \nwith scene hierarchy, giving the model explicit awareness of object-level relations  \nand room-level connectivity. It enriches object tokens with room-local relational  \ncontext via a graph neural network, introduces learned room tokens for room-level  \nabstraction, and applies a hierarchical attention mask with geometric bias to route  \ninformation according to scene topology. CAIRN is developed on CAIRN-MR, a  \nbenchmark we introduce on HM3D for multi-room 3D scene understanding, cover  \ning grounding, captioning, and four question-answering tasks that progressively  \nevaluate from intra-room perception to cross-room reasoning. Experiments show  \nthat CAIRN outperforms prior 3D-LLMs by a large margin across all CAIRN-MR  \ntasks while remaining competitive on five single-room benchmarks.  \n1 Introduction  \nA capable home assistant agent should recall where an object was last seen in the kitchen and retrieve it from another room, or answer queries that relate objects across spaces (e.g., comparing furniture arrangements across bedrooms) . This requires understanding the full 3D scene structure of a household and performing consistent reasoning over the entire multi-room environment.  \nThere has been a rich history of prior work on developing agents that answer questions grounded in 3D scenes, using large-language models (LLMs) augmented with structured 3D scene representations (i.e., 3D-LLMs) as the reasoning backbone [28, 29, 31, 32, 57, 63] . However, these approaches largely restrict 3D scene understanding and spatial reasoning to a single room [8, 10, 20–22, 43, 60], representing scenes as flat sequences of object tokens without modeling higher-level room structure. In this work, we aim to move beyond the single-room paradigm and develop 3D-LLMs that can understand full household settings [42], where multiple interconnected rooms contain hundreds of objects. This unlocks more challenging tasks such as grounding across spatial contexts, tracing relationships between rooms, and reasoning over a scene’s topological structure.  \nTo enable the study of multi-room 3D scene understanding, we first introduce CAIRN-MR, a benchmark for structured multi-room 3D scene understanding built on HM3D [44], as existing 3D  \n∗ Part of this work was done during an internship at Microsoft Research.†Corresponding author  \nProject Page: [https://oceansdepp.github.io/cairn_web/](https://oceansdepp.github.io/cairn_web/)  \nPreprint.  \nFigure 1: 3D-LLMs for multi-room scene understanding. We introduce CAIRN, a topology-aware 3D-LLM for multi-room scene understanding, along with CAIRN-MR, a multi-room 3D scene understanding benchmark with diverse tasks. By representing each scene as a hierarchical scene graph and aligning attention with scene topology through structured masking and geometric bias, CAIRN achieves substantial gains over prior 3D-LLMs.  \nscene-language datasets [1, 5, 11, 37, 61] are mostly limited to single-room scans. Unlike prior benchmarks where reasoning is confined to a single room with a limited candidate set, CAIRN-MR shi","cbCail0S0oQEDpl8","https://ap.wps.com/l/cbCail0S0oQEDpl8","pdf",3117204,2,1,19,"English","en",105,"# Introduction\n## Background and motivation\n## Proposed approach and benchmark","[{\"question\":\"What limitation do existing 3D scene-grounded 3D-LLMs have in household environments?\",\"answer\":\"They focus on simplified single-room scenes and lack the ability to reason over multiple interconnected rooms with diverse object categories in real households.\"},{\"question\":\"How does CAIRN handle multi-room 3D scene understanding differently from prior models?\",\"answer\":\"CAIRN models a hierarchical 3D scene graph and aligns transformer attention with scene topology using structured masking and geometric bias, enabling explicit room connectivity and object-level relations.\"},{\"question\":\"What is CAIRN-MR and how does it evaluate the model?\",\"answer\":\"CAIRN-MR is a benchmark for structured multi-room 3D scene understanding on HM3D, providing grounding, captioning, and four progressively harder question-answering tasks that span from room-grounded perception to cross-room reasoning.\"}]",1784199393,48,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"cairn-cross-room-3d-scene-understanding-with-topology-aware-large-multimodal-models","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/cairn-cross-room-3d-scene-understanding-with-topology-aware-large-multimodal-models/84928/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What limitation do existing 3D scene-grounded 3D-LLMs have in household environments?","Question",{"text":75,"@type":76},"They focus on simplified single-room scenes and lack the ability to reason over multiple interconnected rooms with diverse object categories in real households.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does CAIRN handle multi-room 3D scene understanding differently from prior models?",{"text":80,"@type":76},"CAIRN models a hierarchical 3D scene graph and aligns transformer attention with scene topology using structured masking and geometric bias, enabling explicit room connectivity and object-level relations.",{"name":82,"@type":73,"acceptedAnswer":83},"What is CAIRN-MR and how does it evaluate the model?",{"text":84,"@type":76},"CAIRN-MR is a benchmark for structured multi-room 3D scene understanding on HM3D, providing grounding, captioning, and four progressively harder question-answering tasks that span from room-grounded perception to cross-room reasoning.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},"General","general"]