[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86376-en":3,"doc-seo-86376-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86376,34359740700684,"Finn","https://ap-avatar.wpscdn.com/avatar/1f400023980c374ae676?_k=1777273430885731487",8,"Research & Report","SplatReasoner: Enhancing Embodied Reasoning and Grounding by Novel View Synthesis","Vision-Language Models (VLMs) show strong reasoning on images and videos, but embodied scene understanding is limited by fixed viewpoints stored in episodic RGB-D memories. Such constraints can hide query-relevant evidence through occlusions, truncations, and narrow fields of view. SplatReasoner introduces novel view synthesis within VLM reasoning using 3D Gaussian Splatting (3DGS). Given a 3D-scene query, it retrieves relevant observations and synthesizes query-conditioned viewpoints that expose evidence and ground referred entities in 3D. Experiments demonstrate improved embodied reasoning and 3D grounding over fixed-view memory and language-embedded 3DGS baselines.","arXiv :2601 . 13132v2 [ cs .CV] 13 Jul 2026  \nSplatReasoner: Enhancing Embodied Reasoning and Grounding by Novel View Synthesis  \nKim Yu-Ji 1 Dahye Lee2 Kim Jun-Seong 1 Nam Hyeon-Woo 1 GeonU Kim2 Yongjin Kwon3 Yu-Chiang Frank Wang4 Jaesung Choe4 Tae-Hyun Oh2  \n1 POSTECH 2 KAIST 3 ETRI 4 NVIDIA  \n[ugkim@postech.ac.kr](ugkim@postech.ac.kr)  \n[https://splatreasoner.github.io/](https://splatreasoner.github.io/)  \nAbstract. Vision-Language Models (VLMs) have demonstrated strong reasoning capabilities over images and videos, yet their application to embodied scene understanding often constrained by the fixed viewpoints stored in episodic RGB-D memories. These observations may fail to capture query-relevant evidence due to occlusions, object truncation, restricted fields of view, or suboptimal view composition. We present SplatReasoner, a framework that introduces novel view synthesis into the VLM reasoning process by leveraging 3D Gaussian Splatting (3DGS) .  \nGiven a user query about a 3D scene, SplatReasoner retrieves relevant observations and synthesizes query-conditioned viewpoints that reveal the visual evidence needed to answer the query and ground the referred entities in 3D. Experiments show that query-conditioned novel view synthesis improves both embodied reasoning and 3D grounding over fixed-view memory and language-embedded 3DGS baselines.  \nKeywords: Embodied Reasoning · 3D Grounding · Novel View Synthesis  \n1 Introduction  \nBuilding embodied agents capable of understanding 3D environments [33] and executing natural-language instructions [6, 66] remains a foundational goal in computer vision and robotics. Recently, Vision-Language Models (VLMs) [3, 7, 16, 22, 34] have demonstrated advanced reasoning and grounding capability for semantic understanding. As these models have proven effective at interpreting image and video data, research has naturally pivoted toward applying their capabilities to embodied tasks in 3D scenes. However, bridging this gap is a critical challenge, as it requires equipping 2D-native VLMs with a cohesive memory of past observations and a robust, structural understanding of 3D space.  \nPrevious approaches [4,7,9,15,16,37,59] demonstrate that VLMs can perform complex reasoning and grounding across multi-view images when provided with structured memory built from RGB-D observations. While these approaches make notable progress toward memory-based embodied reasoning, they typically operate on a set of pre-captured image sequences as a structured memory and  \n2 K. Yu-Ji et al.  \nNovel view synthesis  \n3D Gaussians  \nEmbodied Reasoning  \nQ: “What color is the car?”  \nA: Blue  \n3D Visual Grounding  \nQ: “The blue car in the garage.”  \nA:  \nFig. 1: SplatReasoner aims to enhance embodied task capacity by integrating VLM with novel view synthesis of the 3DGS representation. Specifically, SplatReasoner retrieves and synthesizes informative camera viewpoints to provide the most relevant visual evidence, thereby facilitating improved query answering and 3D entity grounding.  \nobject-level abstractions, which naturally constrain the investigation of 3D environments into a fixed set of 2D views. Relying on a fixed set of pre-recorded images inherently restricts a VLM’s reasoning and grounding capabilities. Consequently, these models often struggle with sub-optimal view compositions, such as occlusions, small objects, and truncations at image boundaries. While the state-of-the-art methods, including recent 3D-Mem [59], demonstrated strong reasoning capabilities, they particularly faces challenges in these cases by design.  \nIn this context, 3D Gaussian Splatting (3DGS) [25] can be an advantageous representation alternative to the fixed-view based memory by virtue of its favorable novel view synthesis functionality. As a separate line of research, language-driven 3D scene understanding methods [1, 11, 18, 24, 32, 39, 43, 50, 55, 56] have been proposed to connect 3D and language with the novel view synthesis functi","cbCaifn5NCcuHeOE","https://ap.wps.com/l/cbCaifn5NCcuHeOE","pdf",4321283,4,1,36,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"What problem does SplatReasoner address in embodied reasoning with VLMs?\",\"answer\":\"SplatReasoner addresses the limitation of fixed-view episodic RGB-D memory, which can miss query-relevant evidence due to occlusions, truncation, restricted fields of view, or poor view composition.\"},{\"question\":\"How does SplatReasoner use novel view synthesis in its reasoning pipeline?\",\"answer\":\"For a user query, SplatReasoner selects initial candidate views by matching simplified queries to 3D Gaussians, then uses a VLM-as-Judge mechanism with synthesized novel views to finalize view selection for downstream tasks.\"},{\"question\":\"What representation and method underpin SplatReasoner’s view synthesis?\",\"answer\":\"SplatReasoner leverages 3D Gaussian Splatting (3DGS) to synthesize query-conditioned viewpoints, improving evidence retrieval and 3D entity grounding.\"}]",1784211273,91,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"splatreasoner-enhancing-embodied-reasoning-and-grounding-by-novel-view-synthesis","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":20},"https://docshare.wps.com/document/splatreasoner-enhancing-embodied-reasoning-and-grounding-by-novel-view-synthesis/86376/",{"url":52,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does SplatReasoner address in embodied reasoning with VLMs?","Question",{"text":75,"@type":76},"SplatReasoner addresses the limitation of fixed-view episodic RGB-D memory, which can miss query-relevant evidence due to occlusions, truncation, restricted fields of view, or poor view composition.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does SplatReasoner use novel view synthesis in its reasoning pipeline?",{"text":80,"@type":76},"For a user query, SplatReasoner selects initial candidate views by matching simplified queries to 3D Gaussians, then uses a VLM-as-Judge mechanism with synthesized novel views to finalize view selection for downstream tasks.",{"name":82,"@type":73,"acceptedAnswer":83},"What representation and method underpin SplatReasoner’s view synthesis?",{"text":84,"@type":76},"SplatReasoner leverages 3D Gaussian Splatting (3DGS) to synthesize query-conditioned viewpoints, improving evidence retrieval and 3D entity grounding.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]