[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82721-en":3,"doc-seo-82721-105":29,"detail-sidebar-cat-0-en-105":90},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},82721,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Fast 3D Foundation Model Initialized Gaussian Splatting","This work presents a fast approach to high-quality 3D Gaussian Splatting (3DGS) reconstruction without relying on traditional Structure-from-Motion (SfM). The method uses 3D Foundation Models (3DFMs) to initialize camera poses and point clouds, then jointly optimizes camera poses and Gaussian primitives with a depth-guided loss. Convergence is achieved quickly, even from rough initialization with about 50–60 input views. In sparse views, an MLP pose refinement module and depth supervision improve quality. Experiments on Mip-NeRF 360, Tanks and Temples, and RobustNeRF report competitive results (23.61 dB PSNR, 0.19 LPIPS) with training time around three minutes per scene, enabling near real-time robotics, VR, and autonomous navigation.","VI. International Conference on Electrical, Computer and Energy Technologies (ICECET 2026)  \n6-8 July 2026, Rome-Italy  \nFast 3D Foundation Model Initialized Gaussian Splatting  \n1st Anurag Dalal  Dep. of Eng. Sciences University of Agder Grimstad, Norway [anurag.dalal@uia.no](anurag.dalal@uia.no)  \n2nd Daniel Hagen  Dep. of Eng. Sciences University of Agder Grimstad, Norway [daniel.hagen@uia.no](daniel.hagen@uia.no)  \n3rd Kjell G. Robbersmyr  Dep. of Eng. Sciences University of Agder Grimstad, Norway [kjell.g.robbersmyr@uia.no](kjell.g.robbersmyr@uia.no)  \n4th Kristian Muri Knausgrd  Top Research Centre Mechatronics University of Agder Grimstad, Norway [kristianmk@ieee.org](kristianmk@ieee.org)  \narXiv :2607 .03209v 1 [ cs .CV] 3 Jul 2026  \nAbstract—This paper introduces a fast method for high-quality 3D Gaussian Splatting (3DGS) reconstruction without traditional Structure-from-Motion (SfM). The proposed approach leverages 3D Foundation Models (3DFMs) for camera pose and pointcloud initialization, then jointly optimizes both camera posesand Gaussian primitives using a depth-guided loss function. This enables fast convergence even from rough initialization with as few as 50–60 input views. To further improve reconstruction quality in sparse-view scenarios, an MLP-based pose refinement module is introduced alongside depth-guided supervision from the foundation model. Extensive experiments on Mip-NeRF 360, Tanks and Temples, and RobustNeRF demonstrate that the proposed method achieves competitive reconstruction quality (23.61 dB PSNR, 0.19 LPIPS) while reducing training time to approximately three minutes per scene. The proposed method produces ready-to-use 3DGS models at a fraction of the time required by existing pipelines, making it suitable for near realtime applications in robotics, VR, and autonomous navigation.  \nI. INTRODUCTION  \n3D Gaussian Splatting (3DGS) [1] has emerged as a powerful technique for high-quality and computationally efficient 3D scene reconstruction, transforming applications ranging from VR and robotics to autonomous navigation. Alongside, 3D Foundation Models (3DFMs) have gained tremendous popularity for geometric understanding tasks. However, most existing 3DGS pipelines rely on precise camera poses obtained from Structure-from-Motion (SfM) systems like COLMAP [2], [3] . SfM is computationally intensive, scales quadratically with the number of input images, and can be unreliable in challenging scenarios such as low-texture scenes or limited viewpoints [4] . This dependency not only increases overall processing time but also limits the use of 3DGS in real-time or large-scale scenarios, while propagating pose estimation errors into thereconstruction.  \nRecent advances in 3DFMs [5]–[7] have shown promise in reducing the reliance on SfM by employing alternative  \nWe would like to extend our sincere thanks to Aust-Agder utviklingsog kompetansefond (AAUKF) for the generous funding of the Arven etter Dannevig (The legacy of Dannevig) project, nr 62/22 which has been instrumental in the completion of this paper.  \nmethods for camera pose estimation and scene reconstruction. These methods often utilize transformer-based deep learning techniques to infer camera poses directly from multi-view images of a scene. By integrating these approaches with 3DGS, it is possible to achieve rapid initialization of the gaussian primitives in order of milliseconds and accurate 3D reconstructions without the overhead of traditional SfM pipelines, which takes a long time depending on the datasetsize. VGGT-X [8], which is based on 3DFM VGGT [5] isone such method that has demonstrated the ability to estimate camera poses along with a point-cloud with improved accuracy and low latency compared to COLMAP, making it a suitable candidate for integration with 3DGS.  \nAnother technique that has shown potential in this domain is 3R-GS [9], which optimizes camera poses along with 3DGS. By jointly optimizing both the camera parameters a","cbCaiaV38ZfCNw5c","https://ap.wps.com/l/cbCaiaV38ZfCNw5c","pdf",6462758,3,1,"English","en",105,"# Introduction\n## Problem with SfM in 3DGS Pipelines\n## Recent 3DFM-Based Pose Estimation\n## Prior Work: VGGT-X and 3R-GS\n## Proposed SfM-Free Fast 3DGS Pipeline","[{\"question\":\"What problem does the paper address in standard 3D Gaussian Splatting pipelines?\",\"answer\":\"Most existing 3DGS pipelines depend on accurate camera poses from SfM systems such as COLMAP, which is computationally intensive, can fail in low-texture or limited-view scenarios, and may propagate pose errors into reconstruction.\"},{\"question\":\"How does the proposed method initialize cameras and point clouds without SfM?\",\"answer\":\"It leverages a 3D Foundation Model (3DFM) to estimate camera poses and initialize point clouds, then performs joint optimization of camera poses and Gaussian primitives.\"},{\"question\":\"How does the method improve reconstruction quality when the number of input views is sparse?\",\"answer\":\"It introduces a depth-guided loss using depth priors from the foundation model and adds an MLP-based pose refinement module, providing stronger geometric constraints than photometric losses alone.\"}]",1784182492,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":85,"head_meta":87,"extra_data":89,"updated_unix":27},"fast-3d-foundation-model-initialized-gaussian-splatting","",{"@graph":35,"@context":84},[36,52,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,49],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":20},"https://docshare.wps.com/document/research-report/",{"item":50,"name":13,"@type":42,"position":51},"https://docshare.wps.com/document/fast-3d-foundation-model-initialized-gaussian-splatting/82721/",4,{"url":50,"name":13,"@type":53,"author":54,"headline":13,"publisher":56,"fileFormat":59,"inLanguage":23,"description":14,"dateModified":60,"datePublished":61,"encodingFormat":59,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":55},"Person",{"url":40,"name":57,"@type":58},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":64,"interactionType":65,"userInteractionCount":20},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70,76,80],{"name":71,"@type":72,"acceptedAnswer":73},"What problem does the paper address in standard 3D Gaussian Splatting pipelines?","Question",{"text":74,"@type":75},"Most existing 3DGS pipelines depend on accurate camera poses from SfM systems such as COLMAP, which is computationally intensive, can fail in low-texture or limited-view scenarios, and may propagate pose errors into reconstruction.","Answer",{"name":77,"@type":72,"acceptedAnswer":78},"How does the proposed method initialize cameras and point clouds without SfM?",{"text":79,"@type":75},"It leverages a 3D Foundation Model (3DFM) to estimate camera poses and initialize point clouds, then performs joint optimization of camera poses and Gaussian primitives.",{"name":81,"@type":72,"acceptedAnswer":82},"How does the method improve reconstruction quality when the number of input views is sparse?",{"text":83,"@type":75},"It introduces a depth-guided loss using depth priors from the foundation model and adds an MLP-based pose refinement module, providing stronger geometric constraints than photometric losses alone.","https://schema.org",{"og:url":50,"og:type":86,"og:title":13,"og:site_name":57,"og:description":14},"article",{"robots":88,"canonical":50},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":91},[92,96,100,104,109,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":97,"show_sort_weight":98,"slug":99},"Literature",80,"literature",{"id":51,"doc_module":4,"doc_module_name":45,"category_name":101,"show_sort_weight":102,"slug":103},"Exam",70,"exam",{"id":105,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},5,"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":28,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":28,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":45,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":45,"category_name":135,"show_sort_weight":105,"slug":136},19,"General","general"]