[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82769-en":3,"doc-seo-82769-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82769,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation","Ego-centric 3D scene generation remains limited by small view overlap and the strong bias of individual perspectives, which obstruct viewpoint-consistent, semantically aligned visual synthesis and the recovery of accurate geometric structures. CGGS proposes a text-to-3D framework that improves 3D content awareness and reduces geometric distortions. It fine-tunes a Multi-View Latent Diffusion Model with consistency-augmented loss for consistent 2D priors, derives dense point clouds using optical flow and correspondence, and refines 3D Gaussians via entropy-based mutual information depth loss with hierarchical optimization, achieving superior results over prior methods.","© 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. This article has been accepted for publication in IEEE Transactions on Image Processing.  \nCGGS: Consistency–Augmented Geometric Gaussian Splatting for  \nEgo-centric 3D Scene Generation  \n1  \nZhenyu Sun, Xiaohan Zhang, Qi Liu†, Senior Member, IEEE, and Huan Wang†, Senior Member, IEEE  \narXiv :2607 .038 19v2 [ cs .GR] 7 Jul 2026  \nAbstract— Challenges remain in ego-centric 3D scene generation due to limited view overlap and the dominant influence of individual perspectives on scene interpretation. These factors hinder the creation of viewpoint-consistent and semantically aligned visual content, as well as the construction of accurate geometric structures. In this paper, we propose CGGS, a textto-3D framework aiming to enhance 3D-content-awareness and address geometric distortions in ego-centric scene generation. Firstly, the Ego-centric Generator is proposed by fine-tuning a Multi-View Latent Diffusion Model with consistency-augmented loss to generate consistent, high-fidelity 2D content aligned with textual descriptions. Then, Layout Decorator leverages optical flow and point-track correspondence to estimate depth, therefore producing dense point clouds as coarse layouts from the egocentric 2D priors. Building on this initialization, Geometric Refiner is proposed to enhance 3D Gaussian reconstruction via an entropy-based Mutual Information Depth Loss (MID) combined with a hierarchical optimization scheme for improving visual quality and geometric structure. Comprehensive experiments demonstrate that CGGS outperforms previous methods in generating coherent and accurate text-driven 3D scenes. Project page: [https://cggs-26.github.io/cggs26/](https://cggs-26.github.io/cggs26/).  \nIndex Terms—3D gaussian splatting, ego-centric generation, semantic alignment, global coherence.  \nI. INTRODUCTION  \n3D scene generation has recently gained significant attention, fueled by advances in generative models and strong image priors. In particular, generating 3D scenes from textual descriptions holds great promise for a wide range of real-world applications in AR/VR, robotics, and autonomous driving. With the rapid development of text-to-image generation [2],[3], [4], progress has been made toward text-to-3D generation. Latent Diffusion Models (LDMs) [4], [5] have been leveraged to optimize Neural Radiance Fields (NeRF) [6] via CLIP  \nThis paper is supported by Young Scientists Fund of the National Natural Science Foundation of China (NSFC) (No. 62506305), Zhejiang Leading Innovative and Entrepreneur Team Introduction Program (No. 2024R01007), Key Research and Development Program of Zhejiang Province (No. 2025C01026), Scientific Research Project of Westlake University (No. WU2025WF003), Chinese Association for Artificial Intelligence (CAAI) & Ant Group Research Fund - AGI Track (No. 2025CAAI-ANT-13) . It is also supported by the research funds of the National Talent Program and Hangzhou Municipal Talent Program. It is also supported in part by the GJYC program of Guangzhou under Grant 2024D01J0081, in part by the ZJ program of Guangdong under Grant 2023QN10X455, and in part by the Fundamental Research Funds for the Central Universities under Grant 2025ZYGXZR053 .  \nZhenyu Sun, Xiaohan Zhang and Qi Liu are with the School of Future Technology, South China University of Technology, Guangzhou 511442, China (email: [ftsunzhenyu@mail.scut.edu.cn](ftsunzhenyu@mail.scut.edu.cn); [ftxiaohanzhang@mail.scut.edu.cn](ftxiaohanzhang@mail.scut.edu.cn); dr  \n[liuqi@scut.edu.cn](liuqi@scut.edu.cn)).  \nHuan Wang is with the School of Engineering, Westlake University, Hangzhou 310030,","cbCaidTub8naFqkr","https://ap.wps.com/l/cbCaidTub8naFqkr","pdf",2856494,3,1,15,"English","en",105,"# Introduction\n## Text-to-3D and multi-view priors\n## Limitations of ego-centric generation\n## CGGS overview and motivation","[{\"question\":\"What key problems does CGGS address in ego-centric 3D scene generation?\",\"answer\":\"Limited view overlap and the dominant influence of individual perspectives reduce viewpoint consistency and semantic alignment, and hinder accurate geometric structure construction.\"},{\"question\":\"How does CGGS generate consistent 2D priors from text?\",\"answer\":\"It fine-tunes a Multi-View Latent Diffusion Model using a consistency-augmented loss in the proposed Ego-centric Generator.\"},{\"question\":\"How does CGGS turn ego-centric 2D priors into 3D geometry and refine it?\",\"answer\":\"Layout Decorator estimates depth using optical flow and point-track correspondences to form dense point clouds as coarse layouts, then Geometric Refiner reconstructs 3D Gaussians using an entropy-based Mutual Information Depth Loss with hierarchical optimization.\"}]",1784182815,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"consistency-augmented-geometric-gaussian-splatting-for-ego-centric-3d-scene-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/consistency-augmented-geometric-gaussian-splatting-for-ego-centric-3d-scene-generation/82769/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What key problems does CGGS address in ego-centric 3D scene generation?","Question",{"text":75,"@type":76},"Limited view overlap and the dominant influence of individual perspectives reduce viewpoint consistency and semantic alignment, and hinder accurate geometric structure construction.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does CGGS generate consistent 2D priors from text?",{"text":80,"@type":76},"It fine-tunes a Multi-View Latent Diffusion Model using a consistency-augmented loss in the proposed Ego-centric Generator.",{"name":82,"@type":73,"acceptedAnswer":83},"How does CGGS turn ego-centric 2D priors into 3D geometry and refine it?",{"text":84,"@type":76},"Layout Decorator estimates depth using optical flow and point-track correspondences to form dense point clouds as coarse layouts, then Geometric Refiner reconstructs 3D Gaussians using an entropy-based Mutual Information Depth Loss with hierarchical optimization.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]