[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85471-en":3,"doc-seo-85471-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85471,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","GeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation","Previous works that use video models for image-to-3D scene generation often produce geometric distortions and blurry results. GeoWorld introduces a two-stage pipeline that revises the generation process by supplying full-frame geometry features. A first-stage video generator, followed by a multi-view geometry model, produces full-frame geometry features that serve as geometric conditions for the second-stage video generation. The method adds a geometric loss to enforce real-world constraints and a geometry adaptation module to utilize geometry features effectively.","arXiv :2511 .23191v3 [ cs .CV] 13 Jul 2026  \nGeoWorld: Providing Full-frame Geometry Features to Facilitate 3D Scene Generation  \nYuhao Wan 1 ,2, Lijuan Liu2 , Jingzhi Zhou 1 , Zihan Zhou3 , Xuying Zhang 1 , Dongbo Zhang2 , Shaohui Jiao2 , Qibin Hou 1 ,4, and Ming-Ming Cheng4 , 1B  \n1VCIP & AAIS, Nankai University 2 ByteDance Inc.  \n3 Renmin University of China 4 NKIARI, Shenzhen Futian  \nInput  \nSee3D ViewCrafter FlexWorld Hunyuan-Voyager GeoWorld(Ours)  \nLimited geometry cond Full-frame geometry constraint  \nBefore full-frame geometry constraint  \nDistorted geometry Blurry content  \nFig. 1: Visual comparisons. Top: Comparison between our GeoWorld and previous methods. By incorporating full-frame geometry constraints, our approach achieves superior visual quality. Bottom left: Results before applying full-frame geometry constraints, which often suffer from geometric distortions and blurry content. Bottom right: Results after applying full-frame geometry constraints. By unlocking the potential of geometry models, our GeoWorld produces clear geometric structures and sharp visual details.  \nAbstract. Previous works that leverage video models for image-to-3D scene generation often suffer from geometric distortions and blurry content. Using video generation models to implicitly maintain geometric consistency according to a single-frame input is ineffective. In this paper, we present a two-stage method, named GeoWorld, that renovates the image-to-3D scene generation pipeline by providing full-frame geometry features. The first-stage video generation model, followed by a multi-view geometry model, produces full-frame geometry features, which are then used as a mental draft of geometric conditions to aid the second-stage video-generation model. A geometric loss is proposed to impose real-world geometric constraints, and a geometry adaptation module is introduced to ensure the effective utilization of geometry features. Thanks to fullframe geometric modeling, the two smaller video models in our two-stage  \n2 Wan et al.  \nmethod can generate higher-fidelity 3D scenes than SOTA methods, while being even faster, e.g . 7.5 × faster than Hunyuan-Voyager. Project page:  \n[https://peaes.github.io/GeoWorld](https://peaes.github.io/GeoWorld).  \nKeywords: 3D scene generation · Video models · Diffusion models  \n1 Introduction  \nGenerating a high-fidelity 3D scene from a single image has become a significant topic in recent years due to its high value in applications such as entertainment, interior and architectural design, and autonomous driving [10–12, 27 , 51 , 63 , 73 , 75 , 78] . Leveraging deep learning methods for this task can significantly advance traditional 3D modeling pipelines. Given the limited information in a single image, a common approach is to use generative model priors to synthesize the scene content. Early methods [5, 6 , 17 , 29 , 43 , 65 , 70 , 76] often rely on 2D generative models [16, 44], which often lead to issues such as structural inconsistency and inconsistencies within the scene content.  \nThanks to the advances in foundational 3D generative models, some recent works [4, 14 , 20 , 21 , 28 , 32 , 36 , 40 , 45–47, 58 , 64 , 67 , 68] use video models [18, 50 , 62] and leverage their implicit 3D priors to alleviate the aforementioned issues. Such methods typically employ video models to synthesize a video under a specified camera trajectory from a single input image, and subsequently reconstruct a 3D scene from the generated video. However, generating high-fidelity videos from a single image remains challenging. As shown at the top of Fig. 1, these methods often suffer from geometric distortions and blurry content, which degrade the quality of the final 3D reconstruction. To address the above issues, a common approach is to provide additional geometric guidance for the model. As shown in Fig. 2(a), some previous works [4, 21 , 67] have utilized estimated monocular depth maps as spatial priors or camera embeddings to a","cbCaifyqhgjKAJsH","https://ap.wps.com/l/cbCaifyqhgjKAJsH","pdf",7186319,3,1,20,"English","en",105,"# Introduction\n## Problem and motivation\n## Related geometric guidance\n## Proposed two-stage pipeline","[{\"question\":\"Why do existing video-model-based methods for image-to-3D often produce distortion and blur?\",\"answer\":\"They typically rely on geometric consistency implicitly maintained from single-frame conditioning, which is ineffective and leads to structural and geometric issues in the reconstructed 3D scene.\"},{\"question\":\"What is the core idea of GeoWorld?\",\"answer\":\"GeoWorld uses a two-stage pipeline where full-frame geometry features are extracted in the first stage and then used as geometric conditions to guide the second-stage video generation for clearer 3D structures.\"},{\"question\":\"How does GeoWorld ensure that geometry features are effectively used?\",\"answer\":\"It proposes a geometric loss to impose real-world geometric constraints and introduces a geometry adaptation module so the second-stage generation can effectively leverage the extracted geometry features.\"}]",1784203840,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"geoworld-providing-full-frame-geometry-features-to-facilitate-3d-scene-generation","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/geoworld-providing-full-frame-geometry-features-to-facilitate-3d-scene-generation/85471/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-24","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do existing video-model-based methods for image-to-3D often produce distortion and blur?","Question",{"text":75,"@type":76},"They typically rely on geometric consistency implicitly maintained from single-frame conditioning, which is ineffective and leads to structural and geometric issues in the reconstructed 3D scene.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the core idea of GeoWorld?",{"text":80,"@type":76},"GeoWorld uses a two-stage pipeline where full-frame geometry features are extracted in the first stage and then used as geometric conditions to guide the second-stage video generation for clearer 3D structures.",{"name":82,"@type":73,"acceptedAnswer":83},"How does GeoWorld ensure that geometry features are effectively used?",{"text":84,"@type":76},"It proposes a geometric loss to impose real-world geometric constraints and introduces a geometry adaptation module so the second-stage generation can effectively leverage the extracted geometry features.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":22,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":22,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]