[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81490-en":3,"doc-seo-81490-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},81490,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","ProSGNeRF Progressive Dynamic Neural Scene Graph with Frequency Modulated Foundation Model in Urban Scenes","Implicit neural representations deliver strong 3D reconstruction and novel view synthesis, yet existing methods struggle with fast-moving objects and cannot reliably handle large-scale camera ego-motion in urban environments. This work addresses both scale and dynamics via a progressive scene-graph network that learns local dynamic objects and global urban context, with a temporally windowed allocation for scalable coverage. A foundation-guided object representation and frequency-progressive regularization improve dynamic accuracy and reduce overfitting under sparse observations.","arXiv :2312 .09076v4 [ cs .CV] 10 Jul 2026  \nProSGNeRF: Progressive Dynamic Neural Scene Graph with Frequency Modulated Foundation Model in Urban Scenes  \nTianchen Deng 1 , Yanbo Wang 1 , Yejia Liu 1 , Chenpeng Su 1 , Jingchuan Wang 1 , Hesheng Wang 1 , Danwei Wang2 , Shao-Yuan Lo3 , Weidong Chen* 1  \n1 Shanghai Jiao Tong University.  \n2 Nanyang Technological University.  \n3 National Taiwan University.  \nAbstract  \nImplicit neural representation has demonstrated promising results in 3D reconstruction in various scenes. However, existing approaches either struggle to model fast-moving objects or are incapable of handling large-scale camera ego-motion in urban environments. This leads to low-quality synthesized views of the large-scale urban scenes. In this paper, we aim to jointly solve the problems caused by large-scale scenes and fast-moving vehicles, which are more practical and challenging. To this end, we propose a progressive scene graph network architecture to learn the local scene representations of dynamic objects and global urban scenes. The progressive learning architecture dynamically allocates anew local scene graph trained on frames within a temporal window, with the window size automatically determined, allowing us to scale up the representation to large-scale scenes. Besides, according to our observations, fast-moving objects are observed only in a few frames, which leads to a significant decline in reconstruction accuracy for dynamic objects. Therfore, We introduce a foundation-guided object representation that extracts object-centric visual priors and conditions the density and color decoders in normalized object coordinates. We further propose a frequency-progressive regularization strategy that gradually exposes high-frequency positional, directional, and pose encodings during training, reducing overfitting to sparse object observations. Experimental results demonstrate that our method achieves state-of-the-art view synthesis accuracy, object manipulation, and scene roaming ability in various scenes. The code will be open-sourced on [https://github.com/dtc111111/prosgnerf](https://github.com/dtc111111/prosgnerf).  \nKeywords: 3D Scene Reconstruction, Foundation Models, Neural Scene Graph, Large-scale Urban Scenes  \n1 Introduction  \nUrban scene reconstruction and novel view synthesis are fundamental tasks for autonomous driving [10], robotic navigation [11], city-scale simulation, and virtual/augmented reality [6, 52] . In autonomous driving, photo-realistic and controllable reconstruction of road scenes can support closed-loop simulation, corner-case generation, and  \nsafety validation at a lower cost than large-scale real-world testing.  \nNeRF [28] has shown promising results for view synthesis on static and object-centric scenes. Some recent works improve the original NeRF and achieve large-scale urban scene reconstruction. Block-nerf [38] pre-divides the scene into multiple blocks with per-block update for virtual drive-through reconstruction. Mega-nerf [39] uses  \nFig. 1 Urban scene reconstruction and editing with ProSGNeRF. We show our view synthesis in different time steps (65,262) and scene decomposition results. Our approach significantly improves the view synthesis performance in real-world urban scenes containing multiple dynamic objects and large-scale camera ego-motion. We highlight and enlarge objects in the first column of images, providing the corresponding object PSNR in the top left corner. Scene PSNR is provided in the second column.  \nexpanded octree architecture for large-scale photorealistic fly-through scenarios. BungeeNeRF [44] designs a growing model with residual block structure for city-level scene reconstruction. [13, 27] focus on incremental scene representation, where multiple MLP modules are employed to model the entire scene, and the representations are fused through an online distillation strategy. NSG [32] introduces a scene graph architecture to represent the scene, while SUD","cbCaia4WTk8I5tOJ","https://ap.wps.com/l/cbCaia4WTk8I5tOJ","pdf",13458144,1,23,"English","en",105,"# Introduction\n## Urban scene reconstruction and novel view synthesis\n## Related work and key challenges\n## Motivation and proposed approach","[{\"question\":\"What core problems does ProSGNeRF target in urban novel view synthesis?\",\"answer\":\"It targets low-quality synthesized views caused by large-scale urban scenes combined with fast-moving objects and significant camera ego-motion, which existing approaches handle poorly.\"},{\"question\":\"How does the proposed progressive scene graph architecture scale to large urban scenes?\",\"answer\":\"It progressively allocates and learns a new local scene graph from frames inside a temporal window, where the window size is automatically determined to expand representation capacity while preserving consistency.\"},{\"question\":\"Why is object reconstruction for dynamic objects difficult, and how does the method improve it?\",\"answer\":\"Fast relative motion and occlusion mean each dynamic object is observed in only a few frames, making geometry and appearance under-constrained. The approach introduces a foundation-guided, object-centric representation and a frequency-progressive regularization that gradually reveals high-frequency encodings to reduce overfitting.\"}]",1784173786,58,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"prosgnerf-progressive-dynamic-neural-scene-graph-with-frequency-modulated-foundation-model-in-urban-scenes","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/prosgnerf-progressive-dynamic-neural-scene-graph-with-frequency-modulated-foundation-model-in-urban-scenes/81490/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What core problems does ProSGNeRF target in urban novel view synthesis?","Question",{"text":75,"@type":76},"It targets low-quality synthesized views caused by large-scale urban scenes combined with fast-moving objects and significant camera ego-motion, which existing approaches handle poorly.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed progressive scene graph architecture scale to large urban scenes?",{"text":80,"@type":76},"It progressively allocates and learns a new local scene graph from frames inside a temporal window, where the window size is automatically determined to expand representation capacity while preserving consistency.",{"name":82,"@type":73,"acceptedAnswer":83},"Why is object reconstruction for dynamic objects difficult, and how does the method improve it?",{"text":84,"@type":76},"Fast relative motion and occlusion mean each dynamic object is observed in only a few frames, making geometry and appearance under-constrained. The approach introduces a foundation-guided, object-centric representation and a frequency-progressive regularization that gradually reveals high-frequency encodings to reduce overfitting.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]