[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-83672-en":3,"doc-seo-83672-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},83672,3848291630094,"Emma Wilson","https://eur-avatar.wpscdn.com/davatar_085a072bc5b1113ac321206ff7593b45",8,"Research & Report","Reliability-Aware Monocular Depth Supervision for Sparse-View Neural Reconstruction","Sparse-view neural reconstruction is difficult in outdoor driving scenarios where camera motion follows a narrow forward trajectory, yielding limited multiview overlap. Monocular depth estimators provide dense geometric priors but produce noisy, region-varying reliability. This work aligns Depth Anything V2 predictions to metric depth via scale-shift fitting, then applies masked depth supervision using photometric errors from an RGB-only baseline. Experiments on Mip-NeRF-360 and Splatfacto on KITTISeq show marginal gains for Mip-NeRF-360, but PSNR and RMSE improve substantially for Splatfacto, attributing gains to selecting reliable low-error regions and moderate weighting.","Reliability-Aware Monocular Depth Supervision for Sparse-View Neural Reconstruction  \nWei-Teng Chu* Yashasvini Gopalan* Changju Yuan*  \nStanford University Stanford University Stanford University [waynechu@stanford.edu](waynechu@stanford.edu) [ygopalan@stanford.edu](ygopalan@stanford.edu) [ycj2003@stanford.edu](ycj2003@stanford.edu)  \n*Equal contribution  \narXiv :2607 .02554v1 [ cs .CV] 27 Jun 2026  \nAbstract  \nSparse-view neural reconstruction is challenging in outdoor driving scenes, where cameras usually move along a narrow forward-facing trajectory and provide limited multiview overlap. Although monocular depth estimators can provide dense geometric priors, their predictions are noisy, and not uniformly reliable across image regions. In this work, we study monocular depth supervision for sparseview neural reconstruction. We use Depth Anything V2 asa dense monocular depth prior, align its predictions to metric depth using scale-shift fitting, and apply depth supervision selectively through photometric masks generated from an RGB-only baseline model. We evaluate this strategy on two representative scene representations: Mip-NeRF-360 and Splatfacto. On KITTISeq02 under an every2 sparseview setting, masked monocular depth supervision gives only marginal rendering gains forMip-NeRF-360 and does not improve metric geometry. In contrast, Splatfacto benefits more clearly, improving PSNR from 14.903 to 15.932 and reducing RMSE from 0.542 to 0.100. Additional KITTISeq05 experiments and matched-ratio mask ablations further show that the gains for Splatfacto come from selecting reliable low-error regions rather than simply reducing the number of depth-supervised pixels. Additional experiments on the Bicycle scene show that depth supervision can improve geometry while hurting RGB rendering quality when multi-view coverage is already strong. Overall, our results suggest that monocular depth priors are useful for underconstrained sparse-view reconstruction, but should be applied selectively and with moderate weighting.  \n1. Introduction  \nNovel view synthesis aims to reconstruct a scene representation from a set of posed images and render photorealistic images from unseen viewpoints. This problem is cen-  \nFigure 1 . Outline of our reliability-aware monocular depth supervision pipeline for sparse-view outdoor reconstruction. Given sparse forward-facing RGB inputs, we first estimate a monocular depth prior using Depth Anything V2 and fit it to metric depth with scale-shift fitting. We train an RGB-only baseline and obtain photometric reconstruction errors from it. These errors are used to build a reliability mask that keeps the lower-error regions. Finally, the aligned depth prior is applied only to the reliable pixels with a masked depth loss, and the reconstruction model is jointly optimized with the photometric loss for 3DGS/NeRF reconstruction.  \ntral to applications such as autonomous driving simulation, robotics, augmented reality, and digital twins. Recent neural scene representations, including Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS), have achieved impressive rendering quality by optimizing differentiable scene representations directly from images [8, 10] . However, these methods still rely heavily on sufficient multiview coverage. When the input views are sparse, the recon-  \nstruction problem becomes under-constrained, often leading to incorrect geometry and floaters [6, 14, 17] .  \nThis limitation is especially severe in outdoor driving scenes [7, 17] . Unlike object-centric datasets, where cameras move around the target object, autonomous-driving sequences are usually captured by a forward-moving camera along a narrow trajectory. As a result, nearby objects may have limited viewpoint coverage, distant regions provide weak parallax, and large sky regions, shadows, reflective surfaces, and dynamic objects further complicate reconstruction [17] . These challenges make sparse-view outdoor reconstruction an i","cbCaiukeT4VTmtKe","https://ap.wps.com/l/cbCaiukeT4VTmtKe","pdf",4405054,6,1,10,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"Why is sparse-view neural reconstruction challenging in outdoor driving scenes?\",\"answer\":\"Outdoor driving typically uses forward-moving cameras along a narrow trajectory, creating limited multiview overlap. This under-constrains geometry and can lead to incorrect structures such as floaters, while difficult regions (sky, shadows, reflections, dynamic objects) further complicate reconstruction.\"},{\"question\":\"How does the proposed method decide when to trust monocular depth supervision?\",\"answer\":\"It uses Depth Anything V2 as a dense monocular depth prior, aligns predictions to metric depth with scale-shift fitting, then selectively applies depth supervision through photometric masks derived from an RGB-only baseline’s reconstruction error.\"},{\"question\":\"What do the experiments show about different scene representations?\",\"answer\":\"On KITTISeq02 with every-2 sparse views, masked monocular depth supervision yields only marginal rendering gains for Mip-NeRF-360 and does not improve metric geometry. In contrast, Splatfacto benefits clearly, improving PSNR and reducing RMSE, with gains linked to supervising only reliable low-error regions rather than reducing the number of supervised pixels.\"}]",1784189647,25,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"reliability-aware-monocular-depth-supervision-for-sparse-view-neural-reconstruction","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/reliability-aware-monocular-depth-supervision-for-sparse-view-neural-reconstruction/83672/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-26","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is sparse-view neural reconstruction challenging in outdoor driving scenes?","Question",{"text":76,"@type":77},"Outdoor driving typically uses forward-moving cameras along a narrow trajectory, creating limited multiview overlap. This under-constrains geometry and can lead to incorrect structures such as floaters, while difficult regions (sky, shadows, reflections, dynamic objects) further complicate reconstruction.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the proposed method decide when to trust monocular depth supervision?",{"text":81,"@type":77},"It uses Depth Anything V2 as a dense monocular depth prior, aligns predictions to metric depth with scale-shift fitting, then selectively applies depth supervision through photometric masks derived from an RGB-only baseline’s reconstruction error.",{"name":83,"@type":74,"acceptedAnswer":84},"What do the experiments show about different scene representations?",{"text":85,"@type":77},"On KITTISeq02 with every-2 sparse views, masked monocular depth supervision yields only marginal rendering gains for Mip-NeRF-360 and does not improve metric geometry. In contrast, Splatfacto benefits clearly, improving PSNR and reducing RMSE, with gains linked to supervising only reliable low-error regions rather than reducing the number of supervised pixels.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":22,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":22,"slug":133},"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":107,"slug":137},19,"General","general"]