[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85118-en":3,"doc-seo-85118-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85118,2336464648746,"Skyler","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Wat3R Underwater 3D Geometry Learning without Annotations","Underwater 3D geometry estimation faces major obstacles from light attenuation, scattering, and the lack of large-scale, high-quality 3D annotations. Wat3R introduces a cross-domain semi-supervised framework that adapts feed-forward 3D reconstruction from air to underwater scenes without any annotated underwater data, using a teacher-student design trained on abundant unlabeled real underwater video. A cross-view consistency loss transfers geometric cues across views to mitigate degradation. To enable evaluation, Water3D is built for geometric task assessment, and results show clear gains in underwater multi-view depth estimation and point cloud reconstruction.","arXiv :2607 .08772v 1 [ cs .CV] 9 Jul 2026  \nWat3R: Underwater 3D Geometry Learning without Annotations  \nJiangwei Ren, Xingyu Jiang†, Zijie Song, Wei Xu, Hongkai Lin,  \nDingkang Liang, and Xiang Bai  \nHuazhong University of Science and Technology {jwren,jiangxy998,dkliang,[xbai}@hust.edu.cn](xbai}@hust.edu.cn)  \nAbstract. Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-quality 3D annotations. Pioneering methods rely on massive dense annotations that are impractical in underwater settings.  \nIn this paper, we propose Wat3R, a cross-domain semi-supervised learning framework designed to adapt feed-forward 3D reconstruction models from air to underwater scenes. Uniquely, our method eliminates the need for any annotated underwater data following a teacher-student architecture, that learns robust geometry representations merely on abundant unlabeled real underwater video footage. We also design a cross-view consistency loss that leverages geometric cues from other views to compensate for the information degradation in the current view caused by water attenuation and scattering. Furthermore, considering the lack of comprehensive evaluation benchmarks, we construct Water3D, a diverse dataset covering various water bodies and underwater scenarios, designed for geometric task evaluation. Experimental results demonstrate that Wat3R outperforms current state-of-the-art methods in underwater multi-view depth estimation and point cloud reconstruction. The dataset and code are available at [https://github.com/LSXI7/Wat3R](https://github.com/LSXI7/Wat3R) .  \nKeywords: Underwater Vision, Geometry Estimation, VGGT, Crossdomain Semi-supervised Learning  \n1 Introduction  \nUnderwater visual geometry estimation aims to recover the 3D structures, including camera poses, depths, and point clouds, from multi-view underwater imagery. This capability underpins practical applications ranging from underwater robot navigation and obstacle avoidance, to marine mapping, terrain modeling, and underwater archaeology [18, 25] . Unlike on-land scenes, underwater environments present unique challenges caused by the light absorption and scattering, as well as view-dependent degradation. These physical difficulties further lead to the critical scarcity of large-scale, high-quality 3D annotations [29] in this domain, posing great challenges in training a well-performing model.  \n† Corresponding author.  \n2 J. Ren et al.  \n\n|  |  |\n| --- | --- |\n|  |  |\n\n\n| \u003Cbr>\u003Cbr> |  |  |  |  |  |  |  |\n| --- | --- | --- | --- | --- | --- | --- | --- |\n|  | Wat3R (ours) |  | 􀟨3 (ICLR'26) |  | DA3 (ICLR'26) |  | VGGT (CVPR'25) |\n\n\n|  | 6 views | \u003Cbr>Multi-view 3D Reconstruction\u003Cbr>Wat3R VGGT DA3 |\n| --- | --- | --- |\n|  |  |  |\n| \u003Cbr>\u003Cbr>…  50 views\u003Cbr>\u003Cbr>\u003Cbr>\u003Cbr>Wat3R VGGT DA3 |  |  |\n\nFig. 1: Wat3R reconstructs from the open-domain underwater images in a feedforward manner without requiring any underwater 3D annotations. Our Wat3R achieves significant enhancement in both single-view and multi-view tasks. Statistic results also reveal the superior performance of our Wat3R against the SOTA.  \nIn recent years, 3D vision has witnessed a paradigm shift from classical multiview geometry pipelines to feed-forward neural reconstruction models. Advanced methods like DUSt3R [41], VGGT [40] and their successors [6, 19 , 39 , 47] have demonstrated remarkable performance in recovering 3D geometry in a single forward pass. These models learn strong geometric priors from massive on-land datasets with dense 3D ground truths, a.k. a. camera poses and depths. However, directly deploying these powerful models in underwater scenarios results in poor generalization due to the significant domain shift. And the difficulty in obtaining dense 3D labels for real underwater data also prevents the direct training of effective, generalizable models tailored to underwater environments.  \nTo relieve the abo","cbCaidx6lsgV19Pw","https://ap.wps.com/l/cbCaidx6lsgV19Pw","pdf",12413159,2,1,27,"English","en",105,"# Introduction\n# Wat3R Framework\n# Cross-view Consistency Loss\n# Water3D Dataset Construction\n# Experiments and Results","[{\"question\":\"What makes underwater 3D geometry estimation difficult compared with on-land scenes?\",\"answer\":\"Underwater environments cause strong light absorption and scattering, leading to view-dependent degradation and severe scarcity of high-quality 3D annotations.\"},{\"question\":\"How does Wat3R avoid the need for annotated underwater data?\",\"answer\":\"Wat3R uses a teacher-student semi-supervised learning setup and trains the model with abundant unlabeled real underwater video while leveraging geometry priors initialized from annotated on-land data.\"},{\"question\":\"What is the role of the cross-view consistency loss in Wat3R?\",\"answer\":\"It compensates for information degradation in the current view by leveraging geometric cues from other views, improving robustness under attenuation and scattering.\"}]",1784201213,68,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"wat3r-underwater-3d-geometry-learning-without-annotations","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/wat3r-underwater-3d-geometry-learning-without-annotations/85118/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What makes underwater 3D geometry estimation difficult compared with on-land scenes?","Question",{"text":75,"@type":76},"Underwater environments cause strong light absorption and scattering, leading to view-dependent degradation and severe scarcity of high-quality 3D annotations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does Wat3R avoid the need for annotated underwater data?",{"text":80,"@type":76},"Wat3R uses a teacher-student semi-supervised learning setup and trains the model with abundant unlabeled real underwater video while leveraging geometry priors initialized from annotated on-land data.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the role of the cross-view consistency loss in Wat3R?",{"text":84,"@type":76},"It compensates for information degradation in the current view by leveraging geometric cues from other views, improving robustness under attenuation and scattering.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]