[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-82375-en":3,"doc-seo-82375-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},82375,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility","A fundamental challenge in 3D reconstruction and robotic localization is co-visibility: identifying which image pairs share overlapping visible surfaces, especially under minimal overlap. VGGT implicitly encodes co-visibility as emergent behavior without task-specific supervision, with early layers forming 3D-aware scene representations and late layers acting as co-visibility reasoners. Layer L17 consistently routes nonco-visible pairs. Based on this, Co-VGGT freezes VGGT and trains a lightweight mixture-of-experts head from RGB to classify co-visibility. On Co-VisiON it exceeds human annotation baselines and improves prior work by over 25% pairwise and 10% multiview, with well-calibrated predictions usable as edge weights for SfM/SLAM pipelines.","arXiv :2607 .09503v1 [ cs .CV] 10 Jul 2026  \nWhat VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility  \nFilippo Ziliotto 1 ,2, Luciano Serafini2, Lamberto Ballan 1, and Tommaso  \nCampari2  \n1 University of Padova  \n2 Fondazione Bruno Kessler (FBK)  \nAbstract. A fundamental challenge in 3D reconstruction and robotic localization is co-visibility: determining which image pairs share overlapping visible surfaces, particularly in scenarios with minimal overlap.  \nWe demonstrate that VGGT implicitly encodes co-visibility as an emergent behavior: without any supervision for this task, its internal representations exhibit a clear hierarchical structure mirroring that of large language models, i.e. early layers build a 3D-aware scene representation, while late layers act as dedicated co-visibility reasoners. In particular, we identify layer L17 as a negative anchor that consistently routes nonco-visible pairs for this backbone, regardless of the evaluation setting, providing task-grounded evidence of layer specialization in a geometrygrounded foundation model. Building on this, we introduce Co-VGGT, which freezes VGGT and trains only a lightweight layer-wise mixtureof-experts head (∼7.5M parameters) to classify co-visibility from RGB alone, treating each layer as a specialized expert whose geometric abstraction is adaptively weighted per input pair. On the Co-VisiON benchmark, Co-VGGT surpasses the human annotation baseline and improves over prior work by more than 25% pairwise and 10% multiview. Pairwise predictions are well-calibrated (ECE = 0.030), enabling direct use as edge weights in visibility graphs for downstream SfM and SLAM pipelines without post-hoc correction. Code and data are available 1 .  \nKeywords: Co-visibility · Multiview Geometry · Embodied perception  \n1 Introduction  \nRobotic perception and 3D reconstruction systems typically operate on sparse, imperfect image sets rather than isolated, perfect views. A fundamental challenge in these pipelines—whether for mapping, localization, or scene reconstruction—is determining co-visibility, the subset of surfaces jointly observed by multiple cameras. High co-visibility yields abundant geometric constraints and stable optimization. Conversely, when spatial overlap is limited or absent, matching becomes ambiguous and pose estimation drifts. In these scenarios, reconstruction pipelines often fail silently, generating plausible but geometrically inconsistent  \n1 [https://github.com/covisibility-probing](https://github.com/covisibility-probing)  \n2 F. Ziliotto et al.  \nFig. 1: Co-visibility Task. We study co-visibility prediction: given multiple RGB views of a scene, the goal is to determine which image pairs share overlapping visible regions. Our approach probes geometric consistency in a frozen VGGT foundation model, extracting layer-wise view embeddings and combining them through a lightweight mixture-of-experts head. Without modifying the backbone, the model aggregates information across transformer layers to produce a co-visibility probability for each pair, enabling efficient construction of scene-level visibility graphs.  \nstructures. Operating in this sparse-view regime is not an edge case but rather a standard condition for embodied agents navigating complex, occluded environments.  \nDespite advances in multiview transformers and learned reconstruction models, performance drops significantly as spatial overlap decreases. Under these conditions, fusion modules misalign features, learned descriptors drift, and global reasoning degrades into spurious correlations. Recent benchmarks, such as CoVisiON [8], explicitly isolate this co-visibility reasoning, exposing a critical performance gap between current models and human baselines in sparse, highly imbalanced scenarios. Bridging this gap is essential for robotics: robust co-visibility estimation dictates which view pairs to match, which constraints to trust, and when to flag reconstruction fail","cbCainsqHhm6T0h9","https://ap.wps.com/l/cbCainsqHhm6T0h9","pdf",10559132,3,1,25,"English","en",105,"# Introduction\n## Co-visibility in sparse-view robotics\n## Geometry-grounded foundation models and VGGT","[{\"question\":\"What problem does the document address?\",\"answer\":\"It addresses co-visibility estimation: determining which pairs of images share overlapping visible surfaces, particularly when spatial overlap is minimal.\"},{\"question\":\"How does VGGT relate to co-visibility?\",\"answer\":\"VGGT implicitly encodes co-visibility as an emergent behavior, where early layers build 3D-aware representations and late layers function as co-visibility reasoners; layer L17 acts as a consistent negative anchor for nonco-visible pairs.\"},{\"question\":\"What is Co-VGGT and how is it trained?\",\"answer\":\"Co-VGGT freezes the VGGT backbone and trains only a lightweight mixture-of-experts head (about 7.5M parameters) to classify co-visibility using RGB alone, treating each layer as a specialized expert with adaptively weighted geometric abstraction.\"}]",1784180009,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"what-vggt-knows-about-overlap-probing-geometric-foundation-models-for-co-visibility","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":20},"https://docshare.wps.com/document/research-report/",{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/what-vggt-knows-about-overlap-probing-geometric-foundation-models-for-co-visibility/82375/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-23","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the document address?","Question",{"text":75,"@type":76},"It addresses co-visibility estimation: determining which pairs of images share overlapping visible surfaces, particularly when spatial overlap is minimal.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does VGGT relate to co-visibility?",{"text":80,"@type":76},"VGGT implicitly encodes co-visibility as an emergent behavior, where early layers build 3D-aware representations and late layers function as co-visibility reasoners; layer L17 acts as a consistent negative anchor for nonco-visible pairs.",{"name":82,"@type":73,"acceptedAnswer":83},"What is Co-VGGT and how is it trained?",{"text":84,"@type":76},"Co-VGGT freezes the VGGT backbone and trains only a lightweight mixture-of-experts head (about 7.5M parameters) to classify co-visibility using RGB alone, treating each layer as a specialized expert with adaptively weighted geometric abstraction.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]