[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86536-en":3,"doc-seo-86536-105":29,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":11,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},86536,962075114765,"Quinn","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Revisiting Matching Response and Swept Feature Volumes for Wide-baseline Omnidirectional Stereo","This paper presents a confidence-estimation training strategy for omnidirectional stereo under wide-baseline ambiguity, targeting frequent erroneous matches caused by global sweeping. By reinterpreting the expectation values of matching responses from a 3D encoder–decoder block as intrinsic confidence signals, the method penalizes ambiguous responses directly. It avoids auxiliary heads, multi-pass inference, and extra modules, improving efficiency and generalization. In addition, swept feature volume resampling uses regressed positive matching indices to enable joint prediction of meta-information such as surface normals, strengthening depth coherence through geometric regularization and contextual cues, with practical deployment for autonomous mobility.","Revisiting Matching Response and Swept Feature Volumes for Wide-baseline Omnidirectional Stereo  \nSeungjin Jeon 1∗ Jongwoo Lim 1 ,2 and Changhee Won 1∗†  \narXiv :2607 . 11097v1 [ cs .CV] 13 Jul 2026  \nAbstract—In this paper, we propose a training strategy for confidence estimation in omnidirectional stereo, targeting the ambiguous matches that frequently occur in wide-baseline setups. Reinterpreting the matching responses produced by the 3D encoder–decoder block, we show that their expectation values provide intrinsic confidence signals. Building on this, our method directly penalizes ambiguous responses without auxiliary heads, multi-pass inference, or additional modules, resulting in more efficient and generalized predictions. Beyond confidence, we introduce swept feature volume resampling, where response features produced by 3D CNNs are resampled using regressed positive matching indices and then processed by 2D CNNs to predict meta-information such as surface normals. This joint learning introduces auxiliary geometric regularization and improves depth coherence by leveraging additional contextual cues during response aggregation stage. Experimental results demonstrate that our approach enhances both confidence estimation and surface normal prediction while maintaining deployment practicality for autonomous mobility applications.  \nI. INTRODUCTION  \nOmnidirectional perception has become an important component in various computer vision applications, particularly in robotics, autonomous driving, and aerial navigation, where accurate depth estimation critically affects system performance and safety. Although LiDAR sensors can be employed, recent advances in deep neural networks (DNNs) have greatly improved stereo matching [1], [2], making multi-camera omnidirectional stereo approaches increasingly attractive. Existing omnidirectional stereo methods typically involve configurations with multiple cameras [3]–[8] or stereo setups using 360◦ cameras [9]–[11] . More recently, approaches that employ wide-baseline cameras with global spherical sweeping have been proposed [3]–[7], [11] .  \nIn wide-baseline setups with global sweeping, occlusions without positive (valid) matches often lead to erroneous depth estimation results. For practical deployment on mobile platforms such as robots and drones, it is therefore essential to detect and handle such ambiguous matches using a confidence measure of the matching estimates. Various confidence estimation methods have been proposed for stereo matching. Poggi and Mattoccia [12], for example, feed the output disparity into separate 2D convolutional neural networks (CNNs) to predict confidence. Similarly, contextual features  \nfrom input images [13], cost volumes [14], or iteratively ∗ Authors contributed equally.  \n†Corresponding author.  \n1UVify Corporate Affiliated Research Institute, Republic of Korea.{seungjin.jeon, [changhee.won](changhee.won}@uvify.com)[}](changhee.won}@uvify.com)[@uvify.com](changhee.won}@uvify.com)  \n[2](2 Seoul National University)[ Seoul National University](2 Seoul National University), [Seoul](Seoul), [Republic of Korea](Republic of Korea).{ [jongwoo.lim](jongwoo.lim}@snu.ac.kr)[}](jongwoo.lim}@snu.ac.kr)[@snu.ac.kr](jongwoo.lim}@snu.ac.kr)  \nFig. 1: From top: four wide field-of-view (FoV) fisheye input images and our multi-camera equipped drone, followed by inversedepth prediction, confidence prediction, and surface normal prediction from the proposed method.  \ncomputed disparity profile [15] have been used as inputs to dedicated confidence estimation networks. Auxiliary heads within stereo matching networks have also been employed to predict confidence [16], [17] . However, such approaches typically incur higher inference time and architectural complexity, since they rely on additional network modules that go beyond the core stereo matching framework, limiting their practicality for real-time deployment. Without additional modules, Won et al. [3] use entropy as ","cbCaiukedsqvTwsc","https://ap.wps.com/l/cbCaiukedsqvTwsc","pdf",3423542,5,1,"English","en",105,"# Introduction\n## Confidence estimation for wide-baseline omnidirectional stereo\n## Limitations of entropy-based uncertainty under occlusions\n## Joint meta-information learning via swept feature volume resampling","[{\"question\":\"What problem does the paper address in wide-baseline omnidirectional stereo?\",\"answer\":\"The approach performs swept feature volume resampling by resampling response features from 3D CNNs using regressed positive matching indices, then processing them with 2D CNNs. This enables joint learning of meta-information such as surface normals and adds geometric regularization for improved depth coherence.\"}]",1784212467,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":27},"revisiting-matching-response-and-swept-feature-volumes-for-wide-baseline-omnidirectional-stereo","",{"@graph":35,"@context":77},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/revisiting-matching-response-and-swept-feature-volumes-for-wide-baseline-omnidirectional-stereo/86536/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in wide-baseline omnidirectional stereo?","Question",{"text":75,"@type":76},"The approach performs swept feature volume resampling by resampling response features from 3D CNNs using regressed positive matching indices, then processing them with 2D CNNs. This enables joint learning of meta-information such as surface normals and adds geometric regularization for improved depth coherence.","Answer","https://schema.org",{"og:url":51,"og:type":79,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":81,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,101,106,111,114,118,121,125],{"id":21,"doc_module":4,"doc_module_name":45,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":45,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":45,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":28,"slug":117},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":119,"show_sort_weight":28,"slug":120},"World Cup","world-cup",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":122,"slug":124},10,"Lifestyle","lifestyle",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":20,"slug":128},19,"General","general"]