[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85815-en":3,"doc-seo-85815-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85815,8796095461564,"Liam","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Label-Free Target-Domain Adaptation for Unconstrained Event-Image Feature Matching via Dual-Stage Distillation","Building pixel-level correspondence between event and image data is a core task for multi-sensor systems, yet existing cross-modal matching methods are constrained by label dependence or strictly aligned hardware. A two-stage approach is proposed: label-agnostic distillation pretraining learns generalizable representations using distribution-based and contrastive losses on large-scale data, then epipolar-guided self-distillation enables label-free adaptation on unlabeled, unaligned targets via consistency verification and epipolar geometric confidence. Experiments establish state-of-the-art results on MVSEC and TUM-VIE pose estimation benchmarks.","Label-Free Target-Domain Adaptation for Unconstrained Event-Image Feature Matching via Dual-Stage Distillation  \nZhonghua Yi 1 , Hao Shi4, 1 , Qi Jiang 1 , Yufan Zhang3 , Kailun Yang2 , and Kaiwei Wang 1, ∗  \n1 Zhejiang University, 2Hunan University, 3National University of Defense Technology, 4Ant Group  \narXiv :2607 . 10082v1 [ cs .CV] 11 Jul 2026  \nAbstract  \nBuilding pixel-level correspondence between event and image data is a fundamental task for multi-sensor systems. However, existing cross-modal matching methods are largely restricted by their reliance on either matching labels or strictly aligned hardware, which limits them to unlabeled and unconstrained real-world scenarios where neither matching ground truth nor prior sensor relationships are available. To address this, we propose a novel two-stage training paradigm. First, we leverage large-scale data to perform label-agnostic distillation pretraining, upgrading optimization objectives with distribution-based and contrastive losses to learn highly generalizable representations. Second, to tackle unlabeled and unconstrained downstream data, we introduce an epipolarguided self-distillation framework. By utilizing consistency verification to isolate robust matches and incorporating geometric confidence derived from an external epipolar prior, our model can effectively self-evolve directly on target domains without any supervision. Furthermore, we introduce a rigorous cross-modal evaluation benchmark based on TUM-VIE, featuring physically separated cameras with distinct intrinsic parameters and resolutions. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on both MVSEC and TUM-VIE pose estimation tasks. The source code and benchmark will be made publicly available at [https://github.com/ZhonghuaYi/nexus2-official](https://github.com/ZhonghuaYi/nexus2-official).  \nKeywords  \nEvent Cameras, Distillation, Cross-modal Feature Matching  \n1 Introduction  \nCross-modal feature matching [11, 23, 30] has recently attracted significant research interest, focusing on establishing pixel-level correspondences between multi-modal data. Among these tasks, event-image feature matching [37] is particularly challenging. Due to the unique nature of event cameras, their data is represented as asynchronous point clouds in 3D spatiotemporal space, making it significantly more difficult to match with standard 2D images.  \nUnlabeled and unconstrained multi-sensor systems [42] are ubiquitous in real-world applications, since they bypass the need for expensive tracking equipment, especially in distributed sensor networks where relative sensor poses are inherently dynamic and unfixed. As illustrated in Fig. 1(a), our target scenario focuses on these physically decoupled, heterogeneous sensors with unknown spatial relationships. In such environments, the model must learn to establish correspondences on target data without access to any ground-truth matching labels through pose and depth (unlabeled), or predefined sensor relationships (unconstrained) .  \nHowever, previous methods struggle to generalize to such setupsand generally fall into two categories. The first category, shown in Fig. 1(b), relies on large-scale end-to-end training using synthetic multimodal data with matching labels. These methods require image datasets with known camera poses and pixel-wise depth to synthesize large amounts of multimodal data, thereby generalizing to event-image matching [23] . The second category, such as EI-Nexus [37](Fig. 1(c)), relaxes the need for synthetic labels by distilling knowledge from aligned event-image pairs. However, these alignment-dependent approaches are fundamentally bottlenecked by their reliance on specialized hardware (e.g., coaxial DAVIS cameras) to provide pixel-perfect spatial correspondence. Consequently, they cannot adapt to target downstream datasets where such strict hardware alignment is unavailable.  \nTo bridge this gap, we propose","cbCaijvZ8HjzuRCZ","https://ap.wps.com/l/cbCaijvZ8HjzuRCZ","pdf",24210322,1,14,"English","en",105,"# Abstract\n# Introduction\n## Problem setting: unlabeled and unconstrained sensors\n## Limitations of prior methods\n## Proposed two-stage training paradigm\n### Stage 1: label-agnostic distillation pretraining\n### Stage 2: epipolar-guided self-distillation","[{\"question\":\"Why is event-image feature matching difficult in practice?\",\"answer\":\"Event cameras produce asynchronous point clouds in 3D spatiotemporal space, making pixel-level correspondence with standard 2D images significantly harder than conventional image-to-image matching.\"},{\"question\":\"What is the purpose of the two-stage training paradigm?\",\"answer\":\"The first stage learns robust, label-agnostic cross-modal representations via distillation pretraining, and the second stage adapts the model to unlabeled, unaligned target domains using epipolar-guided self-distillation.\"},{\"question\":\"How does epipolar-guided self-distillation work without ground-truth matching labels?\",\"answer\":\"A teacher-student setup extracts high-consistency matches by comparing teacher and student predictions, while geometric confidence derived from an epipolar prior guides the self-distillation so the network focuses on matches consistent with underlying epipolar geometry.\"}]",1784206415,35,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"label-free-target-domain-adaptation-for-unconstrained-event-image-feature-matching-via-dual-stage-distillation","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/label-free-target-domain-adaptation-for-unconstrained-event-image-feature-matching-via-dual-stage-distillation/85815/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is event-image feature matching difficult in practice?","Question",{"text":75,"@type":76},"Event cameras produce asynchronous point clouds in 3D spatiotemporal space, making pixel-level correspondence with standard 2D images significantly harder than conventional image-to-image matching.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the purpose of the two-stage training paradigm?",{"text":80,"@type":76},"The first stage learns robust, label-agnostic cross-modal representations via distillation pretraining, and the second stage adapts the model to unlabeled, unaligned target domains using epipolar-guided self-distillation.",{"name":82,"@type":73,"acceptedAnswer":83},"How does epipolar-guided self-distillation work without ground-truth matching labels?",{"text":84,"@type":76},"A teacher-student setup extracts high-consistency matches by comparing teacher and student predictions, while geometric confidence derived from an epipolar prior guides the self-distillation so the network focuses on matches consistent with underlying epipolar geometry.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]