[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-81598-en":3,"doc-seo-81598-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},81598,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","Robust Self-Supervised Cross-Modal Super-Resolution against Real-World Misaligned Observations","Cross-modal super-resolution (SR) on real-world misaligned data is difficult because only unlabeled low-resolution (LR) sources and high-resolution (HR) guide images are available, often with complex spatial misalignment across modalities and resolutions. Prior approaches depend on simulated training or use weaker alignment that ignores cross-modal dependencies, limiting practical results. RobSelf is a self-supervised framework that jointly optimizes a misalignment-aware feature translator and a content-aware reference filter online to produce high-resolution, high-fidelity predictions, achieving state-of-the-art accuracy and up to 15.3× faster efficiency.","Robust Self-Supervised Cross-Modal Super-Resolution against Real-World Misaligned  \nObservations  \n18822v3 [ cs .CV] 10 Jul 2026  \nXiaoyu Dong 1 , Jiahuan Li 1 ,2 , Ziteng Cui 1 , and Naoto Yokoya 1 ,2 􀀀  \n1 The University of Tokyo, Japan  \n2 RIKEN AIP, Japan  \nAbstract. Cross-modal super-resolution (SR) on real-world misaligned data is challenging, as only unlabeled low-resolution (LR) source and high-resolution (HR) guide images with complex spatial misalignment are available. Previous methods either rely on simulated training data or adopt suboptimal alignment strategies that overlook cross-modal dependencies, limiting their practical performance. To address these issues, we propose RobSelf, a self-supervised model that jointly optimizesa misalignment-aware feature translator and a content-aware reference filter online. The translator resolves unsupervised cross-modal and crossresolution alignment via weakly-supervised, misalignment-aware translation, yielding an aligned guide feature. Guided by this feature, the filter performs reference-based discriminative self-enhancement on the source, enabling SR prediction with high resolution and high fidelity. Experiments on synthesized data and collected real-world data demonstrate that RobSelf achieves state-of-the-art performance, outperforming existing self-supervised and supervised methods. Moreover, it achieves superior efficiency, being up to 15.3 × faster than prior self-supervised methods. [https://github.com/palmdong/RobSelf](https://github.com/palmdong/RobSelf)  \nKeywords: Self-supervised super-resolution · Multi-modal vision  \nFig. 1: Real-world misaligned RGB-guided depth SR (×4) . Our model achieves stateof-the-art performance, requiring no training data, ground-truth supervision, or prealignment. (a) LR source; (b) HR guide; (c) pre-aligned guide by MINIMA [36]; (d) SSGNet [40] + pre-alignment; (e) SGNet [49] + pre-alignment; (f) RobSelf-Re (Ours) .  \narXiv :2602 .  \n2 X. Dong et al.  \n1 Introduction  \nMulti-modal images, e.g ., RGB, depth, and near-infrared (NIR), capture complementary properties of objects and environments [11,57,66,67] . However, nonvisible modalities generally suffer from lower spatial resolution than RGB due to sensing and hardware limits, which hinders downstream tasks and necessitates cross-modal super-resolution (SR) [56, 71, 72] .  \nCross-modal SR enhances a low-resolution (LR) source image using structural cues from a high-resolution (HR) guide image of another modality. Modern methods fall into supervised and self-supervised categories. Supervised methods rely on large-scale domain-specific training data and ground truth [19, 47, 49, 50, 71, 72] . This reliance hinders their practical generalization, as building such datasets is costly and labor-intensive. Self-supervised methods require neither training data nor ground truth, but instead optimize online on each test pair, making them more data-efficient and generalizable [12, 30, 38, 40] .  \nWhile promising progress has been achieved, most existing methods assume that the source and guide images are well-aligned [12,19,30,40,47,49,50,71,72] . In real-world scenarios, however, spatial misalignment is inevitable in multimodal images due to inherent cross-sensor discrepancies [28,51](e.g ., differences in lens distortion, field of view, and physical position) and environmental factors, such as platform-induced viewpoint variation [5,21] and object motion overtime [27, 46] . Although several methods consider misalignment, they either rely on simulated training data [48] or adopt suboptimal alignment strategies [15,38], limiting their practical performance.  \nWe identify two difficulties in cross-modal SR on real-world misaligned data:(I) The scarcity of training data and the absence of SR ground truth for the source [12, 30]; (II) Complex misalignments across modalities and resolutions, along with the lack of alignment ground truth for the guide [7,61] . These hinder the development of reliab","cbCaiuLTrCVAdW6P","https://ap.wps.com/l/cbCaiuLTrCVAdW6P","pdf",22052324,2,1,18,"English","en",105,"# Introduction\n## Problem background: cross-modal SR and misalignment\n## Proposed approach: RobSelf\n## Contributions and evaluation setup","[{\"question\":\"What makes cross-modal super-resolution on real-world misaligned data challenging?\",\"answer\":\"Real-world settings provide only unlabeled LR source images and HR guide images, while spatial misalignment across modalities and resolutions is complex and often lacks alignment ground truth. This reduces the effectiveness of both supervised and self-supervised methods.\"},{\"question\":\"How does RobSelf address cross-modal and cross-resolution misalignment without ground-truth supervision?\",\"answer\":\"RobSelf uses a misalignment-aware feature translator that performs weakly-supervised, misalignment-aware translation to generate an aligned guide feature, and a content-aware reference filter that performs reference-based discriminative self-enhancement on the source guided by that feature.\"},{\"question\":\"What evidence shows that RobSelf works well on complex real-world misaligned observations?\",\"answer\":\"Experiments on both synthesized data and collected real-world RGB-depth and RGB-NIR data demonstrate state-of-the-art performance. The method also shows superior efficiency, up to 15.3× faster than prior self-supervised methods.\"}]",1784174620,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"robust-self-supervised-cross-modal-super-resolution-against-real-world-misaligned-observations","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/robust-self-supervised-cross-modal-super-resolution-against-real-world-misaligned-observations/81598/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-25","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What makes cross-modal super-resolution on real-world misaligned data challenging?","Question",{"text":75,"@type":76},"Real-world settings provide only unlabeled LR source images and HR guide images, while spatial misalignment across modalities and resolutions is complex and often lacks alignment ground truth. This reduces the effectiveness of both supervised and self-supervised methods.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does RobSelf address cross-modal and cross-resolution misalignment without ground-truth supervision?",{"text":80,"@type":76},"RobSelf uses a misalignment-aware feature translator that performs weakly-supervised, misalignment-aware translation to generate an aligned guide feature, and a content-aware reference filter that performs reference-based discriminative self-enhancement on the source guided by that feature.",{"name":82,"@type":73,"acceptedAnswer":83},"What evidence shows that RobSelf works well on complex real-world misaligned observations?",{"text":84,"@type":76},"Experiments on both synthesized data and collected real-world RGB-depth and RGB-NIR data demonstrate state-of-the-art performance. The method also shows superior efficiency, up to 15.3× faster than prior self-supervised methods.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]