[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-84495-en":3,"doc-seo-84495-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},84495,549758146520,"Patrick","https://ap-avatar.wpscdn.com/avatar/80002397d8c0411e94?_k=1775819394049821470",8,"Research & Report","SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval","SOLAR targets symmetric multimodal-to-multimodal (MM2MM) retrieval, where queries and contexts are interchangeable, addressing a gap left by universal multimodal methods that depend on labeled asymmetric datasets. SOLAR is a two-stage self-supervised framework using unlabeled web-scale image–text pairs. Stage one learns an intersection mask to align shared semantics while preserving modality discrepancies. Stage two uses the mask to build positive and hard-negative samples via masking, enabling multimodal embedding learning. A human-verified benchmark and pipeline evaluate realistic conditions.","SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval  \nWenjie Yang 1 Hang Yu 1 † Yuyu Guo 1 Peng Di 1  \narXiv :2605 . 15868v2 [ cs .CV] 13 Jul 2026  \nAbstract  \nIn this work, we address the critical yet underexplored challenge of symmetric multimodal-tomultimodal (MM2MM) retrieval, where queries and contexts are interchangeable. Existing universal multimodal retrieval works struggle with this task, as they are constrained by the labeled asymmetric datasets used. We produce SOLAR (Self-supervised jOint LeArning for symmetric multimodal Retrieval), a novel two-stage selfsupervised framework that leverages readily available unlabeled web-scale image-text pairs. Based on the observation that both semantic alignment and discrepancies exist between two modalities, in the first stage, we learn the intersection mask of image-text pair, allowing us to align intersection while preserving semantic of difference. In the second stage, the learned mask is further utilized to construct positive and hardnegative samples via masking different parts of image/text, which enable us to conduct self-supervised multimodal embedding learning. Complementing this framework, we present a new benchmark featuring highquality human-verified positive and hard-negative pairs to evaluate symmetric MM2MM retrieval under realistic conditions, as well as the corresponding pipeline. Extensive experiments against ten SOTA methods show SOLAR surpasses the strongest supervised VLM by 7.08 points on this benchmark, with over 50x fewer model parameters and a 5x smaller embedding dimension. Code, model and benchmark are available at [https:](https:)//[github.com/codefuse-ai/SOLAR](github.com/codefuse-ai/SOLAR).  \n1. Introduction  \nIn information retrieval, the ability to combine different modalities for search is both essential and beneficial, as  \n1Ant Group. Correspondence to: Hang Yu \u003C[hyu1@e.ntu.edu.sg](hyu1@e.ntu.edu.sg)>.  \nProceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026 . Copyright 2026 by the author(s) .  \nAsymmetric Symmetric  \n| \u003Cbr>CIRR/Fashioni (a) | Barack Obama with Germany’s chancellor\u003Cbr>Angela Merkel at the Brandenburg Gate Berlin on 19 June. |  |\n| --- | --- | --- |\n|  | EDIS (b) |  |\n|  |  |  |\n\nUM2MM /  \nMM2UM  \nMM2MM  \nWho manufactured the plane?  \nOven/Infoseek (c)  \nT-shirt for boy, with letter ‘S’ printed on the back.  \nOurs (d)  \nFigure 1. A comparison of existing multimodal retrieval paradigms with the symmetric MM2MM task addressed in this paper. Tasks are categorized based on two properties: whether the retrieval is symmetric (query and content are interchangeable) and whether both are multimodal (MM2MM vs. UM2MM/MM2UM) .  \nthe cross-modal fusion provides more complete and nuanced representations. Yet, multimodal retrieval has received little attention compared to single-modal tasks. One particularly challenging and overlooked task is symmetric multimodality-to-multimodality (MM2MM) retrieval. As illustrated in Fig. 1, current universal multimodal retrieval works rely on asymmetric training dataset, where the query and content have distinct roles. In contrast, symmetric retrieval, where the query and content are semantically equivalent and interchangeable, is critical for many real-world applications. Consider an e-commerce scenario (Fig. 1d): a user searches with an image of a T-shirt’s front and a description of its back. The desired result is an image of the back paired with a description of the front. To succeed, a model must grasp the holistic compositional meaning, recognizing that these two different multimodal pairs represent a single, coherent product. This requires inferring latent attributes not explicitly present in each modality-such as the T-shirt’s color (white) or intended demographic (a boy) -to understand the full context. Beyond e-commerce, symmetric MM2MM retrieval has potential applications in areas such as news article retrieval, recipe recomm","cbCaisuXTH2OcMlT","https://ap.wps.com/l/cbCaisuXTH2OcMlT","pdf",9603607,1,32,"English","en",105,"# Introduction\n## Symmetric MM2MM Retrieval Motivation\n## Limitations of Supervised and Labeled Datasets\n## SOLAR Self-supervised Approach","[{\"question\":\"What makes symmetric MM2MM retrieval different from existing multimodal retrieval tasks?\",\"answer\":\"In symmetric MM2MM retrieval, the query and the context are semantically equivalent and interchangeable. Many universal multimodal retrieval methods instead rely on asymmetric datasets where query and content play different roles.\"},{\"question\":\"How does SOLAR avoid the need for costly labeled positive and hard-negative pairs?\",\"answer\":\"SOLAR uses unlabeled web-scale image–text pairs in a self-supervised two-stage framework, bypassing manual annotation by exploiting the inherent structure of the symmetric task.\"},{\"question\":\"What is the role of the intersection mask in SOLAR’s two-stage training?\",\"answer\":\"The first stage learns an intersection mask for image–text pairs, aligning shared semantics while preserving differences between modalities. The learned mask then guides construction of positives and hard-negatives in the second stage via masking.\"}]",1784196047,81,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"solar-self-supervised-joint-learning-for-symmetric-multimodal-retrieval","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/solar-self-supervised-joint-learning-for-symmetric-multimodal-retrieval/84495/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What makes symmetric MM2MM retrieval different from existing multimodal retrieval tasks?","Question",{"text":75,"@type":76},"In symmetric MM2MM retrieval, the query and the context are semantically equivalent and interchangeable. Many universal multimodal retrieval methods instead rely on asymmetric datasets where query and content play different roles.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does SOLAR avoid the need for costly labeled positive and hard-negative pairs?",{"text":80,"@type":76},"SOLAR uses unlabeled web-scale image–text pairs in a self-supervised two-stage framework, bypassing manual annotation by exploiting the inherent structure of the symmetric task.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the role of the intersection mask in SOLAR’s two-stage training?",{"text":84,"@type":76},"The first stage learns an intersection mask for image–text pairs, aligning shared semantics while preserving differences between modalities. The learned mask then guides construction of positives and hard-negatives in the second stage via masking.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":45,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":45,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":45,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":45,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":45,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":45,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]