[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-86011-en":3,"doc-seo-86011-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},86011,1099514067415,"Rowan","https://ap-avatar.wpscdn.com/avatar/100002539d78ffe74a7?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779092875211072502",8,"Research & Report","On the Modality Gap and the Contrastive Loss in Multi-Modal Representation Learning","研究聚焦于CLIP式双编码器对比学习中的模态间隙：即使在共享表示空间中训练，图像与文本嵌入仍然难以对齐。作者指出，该现象由独立编码器条件下InfoNCE形式的失效所诱发。通过单模态实验与理论分析，证明在低温下InfoNCE会主动制造该间隙。提出xNCE：同时使用跨模态与模内的负样本对比对。xNCE在MS-COCO检索上匹配性能，并在低温持续缩小间隙；同时提升零样本分类，且不牺牲迁移所需的判别几何结构。","arXiv :2607 . 10698v 1 [ cs .LG] 12 Jul 2026  \nOn the modality gap and the contrastive loss in multi-modal representation learning  \nFabian Mager [fmager@dtu. dk](fmager@dtu. dk)  \nDepartment of Applied Mathematics and Computer Science Technical University of Denmark  \nHiba Nassar [hibna@dtu. dk](hibna@dtu. dk)  \nDepartment of Applied Mathematics and Computer Science Technical University of Denmark  \nLars Kai Hansen [lkai@dtu. dk](lkai@dtu. dk)  \nDepartment of Applied Mathematics and Computer Science Technical University of Denmark  \nAbstract  \nWe study the modality gap in CLIP-style dual-encoder contrastive learning, where image and text embeddings remain misaligned despite being trained in a shared space. We argue that the gap is induced by a failure of the InfoNCE formulation with independent encoders. We conduct a uni-modal experiment with two independent encoders and identical initialization conditions and find that InfoNCE actively generates a gap at low temperatures. We provide a theoretical analysis of this phenomenon and show that the modality gap is indeed a modefailure of InfoNCE, but only at low temperatures. We propose a simple modification called xNCE, which uses intermodal as well as intra-modality negative contrastive pairs. xNCE matches retrieval performance on MS-COCO while consistently reducing the gap even at low temperatures. Notably, xNCE improves zero-shot classification over the InfoNCE baseline across all benchmarks, whereas high-temperature InfoNCE and regularized InfoNCE both fail to do so, demonstrating that xNCE reduces the modality gap without sacrificing the discriminative geometry needed for transfer. Code availability: ...  \n1 Introduction  \nContrastive learning has emerged as a central paradigm for self-supervised representation learning. Wu et al. introduced the view-invariant instance discrimination objective that underpins modern contrastive methods. Subsequent advances such as SimCLR (Chen et al.) demonstrated that strong data augmentation, large batch sizes, and a contrastive loss are sufficient to produce high-quality visual representations, while Bachman et al. broadened the theoretical foundation by defining contrastive learning as maximizing mutual information between augmentations. Parallel developments, including CPC (Oord et al.) and MoCo (He et al.), further refined contrastive objectives through predictive coding and momentum-based negative sampling. Building on these foundations, CLIP (Contrastive Language Image Pretraining) extended contrastive learning to the multimodal domain (Radford et al.) .  \nMultimodal learning aims to build representations that jointly model information from heterogeneous data sources (e.g., images and language), enabling systems to exploit complementary cues and reason about shared semantics across modalities (Radford et al.) . They furthermore showed that pairing images with natural-language descriptions in a contrastive learning paradigm yields highly transferable and zero-shot-capable models at scale.  \nCLIP uses one encoder per modality, trained using a contrastive loss function called Information NoiseContrastive Estimation (InfoNCE) . For a batch of images and captions, e.g., the loss aims to predict the correct caption for a given image and vice versa, based on pairwise cosine similarities. The authors show  \nthat such a pretraining objective produces embeddings with superior zero-shot performance compared to purely vision-based self-supervised models (Radford et al.) .  \nDespite this success, CLIP-style dual-encoder models often exhibit a pronounced modality gap: image and text embeddings, while comparable via cosine similarity, tend to occupy distinct regions (or submanifolds) of the joint representation space rather than forming a fully mixed distribution. Previous analyses show that this separation has geometric and optimization roots, including inductive biases in initialization and training dynamics in contrastive learning that maintain a ","cbCailwBvmTXPmTC","https://ap.wps.com/l/cbCailwBvmTXPmTC","pdf",6476863,5,1,16,"English","en",105,"# Introduction\n# Related Work","[{\"question\":\"模态间隙在CLIP式双编码器对比学习中表现为怎样的现象？\",\"answer\":\"即使图像与文本嵌入通过余弦相似度可比较，它们仍倾向于分布在联合表示空间的不同区域（或子流形），而非形成充分混合的分布。\"},{\"question\":\"作者认为模态间隙的诱因是什么？\",\"answer\":\"作者认为，在独立编码器的设定下，InfoNCE公式会失效，从而诱发模态间隙。\"},{\"question\":\"xNCE如何缓解模态间隙，且其在低温下表现如何？\",\"answer\":\"xNCE同时使用跨模态与模内的负对比样本对。实验显示它在低温下仍能持续降低模态间隙，并在MS-COCO检索与零样本分类上优于InfoNCE基线。\"}]",1784207778,40,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"on-the-modality-gap-and-the-contrastive-loss-in-multi-modal-representation-learning","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/on-the-modality-gap-and-the-contrastive-loss-in-multi-modal-representation-learning/86011/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-07-27","2026-07-16",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"模态间隙在CLIP式双编码器对比学习中表现为怎样的现象？","Question",{"text":76,"@type":77},"即使图像与文本嵌入通过余弦相似度可比较，它们仍倾向于分布在联合表示空间的不同区域（或子流形），而非形成充分混合的分布。","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"作者认为模态间隙的诱因是什么？",{"text":81,"@type":77},"作者认为，在独立编码器的设定下，InfoNCE公式会失效，从而诱发模态间隙。",{"name":83,"@type":74,"acceptedAnswer":84},"xNCE如何缓解模态间隙，且其在低温下表现如何？",{"text":85,"@type":77},"xNCE同时使用跨模态与模内的负对比样本对。实验显示它在低温下仍能持续降低模态间隙，并在MS-COCO检索与零样本分类上优于InfoNCE基线。","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,110,115,119,122,127,130,134],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":29,"slug":118},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":20,"slug":137},19,"General","general"]