[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"detail-sidebar-cat-0-en-105":3,"doc-seo-203749-105":59,"doc-detail-203749-en":130},{"code":4,"msg":5,"data":6},0,"success",[7,13,18,23,28,33,38,43,48,51,55],{"id":8,"doc_module":4,"doc_module_name":9,"category_name":10,"show_sort_weight":11,"slug":12},1,"Document","Story & Novel",90,"story-novel",{"id":14,"doc_module":4,"doc_module_name":9,"category_name":15,"show_sort_weight":16,"slug":17},2,"Literature",80,"literature",{"id":19,"doc_module":4,"doc_module_name":9,"category_name":20,"show_sort_weight":21,"slug":22},4,"Exam",70,"exam",{"id":24,"doc_module":4,"doc_module_name":9,"category_name":25,"show_sort_weight":26,"slug":27},5,"Comic",60,"comic",{"id":29,"doc_module":4,"doc_module_name":9,"category_name":30,"show_sort_weight":31,"slug":32},6,"Technology",50,"technology",{"id":34,"doc_module":4,"doc_module_name":9,"category_name":35,"show_sort_weight":36,"slug":37},7,"Healthcare",40,"healthcare",{"id":39,"doc_module":4,"doc_module_name":9,"category_name":40,"show_sort_weight":41,"slug":42},8,"Research & Report",30,"research-report",{"id":44,"doc_module":4,"doc_module_name":9,"category_name":45,"show_sort_weight":46,"slug":47},9,"Religion & Spirituality",20,"religion-spirituality",{"id":46,"doc_module":4,"doc_module_name":9,"category_name":49,"show_sort_weight":46,"slug":50},"World Cup","world-cup",{"id":52,"doc_module":4,"doc_module_name":9,"category_name":53,"show_sort_weight":52,"slug":54},10,"Lifestyle","lifestyle",{"id":56,"doc_module":4,"doc_module_name":9,"category_name":57,"show_sort_weight":24,"slug":58},19,"General","general",{"code":4,"msg":60,"data":61},"ok",{"site_id":62,"language":63,"slug":64,"title":65,"keywords":66,"description":67,"schema_data":68,"social_meta":123,"head_meta":125,"extra_data":127,"updated_unix":129},105,"en","when-and-how-does-clip-enable-domain-and-compositional-generalization-paper","When and How Does CLIP Enable Domain and Compositional Generalization - Paper","","CLIP’s strong generalization is commonly linked to diverse training distributions, but the conditions and mechanisms remain unclear. This work studies whether CLIP can generalize to entirely unseen domains and to unseen classes within partially seen domains. Training distributions are systematically constructed with controlled domain diversity and object-class exposure. Results show domain diversity is essential for both tasks, while compositional generalization can be weaker under suboptimal domain mixtures, requiring shared intermediate representations.",{"@graph":69,"@context":122},[70,84,105],{"@type":71,"itemListElement":72},"BreadcrumbList",[73,77,79,82],{"item":74,"name":75,"@type":76,"position":8},"https://docshare.wps.com","Home","ListItem",{"item":78,"name":9,"@type":76,"position":14},"https://docshare.wps.com/document/",{"item":80,"name":40,"@type":76,"position":81},"https://docshare.wps.com/document/research-report/",3,{"item":83,"name":65,"@type":76,"position":19},"https://docshare.wps.com/document/when-and-how-does-clip-enable-domain-and-compositional-generalization-paper/203749/",{"url":83,"name":65,"@type":85,"image":86,"author":91,"headline":65,"publisher":94,"fileFormat":97,"inLanguage":63,"description":67,"dateModified":98,"datePublished":99,"encodingFormat":97,"isAccessibleForFree":100,"interactionStatistic":101},"DigitalDocument",{"url":87,"@type":88,"width":89,"height":90},"https://docshare.wps.com/thumbnails/when-and-how-does-clip-enable-domain-and-compositional-generalization-paper/203749.png","ImageObject",300,407,{"name":92,"@type":93},"Dipper","Person",{"url":74,"name":95,"@type":96},"DocShare","Organization","application/pdf","2026-10-07","2026-09-04",true,{"@type":102,"interactionType":103,"userInteractionCount":44},"InteractionCounter",{"@type":104},"ViewAction",{"@type":106,"mainEntity":107},"FAQPage",[108,114,118],{"name":109,"@type":110,"acceptedAnswer":111},"What question does the paper focus on regarding CLIP’s generalization?","Question",{"text":112,"@type":113},"It asks how mixtures of diverse domains in CLIP’s training data affect generalization to unseen domains and unseen classes within partially seen domains.","Answer",{"name":115,"@type":110,"acceptedAnswer":116},"How are the training conditions controlled in the experiments?",{"text":117,"@type":113},"The authors construct training distributions with controlled domain diversity and controlled object-class exposure, while keeping other variables such as model type, training process, and class distribution constant.",{"name":119,"@type":110,"acceptedAnswer":120},"What do the findings say about domain diversity and compositional generalization?",{"text":121,"@type":113},"Domain diversity is essential for both domain and compositional generalization, but compositional generalization can be weaker than domain generalization when the training distribution includes an unfavorable subset of the test domain.","https://schema.org",{"og:url":83,"og:type":124,"og:title":65,"og:site_name":95,"og:description":67},"article",{"robots":126,"canonical":83},"index,follow",{"doc_id":128,"site_id":62},203749,1788563195,{"code":4,"msg":5,"data":131},{"doc_id":128,"user_id":132,"nickname":92,"user_avatar":133,"doc_module":4,"category_id":39,"category_name":40,"doc_title":65,"doc_description":67,"doc_content":134,"file_id":135,"file_url":136,"file_type":137,"file_size":138,"view_count":44,"is_deleted":4,"is_public":8,"is_downloadable":8,"audit_status":8,"page_count":139,"language":140,"language_code":63,"site_id":62,"html_lang":63,"table_of_contents":141,"faqs":142,"seo_title":143,"seo_description":67,"update_tm":129,"read_time":144},1374404997633,"https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd","When and How Does CLIP Enable Domain and Compositional Generalization?  \nElias Kempf * 1 Simon Schrodi * 1 Max Argus 1 Thomas Brox 1  \nAbstract  \nThe remarkable generalization performance of contrastive vision-language models like CLIP is often attributed to the diversity of their training distributions. However, key questions remain unanswered: Can CLIP generalize to an entirely unseen domain when trained on a diverse mixture of domains (domain generalization)? Can it generalize to unseen classes within partially seen domains (compositional generalization)? What factors affect such generalization? To answer these questions, we trained CLIP models on systematically constructed training distributions with controlled domain diversity and object class exposure. Our experiments show that domain diversity is essential for both domain and compositional generalization, yet compositional generalization can be surprisingly weaker than domain generalization when the training distribution contains a suboptimal subset of the test domain. Through data-centric and mechanistic analyses, we find that successful generalization requires the learning of sufficiently shared representations in intermediate layers and circuits.  \n1. Introduction  \nFoundation models are considered a decisive step towards more generic AI models (Bommasani et al., 2021) . For example, CLIP scaled the alignment of image-text pairs via a contrastive loss to millions of samples (Radford et al., 2021 ; Jia et al., 2021 ; Zhai et al., 2023) . Unlike traditional classifiers from the ImageNet era, which often experience substantial performance drops under distribution shifts, CLIP demonstrates unprecedented generalization to“Out-of-Distribution (OOD)” data (Radford et al., 2021) . However, what drives this improved OOD generalization?  \n*Equal contribution 1University of Freiburg. Correspondence to: Elias Kempf \u003C[kempfe@cs.uni-freiburg.de](kempfe@cs.uni-freiburg.de) >, Simon Schrodi \u003C[schrodi@cs.uni-freiburg.de](schrodi@cs.uni-freiburg.de) >.  \nProceedings of the 42 nd International Conference on Machine Learning, Vancouver, Canada. PMLR 267, 2025 . Copyright 2025 by the author(s) .  \nRecent work has converged on the conclusion that CLIP’s diverse training distribution is the primary factor driving its unprecedented generalization performance. For example, Fang et al. (2022) found that other factors such as language supervision, training data size, or the contrastive loss play only a minor role, while Nguyen et al. (2022) showed that data quality is more important than quantity. More recently, Mayilvahanan et al. (2025) demonstrated that CLIP’s “generalization performance [...] drops to levels similar to what has been observed for ImageNet-trained models”(Mayilvahanan et al., 2025, p. 10) by limiting the diversity of (visual) domains 1 to a minimum, i.e., by removing all nonnatural samples. This shows that the mixture of various (non-natural) domains plays an important role for CLIP’s generalization, yet the underlying mechanisms remain unexplored. This brings us to our core research question:  \nHow does the mixture of diverse (visual) domains in the training data affect CLIP’s generalization performance?  \nIn particular, we investigate under which circumstances CLIP can learn the object class invariances across the training domains with the aim to generalize to entirely unseen domains—a fundamental question about its domain generalization capability (Blanchard et al., 2011 ; Muandet et al., 2013 ; Gulrajani & Lopez-Paz, 2021) . We also study questions about CLIP’s compositional generalization (Hupkeset al., 2020 ; Wiedemer et al., 2023), which is believed to bean important factor of its generalization performance (Mayilvahanan et al., 2024 ; Udandarao et al., 2024) and a longstanding challenge of machine learning research. Adapting Szab’s (2012) classical example, we ask pictorially: Can CLIP, trained on natural images of cats and dogs along with sketches of cats, generaliz","cbCaiqHiT34h6G0e","https://ap.wps.com/l/cbCaiqHiT34h6G0e","pdf",3415516,27,"English","# Abstract\n# 1. Introduction\n## Core research questions\n## Experimental design and training data setups","[{\"question\":\"What question does the paper focus on regarding CLIP’s generalization?\",\"answer\":\"It asks how mixtures of diverse domains in CLIP’s training data affect generalization to unseen domains and unseen classes within partially seen domains.\"},{\"question\":\"How are the training conditions controlled in the experiments?\",\"answer\":\"The authors construct training distributions with controlled domain diversity and controlled object-class exposure, while keeping other variables such as model type, training process, and class distribution constant.\"},{\"question\":\"What do the findings say about domain diversity and compositional generalization?\",\"answer\":\"Domain diversity is essential for both domain and compositional generalization, but compositional generalization can be weaker than domain generalization when the training distribution includes an unfavorable subset of the test domain.\"}]","When and How Does CLIP Enable Domain and Compositional Generalization - Paper | PDF",68]