[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-120556-en":3,"doc-seo-120556-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},120556,687197207919,"Theodora","https://ap-avatar.wpscdn.com/avatar/a000253d6f5f7c60be?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779446848396160552",8,"Research & Report","Generalization Analysis for Contrastive Representation Learning - Research paper","Contrastive learning has achieved strong advances across many machine learning tasks, yet existing generalization analyses are limited and often lose usefulness when the number of negative samples k grows. This work derives new generalization bounds for contrastive learning that are essentially independent of k, aside from logarithmic factors, using empirical covering-number structure and Rademacher complexity. By leveraging Lipschitz continuity of the loss, it further establishes optimistic results for self-bounding Lipschitz losses that enable fast rates in low-noise regimes, and it evaluates linear and deep nonlinear representations.","Generalization Analysis for Contrastive Representation Learning  \nYunwen Lei 1 Tianbao Yang 2 Yiming Ying 3 Ding-Xuan Zhou 4  \nAbstract  \nRecently, contrastive learning has found impressive success in advancing the state of the art in solving various machine learning tasks. However, the existing generalization analysis is very limited or even not meaningful. In particular, the existing generalization error bounds depend linearly on the number k of negative examples while it was widely shown in practice that choosing a large k is necessary to guarantee good generalization of contrastive learning in downstream tasks. In this paper, we establish novel generalization bounds for contrastive learning which do not depend on k, up to logarithmic terms. Our analysis uses structural results on empirical covering numbers and Rademacher complexities to exploit the Lipschitz continuity of loss functions.  \nFor self-bounding Lipschitz loss functions, we further improve our results by developing optimistic bounds which imply fast rates in a low noise condition. We apply our results to learning with both linear representation and nonlinear representation by deep neural networks, for both of which we derive Rademacher complexity bounds to get improved generalization bounds.  \n1. Introduction  \nThe performance of machine learning (ML) models often depends largely on the representation of data, which motivates a resurgence of contrastive representation learning (CRL) to learn a representation function f : X 7→ Rd from unsupervised data (Chen et al., 2020 ; Khosla et al., 2020 ; He et al., 2020) . The basic idea is to pull together similar pairs (x, x+ ) and push apart disimilar pairs (x, x − ) in an  \n1Department of Mathematics, The University of Hong Kong 2Department of Computer Science and Engineering, Texas A&M University 3Department of Mathematics and Statistics, State University of New York at Albany 4 School of Mathematics and Statistics, University of Sydney. Correspondence to: Yiming Ying \u003C[yying@albany.edu](yying@albany.edu) >.  \nProceedings of the 40 th International Conference on Machine Learning, Honolulu, Hawaii, USA. PMLR 202, 2023 . Copyright 2023 by the author(s) .  \nembedding space, which can be formulated as minimizing the following objective (Chen et al., 2020 ; Oord et al., 2018)  \nk  \nEx ,x+ , {x−i}ki=1 log 􀀐 1+ X exp 􀀐−f(x)⊤ 􀀀f(x+ )−f(x−i)􀀁􀀑􀀑 , i=1  \nwhere k is the number of negative examples. The hope is that the learned representation f(x) would capture the latent structure and be beneficial to other downstream learning tasks (Arora et al., 2019 ; Tosh et al., 2021a) . CRL has achieved impressive empirical performance in advancing the state-of-the-art performance in various domains such as computer vision (He et al., 2020 ; Caron et al., 2020 ; Chen et al., 2020 ; Caron et al., 2020) and natural language processing (Brown et al., 2020 ; Gao et al., 2021 ; Radford et al., 2021) .  \nThe empirical success of CRL motivates a natural question on theoretically understanding how the learned representation adapts to the downstream tasks, i.e.,  \nHow would the generalization behavior of downstream ML models benefit from the representation function built from positive and negative pairs? Especially, how would the number of negative examples affect the learning performance?  \nArora et al. (2019) provided an attempt to answer the above questions by developing a theoretical framework to study CRL. They first gave generalization bounds for a learned representation function in terms of Rademacher complexities. Then, they showed that this generalization behavior measured by an unsupervised loss guarantees the generalization behavior of a linear classifier in the downstream classification task. However, the generalization bounds there enjoy a linear dependency on k, which would not be effective if k is large. Moreover, this is not consistent with many studies which show a large number of negative examples (Chen et al., 2020 ; Tian et al.,","cbCait6HyKqiOZsP","https://ap.wps.com/l/cbCait6HyKqiOZsP","pdf",665638,1,28,"English","en",105,"# Introduction\n## Motivation and background\n## Main contributions\n### Loss-function-based bounds","[{\"question\":\"Why do existing generalization bounds for contrastive learning become ineffective when k is large?\",\"answer\":\"They often depend linearly on the number of negative examples k, so large k can make the bounds vacuous or less useful.\"},{\"question\":\"What is the key improvement introduced in this paper’s generalization analysis?\",\"answer\":\"The paper establishes new generalization bounds that do not depend on k, up to logarithmic terms, by exploiting structural covering-number results and Rademacher complexities together with Lipschitz continuity.\"},{\"question\":\"How do the results change for different types of Lipschitz loss functions?\",\"answer\":\"For ℓ2-Lipschitz losses, the bounds improve with a square-root dependence on k; for ℓ∞-Lipschitz losses, the bounds become essentially independent of k up to log factors; for self-bounding Lipschitz losses, optimistic bounds yield fast rates under low-noise conditions.\"}]","Generalization Analysis for Contrastive Representation Learning - Research paper | PDF",1785730637,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"generalization-analysis-for-contrastive-representation-learning-research-paper","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/generalization-analysis-for-contrastive-representation-learning-research-paper/120556/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do existing generalization bounds for contrastive learning become ineffective when k is large?","Question",{"text":76,"@type":77},"They often depend linearly on the number of negative examples k, so large k can make the bounds vacuous or less useful.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the key improvement introduced in this paper’s generalization analysis?",{"text":81,"@type":77},"The paper establishes new generalization bounds that do not depend on k, up to logarithmic terms, by exploiting structural covering-number results and Rademacher complexities together with Lipschitz continuity.",{"name":83,"@type":74,"acceptedAnswer":84},"How do the results change for different types of Lipschitz loss functions?",{"text":85,"@type":77},"For ℓ2-Lipschitz losses, the bounds improve with a square-root dependence on k; for ℓ∞-Lipschitz losses, the bounds become essentially independent of k up to log factors; for self-bounding Lipschitz losses, optimistic bounds yield fast rates under low-noise conditions.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]