[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118722-en":3,"doc-seo-118722-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},118722,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Instance-Dependent Generalization Bounds via Optimal Transport","Existing generalization bounds do not explain key drivers behind how modern neural networks generalize. Because many bounds hold uniformly over all parameters, they suffer from over-parameterization and ignore inductive biases induced by initialization and stochastic gradient descent. This work develops an optimal-transport perspective that yields instance-dependent bounds determined by local Lipschitz regularity of the learned predictor. The resulting guarantees are parametrization-agnostic, apply well in small-sample regimes, accelerate rates on low-dimensional manifolds, and address distribution shifts. Experiments on neural networks show the bound values are informative and reflect regularization effects.","arXiv :2211 .01258v2 [ stat .ML] 7 Nov 2022  \nInstance-Dependent Generalization Bounds via Optimal Transport  \nInstance-Dependent Generalization Bounds via Optimal Transport  \nSongyan Hou 􀀃  \nDepartment of Mathematics, ETH Zurich Parnian Kassraie 􀀃  \nDepartment of Computer Science, ETH Zurich Anastasis Kratsios 􀀃  \nDepartment of Mathematics, McMaster University Jonas Rothfuss 􀀃  \nDepartment of Computer Science, ETH Zurich Andreas Krause  \nDepartment of Computer Science, ETH Zurich  \n[songyan.hou@ethz.ch](songyan.hou@ethz.ch)[pkassraie@ethz.ch](pkassraie@ethz.ch)[ ](pkassraie@ethz.ch)[kratsioa@mcmaster.ca](kratsioa@mcmaster.ca)[ ](kratsioa@mcmaster.ca)[jonas.rothfuss@inf.ethz.ch](jonas.rothfuss@inf.ethz.ch)[ ](jonas.rothfuss@inf.ethz.ch)[krausea@ethz.ch](krausea@ethz.ch)  \nAbstract  \nExisting generalization bounds fail to explain crucial factors that drive generalization of modern neural networks. Since such bounds often hold uniformly over all parameters, they suﬀer from over-parametrization, and fail to account for the strong inductive bias of initialization and stochastic gradient descent. As an alternative, we propose a novel optimal transport interpretation of the generalization problem. This allows us to derive instance-dependent generalization bounds that depend on the local Lipschitz regularity of the learned prediction function in the data space. Therefore, our bounds are agnostic to the parametrization of the model and work well when the number of training samples is much smaller than the number of parameters. With small modiﬁcations, our approach yields accelerated rates for data on low-dimensional manifolds, and guarantees under distribution shifts. We empirically analyze our generalization bounds for neural networks, showing that the bound values are meaningful and capture the eﬀect of popular regularization methods during training.  \nKeywords: Generalization Bound, Instance-Dependent, Optimal Transport, Local Lipschitz Regularity  \n1. Introduction  \nA core challenge in machine learning is to generalize well beyond the training data. We want to choose a hypothesis f 2 F that not only gives small training error but also yields good predictions for previously unseen data points. Accordingly, statistical learning theory aims to provide generalization guarantees and understand the factors that drive it. Generalization is typically described through the discrepancy between two key quantities: The empirical risk ^R(f ), i.e. , the prediction error of f on the training data and the expected risk R (f ) , i.e. , the expected error under the unknown data-distribution. A common type of guarantees are uniform bounds which control the generalization gap R (f ) 􀀀 ^R(f ) with high probability, simultaneously for all hypotheses f 2 F (e.g. , Vapnik and Chervonenkis, 1971 ; Bartlett and Mendelson, 2002) . Such bounds include terms that quantify the complexity of the hypothesis  \n∗ . Equal contribution, alphabetic order.  \nHou, Kassraie, Kratsios, Rothfuss, and Krause  \nf or hypothesis space F. For neural networks (NNs), this complexity term grows rapidly with the number of parameters (e.g. , Bartlett et al. , 2017 ; Neyshabur et al. , 2015 ; Harvey et al. , 2017) . While the parameter space of NNs is vast, regular networks which are used in practice only seem to populate a small subset of the parameter space. This subset seemingly generalizes well, and depends on model structure, initialization scheme and optimization method in a complex manner. In addition, there are many NN parameter conﬁgurations that correspond to the same neural network mapping, artiﬁcially inﬂating the complexity of the parametric hypothesis space. Thus, such uniform bounds in the parameter space fail to explain the empirical generalization behavior of neural networks in the over-parameterized setting where the the number of training examples is much smaller than the number of parameters (Belkin et al. , 2019) .  \nAddressing this issue, we base our analysis ","cbCaikPEbGOAxLih","https://ap.wps.com/l/cbCaikPEbGOAxLih","pdf",1376899,1,50,"English","en",105,"# Introduction\n## Instance-dependent view of generalization\n## Geometric characterization via local Lipschitz constants\n## Optimal transport interpretation of generalization gap\n## Properties of the main generalization bound","[{\"question\":\"Why do uniform generalization bounds fail for over-parameterized neural networks?\",\"answer\":\"Uniform bounds typically control the generalization gap over all hypotheses and grow with the parameter count, which becomes overly large in over-parameterized regimes. They also do not reflect the specific inductive bias from initialization and stochastic gradient descent.\"},{\"question\":\"What is the main idea of the proposed optimal transport interpretation?\",\"answer\":\"The approach treats the generalization gap as the worst-case loss impact when probability mass is transported from the empirical distribution to the true data distribution. This impact depends on the predictor’s local regularity and the transport cost.\"},{\"question\":\"How do the derived bounds differ from global Lipschitz-based bounds?\",\"answer\":\"The bounds use fine-grained local analysis, so they incorporate changes in regularity across the input domain. This makes them tighter than bounds relying only on global Lipschitz properties.\"}]","Instance-Dependent Generalization Bounds via Optimal Transport | PDF",1785719913,126,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"instance-dependent-generalization-bounds-via-optimal-transport","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/instance-dependent-generalization-bounds-via-optimal-transport/118722/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04","2026-08-03",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why do uniform generalization bounds fail for over-parameterized neural networks?","Question",{"text":76,"@type":77},"Uniform bounds typically control the generalization gap over all hypotheses and grow with the parameter count, which becomes overly large in over-parameterized regimes. They also do not reflect the specific inductive bias from initialization and stochastic gradient descent.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the main idea of the proposed optimal transport interpretation?",{"text":81,"@type":77},"The approach treats the generalization gap as the worst-case loss impact when probability mass is transported from the empirical distribution to the true data distribution. This impact depends on the predictor’s local regularity and the transport cost.",{"name":83,"@type":74,"acceptedAnswer":84},"How do the derived bounds differ from global Lipschitz-based bounds?",{"text":85,"@type":77},"The bounds use fine-grained local analysis, so they incorporate changes in regularity across the input domain. This makes them tighter than bounds relying only on global Lipschitz properties.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":21,"slug":114},6,"Technology","technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]