[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-seo-203849-105":3,"detail-sidebar-cat-0-en-105":73,"doc-detail-203849-en":122},{"code":4,"msg":5,"data":6},0,"ok",{"site_id":7,"language":8,"slug":9,"title":10,"keywords":11,"description":12,"schema_data":13,"social_meta":66,"head_meta":68,"extra_data":70,"updated_unix":72},105,"en","grounding-representation-similarity-with-statistical-testing","Grounding Representation Similarity with Statistical Testing","","The paper addresses the inconsistency of widely used neural representation similarity and dissimilarity metrics such as CCA, CKA, and related measures. It proposes that valid metrics must be sensitive to changes that alter functional behavior while remaining specific against changes that do not. Functional behavior is quantified via probing accuracy and robustness under distribution shift. Benchmarks vary random initialization and principal component deletion, revealing differing metric weaknesses, strong performance of a classical baseline, and challenging settings where all metrics fail.",{"@graph":14,"@context":65},[15,34,56],{"@type":16,"itemListElement":17},"BreadcrumbList",[18,23,27,31],{"item":19,"name":20,"@type":21,"position":22},"https://docshare.wps.com","Home","ListItem",1,{"item":24,"name":25,"@type":21,"position":26},"https://docshare.wps.com/document/","Document",2,{"item":28,"name":29,"@type":21,"position":30},"https://docshare.wps.com/document/research-report/","Research & Report",3,{"item":32,"name":10,"@type":21,"position":33},"https://docshare.wps.com/document/grounding-representation-similarity-with-statistical-testing/203849/",4,{"url":32,"name":10,"@type":35,"image":36,"author":41,"headline":10,"publisher":44,"fileFormat":47,"inLanguage":8,"description":12,"dateModified":48,"datePublished":49,"encodingFormat":47,"isAccessibleForFree":50,"interactionStatistic":51},"DigitalDocument",{"url":37,"@type":38,"width":39,"height":40},"https://docshare.wps.com/thumbnails/grounding-representation-similarity-with-statistical-testing/203849.png","ImageObject",300,407,{"name":42,"@type":43},"Adam","Person",{"url":19,"name":45,"@type":46},"DocShare","Organization","application/pdf","2026-09-27","2026-09-04",true,{"@type":52,"interactionType":53,"userInteractionCount":55},"InteractionCounter",{"@type":54},"ViewAction",5,{"@type":57,"mainEntity":58},"FAQPage",[59],{"name":60,"@type":61,"acceptedAnswer":62},"Which functional behaviors are used to evaluate the metrics?","Question",{"text":63,"@type":64},"The paper uses probing accuracy and out-of-distribution performance to score how well dissimilarity measures track functional differences.","Answer","https://schema.org",{"og:url":32,"og:type":67,"og:title":10,"og:site_name":45,"og:description":12},"article",{"robots":69,"canonical":32},"index,follow",{"doc_id":71,"site_id":7},203849,1788564433,{"code":4,"msg":74,"data":75},"success",[76,80,84,88,92,97,102,106,111,114,118],{"id":22,"doc_module":4,"doc_module_name":25,"category_name":77,"show_sort_weight":78,"slug":79},"Story & Novel",90,"story-novel",{"id":26,"doc_module":4,"doc_module_name":25,"category_name":81,"show_sort_weight":82,"slug":83},"Literature",80,"literature",{"id":33,"doc_module":4,"doc_module_name":25,"category_name":85,"show_sort_weight":86,"slug":87},"Exam",70,"exam",{"id":55,"doc_module":4,"doc_module_name":25,"category_name":89,"show_sort_weight":90,"slug":91},"Comic",60,"comic",{"id":93,"doc_module":4,"doc_module_name":25,"category_name":94,"show_sort_weight":95,"slug":96},6,"Technology",50,"technology",{"id":98,"doc_module":4,"doc_module_name":25,"category_name":99,"show_sort_weight":100,"slug":101},7,"Healthcare",40,"healthcare",{"id":103,"doc_module":4,"doc_module_name":25,"category_name":29,"show_sort_weight":104,"slug":105},8,30,"research-report",{"id":107,"doc_module":4,"doc_module_name":25,"category_name":108,"show_sort_weight":109,"slug":110},9,"Religion & Spirituality",20,"religion-spirituality",{"id":109,"doc_module":4,"doc_module_name":25,"category_name":112,"show_sort_weight":109,"slug":113},"World Cup","world-cup",{"id":115,"doc_module":4,"doc_module_name":25,"category_name":116,"show_sort_weight":115,"slug":117},10,"Lifestyle","lifestyle",{"id":119,"doc_module":4,"doc_module_name":25,"category_name":120,"show_sort_weight":55,"slug":121},19,"General","general",{"code":4,"msg":74,"data":123},{"doc_id":71,"user_id":124,"nickname":42,"user_avatar":125,"doc_module":4,"category_id":103,"category_name":29,"doc_title":10,"doc_description":12,"doc_content":126,"file_id":127,"file_url":128,"file_type":129,"file_size":130,"view_count":55,"is_deleted":4,"is_public":22,"is_downloadable":22,"audit_status":22,"page_count":131,"language":132,"language_code":8,"site_id":7,"html_lang":8,"table_of_contents":133,"faqs":134,"seo_title":135,"seo_description":12,"update_tm":72,"read_time":136},1374404737137,"https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d","arXiv :2108 .01661v2 [ cs .LG] 3 Nov 2021  \nGrounding Representation Similarity with Statistical  \nTesting  \nFrances Ding, Jean-Stanislas Denain, Jacob Steinhardt  \nUniversity of California Berkeley  \n{frances, js_denain, [jsteinhardt}@berkeley.edu](jsteinhardt}@berkeley.edu)  \nAbstract  \nTo understand neural network behavior, recent works quantitatively compare different networks' learned representations using canonical correlation analysis (CCA), centered kernel alignment (CKA), and other dissimilarity measures. Unfortunately, these widely used measures often disagree on fundamental observations, such as whether deep networks differing only in random initialization learn similar representations. These disagreements raise the question: which, if any, of these dissimilarity measures should we believe? We provide a framework to ground this question through a concrete test: measures should have sensitivity to changes that affect functional behavior, and speciﬁcity against changes that do not. We quantify this through a variety of functional behaviors including probing accuracy and robustness to distribution shift, and examine changes such as varying random initialization and deleting principal components. We ﬁnd that current metrics exhibit different weaknesses, note that a classical baseline performs surprisingly well, and highlight settings where all metrics appear to fail, thus providing a challenge set for further improvement.  \n1 Introduction  \nUnderstanding neural networks is not only scientiﬁcally interesting, but critical for applying deep networks in high-stakes situations. Recent work has highlighted the value of analyzing not just the ﬁnal outputs of a network, but also its intermediate representations [20, 29] . This has motivated the development of representation similarity measures, which can provide insight into how different training schemes, architectures, and datasets affect networks' learned representations.  \nA number of similarity measures have been proposed, including centered kernel alignment (CKA)  \n[13], ones based on canonical correlation analysis (CCA) [24, 30], single neuron alignment [20], vector space alignment [3, 6, 32], and others [2, 9, 16, 18, 21, 39] . Unfortunately, these different measures tell different stories. For instance, CKA and projection weighted CCA disagree on which layers of different networks are most similar [13] . This lack of consensus is worrying, as measures are often designed according to different and incompatible intuitive desiderata, such as whether ﬁnding a one-to-one assignment, or ﬁnding few-to-one mappings, between neurons is more appropriate [20] . As a community, we need well-chosen formal criteria for evaluating metrics to avoid over-reliance on intuition and the pitfalls of too many researcher degrees of freedom [17] .  \nIn this paper we view representation dissimilarity measures as implicitly answering a classiﬁcation question–whether two representations are essentially similar or importantly different. Thus, in analogy to statistical testing, we can evaluate them based on their sensitivity to important change and speciﬁcity (non-responsiveness) against unimportant changes or noise.  \nAs a warm-up, we ﬁrst initially consider two intuitive criteria: ﬁrst, that metrics should have speciﬁcity against random initialization; and second, that they should be sensitive to deleting important principal  \n35th Conference on Neural Information Processing Systems (NeurIPS 2021) .  \ncomponents (those that affect probing accuracy) . Unfortunately, popular metrics fail at least one of these two tests. CCA is not speciﬁc – random initialization noise overwhelms differences between even far-apart layers in a network (Section 3.1) . CKA on the other hand is not sensitive, failing to detect changes in all but the top 10 principal components of a representation (Section 3.2) .  \nWe next construct quantitative benchmarks to evaluate a dissimilarity measure's quality. To move beyond o","cbCaif1bQIJcriJL","https://ap.wps.com/l/cbCaif1bQIJcriJL","pdf",2049806,25,"English","# Introduction\n## Problem Setup: Metrics and Models\n## Benchmarks and Evaluation\n## Results and Findings","[{\"question\":\"Which functional behaviors are used to evaluate the metrics?\",\"answer\":\"The paper uses probing accuracy and out-of-distribution performance to score how well dissimilarity measures track functional differences.\"}]","Grounding Representation Similarity with Statistical Testing | PDF",63]