[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125934-en":3,"doc-seo-125934-105":31,"detail-sidebar-cat-0-en-105":96},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},125934,2336474459895,"Aria","https://ap-avatar.wpscdn.com/avatar/22000baeef7a5ed0655?x-image-process=image/resize,m_fixed,w_180,h_180&k=1786071322749376916",8,"Research & Report","Re-examining Metrics for Success in Machine Learning - from Fairness and Interpretability to Protein Design","Quantitative metrics and the datasets used to evaluate them drive rapid advances in machine learning by setting priorities and enabling efficient discovery of models. Yet metrics must reflect real-world goals so that measurable gains transfer to practical tasks, an external-validity challenge that evolves as new issues surface. This dissertation develops tests for representation similarity metrics, new data for fair classification, and protein-language-model metrics capturing evolutionary-taxa bias, including a mitigation method.","UC Berkeley  \nUC Berkeley Electronic Theses and Dissertations  \nTitle  \nRe-examining Metrics for Success in Machine Learning, from Fairness and Interpretability to Protein Design  \nPermalink  \n[https://escholarship.org/uc/item/4637w8xs](https://escholarship.org/uc/item/4637w8xs)  \nAuthor  \nDing, Frances  \nPublication Date  \n2024  \nPeer reviewed|Thesis/dissertation  \n[eScholarship.org](eScholarship.org) Powered by the California Digital Library  \nUniversity of California  \nRe-examining Metrics for Success in Machine Learning, from Fairness and Interpretability to  \nProtein Design  \nBy  \nFrances Ding  \nA dissertation submitted in partial satisfaction of the requirements for the degree of  \nDoctor of Philosophy  \nin  \nComputer Science  \nin the  \nGraduate Division  \nof the  \nUniversity of California, Berkeley  \nCommittee in charge:  \nAssistant Professor Jacob Steinhardt, Co-chair Associate Professor Moritz Hardt, Co-chair  \nProfessor Yun Song  \nProfessor Jeremy Reiter  \nSummer 2024  \nRe-examining Metrics for Success in Machine Learning, from Fairness and Interpretability to  \nProtein Design  \nCopyright 2024  \nBy  \nFrances Ding  \n1  \nAbstract  \nRe-examining Metrics for Success in Machine Learning, from Fairness and Interpretability to  \nProtein Design  \nBy  \nFrances Ding  \nDoctor of Philosophy in Computer Science  \nUniversity of California, Berkeley  \nAssistant Professor Jacob Steinhardt, Co-chair Associate Professor Moritz Hardt, Co-chair  \nQuantitative metrics, along with datasets to assess them with, are key ingredients that have fueled rapid progress in machine learning (ML) in recent years. These metrics, datasets, and benchmarks deﬁne priorities and facilitate eﬃcient discovery of model designs that make progress on those priorities. Ideally, metrics track real world goals, such that improvement on them translates to improvement in related, real tasks. Creating metrics that achieve this external validity is an ever-present challenge in ML. Thus, the science of metrics is an iterative one, as identifying and resolving one issue allows other, more subtle ones, to become apparent.  \nIn this thesis, we describe a series of works that highlight limitations in metrics across diﬀerent subﬁelds of ML and design new metrics to ﬁll these gaps. We ﬁrst examine representation similarity metrics used in the interpretability subﬁeld to compare neural network representations. We show that current popular metrics often disagree on fundamental observations, making it unclear which one we should believe. We develop practical, statistically grounded tests to evaluate these metrics and ﬁnd diﬀerent weaknesses in each. We next examine metrics and benchmarks for fair classiﬁcation. We highlight idiosyncrasies in the popular UCI Adult dataset that limit its external validity, and we contribute a suite of new datasets derived from US Census surveys that extend the existing data ecosystem for research on fair machine learning. Finally, we examine the subﬁeld of protein modeling with ML. We develop metrics to quantify a novel type of bias present in popular protein language models–bias towards sequences from certain evolutionary taxa. We additionally introduce a method to mitigate this bias. Across these works in diverse subﬁelds, we demonstrate the challenges and opportunities present in developing metrics that advance technical capabilities in alignment with real world needs.  \ni  \nTo my parents, Christine and George.  \nii  \nContents  \nContents ii  \nList of Figures iv  \nList of Tables vi  \n1 Introduction 1  \n2 Evaluating Representation Similarity Metrics 4  \n2.1 Introduction .................................... 4  \n2.2 Problem Setup: Metrics and Models ....................... 5  \n2.3 Warm-up: Intuitive Tests for Sensitivity and Speciﬁcity ............ 7  \n2.4 Rigorously Evaluating Dissimilarity Metrics .................. 10  \n2.5 Discussion ..................................... 16  \n2.6 Supplementary Materials ............................. 18  \n3 As","cbCaibBCTc7vorrx","https://ap.wps.com/l/cbCaibBCTc7vorrx","pdf",8520207,4,1,144,"English","en",105,"# Introduction\n# Evaluating Representation Similarity Metrics\n## Problem Setup: Metrics and Models\n## Warm-up: Intuitive Tests for Sensitivity and Speciﬁcity\n## Rigorously Evaluating Dissimilarity Metrics\n# Assessing Fair Machine Learning with New Datasets\n## Archaeology of UCI Adult: Origin, Impact, Limitations\n## New datasets for algorithmic fairness\n# Identifying Biases in Protein Language Models\n## Bias mitigation","[{\"question\":\"Why are metrics and benchmarks central to progress in machine learning?\",\"answer\":\"They define priorities and enable efficient discovery of model designs that advance those priorities. The dissertation emphasizes that the ultimate goal is external validity: metric improvements should translate to improvements on real tasks.\"},{\"question\":\"What limitations are found in representation similarity metrics for interpretability?\",\"answer\":\"The work shows popular metrics can disagree on fundamental observations, making it unclear which metric to trust. It then introduces practical, statistically grounded tests to evaluate these metrics and uncover distinct weaknesses.\"},{\"question\":\"How does the dissertation address fairness in machine learning evaluation?\",\"answer\":\"It highlights limitations in the UCI Adult dataset that restrict external validity, and it contributes new datasets derived from US Census surveys to extend research on fair machine learning.\"},{\"question\":\"What bias is identified in protein language models, and how is it mitigated?\",\"answer\":\"The dissertation develops metrics for a bias toward sequences from certain evolutionary taxa in popular protein language models. It also introduces a method to mitigate this bias and evaluate its impact on protein design.\"}]","Re-examining Metrics for Success in Machine Learning - from Fairness and Interpretability to Protein Design | PDF",1785902116,363,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":91,"head_meta":93,"extra_data":95,"updated_unix":29},"re-examining-metrics-for-success-in-machine-learning-from-fairness-and-interpretability-to-protein-design","",{"@graph":37,"@context":90},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/re-examining-metrics-for-success-in-machine-learning-from-fairness-and-interpretability-to-protein-design/125934/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-16","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82,86],{"name":73,"@type":74,"acceptedAnswer":75},"Why are metrics and benchmarks central to progress in machine learning?","Question",{"text":76,"@type":77},"They define priorities and enable efficient discovery of model designs that advance those priorities. The dissertation emphasizes that the ultimate goal is external validity: metric improvements should translate to improvements on real tasks.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What limitations are found in representation similarity metrics for interpretability?",{"text":81,"@type":77},"The work shows popular metrics can disagree on fundamental observations, making it unclear which metric to trust. It then introduces practical, statistically grounded tests to evaluate these metrics and uncover distinct weaknesses.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the dissertation address fairness in machine learning evaluation?",{"text":85,"@type":77},"It highlights limitations in the UCI Adult dataset that restrict external validity, and it contributes new datasets derived from US Census surveys to extend research on fair machine learning.",{"name":87,"@type":74,"acceptedAnswer":88},"What bias is identified in protein language models, and how is it mitigated?",{"text":89,"@type":77},"The dissertation develops metrics for a bias toward sequences from certain evolutionary taxa in popular protein language models. It also introduces a method to mitigate this bias and evaluate its impact on protein design.","https://schema.org",{"og:url":53,"og:type":92,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":94,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":97},[98,102,106,110,115,120,125,128,133,136,140],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":107,"show_sort_weight":108,"slug":109},"Exam",70,"exam",{"id":111,"doc_module":4,"doc_module_name":47,"category_name":112,"show_sort_weight":113,"slug":114},5,"Comic",60,"comic",{"id":116,"doc_module":4,"doc_module_name":47,"category_name":117,"show_sort_weight":118,"slug":119},6,"Technology",50,"technology",{"id":121,"doc_module":4,"doc_module_name":47,"category_name":122,"show_sort_weight":123,"slug":124},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":126,"slug":127},30,"research-report",{"id":129,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":131,"slug":132},9,"Religion & Spirituality",20,"religion-spirituality",{"id":131,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":131,"slug":135},"World Cup","world-cup",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":137,"slug":139},10,"Lifestyle","lifestyle",{"id":141,"doc_module":4,"doc_module_name":47,"category_name":142,"show_sort_weight":111,"slug":143},19,"General","general"]