[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117168-en":3,"doc-seo-117168-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117168,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",8,"Research & Report","Correctly Evaluating Machine Learning Algorithms - A Dissertation - Doctor of Philosophy","Correct evaluation of machine learning algorithms determines whether published methods reflect true performance and prevents misleading follow-on research. This dissertation examines two actively studied yet inconsistently scrutinized areas: metric learning and unsupervised domain adaptation. It identifies flawed evaluation methodologies that can persist unchallenged, producing distorted experimental results and slowing progress. The work then proposes improved evaluation approaches, including fair comparisons, reproducible protocols, informative accuracy metrics, and hyperparameter search via cross-validation, validated through controlled experiments and benchmark analysis.","CORRECTLY EVALUATING MACHINE LEARNING ALGORITHMS  \nA Dissertation  \nPresented to the Faculty of the Graduate School of Cornell University  \nin Partial Fulfillment of the Requirements for the Degree of Doctor of Philosophy  \nby  \nKevin Musgrave  \nMay 2023  \n© 2023 Kevin Musgrave  \nALL RIGHTS RESERVED  \nCORRECTLY EVALUATING  \nMACHINE LEARNING ALGORITHMS  \nKevin Musgrave, Ph.D.  \nCornell University 2023  \nMachine-learning research is progressing rapidly thanks to advances in model architectures, optimization, training algorithms, and scaling. There are many ingredients that allow this rapid progress. One is the ability to quickly build on prior papers, which depends on the clarity of the papers, and the availability of accompanying research code. Another essential ingredient is the correct evaluation of algorithms, which ensures that the papers accurately portray their algorithm’s performance. This is important because researchers think of new ideas under the assumption that prior works contain accurate information. If papers do not evaluate algorithms correctly, then the inaccurate results could lead future researchers astray.  \nSome machine learning sub-fields are studied very actively. With so many critical eyes, it is unlikely for incorrect evaluation methodologies to propagate through the literature. However, some other sub-fields attract relatively little scrutiny. As a result, various flawed evaluation methodologies can go unchallenged for years. This leads to distortions in experiment results, which misleads researchers and hinders progress. In this dissertation, we investigate two machine-learning sub-fields: metric learning and unsupervised domain adaptation. We reveal the flaws in the evaluation methodologies of both fields, and present improvements.  \nBIOGRAPHICAL SKETCH  \nKevin Musgrave is a PhD candidate working with Professor Serge Belongie. Kevin’s research focuses on the correct evaluation of machine learning algorithms. Prior to attending Cornell, he received his BEng in Electrical Engineering at McGill University.  \nACKNOWLEDGEMENTS  \nThank you Serge Belongie for your insightful discussions and brainstorming, for guiding me through my PhD, and for giving me the freedom to work on subjects that interest me. Thank you Ser-Nam Lim for helping me refine my ideas, for running my scripts, and for funding a large part of my degree through Meta (Facebook AI) . I also thank everyone in the lab for your feedback and suggestions during our group meetings.  \nTABLE OF CONTENTS  \nBiographical Sketch .............................. iii  \nAcknowledgements .............................. iv  \nTable of Contents . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . v  \nList of Tables .................................. vii  \nList of Figures ................................. xii  \n1 Introduction 1  \n2 Metric Learning 3  \n2.1 Metric Learning Overview ....................... 3  \n2.1.1 What is metric learning? .................... 3  \n2.1.2 How does metric learning work? ............... 4  \n2.1.3 How much has metric learning improved over time? ... 11  \n2.1.4 Related work .......................... 12  \n2.2 Flaws in the Literature from 2016 to 2020 .............. 13  \n2.2.1 Unfair comparisons ...................... 13  \n2.2.2 Weakness of commonly used accuracy metrics ....... 17  \n2.2.3 Training with test set feedback ................ 18  \n2.3 Proposed Evaluation Method ..................... 19  \n2.3.1 Fair comparisons and reproducibility ............ 19  \n2.3.2 Informative accuracy metrics ................. 20  \n2.3.3 Hyperparameter search via cross validation ........ 21  \n2.4 Experiments ............................... 24  \n2.4.1 Losses and datasets ....................... 24  \n2.4.2 Papers versus reality ...................... 24  \n2.5 Conclusion ................................ 29  \n3 Unsupervised Domain Adaptation 30  \n3.1 Unsupervised Domain Adaptation Overview ............ 30  \n3.1.1 What is unsupervised domain adaptatio","cbCaia5nhcuLkS87","https://ap.wps.com/l/cbCaia5nhcuLkS87","pdf",5667361,1,126,"English","en",105,"# Introduction\n# Metric Learning\n## Metric Learning Overview\n## Flaws in the Literature from 2016 to 2020\n## Proposed Evaluation Method\n## Experiments\n## Conclusion\n# Unsupervised Domain Adaptation\n## Unsupervised Domain Adaptation Overview\n## UDA Validators Overview\n## Experiment Methodology\n## Results\n## Conclusion\n# Conclusion\n# Appendix\n## Open Source Code","[{\"question\":\"Why is correctly evaluating machine learning algorithms important for the research community?\",\"answer\":\"It ensures papers portray their algorithms’ performance accurately. Correct evaluation prevents incorrect results from steering future researchers toward flawed assumptions.\"},{\"question\":\"Which two machine-learning sub-fields are investigated in this dissertation?\",\"answer\":\"The dissertation investigates metric learning and unsupervised domain adaptation. It analyzes evaluation flaws in both areas and proposes improvements.\"},{\"question\":\"What kinds of improvements does the dissertation propose for evaluation?\",\"answer\":\"It emphasizes fair comparisons and reproducibility, informative accuracy metrics, and hyperparameter search via cross-validation. These changes aim to reduce distorted experiment results and better reflect real performance.\"}]","Correctly Evaluating Machine Learning Algorithms - A Dissertation - Doctor of Philosophy | PDF",1785674195,318,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"correctly-evaluating-machine-learning-algorithms-a-dissertation-doctor-of-philosophy","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/correctly-evaluating-machine-learning-algorithms-a-dissertation-doctor-of-philosophy/117168/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is correctly evaluating machine learning algorithms important for the research community?","Question",{"text":76,"@type":77},"It ensures papers portray their algorithms’ performance accurately. Correct evaluation prevents incorrect results from steering future researchers toward flawed assumptions.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which two machine-learning sub-fields are investigated in this dissertation?",{"text":81,"@type":77},"The dissertation investigates metric learning and unsupervised domain adaptation. It analyzes evaluation flaws in both areas and proposes improvements.",{"name":83,"@type":74,"acceptedAnswer":84},"What kinds of improvements does the dissertation propose for evaluation?",{"text":85,"@type":77},"It emphasizes fair comparisons and reproducibility, informative accuracy metrics, and hyperparameter search via cross-validation. These changes aim to reduce distorted experiment results and better reflect real performance.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]