[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126979-en":3,"doc-seo-126979-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126979,687207024643,"Oliver","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","A review of model evaluation metrics for machine learning in genetics and genomics","Machine learning has strong potential in genetics and genomics, where large, complex datasets can inform disease risk, pathogenesis, and health and wellbeing prediction. Model evaluation, however, can be affected by bias and result inflation, creating unintended harm. This review explains how metrics shape interpretation by surveying clustering, classification, and regression evaluation measures, comparing their advantages and disadvantages, and outlining common evaluation pitfalls, with genomics-focused examples.","TYPE Review  \nPUBLISHED 10 September 2024 DOI 10.3389/fbinf.2024.1457619  \nOPEN ACCESS  \nEDITED BY  \nKeith A. Crandall,  \nGeorge Washington University, United States  \nREVIEWED BY  \nPiyali Basak,  \nMerck (United States), United States Ali Taheriyoun,  \nGeorge Washington University, United States  \n*CORRESPONDENCE  \nCatriona Miller,  \n [catriona.miller@auckland.ac. nz](catriona.miller@auckland.ac. nz)[ ](catriona.miller@auckland.ac. nz)Justin O’Sullivan,  \n [justin.osullivan@auckland.ac. nz](justin.osullivan@auckland.ac. nz)  \nRECEIVED 01 July 2024  \nACCEPTED 27 August 2024  \nPUBLISHED 10 September 2024  \nCITATION  \nMiller C, Portlock T, Nyaga DM and O’Sullivan JM (2024) A review of model evaluation metrics for machine learning in genetics and genomics.  \nFront. Bioinform. 4:1457619 .  \ndoi: 10.3389/fbinf.2024.1457619  \nCOPYRIGHT  \n© 2024 Miller, Portlock, Nyaga and O’Sullivan. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY) . The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.  \nA review of model evaluation metrics for machine learning in genetics and genomics  \nCatriona Miller 1*, Theo Portlock 1, Denis M. Nyaga 1 and Justin M. O’Sullivan 1,2,3,4*  \n1The Liggins Institute, The University of Auckland, Auckland, New Zealand, 2The Maurice Wilkins Centre, The University of Auckland, Auckland, New Zealand, 3MRC Lifecourse Epidemiology Unit, University of Southampton, Southampton, United Kingdom, 4Singapore Institute for Clinical Sciences, Agency for Science Technology and Research, Singapore, Singapore  \nMachine learning (ML) has shown great promise in genetics and genomics where large and complex datasets have the potential to provide insight into many aspects of disease risk, pathogenesis of genetic disorders, and prediction of health and wellbeing. However, with this possibility there is a responsibility to exercise caution against biases and inﬂation of results that can have harmful unintended impacts. Therefore, researchers must understand the metrics used to evaluate ML models which can inﬂuence the critical interpretation of results. In this review we provide an overview of ML metrics for clustering, classiﬁcation, and regression and highlight the advantages and disadvantages of each. We also detail common pitfalls that occur during model evaluation. Finally, we provide examples of how researchers can assess and utilise the results of ML models, speciﬁcally from a genomics perspective.  \nKEYWORDS  \nmetrics, machine learning, genomics prediction, clustering, classiﬁcation, regression, disease prediction  \n1 Introduction  \nThe general hype around the generative artiﬁcial intelligence (AI) era has increased the popularity of machine learning (ML) for a range of applications. Alongside this, the advent of “plug and play” style ML tools, such as PyCaret, has dramatically increased the accessibility of ML to scientists and researchers without a traditional computational background (Ali, 2020; Manduchi et al., 2022; Whig et al., 2023) . In genomics, ML is becoming increasingly used to analyse large and complex datasets, including sequencing data (Caudai et al., 2021; Chafai et al., 2024). Therefore, it is increasingly important that “all”researchers understand what happens after an ML model has been deployed. This is particularly true for the choice of performance metrics and how to interpret the validity of the results. As such, without understanding the common metrics used in ML, together with an awareness of the inherent strengths and weaknesses of such metrics, there is a possible risk of result inﬂation (Kapoor and Narayanan, 2023) . Therefore, understanding the potential biases withi","cbCaivd3yM8cfDQA","https://ap.wps.com/l/cbCaivd3yM8cfDQA","pdf",1877565,1,13,"English","en",105,"# Introduction\n## Types of ML typically used in genomics\n# Model evaluation metrics in genomics\n## Clustering metrics\n## Classification metrics\n## Regression metrics\n## Common pitfalls and result inflation risks\n# Practical examples for interpreting ML results","[{\"question\":\"Why are model evaluation metrics especially important in genetics and genomics?\",\"answer\":\"Metrics influence how results are interpreted. In genetics and genomics, improper evaluation can amplify bias and inflate performance estimates, leading to harmful or misleading conclusions.\"},{\"question\":\"Which machine learning tasks are covered for model evaluation in this review?\",\"answer\":\"The review focuses on clustering, classification, and regression, which correspond to unsupervised and supervised learning subcategories.\"},{\"question\":\"What kinds of pitfalls are highlighted during model evaluation?\",\"answer\":\"Common pitfalls that can bias model performance and inflate reported metrics are detailed, emphasizing cautious interpretation of evaluation outcomes.\"}]","A review of model evaluation metrics for machine learning in genetics and genomics | PDF",1785936015,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-review-of-model-evaluation-metrics-for-machine-learning-in-genetics-and-genomics","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-review-of-model-evaluation-metrics-for-machine-learning-in-genetics-and-genomics/126979/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why are model evaluation metrics especially important in genetics and genomics?","Question",{"text":75,"@type":76},"Metrics influence how results are interpreted. In genetics and genomics, improper evaluation can amplify bias and inflate performance estimates, leading to harmful or misleading conclusions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which machine learning tasks are covered for model evaluation in this review?",{"text":80,"@type":76},"The review focuses on clustering, classification, and regression, which correspond to unsupervised and supervised learning subcategories.",{"name":82,"@type":73,"acceptedAnswer":83},"What kinds of pitfalls are highlighted during model evaluation?",{"text":84,"@type":76},"Common pitfalls that can bias model performance and inflate reported metrics are detailed, emphasizing cautious interpretation of evaluation outcomes.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]