[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-122476-en":3,"doc-seo-122476-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},122476,1374391974585,"Genevieve","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Evaluation metrics and statistical tests for machine learning - Practical guide to model comparison","Evaluation metrics and statistical tests support reliable comparison of machine learning (ML) models, especially in supervised tasks such as binary, multi-class, and multi-label classification, regression, image segmentation, object detection, and information retrieval. The work explains how to select appropriate metrics, gather enough metric values for testing, and perform statistical tests with correct interpretation. Practical examples compare convolutional neural networks for X-ray lung infection classification and segmentation in positron emission tomography images.","[www. nature.com/scientificreports](www. nature.com/scientificreports)  \nOPEN  \nEvaluation metrics and statistical tests for machine learning  \nOona Rainio*, JarmoTeuho & Riku Klén  \nResearch on different machine learning (ML) has become incredibly popular during the past few decades. However, for some researchers not familiar with statistics, it might be difficult to understand how to evaluate the performance of ML models and compare them with each other. Here, we introduce the most common evaluation metrics used for the typical supervised ML tasks including binary, multi-class, and multi-label classification, regression, image segmentation, object detection, and information retrieval. We explain how to choose a suitable statistical test for comparing models, how to obtain enough values of the metric for testing, and how to perform the test and interpret its results. We also present a few practical examples about comparing convolutional neural networks used to classify X-rays with different lung infections and detect cancer tumors in positron emission tomography images.  \nKeywords Evaluation metrics, Machine learning, Medical images, Statistical testing  \nDue to our developed technology and access to huge amounts of digitized data, the number of different applications using machine learning (ML) has increased dramatically during the past few decades 1. Whereas ML techniques initially included only statistical methods and simple algorithms2, ML is currently used for different purposes across the fields of engineering, medicine, public health, finance, politics, and natural sciences, both in academia and industry3. However, because of this immerse interdisciplinary interest, some of the new ML researchers might not have a good grasp of basic statistical concepts. This prompts need for ongoing education about the proper use of statistics and appropriate metrics for evaluation of performance of ML algorithms.  \nWhen new ML models are created, it is necessary to compare their performance to the already existing ones4. Evaluation serves two purposes: methods that do not perform well can be discarded, and the ones that seem promising can be further optimized. Also, especially in medicine, it is often useful to know whether an ML model outperforms an educated professional or not5–7. In supervised ML, we first divide our data for training and test sets, use the training data for training and validation of the model, predict all the instances of the test data, and compare the obtained predictions to the corresponding ground-truth values of the test set8. In this way, we can estimate whether the predictions of a new ML model are better than the predictions of a human or existing models in our test set.  \nDespite complexity of final applications, ML models typically consists of relatively simple sub-tasks, such as binary or multi-class classification and regression. In addition, a special image processing ML technique called a convolutional neural network (CNN) can be used to perform image segmentation9 and object detectors are used to find desired targets in images or video footage10. Depending on the task in question, there are certain choices of evaluation metrics that can be used to assess the performance of supervised ML models11. There are also established statistical testing practices, especially for metrics used in binary classification8, 12. Nonetheless, the misuse of certain well-known tests, such as the paired t-test, is common4, and the required assumptions of the tests are often ignored11.  \nOur aim here is to introduce the most common metrics for binary and multi-class classification, regression, image segmentation, and object detection. We explain the basics of statistical testing and what tests should be used in different situations related to supervised ML. At the end, we also give three examples about comparing the performance of CNNs for classifying X-rays related to lung infections and performing image segmentation fo","cbCaipk4X4IvLDJo","https://ap.wps.com/l/cbCaipk4X4IvLDJo","pdf",1487578,1,14,"English","en",105,"# Different machine learning tasks\n## Binary classification\n## Confusion matrix and core metrics\n## Statistical testing for comparing models\n## Practical examples with CNNs and medical imaging","[{\"question\":\"Which supervised ML tasks are covered by the evaluation metrics and statistical testing approach?\",\"answer\":\"The document covers common supervised tasks including binary, multi-class, and multi-label classification, regression, image segmentation, object detection, and information retrieval.\"},{\"question\":\"How does the document recommend comparing a new ML model against existing models?\",\"answer\":\"It describes splitting data into training and test sets, training and validating with training data, predicting the test set, and comparing predictions to ground-truth labels using evaluation metrics.\"},{\"question\":\"What statistical testing guidance is emphasized for model comparison?\",\"answer\":\"It explains how to choose suitable statistical tests, obtain enough metric values for testing, and interpret results while noting that assumptions and correct usage of tests (e.g., paired t-test misuse) matter.\"}]","Evaluation metrics and statistical tests for machine learning - Practical guide to model comparison | PDF",1785810851,35,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"evaluation-metrics-and-statistical-tests-for-machine-learning-practical-guide-to-model-comparison","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/evaluation-metrics-and-statistical-tests-for-machine-learning-practical-guide-to-model-comparison/122476/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Which supervised ML tasks are covered by the evaluation metrics and statistical testing approach?","Question",{"text":75,"@type":76},"The document covers common supervised tasks including binary, multi-class, and multi-label classification, regression, image segmentation, object detection, and information retrieval.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the document recommend comparing a new ML model against existing models?",{"text":80,"@type":76},"It describes splitting data into training and test sets, training and validating with training data, predicting the test set, and comparing predictions to ground-truth labels using evaluation metrics.",{"name":82,"@type":73,"acceptedAnswer":83},"What statistical testing guidance is emphasized for model comparison?",{"text":84,"@type":76},"It explains how to choose suitable statistical tests, obtain enough metric values for testing, and interpret results while noting that assumptions and correct usage of tests (e.g., paired t-test misuse) matter.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]