[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-118313-en":3,"doc-seo-118313-105":29,"detail-sidebar-cat-0-en-105":82},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":11,"language":21,"language_code":22,"site_id":23,"html_lang":22,"table_of_contents":24,"faqs":25,"seo_title":26,"seo_description":14,"update_tm":27,"read_time":28},118313,7971461740909,"Levi","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","An Entropic Metric for Measuring Calibration of Machine Learning Models","Machine learning model calibration determines whether predicted confidence probabilities reflect true correctness rates rather than only accuracy. The paper introduces Entropic Calibration Difference (ECD), a new calibration metric grounded in state-estimation research on target tracking. ECD separates and quantifies under-confidence and over-confidence instead of conflating them. The work analyzes why safer under-confident behavior can trade accuracy for statistical inefficiency, then evaluates ECD on simulated and real data. Results are compared against Expected Calibration Error (ECE) and Expected Signed Calibration Error (ESCE).","An Entropic Metric for Measuring Calibration of  \nMachine Learning Models  \nDaniel James Sumler∗ , Lee Devlin∗ , Simon Maskell∗ , and Richard O. Lane†  \n∗ The University of Liverpool, UK.  \n†QinetiQ, Great Malvern, UK.  \narXiv :2502 . 14545v1 [ cs .LG] 20 Feb 2025  \nAbstract—Understanding the confidence with which a machine learning model classifies an input datum is an important, and perhaps under-investigated, concept. In this paper, we propose a new calibration metric, the Entropic Calibration Difference (ECD). Based on existing research in the field of state estimation, specifically target tracking (TT), we show how ECD may be applied to binary classification machine learning models. We describe the relative importance of under- and over-confidence and how they are not conflated in the TT literature. Indeed, our metric distinguishes under- from over-confidence. We consider this important given that algorithms that are under-confident are likely to be “safer” than algorithms that are over-confident, albeit at the expense of also being over-cautious and so statistically inefficient. We demonstrate how this new metric performs on real and simulated data and compare with other metrics for machine learning model probability calibration, including the Expected Calibration Error (ECE) and its signed counterpart, the Expected Signed Calibration Error (ESCE).  \nI. INTRODUCTION  \nCalibration of probabilities is an important and oftenoverlooked concept when developing machine learning (ML) models. Usually, accuracy is the main metric used to calculate how well an ML model performs in terms of predicting a class for unseen data. Generally speaking, the closer the accuracy is to 100%, the better the model is deemed to be. However, this does not take into account the probability of predictions that the model outputs, which can be just as important, if not more, than the accuracy.  \nIn binary classification, a probability greater than a threshold, typically 0.5, is enough to decide whether an input belongs to one of two classes. While accuracy informs whether a classification is correct, a calibration metric informs how well the confidence probabilities match the true proportions of correct decisions. For example, a model that always outputsa probability of 0.6 for class label 1, but always gets this classification correct, should produce a poor calibration score, as even though the model has classified the output correctly, it has low confidence in that decision.  \nCalibration has become even more important as of late, asthe research of Guo et al. [1] reveals that while modern neural networks are more accurate than ever, they are also badly calibrated. This could be attributed to the over-confidence of said networks due to the large amount of data they are able to be trained on.  \nA well-calibrated model can be defined as one that outputs probabilities that are representative of the real-life occurrences from the unseen data. For example if, on average, 70% of  \npeople are correctly predicted to contract a certain disease, then one would expect the average probability outputted by a diagnosis model to be 0.7 . A mathematical representation of calibration can be seen in (1), where x ∈ {0, 1} is the true label, y ∈ Rd is an observed data sample of dimension d belonging to a binary class k ∈ {0, 1}, and pk is the confidence in class k, while P is the true probability.  \nP (x = k|y) = pk (1)  \nIn this paper, we present a novel calibration metric that addresses some weaknesses of some of the most commonly-used existing metrics. In section II, we discuss existing calibration metrics that are widely discussed throughout the literature. In section III, we detail our motivations for “safe” calibration and why we feel our metric is necessary. Section IV explains our new metric and details how the results can be interpreted. In section V, we test our metric on simulated and real data, and compare it with other popular calibration metrics. Finally,","cbCaijEmf0xE0Qea","https://ap.wps.com/l/cbCaijEmf0xE0Qea","pdf",1462669,1,"English","en",105,"# Abstract\n# Introduction\n## Calibration of probabilities and limitations of accuracy\n## Motivation from modern neural networks being badly calibrated\n# Existing Metrics for Model Calibration\n## Brier Score\n## Probability adjustment and evaluation metrics\n# Proposed Entropic Calibration Difference (ECD)\n## Safe calibration and interpretation\n# Experiments and Results\n## Simulated and real data evaluation\n## Comparison with ECE and ESCE\n# Conclusion","[{\"question\":\"Why does the paper discuss “safe” calibration?\",\"answer\":\"It argues that under-confident models are often safer than over-confident ones, though this safety can come with over-cautious and statistically inefficient predictions.\"}]","An Entropic Metric for Measuring Calibration of Machine Learning Models | PDF",1785682991,20,{"code":4,"msg":30,"data":31},"ok",{"site_id":23,"language":22,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":77,"head_meta":79,"extra_data":81,"updated_unix":27},"an-entropic-metric-for-measuring-calibration-of-machine-learning-models","",{"@graph":35,"@context":76},[36,53,67],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/an-entropic-metric-for-measuring-calibration-of-machine-learning-models/118313/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":22,"description":14,"dateModified":61,"datePublished":61,"encodingFormat":60,"isAccessibleForFree":62,"interactionStatistic":63},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":64,"interactionType":65,"userInteractionCount":4},"InteractionCounter",{"@type":66},"ViewAction",{"@type":68,"mainEntity":69},"FAQPage",[70],{"name":71,"@type":72,"acceptedAnswer":73},"Why does the paper discuss “safe” calibration?","Question",{"text":74,"@type":75},"It argues that under-confident models are often safer than over-confident ones, though this safety can come with over-cautious and statistically inefficient predictions.","Answer","https://schema.org",{"og:url":51,"og:type":78,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":80,"canonical":51},"index,follow",{"doc_id":7,"site_id":23},{"code":4,"msg":5,"data":83},[84,88,92,96,101,106,111,114,118,121,125],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":85,"show_sort_weight":86,"slug":87},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":89,"show_sort_weight":90,"slug":91},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":93,"show_sort_weight":94,"slug":95},"Exam",70,"exam",{"id":97,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},5,"Comic",60,"comic",{"id":102,"doc_module":4,"doc_module_name":45,"category_name":103,"show_sort_weight":104,"slug":105},6,"Technology",50,"technology",{"id":107,"doc_module":4,"doc_module_name":45,"category_name":108,"show_sort_weight":109,"slug":110},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":112,"slug":113},30,"research-report",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":28,"slug":117},9,"Religion & Spirituality","religion-spirituality",{"id":28,"doc_module":4,"doc_module_name":45,"category_name":119,"show_sort_weight":28,"slug":120},"World Cup","world-cup",{"id":122,"doc_module":4,"doc_module_name":45,"category_name":123,"show_sort_weight":122,"slug":124},10,"Lifestyle","lifestyle",{"id":126,"doc_module":4,"doc_module_name":45,"category_name":127,"show_sort_weight":97,"slug":128},19,"General","general"]