[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123200-en":3,"doc-seo-123200-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123200,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Reassessing How to Compare and Improve the Calibration of Machine Learning Models - Research paper analysis and visualization","Machine learning model calibration is defined by whether predicted outcome probabilities match observed frequencies conditioned on those predictions, and it has become critical as ML systems expand into high-stakes domains. This work reassesses calibration reporting in recent deep learning literature, showing that trivial recalibration methods can look state of the art when calibration and prediction metrics are not paired with generalization metrics such as negative log-likelihood. A Bregman-divergence decomposition is derived to motivate calibration-metric selection, detect trivial calibration, and power an extended reliability-diagram approach for joint visualization of calibration and estimated generalization error.","Reassessing How to Compare and Improve the Calibration of  \nMachine Learning Models  \narXiv :2406 .04068v 1 [ cs .LG] 6 Jun 2024  \nMuthu Chidambaram Duke University [muthu@cs. duke. edu](muthu@cs. duke. edu)  \nRong Ge Duke University [rongge@cs. duke. edu](rongge@cs. duke. edu)  \nJune 7, 2024  \nAbstract  \nA machine learning model is calibrated if its predicted probability for an outcome matches the observed frequency for that outcome conditional on the model prediction. This property has become increasingly important as the impact of machine learning models has continued to spread to various domains. As a result, there are now a dizzying number of recent papers on measuring and improving the calibration of (specifically deep learning) models. In this work, we reassess the reporting of calibration metrics in the recent literature. We show that there exist trivial recalibration approaches that can appear seemingly state-of-the-art unless calibration and prediction metrics (i.e. test accuracy) are accompanied by additional generalization metrics such as negative log-likelihood. We then derive a calibration-based decomposition of Bregman divergences that can be used to both motivate a choice of calibration metric based on a generalization metric, and to detect trivial calibration. Finally, we apply these ideas to develop a new extension to reliability diagrams that can be used to jointly visualize calibration as well as the estimated generalization error of a model.  \n1 Introduction  \nStandard machine learning models are trained to predict probability distributions over a set of possible actions or outcomes. Model-based decision-making is then typically done by using the action or outcome associated with the highest probability, and ideally one would like to interpret the model-predicted probability as a notion of confidence in the predicted action/outcome.  \nIn order for this confidence interpretation to be valid, it is crucial that the predicted probabilities are calibrated (Lichtenstein et al. , 1982 ; Dawid, 1982 ; DeGroot & Fienberg, 1983), or accurately reflect the true frequencies of the outcome conditional on the prediction. As an informal (classic) example, a calibrated weather prediction model would satisfy the property that we observe rain 80% of the time on days for which our model predicted a 0.8 probability of rain.  \nAs the applications of machine learning models-particularly deep learning models-continue to expand to include high-stakes areas such as medical image diagnoses (Mehrtash et al. , 2019 ; Elmarakeby et al. , 2021 ; Nogales et al. , 2021) and self-driving cars (Hu et al. , 2023), so too does the importance of having calibrated model probabilities. Unfortunately, the seminal empirical investigation of Guo et al. (2017) demonstrated that deep learning models can be poorly calibrated, largely due to overconfidence.  \nThis observation has led to a number of follow-up works intended to improve model calibration using both training-time (Thulasidasan et al. , 2019 ; Müller et al. , 2020 ; Wang et al. , 2021) and post-training methods (Joy et al. , 2022 ; Gupta & Ramdas, 2022) . Comparing these proposed improvements, however, is non-trivial due to the fact that the measurement of calibration in practice is itself an active area of research (Nixon et al. ,  \n2019 ; Kumar et al. , 2019 ; Błasiok et al. , 2023), and improvements with respect to one calibration measure do not necessarily indicate improvements with respect to another.  \nEven if we fix a choice of calibration measure, the matter is further complicated by the existence of trivially calibrated models, such as models whose confidence is always their test accuracy (see Section 3) . Most works on improving calibration have approached these issues by choosing to report a number of different calibration metrics along with generalization metrics (e.g. negative log-likelihood or mean-squared error), but these choices vary greatly across works and are always no","cbCailj62BOkvojP","https://ap.wps.com/l/cbCailj62BOkvojP","pdf",2069394,1,20,"English","en",105,"# Introduction\n## Summary of Main Contributions and Takeaways","[{\"question\":\"What does calibration mean for a machine learning model?\",\"answer\":\"A model is calibrated when its predicted probability for an outcome matches the observed frequency of that outcome conditional on the model’s prediction.\"},{\"question\":\"Why can some recalibration approaches appear state-of-the-art?\",\"answer\":\"Some trivial recalibration strategies can seem strong when papers report calibration and prediction metrics without adding generalization metrics such as negative log-likelihood.\"},{\"question\":\"How does the paper use Bregman divergences for calibration?\",\"answer\":\"It derives a decomposition of Bregman divergences into a calibration error term and a sharpness term, which helps detect trivial calibration and connect calibration notions to proper scoring rules.\"}]","Reassessing How to Compare and Improve the Calibration of Machine Learning Models - Research paper analysis and visualization | PDF",1785815187,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"reassessing-how-to-compare-and-improve-the-calibration-of-machine-learning-models-research-paper-analysis-and-visualization","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/reassessing-how-to-compare-and-improve-the-calibration-of-machine-learning-models-research-paper-analysis-and-visualization/123200/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What does calibration mean for a machine learning model?","Question",{"text":75,"@type":76},"A model is calibrated when its predicted probability for an outcome matches the observed frequency of that outcome conditional on the model’s prediction.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why can some recalibration approaches appear state-of-the-art?",{"text":80,"@type":76},"Some trivial recalibration strategies can seem strong when papers report calibration and prediction metrics without adding generalization metrics such as negative log-likelihood.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper use Bregman divergences for calibration?",{"text":84,"@type":76},"It derives a decomposition of Bregman divergences into a calibration error term and a sharpness term, which helps detect trivial calibration and connect calibration notions to proper scoring rules.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":29,"slug":113},6,"Technology","technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]