[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126740-en":3,"doc-seo-126740-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126740,962084928432,"Emma Wilson","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Calibration in Machine Learning Uncertainty Quantification - beyond consistency to target adaptivity","Reliable uncertainty quantification for machine-learning regression is increasingly critical in materials and chemical sciences, where uncertainty must reflect a range of plausible predicted values. Average calibration and conditional calibration via consistency (often assessed with reliability diagrams) are insufficient to ensure dependable individual predictions. This work argues that consistency and adaptivity are complementary validation targets and proposes an integrated validation framework that evaluates variance-based UQ metrics for reliability across the input feature space.","arXiv :2309 .06240v2 [ stat .ML] 7 Dec 2023  \nCalibration in Machine Learning Uncertainty Quanti􀀜cation: beyond consistency to target adaptivity  \nPascal PERNOT 1  \nInstitut de Chimie Physique, UMR8000 CNRS,  \nUniversité Paris-Saclay, 91405 Orsay, Francea)  \nReliable uncertainty quanti􀀂cation (UQ) in machine learning (ML) regression tasks is becoming the focus of many studies in materials and chemical science. It is now well understood that average calibration is insuf􀀂cient, and most studies implement additional methods testing the conditional calibration with respect to uncertainty, i.e. consistency. Consistency is assessed mostly by so-called reliability diagrams. There exists however another way beyond average calibration, which is conditional calibration with respect to input features, i.e. adaptivity. In practice, adaptivity is the main concern of the 􀀂nal users of a ML-UQ method, seeking for the reliability of predictions and uncertainties for any point in features space. This article aims to show that consistency and adaptivity are complementary validation targets, and that a good consistency does not imply a good adaptivity. An integrated validation framework is proposed and illustrated on a representative example.  \na)[Electronic mail:](Electronic mail: pascal.pernot@cnrs.fr)[ pascal.pernot@cnrs.fr](Electronic mail: pascal.pernot@cnrs.fr)  \nI. INTRODUCTION  \nThe quest for trust or con􀀂dence in the predictions of data-based algorithms1–4 has led to a profusion of uncertainty quanti􀀂cation (UQ) methods in machine learning (ML) .5–19 However, not all of these UQ methods provide uncertainties that can be relied upon,20,21 notably if, as in metrology, one expects uncertainty to inform us on a range of plausible values for a predicted property.22,23  \nIn pre-ML computational chemistry, UQ metrics consisted essentially in standard uncertainty, i.e. the standard deviation of the distribution of plausible values (a variance-based metric), or expanded uncertainty, i.e. the half-range of a prediction interval, typically at the 95 % level (an interval-based metric) .23,24 The advent of ML methods provided UQ metrics beyond this standard setup, for instance distances in feature or latent space13,25,26 or the ∆-metric,27 which have no direct statistical or probabilistic meaning. These metrics might however be converted to variance-based metrics by post hoc recalibration methods such as temperature scaling28–30 and isotonic regression31 , or to interval-based metrics by conformal inference.32–35 Nevertheless, all UQ metrics need to be validated to ensure that they are adapted to their intended use. In this study, I focus on the reliability of variance-based UQ metrics for the prediction of properties at the individual level.36  \nThe validation of UQ metrics is based on the concept of calibration. A handful of validation methods exist that explore more or less complementary aspects of calibration. A trio of methods seems to have recently taken the center stage: the reliability diagram28,30, the calibration curve29 and the con􀀂dence curve37,38. They implement three different approaches to calibration which are not necessarily independent, but it is essential to realize that they do not cover the full spectrum of calibration requirements. In particular, none of these methods addresses the essential reliability of predicted uncertainties with respect to the input features, sometimes called individual calibration39,40.  \nA. Scope and limitations of the study  \nThe aim of this article is to propose a complete validation framework for variance-based UQ metrics, based on the concept of conditional calibration and its complementary aspects of consistency (conditional calibration with respect to uncertainty) and adaptivity (conditional calibration with respect to input features) .  \nIt is well known that average calibration is not suf􀀂cient to establish the reliability of ML-UQ predictions. This study goes one step further and is designed to","cbCailf0WLzsi0rg","https://ap.wps.com/l/cbCailf0WLzsi0rg","pdf",1976403,1,25,"English","en",105,"# Introduction\n## Scope and limitations of the study\n## Structure of the article\n# Validation of variance-based UQ metrics","[{\"question\":\"Why is average calibration not enough for ML uncertainty quantification?\",\"answer\":\"Average calibration does not guarantee that predictions and uncertainties remain reliable at the individual level. Conditional reliability with respect to uncertainty is commonly tested, but it still does not ensure reliability with respect to input features.\"},{\"question\":\"What is the difference between consistency and adaptivity in calibration?\",\"answer\":\"Consistency refers to conditional calibration with respect to uncertainty, typically evaluated using reliability diagrams. Adaptivity refers to conditional calibration with respect to input features, ensuring reliability throughout the feature space.\"},{\"question\":\"What does the article propose for validating variance-based UQ metrics?\",\"answer\":\"It proposes an integrated validation framework that treats consistency and adaptivity as complementary validation targets, illustrated on a representative computational chemistry example.\"}]","Calibration in Machine Learning Uncertainty Quantification - beyond consistency to target adaptivity | PDF",1785934546,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"calibration-in-machine-learning-uncertainty-quantification-beyond-consistency-to-target-adaptivity","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/calibration-in-machine-learning-uncertainty-quantification-beyond-consistency-to-target-adaptivity/126740/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is average calibration not enough for ML uncertainty quantification?","Question",{"text":75,"@type":76},"Average calibration does not guarantee that predictions and uncertainties remain reliable at the individual level. Conditional reliability with respect to uncertainty is commonly tested, but it still does not ensure reliability with respect to input features.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the difference between consistency and adaptivity in calibration?",{"text":80,"@type":76},"Consistency refers to conditional calibration with respect to uncertainty, typically evaluated using reliability diagrams. Adaptivity refers to conditional calibration with respect to input features, ensuring reliability throughout the feature space.",{"name":82,"@type":73,"acceptedAnswer":83},"What does the article propose for validating variance-based UQ metrics?",{"text":84,"@type":76},"It proposes an integrated validation framework that treats consistency and adaptivity as complementary validation targets, illustrated on a representative computational chemistry example.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]