[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119341-en":3,"doc-seo-119341-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119341,1374391975076,"Riley","https://ap-avatar.wpscdn.com/avatar/14000253ca4ec9f6853?x-image-process=image/resize,m_fixed,w_180,h_180&k=1783305029341752051",8,"Research & Report","Characterizing Uncertainty in Machine Learning for Chemistry - Article Abstract and Key Findings","Characterizing uncertainty in machine learning models supports reliability, robustness, safety, and active learning in chemistry. The work decomposes total uncertainty into aleatoric noise from data and epistemic shortcomings of the model, then further distinguishes epistemic contributions from model bias and variance. It examines how noise level, dataset size, architecture, molecular representation, ensemble size, and data splitting affect chemical property predictions across diverse target properties and chemical space. Controlled experiments show that test-set noise can depress observed performance, size-extensive aggregation is vital for extensive properties, and ensembling improves uncertainty quantification through variance reduction. Practical guidelines address improving underperforming models under different uncertainty regimes.","[pubs.acs.org/jcim](pubs.acs.org/jcim)  Article   \nCharacterizing Uncertainty in Machine Learning for Chemistry  \nEsther Heid, Charles J. McGill, Florence H. Vermeire, and William H. Green*  \n Cite This: J. Chem. Inf. Model. 2023, 63, 4012−4029  \nRead Online  \nACCESS  \n Metrics & More  \n Article Recommendations  \nDownloaded via TU WIEN on January 8, 2024 at 13:52:06 (UTC) . See [https://pubs.acs.org/sharingguidelines](https://pubs.acs.org/sharingguidelines) for options on how to legitimately share published articles.  \nABSTRACT: Characterizing uncertainty in machine learning models has recently gained interest in the context of machine learning reliability, robustness, safety, and active learning. Here, we separate the total uncertainty into contributions from noise in the data (aleatoric) and shortcomings of the model (epistemic), further dividing epistemic uncertainty into model bias and variance contributions. We systematically address the influence of noise, model bias, and model variance in the context of chemical property predictions, where the diverse nature of target properties and the vast chemical chemical space give rise to many different distinct sources of prediction error. We demonstrate that different sources of error can each be significant in different contexts and must be  \nindividually addressed during model development. Through controlled experiments on data sets of molecular properties, we show important trends in model performance associated with the level of noise in the data set, size of the data set, model architecture, molecule representation, ensemble size, and data set splitting. In particular, we show that 1) noise in the test set can limit a model’s observed performance when the actual performance is much better, 2) using size-extensive model aggregation structures is crucial for extensive property prediction, and 3) ensembling is a reliable tool for uncertainty quantification and improvement specifically for the contribution of model variance. We develop general guidelines on how to improve an underperforming model when falling into different uncertainty contexts.  \n■ INTRODUCTION  \nMachine learning models for chemical applications such as predicting molecular and reaction properties are becoming not only increasingly popular but also increasingly accurate, for example for quantum-mechanical properties, 1−3 biological effects,4−6 physicochemical properties, 7 − 11 reaction yields,12−14 or reaction rates and barriers. 15−19 Also, promising developments in the fields of retrosynthesis20−24 and forward reaction prediction25−28 have been made.  \nHowever, despite the increase in accuracy, many machine learning models fail in real-world applications.29,30 This can be due to a lack of generalization, lack of ability to filter out erroneous predictions for edge cases, or because the employed training and test sets are simply not reflective of the application of interest, so that the developed model is suboptimal for the proposed task. Poor choice of a test set can overestimate or, more commonly, underestimate the actual errors that a user will encounter when the model is applied. Optimizing a mediocre model can be tedious, time-consuming, and often unfruitful. Moreover, the model architectures, input representations, and data set characteristics for chemical applications differ considerably from other fields of research, so that following general guidelines for optimizing machine learning models often fails to produce accurate models for molecular and reaction properties. To optimize a model in a targeted and  \nefficient manner, it is imperative to understand and identify possible sources of error and uncertainty in a model.  \nThe separation of the total uncertainty into aleatoric (datadependent, noise-induced, irreducible) and epistemic (modeldependent, reducible) contributions31 has recently received increasing attention.32−34 The aleatoric uncertainty is often referred to as the irreducible component ","cbCaiu5iOgzMekL0","https://ap.wps.com/l/cbCaiu5iOgzMekL0","pdf",3889707,1,18,"English","en",105,"# Abstract\n## Uncertainty Decomposition (Aleatoric vs. Epistemic)\n## Sources of Prediction Error in Chemistry\n## Controlled Experiments and Observed Trends\n## Ensembling and Model Performance Improvements\n## Guidelines for Improving Underperforming Models","[{\"question\":\"How is total uncertainty decomposed in machine learning models for chemistry?\",\"answer\":\"Total uncertainty is separated into aleatoric uncertainty from noise in the data and epistemic uncertainty from model shortcomings. Epistemic uncertainty is further divided into bias and variance contributions.\"},{\"question\":\"Why can test-set noise reduce a model’s observed performance?\",\"answer\":\"Noise in the test set can limit the performance measured during evaluation, even when the model’s true underlying performance would be much better.\"},{\"question\":\"What role does ensembling play in uncertainty quantification and improvement?\",\"answer\":\"Ensembling is presented as a reliable approach for uncertainty quantification, specifically improving the variance-related contribution of epistemic uncertainty and thereby supporting better model behavior.\"}]","Characterizing Uncertainty in Machine Learning for Chemistry - Article Abstract and Key Findings | PDF",1785723789,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"characterizing-uncertainty-in-machine-learning-for-chemistry-article-abstract-and-key-findings","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/characterizing-uncertainty-in-machine-learning-for-chemistry-article-abstract-and-key-findings/119341/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"How is total uncertainty decomposed in machine learning models for chemistry?","Question",{"text":75,"@type":76},"Total uncertainty is separated into aleatoric uncertainty from noise in the data and epistemic uncertainty from model shortcomings. Epistemic uncertainty is further divided into bias and variance contributions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why can test-set noise reduce a model’s observed performance?",{"text":80,"@type":76},"Noise in the test set can limit the performance measured during evaluation, even when the model’s true underlying performance would be much better.",{"name":82,"@type":73,"acceptedAnswer":83},"What role does ensembling play in uncertainty quantification and improvement?",{"text":84,"@type":76},"Ensembling is presented as a reliable approach for uncertainty quantification, specifically improving the variance-related contribution of epistemic uncertainty and thereby supporting better model behavior.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]