[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-121310-en":3,"doc-seo-121310-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},121310,1099514067438,"River Wang","https://ap-avatar.wpscdn.com/avatar/100002539ee87300030?x-image-process=image/resize,m_fixed,w_180,h_180&k=1780474512215547542",8,"Research & Report","Denoising ESG - quantifying data uncertainty from missing data with Machine Learning and prediction intervals","Environmental, Social, and Governance (ESG) datasets often contain substantial missing values, which can create inconsistencies in ESG ratings when different imputation strategies are used. This study investigates missing-data handling in ESG by comparing K-Nearest Neighbors, Gradient Boosting, Multiple Imputation by Chained Equations (MICE), and Neural Networks. The work quantifies risk from data anomalies and evaluates how this risk affects score variability. Prediction-uncertainty methods such as Predictive Mean Matching and Local Residual Draw provide confidence measures for individual predictions, improving imputation accuracy and enabling more reliable ESG scoring in banking and finance.","arXiv :2407 .20047v 1 [ cs .LG] 29 Jul 2024  \nDenoising ESG: quantifying data uncertainty from missing data with Machine Learning and prediction intervals  \nSergio Caprioli 1 , Jacopo Foschi2 , Riccardo Crupi2[0009−0005−6714−5161], and  \nAlessandro Sabatino2[0000−0002−1336−2057]  \n1 Intesa Sanpaolo S.P.A. , Milano MI 20121, Italy [sergio.caprioli@intesasanpaolo.com](sergio.caprioli@intesasanpaolo.com)  \n2 Intesa Sanpaolo S.P.A. , Torino TO 10121, Italy {[name.surname](name.surname}@intesasanpaolo.com)[}](name.surname}@intesasanpaolo.com)[@intesasanpaolo.com](name.surname}@intesasanpaolo.com)  \nAbstract. Environmental, Social, and Governance (ESG) datasets are frequently plagued by significant data gaps, leading to inconsistencies in ESG ratings due to varying imputation methods. This study addresses the missing data issues in ESG datasets using machine learning techniques, comparing K-Nearest Neighbors, Gradient Boosting, Multiple Imputation by Chained Equations (MICE) and Neural Networks. We focus on quantifying the risk induced by data anomalies and provide tools to assess the impacts of this risk on the variability of the scores.  \nBy introducing prediction uncertainty using methods such as Predictive Mean Matching and Local Residual Draw, in order to assign confidence measures to individual predictions, we provide a nuanced understanding of prediction uncertainty. Empirical analyses show that these methods improve imputation accuracy and quantify uncertainty, which is required for reliable ESG scoring in banking and finance.  \nKeywords: ESG · Multiple Imputation · MICE · Machine Learning  \n1 Introduction  \nThe growing attention towards the ramifications of climate change has bolstered international cooperation on sustainable finance. This collaboration involves initiatives from industry and institutions, as well as recommendations from regulators and legislators. There is a growing emphasis on integrating ESG criteria into the strategies and operations of banks, accompanied by the requisite development of adequate support tools.  \nEuropean regulators acknowledge that ESG ratings play an important role in global capital markets, as investors, borrowers and issuers increasingly use those ESG ratings as part of the process of making informed, sustainable investment and financing decisions [21] . For this reason  \n“Better comparability and increased reliability of ESG ratings would enhance the efficiency of that fast-growing market, thereby facilitating progress towards the objectives of the Green Deal”.  \n2 Sergio Caprioli , Jacopo Foschi , Riccardo Crupi , and Alessandro Sabatino  \nThis aspect reverberated also on the banks, given that, as stated in [4],  \n“Institutions should embed ESG risks in their regular processes including risk appetite, internal controls and ICAAP. Besides, institutions should monitor ESG risks through effective internal reporting frameworks anda range of backward and forward-looking ESG risks metrics and indicators.”  \nGiven the importance attributed to the ESG metrics for the assessment of firms’ performances on a given ESG issue, the topic of the reliability of such metrics plays a crucial role. Berg [et.al](et.al) [5] analyzed the divergence of ESG ratings based on data from six agencies, decomposing the divergence into contributions of scope, measurement and weight. They estimated that measurement contributes 56% of the divergence. The issue assumes even a greater impact when considering that Banks should measure ESG risk for a variety of counterparties aggregating information derived from different sources: e.g. raw data from different rating agencies, internal sources, questionnaires.  \nAs part of the measurement divergence issue, datasets used to calculate ESG Scores are affected by vast data gaps, which are imputed by researchers and analysts with different methods, contributing to the inconsistency of ESG ratings ([5], [6]) . There is lack of regulation and standardization on ESG data disclo","cbCaijSf2LGyJJ4e","https://ap.wps.com/l/cbCaijSf2LGyJJ4e","pdf",1043974,1,15,"English","en",105,"# 1 Introduction\n## ESG ratings, sustainable finance, and regulatory context\n## Divergence in ESG ratings and measurement contributions\n## Missing data, imputation methods, and inconsistency\n## Prediction-interval uncertainty and confidence measures\n# 2 Methods and Models\n## K-Nearest Neighbors and Gradient Boosting\n## MICE and Neural Networks","[{\"question\":\"Why do missing values matter for ESG ratings?\",\"answer\":\"Missing values in ESG datasets can lead to inconsistencies in ESG ratings when different imputation methods are applied. This affects comparability and reliability of ESG scoring.\"},{\"question\":\"Which machine learning and imputation approaches are compared in the study?\",\"answer\":\"The study compares K-Nearest Neighbors, Gradient Boosting, Multiple Imputation by Chained Equations (MICE), and Neural Networks. It also incorporates uncertainty-aware techniques for prediction.\"},{\"question\":\"How does the study quantify prediction uncertainty from missing data?\",\"answer\":\"It introduces prediction uncertainty using methods such as Predictive Mean Matching and Local Residual Draw. These provide confidence measures for individual predictions to assess impacts on score variability.\"}]","Denoising ESG - quantifying data uncertainty from missing data with Machine Learning and prediction intervals | PDF",1785735015,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"denoising-esg-quantifying-data-uncertainty-from-missing-data-with-machine-learning-and-prediction-intervals","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/denoising-esg-quantifying-data-uncertainty-from-missing-data-with-machine-learning-and-prediction-intervals/121310/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do missing values matter for ESG ratings?","Question",{"text":75,"@type":76},"Missing values in ESG datasets can lead to inconsistencies in ESG ratings when different imputation methods are applied. This affects comparability and reliability of ESG scoring.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which machine learning and imputation approaches are compared in the study?",{"text":80,"@type":76},"The study compares K-Nearest Neighbors, Gradient Boosting, Multiple Imputation by Chained Equations (MICE), and Neural Networks. It also incorporates uncertainty-aware techniques for prediction.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the study quantify prediction uncertainty from missing data?",{"text":84,"@type":76},"It introduces prediction uncertainty using methods such as Predictive Mean Matching and Local Residual Draw. These provide confidence measures for individual predictions to assess impacts on score variability.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]