[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123994-en":3,"doc-seo-123994-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123994,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","The harms of class imbalance corrections for machine learning based prediction models - a simulation study","Risk prediction models increasingly support healthcare decision making, where reliable calibration of risk estimates is essential for clinical use. Development data are often class-imbalanced, so researchers frequently apply imbalance corrections to balance event vs non-event observations. The study evaluates how imbalance corrections affect calibration across multiple machine learning algorithms. Extensive Monte Carlo simulations compare out-of-sample predictive performance with and without correction under varying sample sizes, predictor counts, and event fractions, and findings are illustrated using MIMIC-III data. Models trained without correction consistently achieve equal or better calibration; correction induces miscalibration via risk overestimation and re-calibration may not fully fix it. Correcting class imbalance is therefore not always necessary and can be harmful for clinically reliable individualized risk estimation.","arXiv :2404 . 19494v1 [ stat .ME] 30 Apr 2024  \nThe harms of class imbalance corrections for machine learning based prediction models: a  \nsimulation study.  \nA Preprint  \nAlex Carriero  \nJulius Center for Health Sciences and Primary Care University Medical Center Utrecht Netherlands [a.j.carriero@umcutrecht.nl](a.j.carriero@umcutrecht.nl)  \nKim Luijken  \nJulius Center for Health Sciences and Primary Care University Medical Center Utrecht Netherlands  \nAnne de Hond  \nJulius Center for Health Sciences and Primary Care University Medical Center Utrecht Netherlands  \nKarel GM Moons  \nJulius Center for Health Sciences and Primary Care University Medical Center Utrecht Netherlands  \nBen van Calster  \nDepartment of Development and Regeneration KU, Leuven  \nBelgium  \nMaarten van Smeden  \nJulius Center for Health Sciences and Primary Care  \nUniversity Medical Center Utrecht  \nNetherlands  \nMay 1, 2024  \nAbstract  \nRisk prediction models are increasingly used in healthcare to aid in clinical decision making. In most clinical contexts, model calibration (i.e., assessing the reliability of risk estimates) is critical. Data available for model development are often not perfectly balanced with respect to the modeled outcome (i.e., individuals with vs. without the event of interest are not equally represented in the data) . It is common for researchers to correct this class imbalance, yet, the effect of such imbalance corrections on the calibration of machine learning models is largely unknown. We studied the effect of imbalance corrections on model calibration for a variety of machine learning algorithms. Using extensive Monte Carlo simulations we  \nA preprint-May 1, 2024  \ncompared the out-of-sample predictive performance of models developed with an imbalance correction to those developed without a correction for class imbalance across different datagenerating scenarios (varying sample size, the number of predictors and event fraction) . Our findings were illustrated in a case study using MIMIC-III data. In all simulation scenarios, prediction models developed without a correction for class imbalance consistently had equal or better calibration performance than prediction models developed with a correction for class imbalance. The miscalibration introduced by correcting for class imbalance was characterized by an over-estimation of risk and was not always able to be corrected with re-calibration.  \nCorrecting for class imbalance is not always necessary and may even be harmful for clinical prediction models which aim to produce reliable risk estimates on an individual basis.  \nKeywords Class Imbalance · Machine Learning · Calibration · Prediction Modeling  \n1 Introduction  \nRisk prediction models are increasingly used in healthcare to aid in clinical decision making; for example, to help decide if a patient is a good candidate for surgery or to communicate a patient’s risk of disease [1, 2 , 3] . As such, the purpose of a clinical prediction model is often to estimate a patient’s risk of experiencing a particular event (e.g., successful surgery, disease) [4, 5] . Due to the rarity of many diseases, data available to train clinical prediction models often exhibit class imbalance i.e. , observations from patients with vs. without the event of interest are not equally represented in the data. In machine learning literature, imbalance correction methods are commonly applied to correct class imbalance by artificially creating data that are more or perfectly balanced [6, 7 , 8 , 9], although the benefit of such corrections for model performance is not always clear.  \nAn abundance of imbalance correction methods exist [7, 8 , 9], yet, information regarding the effect of these imbalance corrections on model calibration is sparse. Model calibration captures the accuracy of risk estimates, relating to the agreement between the estimated (predicted) and observed number of events [3] . In clinical applications where a patient’s predicted risk is the e","cbCaifSDnBv8KzMd","https://ap.wps.com/l/cbCaifSDnBv8KzMd","pdf",24049447,1,25,"English","en",105,"# Abstract\n# Introduction\n## Purpose of clinical prediction models\n## Class imbalance and common correction practices\n## Calibration and consequences of miscalibration\n## Prior evidence and the research gap\n# Study approach and scope","[{\"question\":\"Why is calibration critical for healthcare risk prediction models?\",\"answer\":\"Calibration measures whether predicted risks match observed event frequencies. Poor calibration can systematically over- or under-estimate risk or produce predictions that are too extreme or too modest, leading to misguided clinical decisions and potentially false reassurance.\"},{\"question\":\"What does this study investigate about class imbalance corrections?\",\"answer\":\"It examines how common imbalance correction methods influence the calibration of machine learning-based risk prediction models, focusing on whether corrected models maintain reliable individualized risk estimates.\"},{\"question\":\"What are the main findings from the simulations and the case study?\",\"answer\":\"Across simulation scenarios, models developed without class imbalance correction consistently show equal or better calibration performance. When correction is applied, miscalibration is characterized by risk overestimation and re-calibration does not always repair it; results are further illustrated using MIMIC-III data.\"}]","The harms of class imbalance corrections for machine learning based prediction models - a simulation study | PDF",1785819712,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-harms-of-class-imbalance-corrections-for-machine-learning-based-prediction-models-a-simulation-study","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-harms-of-class-imbalance-corrections-for-machine-learning-based-prediction-models-a-simulation-study/123994/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is calibration critical for healthcare risk prediction models?","Question",{"text":75,"@type":76},"Calibration measures whether predicted risks match observed event frequencies. Poor calibration can systematically over- or under-estimate risk or produce predictions that are too extreme or too modest, leading to misguided clinical decisions and potentially false reassurance.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does this study investigate about class imbalance corrections?",{"text":80,"@type":76},"It examines how common imbalance correction methods influence the calibration of machine learning-based risk prediction models, focusing on whether corrected models maintain reliable individualized risk estimates.",{"name":82,"@type":73,"acceptedAnswer":83},"What are the main findings from the simulations and the case study?",{"text":84,"@type":76},"Across simulation scenarios, models developed without class imbalance correction consistently show equal or better calibration performance. When correction is applied, miscalibration is characterized by risk overestimation and re-calibration does not always repair it; results are further illustrated using MIMIC-III data.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]