[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126008-en":3,"doc-seo-126008-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},126008,687207024643,"Oliver","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","The impact of imputation quality on machine learning classifiers for datasets with missing values","Machine learning classification in incomplete datasets is a widespread but difficult task because real-world data often contains missing values that must be imputed before models can be trained and evaluated. Using three simulated and three clinical datasets, the study analyzes how classifier choice, imputation method, and missingness patterns affect downstream performance via quantitative ANOVA. It also compares common imputation-quality metrics against new discrepancy scores based on sliced Wasserstein distance, examining stability and interpretability.","ARTICLE  \n [https://doi.org/10.1038/s43856-023-00356-z](https://doi.org/10.1038/s43856-023-00356-z)  OPEN  \nThe impact of imputation quality on machine  \nlearning classiﬁers for datasets with missing values  \nTolou Shadbahr  1,23, Michael Roberts  2,3,23✉ , Jan Stanczuk2,23, Julian Gilbey  2,23, Philip Teare3,23, Sören Dittmer2,4, Matthew Thorpe5, Ramon Viñas Torné6, Evis Sala  7, Pietro Lió  5, Mishal Patel3,8, Jacobus Preller  9, AIX-COVNET Collaboration*, James H. F. Rudd  10, Tuomas Mirtti  1,11,12, Antti Sakari Rannikko1,12,13, John A. D. Aston14, Jing Tang  1 & Carola-Bibiane Schönlieb2  \nAbstract  \nBackground Classifying samples in incomplete datasets is a common aim for machine learning practitioners, but is non-trivial. Missing data is found in most real-world datasets and these missing values are typically imputed using established methods, followed by classiﬁcation of the now complete samples. The focus of the machine learning researcher is tooptimise the classiﬁer’s performance.  \nMethods We utilise three simulated and three real-world clinical datasets with different feature types and missingness patterns. Initially, we evaluate how the downstream classiﬁer performance depends on the choice of classiﬁer and imputation methods. We employ ANOVA to quantitatively evaluate how the choice of missingness rate, imputation method, and classiﬁer method inﬂuences the performance. Additionally, we compare commonly used methods for assessing imputation quality and introduce a class of discrepancy scores based on the sliced Wasserstein distance. We also assess the stability of the imputations and the interpretability of model built on the imputed data.  \nResults The performance of the classiﬁer is most affected by the percentage of missingness in the test data, with a considerable performance decline observed as the test missingness rate increases. We also show that the commonly used measures for assessing imputation quality tend to lead to imputed data which poorly matches the underlying data distribution, whereas our new class of discrepancy scores performs much better on this measure. Furthermore, we show that the interpretability of classiﬁer models trained using poorly imputed data is compromised.  \nConclusions It is imperative to consider the quality of the imputation when performing downstream classiﬁcation as the effects on the classiﬁer can be considerable.  \nPlain language summary  \nMany artiﬁcial intelligence (AI) methods aim to classify samples of data into groups, e.g., patients with disease vs. those without. This often requires datasets to be complete, i. e., that all data has been collected for all samples. However, in clinical practice this is often not the case and some data can be missing. One solution is to ‘complete’ the dataset using a technique called imputation to replace those missing values. However, assessing how well the imputation method performs is challenging. In this work, we demonstrate why people should care about imputation, develop a new method for assessing imputation quality, and demonstrate that if we build AI models on poorly imputed data, the model can give different results to those we would hope for. Our ﬁndings may improve the utility and quality of AI models in the clinic.  \n1 Research Program in Systems Oncology, Faculty of Medicine, University of Helsinki, Helsinki, Finland. 2 Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Cambridge, UK. 3 Data Science & Artiﬁcial Intelligence, AstraZeneca, Cambridge, UK. 4 ZeTeM, University of Bremen, Bremen, Germany. 5 Department of Mathematics, University of Manchester, Manchester, UK. 6 Department of Computer Science and Technology, University of Cambridge, Cambridge, UK. 7 Department of Radiology, University of Cambridge, Cambridge, UK. 8 Clinical Pharmacology & Safety Sciences, AstraZeneca, Cambridge, UK. 9 Addenbrooke’s Hospital, Cambridge University Hospitals NHS Trust, Cambridge, UK. 10 Department of M","cbCair1UJfYPyY3m","https://ap.wps.com/l/cbCair1UJfYPyY3m","pdf",2329735,4,1,15,"English","en",105,"# Abstract\n## Background\n## Methods\n## Results\n## Conclusions\n# Plain language summary","[{\"question\":\"Why is imputation quality important for machine learning classification with missing values?\",\"answer\":\"Imputation directly affects downstream classifier performance. Poorly imputed data can cause performance declines and can compromise model interpretability.\"},{\"question\":\"How does the study evaluate the effect of missingness and methods?\",\"answer\":\"It uses three simulated and three real-world clinical datasets and applies ANOVA to quantify how missingness rate, imputation method, and classifier method influence performance.\"},{\"question\":\"What new idea does the paper introduce for assessing imputation quality?\",\"answer\":\"It introduces a class of discrepancy scores based on sliced Wasserstein distance, which better matches the underlying data distribution than commonly used measures.\"}]","The impact of imputation quality on machine learning classifiers for datasets with missing values | PDF",1785902530,38,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"the-impact-of-imputation-quality-on-machine-learning-classifiers-for-datasets-with-missing-values-126008","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/the-impact-of-imputation-quality-on-machine-learning-classifiers-for-datasets-with-missing-values-126008/126008/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-22","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is imputation quality important for machine learning classification with missing values?","Question",{"text":76,"@type":77},"Imputation directly affects downstream classifier performance. Poorly imputed data can cause performance declines and can compromise model interpretability.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the study evaluate the effect of missingness and methods?",{"text":81,"@type":77},"It uses three simulated and three real-world clinical datasets and applies ANOVA to quantify how missingness rate, imputation method, and classifier method influence performance.",{"name":83,"@type":74,"acceptedAnswer":84},"What new idea does the paper introduce for assessing imputation quality?",{"text":85,"@type":77},"It introduces a class of discrepancy scores based on sliced Wasserstein distance, which better matches the underlying data distribution than commonly used measures.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]