[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-124669-en":3,"doc-seo-124669-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},124669,962075006959,"Anda","https://ap-avatar.wpscdn.com/avatar/e0002397efbe92a78e?_k=1776741047341049297",8,"Research & Report","The Impact of Imputation Quality on Machine Learning Classifiers for Datasets with Missing Values - Research findings and evaluation methods","Machine-learning classification on incomplete datasets requires imputation, yet imputation quality strongly shapes downstream predictive behavior. Using three simulated and three real-world clinical datasets covering distinct feature types and missingness patterns, the work evaluates how classifier choice and imputation strategy interact. ANOVA quantifies the influence of missingness rate, imputation method, and classifier method. Common imputation-quality metrics are benchmarked against new class-discrepancy scores based on sliced Wasserstein distance, with additional checks on imputation stability and model interpretability. Results show test-missingness percentage drives performance decline and poorly imputed data harms interpretability.","The Impact of Imputation Quality on Machine  \nLearning Classifiers for Datasets with Missing Values Tolou Shadbahr 1,+ , Michael Roberts2,3,+,* , Jan S˜ tanczuk2,+ , Julian Gilbey2,+ , Philip Teare3,+ ,  \nSren Dittmer2,4 , Matthew Thorpe5 , Ramon Vinas Torn6 , Evis Sala7 , Pietro Li5 , Mishal Patel3,8 , Jacobus Preller9 , AIX-COVNET Collaboration‡, James H. F. Rudd 10 , Tuomas Mirtti 1,11,12 , Antti Sakari Rannikko 1,12,13 , John A. D. Aston 14 , Jing Tang 1 , and  \nCarola-Bibiane Schnlieb2  \n1 Research Program in Systems Oncology, Faculty of Medicine, University of Helsinki, Helsinki, Finland  \n2 Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Cambridge, UK  \n3 Data Science & Artificial Intelligence, AstraZeneca, Cambridge, UK  \n4 ZeTeM, University of Bremen, Bremen, Germany  \n5 Department of Mathematics, University of Manchester, Manchester, UK  \n6 Department of Computer Science and Technology, University of Cambridge, Cambridge, UK  \n7 Department of Radiology, University of Cambridge, Cambridge, UK  \n8 Clinical Pharmacology & Safety Sciences, AstraZeneca, Cambridge, UK  \n9Addenbrooke’s Hospital, Cambridge University Hospitals NHS Trust, Cambridge, UK  \n10 Department of Medicine, University of Cambridge, Cambridge, UK  \n11 Department of Pathology, University of Helsinki and Helsinki University Hospital, Helsinki, Finland.  \n12 iCAN-Digital Precision Cancer Medicine Flagship, Helsinki, Finland.  \n13 Department of Urology, University of Helsinki and Helsinki University Hospital, Helsinki, Finland  \n14 Department of Pure Mathematics and Mathematical Statistics, University of Cambridge, Cambridge, UK ‡A list of authors and their affiliations appears at the end of the paper  \n* corresponding author, [email: michael.roberts@maths.cam.ac.uk](email: michael.roberts@maths.cam.ac.uk)  \n+these authors contributed equally to this work  \nABSTRACT  \nBackground  \nClassifying samples in incomplete datasets is a common aim for machine learning practitioners, but is non-trivial. Missing data is found in most real-world datasets and these missing values are typically imputed using established methods, followed by classification of the now complete samples. The focus of the machine learning researcher is to optimise the classifier’s performance.  \nMethods  \nWe utilise three simulated and three real-world clinical datasets with different feature types and missingness patterns. Initially, we evaluate how the downstream classifier performance depends on the choice of classifier and imputation methods. We employ ANOVA to quantitatively evaluate how the choice of missingness rate, imputation method, and classifier method influences the performance. Additionally, we compare commonly used methods for assessing imputation quality and introduce a class of discrepancy scores based on the sliced Wasserstein distance. We also assess the stability of the imputations and the interpretability of model built on the imputed data.  \nResults  \nThe performance of the classifier is most affected by the percentage of missingness in the test data, with a considerable performance decline observed as the test missingness rate increases. We also show that the commonly used measures for assessing imputation quality tend to lead to imputed data which poorly matches the underlying data distribution, whereas our new class of discrepancy scores performs much better on this measure. Furthermore, we show that the interpretability ofclassifier models trained using poorly imputed data is compromised.  \nConclusions  \nIt is imperative to consider the quality of the imputation when performing downstream classification as the effects on the classifier can be considerable.  \nPlain Language Summary  \nMany artificial intelligence (AI) methods aim to classify samples of data into groups, e.g. patients with disease vs. those without. This often requires datasets to be complete, i.e. that all data has been collected for all samples. However, in clinical","cbCaibW7B2ZIvm4u","https://ap.wps.com/l/cbCaibW7B2ZIvm4u","pdf",216096,1,15,"English","en",105,"# Abstract\n## Background\n## Methods\n## Results\n## Conclusions\n# Plain Language Summary\n# Introduction","[{\"question\":\"Why is imputation quality important for machine learning classifiers with missing values?\",\"answer\":\"Imputation quality can substantially affect classifier performance and can also compromise the interpretability of models trained on imputed data.\"},{\"question\":\"How did the study evaluate the effect of imputation and classifier choices?\",\"answer\":\"It used three simulated and three real-world clinical datasets, then applied ANOVA to quantify how missingness rate, imputation method, and classifier method influence performance.\"},{\"question\":\"What improvement does the paper propose for assessing imputation quality?\",\"answer\":\"It introduces class discrepancy scores based on the sliced Wasserstein distance, which better match the underlying data distribution than commonly used imputation-quality measures.\"}]","The Impact of Imputation Quality on Machine Learning Classifiers for Datasets with Missing Values - Research findings and evaluation methods | PDF",1785893826,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"the-impact-of-imputation-quality-on-machine-learning-classifiers-for-datasets-with-missing-values-research-findings-and-evaluation-methods","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/the-impact-of-imputation-quality-on-machine-learning-classifiers-for-datasets-with-missing-values-research-findings-and-evaluation-methods/124669/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is imputation quality important for machine learning classifiers with missing values?","Question",{"text":75,"@type":76},"Imputation quality can substantially affect classifier performance and can also compromise the interpretability of models trained on imputed data.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How did the study evaluate the effect of imputation and classifier choices?",{"text":80,"@type":76},"It used three simulated and three real-world clinical datasets, then applied ANOVA to quantify how missingness rate, imputation method, and classifier method influence performance.",{"name":82,"@type":73,"acceptedAnswer":83},"What improvement does the paper propose for assessing imputation quality?",{"text":84,"@type":76},"It introduces class discrepancy scores based on the sliced Wasserstein distance, which better match the underlying data distribution than commonly used imputation-quality measures.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]