[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123510-en":3,"doc-seo-123510-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123510,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Common Steps in Machine Learning Might Hinder The Explainability Aims in Medicine - Paper","Data pre-processing improves machine-learning performance and reduces runtime, but it can weaken explainability and interpretability when applied without careful consideration in medicine. The paper covers common steps such as missing-value handling, outlier detection and removal, normalization and standardization, dimensionality reduction, data augmentation for imbalanced data, and confounding-variable treatment. Misapplied preprocessing may prevent new findings, reduce fairness across groups, distort clinically meaningful features, and turn them unitless or non-explainable, and it offers possible solutions that preserve both accuracy and XAI goals.","arXiv :2409 .00155v1 [ cs .LG] 30 Aug 2024  \nCommon Steps in Machine Learning Might Hinder The Explainability Aims in Medicine  \nAhmed M Salih 1,2,3,4  \n1Department of Population Health Sciences, University of Leicester, University Rd, LE1 7RH, Leicester, UK  \n2William Harvey Research Institute, NIHR Barts Biomedical Research Centre, Queen Mary University of  \nLondon, Charterhouse Square, London, EC1M 6BQ, London, UK  \n3Barts Heart Centre, St Bartholomew’s Hospital, Barts Health NHS Trust, West Smithfield, London, EC1A 7BE, UK  \n4Department of Computer Science, University of Zakho, Duhok road, Zakho, Kurdistan, Iraq  \nAbstract  \nData pre-processing is a significant step in machine learning to improve the performance of the model and decreases the running time. This might include dealing with missing values, outliers detection and removing, data augmentation, dimensionality reduction, data normalization and handling the impact of confounding variables.  \nAlthough it is found the steps improve the accuracy of the model, but they might hinder the explainability of the model if they are not carefully considered especially in medicine. They might block new findings when missing values and outliers removal are implemented inappropriately. In addition, they might make the model unfair against all the groups in the model when making the decision. Moreover, they turn the features into unitless and clinically meaningless and consequently not explainable. This paper discusses the common steps of the data preprocessing in machine learning and their impacts on the explainability and interpretability of the model. Finally, the paper discusses some possible solutions that improve the performance of the model while not decreasing its explainability.  \nKeywords—Pre-processing, XAI, medicine  \n1 Introduction  \nData pre-processing is an initial indispensable step in machine learning and data science. It consists of several components and processing steps that are performed on raw data to ensure its quality before fitting them to any model. It might involves dealing with missing values, outliers detection and removal, normalization and standardization, dimensionality reduction (e.g., feature selection, principal component analysis), data augmentation of imbalanced data, and dealing with confounding variables [1] . All these steps and more are conducted to ensure the quality of the data which might improve the performance of the model and decrease the running time.  \nOn the other hand, explainable artificial intelligence (XAI) as an emerging topic aims to understand how the model works. Its direct aims are more to do with improving the interpretability of the model than improving the performance of the model. It involves quantifying the uncertainty of a machine learning model when making a decision. Moreover, it helps to identify and highlight the most informative features (pixels in an image) in the model that drive its outcome [2] . Furthermore, it aims to improve model fairness toward all groups (e.g., sex, ethnicity, unemployment, etc) in the model. The ultimate aim of the XAI is to make the model more transparent and consequently trustfully.  \nData pre-processing and XAI should not hinder each other. They should rather work together to improve the performance of the model and its explainability simultaneously. In other words, any step or steps to improve the performance of the model should not affect its explainability negatively. However, data-preprocessing might hinder the aims of the XAI and eventually its interpretability. This include how to deal with missing values and which method should be considered for imputation. Moreover, outliers should not be ignored and removed from the model in the medicine domain as they might represent a new case or at least should be explained why they are with extreme values. In addition, dimensionality reduction methods might lead to remove clinical significant features or transfer them into unitless whic","cbCailgArZy66TlA","https://ap.wps.com/l/cbCailgArZy66TlA","pdf",967878,1,11,"English","en",105,"# Introduction\n## Data preprocessing steps\n### Missing values","[{\"question\":\"为什么数据预处理在机器学习中很关键？\",\"answer\":\"数据预处理用于提升数据质量并为模型训练做准备，通常包括缺失值处理、异常值检测/移除、归一化与标准化、降维、数据增强以及混杂变量处理等，从而提高性能并降低运行时间。\"},{\"question\":\"哪些预处理操作可能会削弱医疗场景下的可解释性（XAI）？\",\"answer\":\"不恰当的缺失值删除或插补、异常值移除可能阻断新的发现；降维可能移除临床关键信息或使特征变得难以解释；归一化与标准化可能让特征失去临床意义。\"},{\"question\":\"论文如何强调预处理与XAI应当如何协同？\",\"answer\":\"论文主张数据预处理不应与XAI目标相互阻碍，而应共同提升模型性能与可解释性。通过选择合适的处理方法与插补/异常值策略，并注意数据结构与混杂变量处理的正负影响，可以在不降低可解释性的前提下改进结果。\"}]","Common Steps in Machine Learning Might Hinder The Explainability Aims in Medicine - Paper | PDF",1785816974,28,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"common-steps-in-machine-learning-might-hinder-the-explainability-aims-in-medicine-paper","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/common-steps-in-machine-learning-might-hinder-the-explainability-aims-in-medicine-paper/123510/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"为什么数据预处理在机器学习中很关键？","Question",{"text":75,"@type":76},"数据预处理用于提升数据质量并为模型训练做准备，通常包括缺失值处理、异常值检测/移除、归一化与标准化、降维、数据增强以及混杂变量处理等，从而提高性能并降低运行时间。","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"哪些预处理操作可能会削弱医疗场景下的可解释性（XAI）？",{"text":80,"@type":76},"不恰当的缺失值删除或插补、异常值移除可能阻断新的发现；降维可能移除临床关键信息或使特征变得难以解释；归一化与标准化可能让特征失去临床意义。",{"name":82,"@type":73,"acceptedAnswer":83},"论文如何强调预处理与XAI应当如何协同？",{"text":84,"@type":76},"论文主张数据预处理不应与XAI目标相互阻碍，而应共同提升模型性能与可解释性。通过选择合适的处理方法与插补/异常值策略，并注意数据结构与混杂变量处理的正负影响，可以在不降低可解释性的前提下改进结果。","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]