[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123140-en":3,"doc-seo-123140-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123140,7971461740886,"Theodore","https://ap-avatar.wpscdn.com/davatar_3d24733baf745e90a7e4bdd5f77d97b2",8,"Research & Report","Transparent Data Preprocessing for Machine Learning","Data preprocessing is a core step in machine learning that can materially improve model outcomes, yet assessing its true impact is often difficult. The document proposes a transparency system that logs preprocessing transformations and processed data using a Python library, then generates pipeline data summaries and change profiles for each step. A companion interface lets users interactively explore the implemented pipeline and understand how changes affect the data. The paper outlines an initial concept and discusses additional challenges and solutions for making preprocessing transparent.","Transparent Data Preprocessing for Machine Learning  \nSebastian Strasser  \n[sebastian.strasser@ur.de](sebastian.strasser@ur.de)[ ](sebastian.strasser@ur.de)University of Regensburg Regensburg, Germany  \nMeike Klettke  \n[meike.klettke@ur.de](meike.klettke@ur.de)[ ](meike.klettke@ur.de)University of Regensburg Regensburg, Germany  \nABSTRACT  \nData preprocessing is an important task in machine learning which can significantly improve model outcomes. However, evaluating the impact of data preprocessing is often difficult. There is a need for tools which make it transparent to the user on how certain transformations conducted in preprocessing affect the data. Thus, we propose a vision of a transparency system for data preprocessing that provides insights into data preparation pipelines. Our envisioned system consists of a Python library which enables users to log transformations and processed data. Subsequently, the system generates summaries of the data which was processed in the pipeline and so-called change profiles which capture the changes conducted in each processing step. These abstractions offer insight into the transformations and their effects on data. Additionally, the system includes an user interface where users can interactively discover the implemented pipeline and the changes made during preprocessing. This paper presents an initial concept of such a system. It also examines further challenges related to making preprocessing transparent and discusses potential solutions to address these challenges.  \nCCS CONCEPTS  \n• Information systems → Data cleaning; Extraction, transformation and loading; • General and reference → Evaluation; Validation.  \nKEYWORDS  \ndata preprocessing, data profiles, change profiles, transparency  \nACM Reference Format:  \nSebastian Strasser and Meike Klettke. 2024. Transparent Data Preprocessing for Machine Learning. In Workshop on Human-In-the-Loop Data Analytics (HILDA 24), June 14, 2024, Santiago, AA, Chile. ACM, New York, NY, USA, 6 pages. [https://doi.org/10.1145/3665939.3665960](https://doi.org/10.1145/3665939.3665960)  \n1 INTRODUCTION  \nData preprocessing is a tedious task which often takes a substantial percentage of time spent on data science projects. This is also due to it not being a streamlined process. It can better be explained as an iterative approach where the optimal processing steps and parameters are found based on trial-and-error. The aim of data preprocessing for machine learning is to find a data representation  \nThis work is licensed under a Creative Commons Attribution International 4.0 License.  \nHILDA 24 , June 14, 2024, Santiago, AA, Chile © 2024 Copyright held by the owner/author(s) .  \nACM ISBN 979-8-4007-0693-6/24/06  \n[https://doi.org/10.1145/3665939.3665960](https://doi.org/10.1145/3665939.3665960)  \nwhich yields the best result, i.e., the data format where the machine learning model can fit the input to the output in the best possible way. Therefore, data scientists conducting the experiments measure the effectiveness of their data preprocessing by drawing upon machine learning metrics yielded when evaluating the model.  \nThis approach has some drawbacks: as multiple data processing steps are performed on the data, it is not always clear which of them leads to an improvement or detoriation of model performance. Also, data preprocessing is only one factor influencing machine learning metrics, next to model and hyperparameter selection. Another problem is that especially non-experts cannot assess the effects of their preprocessing pipeline accurately. This is also due to the experimental nature of the model training process. Many decisions have to be made in a data science project, e.g., which preprocessing algorithms to use, what kind of machine learning model is suitable for the task, and what parameters are suitable for both preprocessing and machine learning algorithms. This can be overwhelming for unexperienced, but in some cases even for professional data s","cbCaitAKh9vN3T8D","https://ap.wps.com/l/cbCaitAKh9vN3T8D","pdf",690989,1,6,"English","en",105,"# Introduction\n## Motivation and evaluation challenges\n## Need for transparency in preprocessing\n# Proposed transparency system\n## Python library for logging and summaries\n## Change profiles and user interface","[{\"question\":\"Why is evaluating data preprocessing impact difficult in machine learning?\",\"answer\":\"Multiple preprocessing steps are applied and model quality depends on many interacting factors, so it is not always clear which step improves or harms performance. Non-experts may also struggle to assess pipeline effects because training is experimental and involves many decisions.\"},{\"question\":\"What does the proposed transparency system do for preprocessing pipelines?\",\"answer\":\"It logs transformations and processed data in a Python library, then produces summaries and change profiles that capture what each processing step changes. This information is presented to users to help them understand and assess the pipeline.\"},{\"question\":\"How does the system help detect issues like erroneous preprocessing bias?\",\"answer\":\"By making changes more explicit through logged transformations and step-level change profiles, users can observe when preprocessing introduces problems such as technical bias caused by imputation or other operations.\"}]","Transparent Data Preprocessing for Machine Learning | PDF",1785814841,15,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"transparent-data-preprocessing-for-machine-learning","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/transparent-data-preprocessing-for-machine-learning/123140/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is evaluating data preprocessing impact difficult in machine learning?","Question",{"text":75,"@type":76},"Multiple preprocessing steps are applied and model quality depends on many interacting factors, so it is not always clear which step improves or harms performance. Non-experts may also struggle to assess pipeline effects because training is experimental and involves many decisions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the proposed transparency system do for preprocessing pipelines?",{"text":80,"@type":76},"It logs transformations and processed data in a Python library, then produces summaries and change profiles that capture what each processing step changes. This information is presented to users to help them understand and assess the pipeline.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the system help detect issues like erroneous preprocessing bias?",{"text":84,"@type":76},"By making changes more explicit through logged transformations and step-level change profiles, users can observe when preprocessing introduces problems such as technical bias caused by imputation or other operations.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]