[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125966-en":3,"doc-seo-125966-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":11,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},125966,687207024478,"Liam","https://ap-avatar.wpscdn.com/davatar_a8503ba1806abce46bf441b54a3ca4cd",8,"Research & Report","Data Cleaning and Machine Learning - A Systematic Literature Review","Machine Learning (ML) is increasingly embedded in real-world systems, and model performance strongly depends on training data quality. Growing research interest targets approaches to detect and repair data errors through data cleaning, while also exploring how ML itself can perform data cleaning. This work conducts a systematic literature review (2016–2022) to summarize approaches in both directions and deliver future research recommendations.","arXiv :2310 .01765v2 [ cs .LG] 31 May 2024  \nAutomated Software Engineering manuscript No.  \n(will be inserted by the editor)  \nData Cleaning and Machine Learning: A Systematic Literature Review  \nPierre-Olivier Côté 1 · Amin Nikanjam 1 · Nafisa Ahmed 1 · Dmytro Humeniuk 1 ·  \nFoutse Khomh 1  \nReceived: date / Accepted: date  \nAbstract Context: Machine Learning (ML) is integrated into a growing number of systems for various applications. Because the performance of an ML model is highly dependent on the quality of the data it has been trained on, there is a growing interest in approaches to detect and repair data errors (i.e. , data cleaning) . Researchers are also exploring how ML can be used for data cleaning; hence creating a dual relationship between ML and data cleaning. To the best of our knowledge, there is no study that comprehensively reviews this relationship. Objective: This paper’s objectives are twofold. First, it aims to summarize the latest approaches for data cleaning for ML and ML for data cleaning. Second, it provides future work recommendations. Method: We conduct a systematic literature review of the papers published between 2016 and 2022 inclusively. We identify different types of data cleaning activities with and for ML: feature cleaning, label cleaning, entity matching, outlier detection, imputation, and holistic data cleaning. Results: We summarize the content of 101 papers covering various data cleaning activities and provide 24 future work recommendations. Our review highlights many promising data cleaning techniques that can be further extended. Conclusion: We believe that our review of the literature will help the community develop better approaches to clean data.  \nKeywords Machine Learning, Data Cleaning, Systematic Literature Review, Survey, Taxonomy  \nThis work is funded by the Fonds de Recherche du Quebec (FRQ), the Canadian Institute for Advanced Research (CIFAR), and the Natural Sciences and Engineering Research Council of Canada (NSERC) .  \n1 Polytechnique Montréal, Québec, Canada  \nE-mail: {pierre-olivier.cote, amin.nikanjam, [nafisa.abdelmutalab-ali-ahmed@polymtl.ca](nafisa.abdelmutalab-ali-ahmed@polymtl.ca), [dmytro.humeniuk@polymtl.ca](dmytro.humeniuk@polymtl.ca), [foutse.khomh}@polymtl.ca](foutse.khomh}@polymtl.ca)  \n1 Introduction  \nNowadays, Machine Learning (ML) integrates a growing number of industries, from transportation to healthcare and education. Recent applications of ML achieved performances similar to humans’ performance on complex tasks, such as taking the bar exam (OpenAI, 2023) or driving a car (Badue et al. , 2021; Gitnux, 2023) . ML models are implemented as software components, and then integrated into other components in Machine Learning Software Systems (MLSSs) . Such systems employ trained ML models to intelligently make decisions or generate output based on learned data-derived knowledge, and similar to any software systems, they need quality assurance (Gezici and Tarhan, 2022) .  \nBehind much of the recent success of ML applications are large amounts of training data and powerful computing infrastructure (Roh et al., 2019; Whanget al., 2021) . While a major part of the research in ML is spent on developing better modeling techniques (Ng, 2021), data preparation often is the most arduous and time-consuming task for practitioners. Indeed, researchers have reported that data scientists sometimes spend over 80% of their time preparing data (Whang et al., 2021; Neutatz et al., 2021; Deng et al., 2017; Agrawalet al., 2019) . As a consequence, there is an emerging trend, referred to as Data-Centric AI (DCAI) to steer research to focus on improving datasets instead of models for ML problems. In the last few years, various initiatives aimed at stimulating involvement in the DCAI trend have been proposed. For example, in August 2021 the first DCAI competition (Ng et al., 2021) was organized. During this competition, participants were challenged to improve a model’s performan","cbCaiuwuErz3s3vD","https://ap.wps.com/l/cbCaiuwuErz3s3vD","pdf",1739345,1,82,"English","en",105,"# Abstract\n## Context and motivation\n## Objective\n## Method\n## Results and conclusion","[{\"question\":\"What is the core relationship investigated in the review?\",\"answer\":\"The review focuses on the dual relationship between data cleaning for ML and using ML to clean data, highlighting how each side can improve the other.\"},{\"question\":\"Which time window and study method does the paper use?\",\"answer\":\"It performs a systematic literature review of papers published from 2016 to 2022 inclusive, following a strict and rigorous methodology.\"},{\"question\":\"What types of data cleaning activities are identified?\",\"answer\":\"The review identifies feature cleaning, label cleaning, entity matching, outlier detection, imputation, and holistic data cleaning for and with ML.\"}]","Data Cleaning and Machine Learning - A Systematic Literature Review | PDF",1785902276,207,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"data-cleaning-and-machine-learning-a-systematic-literature-review","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/data-cleaning-and-machine-learning-a-systematic-literature-review/125966/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-24","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":11},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What is the core relationship investigated in the review?","Question",{"text":76,"@type":77},"The review focuses on the dual relationship between data cleaning for ML and using ML to clean data, highlighting how each side can improve the other.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which time window and study method does the paper use?",{"text":81,"@type":77},"It performs a systematic literature review of papers published from 2016 to 2022 inclusive, following a strict and rigorous methodology.",{"name":83,"@type":74,"acceptedAnswer":84},"What types of data cleaning activities are identified?",{"text":85,"@type":77},"The review identifies feature cleaning, label cleaning, entity matching, outlier detection, imputation, and holistic data cleaning for and with ML.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]