[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119199-en":3,"doc-seo-119199-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119199,1099514068035,"Ezra","https://ap-avatar.wpscdn.com/davatar_276721f389ce27ea32af1340a28f341c",8,"Research & Report","Collaborative Data Cleaning Framework - a Pilot Case Study for Machine Learning Development","This study experiments with collaborative data cleaning as a pivotal step in data preparation for both analysis and machine learning development. A provenance Data Cleaning Model (DCM) is employed to support multi-user scenarios, tracking dataset changes while running experiments that simulate multiple data curators collaborating on the same data. The work evaluates how distinct cleaning scenarios influence dataset completeness and correctness metrics, and how those quality changes propagate into downstream machine learning modeling performance.","IJDC | Conference Paper  \nCollaborative Data Cleaning Framework: a Pilot Case Study for Machine Learning Development  \nNikolaus Nova Parulian University of Illinois at Urbana Champaign  \nBertram Ludäscher University of Illinois at Urbana Champaign  \nAbstract  \nThis study experiments with collaborative data cleaning, a pivotal phase in data preparation for both analysis and machine learning. We used a provenance Data Cleaning Model (DCM) for multi-user scenarios to track changes on a dataset and conduct comprehensive experiments that simulate multiple data curators working collaboratively on a dataset. Furthermore, we analyzed how different data-cleaning scenarios to improve quality metrics of completeness and correctness of a dataset can affect the downstream machine learning modeling performance.  \nSubmitted 10 February 2024 ~ Accepted 22 February 2024  \nCorrespondence should be addressed to Nikolas Parulian, [Email: nnp2@illinois.edu](Email: nnp2@illinois.edu)  \nThis paper was presented at the International Digital Curation Conference IDCC25, 19-21 February 2024 The International Journal of Digital Curation is an international journal committed to scholarly excellence and dedicated to the advancement of digital curation across a wide range of sectors. The IJDC is published by the University of  \nEdinburgh on behalf of the Digital Curation Centre. ISSN: 1746-8256. URL: [http://www.ijdc.net/](http://www.ijdc.net/)  \nCopyright rests with the authors. This work is released under a Creative Commons Attribution License, version 4.0. For details please see [https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/)  \nInternational Journal of Digital Curation 2024, Vol. 18, Iss. 1, pp. 12  \n1 [http://dx.doi.org/10.2218/ijdc.v18i1.924](http://dx.doi.org/10.2218/ijdc.v18i1.924)[ ](http://dx.doi.org/10.2218/ijdc.v18i1.924)DOI: 10.2218/ijdc.v 18i1.924  \nIntroduction and Overview  \nData cleaning is a critical step the data preparation and aims to ensure the accuracy, consistency, and integrity of data used for analysis or machine learning. However, data cleaning is time-consuming and labor-intensive, often requiring domain expertise and manual effort. With the increasing availability of complex datasets, collaborative data cleaning has emerged as a promising approach to leverage multiple data curators’collective knowledge and skills to clean data more efficiently and effectively.  \nBuilding upon our previous work on providing a provenance model for Collaborative Data Cleaning (CDCM) (Parulian et al., 2021), (Parulian and Ludäscher, 2022), in this work we describe the use case and an example of implementation of the model for a machine learning development. We aim to address the challenges and opportunities associated with collaborative data cleaning and investigate how different data-cleaning scenarios can impact the cleaned datasets and outcomes for the downstream tasks—increasing the transparency by recording provenance on each cleaning step and analyzing the effect on the data quality (DQ) metrics. To this end, we propose an experiment that involves applying different variations of datacleaning steps to simulate multiple curators working together on cleaning a dataset. Our provenance model can also support reporting queries for multi-curator scenarios such as: Who changed the values in the dataset? ; Who contributed to the data-cleaning workflow development?  \nThe findings from our experiment use case contribute to our understanding of practical use cases for collaborative data cleaning. The outcomes ofour analysis shed light on how a collaborative data cleaning process can impact the quality and reliability of the cleaned data and how it can influence downstream data analysis tasks or modeling outcomes. This study aims to advance the use of provenance for trust and transparency in collaborative data cleaning and provide valuable insights for data curators working on machine learning development.  \nBac","cbCaidNd7pZeyVSM","https://ap.wps.com/l/cbCaidNd7pZeyVSM","pdf",897682,1,15,"English","en",105,"# Abstract\n# Introduction and Overview\n## Background and Motivation\n# Machine Learning Pipeline and Data Cleaning","[{\"question\":\"What is the main goal of the collaborative data cleaning framework in this paper?\",\"answer\":\"The paper aims to study collaborative data cleaning as a key phase in preparing data for analysis and machine learning, using a provenance model to improve transparency and understand impacts on results.\"},{\"question\":\"How does the provenance Data Cleaning Model (DCM) support multi-user collaboration?\",\"answer\":\"The DCM tracks changes at the dataset level across multiple curators and enables experiments that simulate collaborative cleaning workflows.\"},{\"question\":\"Which dataset quality aspects are evaluated, and why do they matter for machine learning?\",\"answer\":\"The study analyzes completeness and correctness metrics, and examines how improvements or differences in these metrics affect downstream machine learning model performance.\"}]","Collaborative Data Cleaning Framework - a Pilot Case Study for Machine Learning Development | PDF",1785723055,38,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"collaborative-data-cleaning-framework-a-pilot-case-study-for-machine-learning-development","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/collaborative-data-cleaning-framework-a-pilot-case-study-for-machine-learning-development/119199/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What is the main goal of the collaborative data cleaning framework in this paper?","Question",{"text":75,"@type":76},"The paper aims to study collaborative data cleaning as a key phase in preparing data for analysis and machine learning, using a provenance model to improve transparency and understand impacts on results.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the provenance Data Cleaning Model (DCM) support multi-user collaboration?",{"text":80,"@type":76},"The DCM tracks changes at the dataset level across multiple curators and enables experiments that simulate collaborative cleaning workflows.",{"name":82,"@type":73,"acceptedAnswer":83},"Which dataset quality aspects are evaluated, and why do they matter for machine learning?",{"text":84,"@type":76},"The study analyzes completeness and correctness metrics, and examines how improvements or differences in these metrics affect downstream machine learning model performance.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]