[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119992-en":3,"doc-seo-119992-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119992,137441390410,"Hazel","https://ap-avatar.wpscdn.com/avatar/2000252f4ab5702993?_k=1776741390130283984",6,"Technology","Machine Learning Based Missing Data Imputation in Categorical Datasets - Ensemble Models with ECOC Framework","This research develops machine learning approaches for predicting and filling gaps in categorical datasets, focusing on ensemble models built with the Error Correction Output Codes (ECOC) framework. The study evaluates classifiers including SVM-based and KNN-based models, plus a hybrid approach combining SVM, KNN, and MLP. Experiments use three datasets—CPU, Hypothyroid, and Breast Cancer—under varying missing-data patterns. Results show ensemble ECOC models improve accuracy and robustness over single models, though deep learning still faces challenges such as large labeled-data needs and potential overfitting.","Received 21 May 2024, accepted 30 May 2024, date of publication 10 June 2024, date of  \ncurrent version 1 July 2024. Digital Object Identifier 10.1109/ACCESS.2024.3411817 Machine Learning Based Missing Data Imputation in Categorical Datasets MUHAMMADISHAQ1, ∗, SANA ZAHIR1, ∗, LAILA IFTIKHAR1, MOHAMMAD FARHAD BULBUL SEUNGMIN RHO 3,ANDMIYOUNGLEE 4,(Member, IEEE)  \n1Institute of Computer Sciences and Information Technology, The University of Agriculture at Peshawar, Peshawar, Khyber Pakhtunkhwa 25000, Pakistan 2Department of Mathematics, Jashore University of Science and Technology, Jashore 7408, Bangladesh 3Department of Industrial Security, Chung-Ang University, Seoul 06974, South Korea 4Department of Research, Chung-Ang University, Seoul 06974, South Korea Corresponding author: Mi Young Lee ( [miylee@cau.ac.kr](miylee@cau.ac.kr)) 2, This work was supported by the Basic Science Research Program through the National Research Foundation of Korea (NRF) funded by the Ministry of Education under Grant 2021R1I1A1A01055652 .  \n∗Muhammad Ishaq and Sana Zahir contributed equally to this work.  \nABSTRACT In order to predict and fill in the gaps in categorical datasets, this research looked into the use of machine learning algorithms. The emphasis was on ensemble models constructed using the Error Correction Output Codes (ECOC) framework, including models based on SVM and KNN as well as a hybrid classifier that combines models based on SVM, KNN, and MLP. Three diverse datasets—the CPU, Hypothyroid, and Breast Cancer datasets—were employed to validate these algorithms. Results indicated that these machine learning techniques provided substantial performance in predicting and completing missing data, with the effectiveness varying based on the specific dataset and missing data pattern. Compared to solo models, ensemble models that made use of the ECOC framework significantly improved prediction accuracy and robustness. Deep learning for missing data imputation has obstacles despite these encouraging results, including the requirement for large amounts of labeled data and the possibility of overfitting. Subsequent research endeavors ought to evaluate the feasibility and efficacy of deep learning algorithms in the context of the imputation of missing data.  \nINDEXTERMS Data cleansing, missing data imputation, classification, regression and categorical datasets.  \nI. INTRODUCTION‘‘Dirty data’’ describes unprocessed or inconsistent, erro neous, or incomplete raw data that has been tampered with. High-quality data is always the foundation for quality decisions. The conclusions drawn from analytical results derived from dirty data are untrustworthy. Consequently, raw data must first be cleaned before being utilized in any analytical process. It is not possible to use raw data directly in  \nanalytical methods. Data cleaning is an important part of information quality management. It aims to enhance the overall quality of data by locating and removing errors, omissions, and inconsistencies. This section provides an overview of the proposed technique and an introduction to its theoretical foundations [1] . The associate editor coordinating the review of this manuscript and approving it for publication was Chun-Wei Tsai . 88332 As a result, preprocessing is required, as illustrated in Figure 1, before machine learning models can be trained or run on raw data. Even though it is necessary and inevitable, data preprocessing is a time-consuming and frustrating procedure. According to industry standards, data scientists typically devote more than half of their analysis time to this task. On the other hand, those who used the software in work were not experts in it [2] . Because of this, data scientists are in high demand for a tool that will assist them in automating the process [3] . Data preprocessing encompasses various tasks such as data cleaning, data integration, and data transformation [3] . It confronts common data challenges like outliers, lost or","cbCaiqFF6R5WRHiN","https://ap.wps.com/l/cbCaiqFF6R5WRHiN","pdf",142174,1,20,"English","en",105,"# Abstract\n# Introduction\n## Dirty data and the need for preprocessing\n## Data cleaning tasks and challenges\n# Methods and models (ECOC ensemble approach)\n# Experimental datasets and evaluation\n# Results and discussion\n# Conclusion and future work","[{\"question\":\"What problem does the paper address?\",\"answer\":\"The paper addresses how to predict and impute missing values in categorical datasets so gaps can be completed for reliable analysis.\"},{\"question\":\"Which modeling strategy is emphasized for imputation?\",\"answer\":\"It emphasizes ensemble models using the Error Correction Output Codes (ECOC) framework, including SVM- and KNN-based ensembles and a hybrid SVM-KNN-MLP classifier.\"},{\"question\":\"How do the results compare between ensemble and single models?\",\"answer\":\"Ensemble ECOC models significantly improve prediction accuracy and robustness compared with solo models, with effectiveness depending on the dataset and missing-data pattern.\"}]","Machine Learning Based Missing Data Imputation in Categorical Datasets - Ensemble Models with ECOC Framework | PDF",1785727548,50,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"machine-learning-based-missing-data-imputation-in-categorical-datasets-ensemble-models-with-ecoc-framework","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/technology/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/machine-learning-based-missing-data-imputation-in-categorical-datasets-ensemble-models-with-ecoc-framework/119992/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address?","Question",{"text":75,"@type":76},"The paper addresses how to predict and impute missing values in categorical datasets so gaps can be completed for reliable analysis.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which modeling strategy is emphasized for imputation?",{"text":80,"@type":76},"It emphasizes ensemble models using the Error Correction Output Codes (ECOC) framework, including SVM- and KNN-based ensembles and a hybrid SVM-KNN-MLP classifier.",{"name":82,"@type":73,"acceptedAnswer":83},"How do the results compare between ensemble and single models?",{"text":84,"@type":76},"Ensemble ECOC models significantly improve prediction accuracy and robustness compared with solo models, with effectiveness depending on the dataset and missing-data pattern.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,112,117,122,126,129,133],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":111},"technology",{"id":113,"doc_module":4,"doc_module_name":46,"category_name":114,"show_sort_weight":115,"slug":116},7,"Healthcare",40,"healthcare",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":119,"show_sort_weight":120,"slug":121},8,"Research & Report",30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":21,"slug":125},9,"Religion & Spirituality","religion-spirituality",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":127,"show_sort_weight":21,"slug":128},"World Cup","world-cup",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":131,"show_sort_weight":130,"slug":132},10,"Lifestyle","lifestyle",{"id":134,"doc_module":4,"doc_module_name":46,"category_name":135,"show_sort_weight":106,"slug":136},19,"General","general"]