[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126339-en":3,"doc-seo-126339-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},126339,962085570644,"Evangeline","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","A method for missing values imputation of machine learning datasets - research","Machine-learning classification pipelines require missing-value handling during preprocessing to maintain reliable training and testing. The class center missing value imputation (CCMVI) method offers strong prediction accuracy with low computing cost, but it assumes class centers are known, limiting its use on test datasets where class labels must be treated as unknown. This study extends CCMVI to impute missing values in test instances using three strategies, including hybridization with literature methods and center-based nearest-label inference. Comparisons show improved accuracy near CCMVI+KNN/mean baselines while reducing time and memory.","A method for missing values imputation of machine learning  \ndatasets  \nYoussef Hanyf1,2, Hassan Silkan3  \n1Research laboratory in management and decision support, AI Data SEED team, Ibn Zohr University, Dakhla, Morocco  \n2 National School of Commerce and Management, Ibn Zohr University, Dakhla, Morocco 3LAROSERI laboratory, Faculty of Sciences, Chouaib Doukali University, EL Jadida, Morocco  \n\n| Article history:\u003Cbr>Received Apr 7, 2023 Revised Sep 30, 2023 Accepted Nov 15, 2023 | In machine learning applications, handling missing data is often required in the pre-processing phase of datasets to train and test models. The class center missing value imputation (CCMVI) is among the best imputation literature methods in terms of prediction accuracy and computing cost. The main drawback of this method is that it is inadequate for test datasets as long as it uses class centers to impute incomplete instances because their classes should be assumed as unknown in real-world classification situations. This work aims to extend the CCMVI method to handle missing values of test datasets. To this end, we propose three techniques: the first technique combines the CCMVI with other literature methods, the second technique imputes incomplete test instances based on their nearest class center, and the third technique uses the mean of centers of classes. The comparison of classification accuracies shows that the second and third proposed techniques ensure accuracy close to that of the combination of CCMVI with literature imputation methods, namely k-nearest neighbors (KNN) and mean methods. Moreover, they significantly decrease the time and memory space required for imputing test datasets.\u003Cbr>This is an open access article under the CC BY-SA license.\u003Cbr> |\n| --- | --- |\n| Keywords:\u003Cbr>Imputation\u003Cbr>Machine learning accuracy Imputation cost\u003Cbr>Missing data |  |\n\nCorresponding Author:  \nYoussef Hanyf  \nNational School of Commerce and Management of Dakhla, Ibn Zohr University Dakhla, Morocco  \n[Email: Youssef.hanyf@gmail.com](Email: Youssef.hanyf@gmail.com); [y.hanyf@uiz.ac.ma](y.hanyf@uiz.ac.ma)  \nArticle Info ABSTRACT  \n1. INTRODUCTION  \nIn the last decade, machine-learning classification methods have become increasingly required and used in various outstanding technologies such as health care [1], social media, and recommendation systems. In consequence, many machine-learning-related problems have attracted the attention of a large community of researchers. Handling missing data is one of the most severe problems of machine-learning classification because it significantly affects classification accuracy [2], [3] . Although the increasing development of data collection and acquisition technologies, various reasons can lead to losing values in datasets like the breakdown of devices, power cuts, and unanswered form questions [4] . Therefore, datasets often require a preprocessing phase to impute the missing values before training and testing classification models.  \nThe intuitive way to deal with missing data is the deletion of the features or instances containing missing values [5]–[7] . However, this method has risks of losing important information in datasets, and it can significantly impact classification accuracy. Many other methods have been used and proposed in the literature to impute missing data for increasing classification accuracy [8], [9] . These methods can be classified into two principal categories; statistical-based methods, such as mean/mod and least squares (LS), and machinelearning-based methods like k-nearest neighbors (KNN), neural networks (NN), and decision tree (DT) [10] .  \nHoque et al. [5] have compared the imputation accuracy of many machine-learning-based methods. They found that adaboost classifier and linear support vector machine (SVM) are better than logistic regression (LR), and random forest (RF) . But this study has been carried out only on one dataset. Thus, these results need to be validated in other datasets.","cbCaiklBvp7StL9R","https://ap.wps.com/l/cbCaiklBvp7StL9R","pdf",942121,4,1,11,"English","en",105,"# Introduction\n## Missing data challenges in ML\n## Imputation approaches: deletion, statistical, and ML-based\n## Trade-off: accuracy vs computational cost\n# Extending CCMVI for test datasets\n## Limitations of CCMVI for unknown test classes\n## Proposed techniques\n### Hybrid method with literature approaches\n### Nearest class center for incomplete test instances\n### Mean of class centers\n# Experimental comparison\n## Classification accuracy results\n## Time and memory efficiency","[{\"question\":\"Why is missing value imputation necessary in machine learning datasets?\",\"answer\":\"Missing values are common due to device breakdowns, power cuts, or unanswered form questions. Preprocessing imputation is needed to avoid degrading classification accuracy before training and testing models.\"},{\"question\":\"What is the main limitation of CCMVI for real-world testing?\",\"answer\":\"CCMVI imputes incomplete instances using class centers, which requires class labels to be known. In real-world use and in test datasets, classes should be treated as unknown.\"},{\"question\":\"How does the proposed approach improve CCMVI for test datasets?\",\"answer\":\"The work extends CCMVI with three techniques, including a hybrid approach with other imputation methods and two center-based strategies. Results indicate accuracy close to the CCMVI+KNN and CCMVI+mean combinations, with lower time and memory consumption for imputing test datasets.\"}]","A method for missing values imputation of machine learning datasets - research | PDF",1785904550,28,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"a-method-for-missing-values-imputation-of-machine-learning-datasets-research","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/a-method-for-missing-values-imputation-of-machine-learning-datasets-research/126339/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is missing value imputation necessary in machine learning datasets?","Question",{"text":76,"@type":77},"Missing values are common due to device breakdowns, power cuts, or unanswered form questions. Preprocessing imputation is needed to avoid degrading classification accuracy before training and testing models.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"What is the main limitation of CCMVI for real-world testing?",{"text":81,"@type":77},"CCMVI imputes incomplete instances using class centers, which requires class labels to be known. In real-world use and in test datasets, classes should be treated as unknown.",{"name":83,"@type":74,"acceptedAnswer":84},"How does the proposed approach improve CCMVI for test datasets?",{"text":85,"@type":77},"The work extends CCMVI with three techniques, including a hybrid approach with other imputation methods and two center-based strategies. Results indicate accuracy close to the CCMVI+KNN and CCMVI+mean combinations, with lower time and memory consumption for imputing test datasets.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":47,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":47,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":47,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":47,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]