[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-116995-en":3,"doc-seo-116995-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},116995,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","Imputation Techniques in Machine Learning - A Survey","Machine learning relies on complete, reliable data, yet missing values commonly appear in real-world datasets due to collection and management errors, intentional omissions, or human mistakes. Since most models cannot process missingness directly, data imputation becomes a prerequisite to maintain distributional integrity and protect predictive performance. This survey examines widely used imputation strategies, including mean, median, K-nearest neighbors (KNN) imputation, linear regression, missForest, and MICE, and discusses how selection impacts model accuracy.","| Imputation Techniques in Machine Learning – A\u003Cbr>Survey\u003Cbr>Y. Angeline Christobel1, R. Jaya Suji2, J. Jeya A Celin3\u003Cbr>1Dean, School of Computational Studies Hindustan College of Arts & Science\u003Cbr>Chennai-603103\u003Cbr>[angelinechristobel5@gmail.com](angelinechristobel5@gmail.com)\u003Cbr>2Assistant Professor, Department of Computer Science\u003Cbr>Hindustan College of Arts & Science\u003Cbr>Chennai-603103\u003Cbr>[jayasuji1981@gmail.com](jayasuji1981@gmail.com)\u003Cbr>3Professor, Department of Information Technology\u003Cbr>Kalasalingam Academy of Research and Education\u003Cbr>Krishnankoil-626126\u003Cbr>[jjeyacelin@gmail.com](jjeyacelin@gmail.com)\u003Cbr>Abstract—Machine learning plays a pivotal role in data analysis and information extraction. However, one common challenge encountered in this process is dealing with missing values. Missing data can find its way into datasets for a variety of reasons. It can result from errors during data collection and management, intentional omissions, or even human errors. It's important to note that most machine learning models are not designed to handle missing values directly. Consequently, it becomes essential to perform data imputation before feeding the data into a machine learning model. Multiple techniques are available for imputing missing values, and the choice of technique should be made judiciously, considering various parameters. An inappropriate choice can disrupt the overall distribution of data values and subsequently impact the model's performance. In this paper, various imputation methods, including Mean, Median, K-nearest neighbors (KNN)-based imputation, Linear Regression, Miss Forest, and MICE are examined.\u003Cbr>Keywords-Missing data, Imputation, Machine learning. |\n| --- |\n|  |\n\nI. INTRODUCTION  \nThe issue of missing values is a widespread challenge encountered across various domains that work with data. It can give rise to a range of problems, including reduced performance, complications in data analysis, and biased results stemming from disparities between missing and complete data. Additionally, the severity of the missing data problem depends on several factors, including the extent of missing data, the missing data pattern, the fundamental mechanism behind data absence, and the underlying mechanism driving the data's missingness. Many studies have been conducted for imputing missing values. Anil Jadhav et al. conducted a comparison of seven data imputation techniques, which included mean imputation, median imputation, kNN imputation, predictive mean matching, Bayesian Linear Regression (norm), Linear Regression, non-Bayesian (norm.nob), and random sample imputation. The findings of their analysis revealed that the kNN imputation method exhibited superior performance when compared to the other methods [1] . According to the findings of Ahmad R Alsaber et [al. in](al. in) [2], the Missing at Random (MAR) technique demonstrated the lowest Root Mean Square Error (RMSE) and Mean Absolute Error (MAE) . Furthermore, the Multiple Imputation (MI) method, specifically employing  \nthe missForest approach, exhibited a high level of accuracy in estimating missing values. Pooja Rani et al. [3], evaluated four imputation techniques—k-nearest neighbor (KNN), multivariate imputation by chained equations (MICE), mean imputation, and mode imputation—utilizing four different classifiers, namely Naive Bayes (NB), support vector machine (SVM), logistic regression (LR), and random forest (RF) . The objective of their study was to compare the root mean square error (RMSE) of these classifiers and identify the most effective imputation method. The results indicate that the MICE imputation method outperformed the other imputation techniques in this context. Vikesh Kumar Gond et al. [4] discussed imputation techniques and compared the merits and drawbacks. Emmanuel et al. [5], introduced and assessed two distinct methods: the k-nearest neighbor approach and an iterative imputation technique known as \"missForest,\" which harnesses the ","cbCaii12T28t9LZT","https://ap.wps.com/l/cbCaii12T28t9LZT","pdf",160519,1,5,"English","en",105,"# Introduction\n## Background on Missing Values\n## Overview of Prior Imputation Studies","[{\"question\":\"Why do missing values cause problems in machine learning workflows?\",\"answer\":\"Missing values can reduce model performance, complicate data analysis, and introduce bias by creating discrepancies between missing and complete data.\"},{\"question\":\"Why is data imputation necessary before training many machine learning models?\",\"answer\":\"Most machine learning models are not built to handle missing values directly, so imputation is needed to supply usable inputs.\"},{\"question\":\"Which imputation techniques are reviewed in the survey?\",\"answer\":\"The paper reviews mean, median, K-nearest neighbors (KNN) imputation, linear regression, missForest, and MICE, comparing their effectiveness and suitability.\"}]","Imputation Techniques in Machine Learning - A Survey | PDF",1785673006,13,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"imputation-techniques-in-machine-learning-a-survey","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/imputation-techniques-in-machine-learning-a-survey/116995/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do missing values cause problems in machine learning workflows?","Question",{"text":75,"@type":76},"Missing values can reduce model performance, complicate data analysis, and introduce bias by creating discrepancies between missing and complete data.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is data imputation necessary before training many machine learning models?",{"text":80,"@type":76},"Most machine learning models are not built to handle missing values directly, so imputation is needed to supply usable inputs.",{"name":82,"@type":73,"acceptedAnswer":83},"Which imputation techniques are reviewed in the survey?",{"text":84,"@type":76},"The paper reviews mean, median, K-nearest neighbors (KNN) imputation, linear regression, missForest, and MICE, comparing their effectiveness and suitability.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":21,"slug":137},19,"General","general"]