[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-127344-en":3,"doc-seo-127344-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},127344,962085564381,"Clementine","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","A survey of missing data imputation techniques - statistical methods, machine learning models, and GAN-based approaches","Efficient missing-data handling is essential for reliable analysis across scientific domains. This survey evaluates traditional statistical, machine learning, and generative adversarial network (GAN) based imputation approaches, comparing strengths, limitations, and suitability for different data types and missing-data mechanisms: MCAR, MAR, and MNAR. GAN models such as GAIN, VIGAN, and SolarGAN are emphasized for complex datasets including images and time series, with attention to computational cost, hyperparameter sensitivity, and overfitting risk. Results indicate GANs better capture non-linear dependencies, while future work targets broader architectures and hybrid methods for improved accuracy and scalability in real-world use.","IAES International Journal of Artificial Intelligence (IJ-AI)  \nVol. 14, No. 4, August 2025, pp. 2876 ✒2888  \nISSN: 2252-8938, DOI: 10.11591/ijai.v14.i4.pp2876-2888 ❒ 2876  \n\n| A survey of missing data imputation techniques: statistical methods, machine learning models, and GAN-based\u003Cbr>approaches\u003Cbr>Rifaa Sadegh, Ahmed Mohameden, Mohamed Lemine Salihi, Mohamedade Farouk Nanne\u003Cbr>Scientific Computing, Computer Science and Data Science, Department of Computer Science, Faculty of Science and Technology,\u003Cbr>University of Nouakchott, Nouakchott, Mauritania |  |  |\n| --- | --- | --- |\n| Article Info\u003Cbr>Article history:\u003Cbr>Received Jun 8, 2024 Revised Jun 11, 2025 Accepted Jul 10, 2025\u003Cbr>Keywords:\u003Cbr>Data imputation\u003Cbr>Generative adversarial networks Machine learning\u003Cbr>Missing data\u003Cbr>Statistical methods |  | ABSTRACT\u003Cbr>Efficiently addressing missing data is critical in data analysis across diverse domains. This study evaluates traditional statistical, machine learning, and generative adversarial network (GAN)-based imputation methods, emphasizing their strengths, limitations, and applicability to different data types and missing data mechanisms (missing completely at random (MCAR), missing at random (MAR), missing not at random (MNAR)) . GAN-based models, including generative adversarial imputation network (GAIN), view imputation generative adversarial network (VIGAN), and SolarGAN, are highlighted for their adaptability and effectiveness in handling complex datasets, such as images and time series. Despite challenges like computational demands, GANs outperform conventional methods in capturing non-linear dependencies. Future work includes optimizing GAN architectures for broader data types and exploring hybrid models to enhance imputation accuracy and scalability in real-world applications.\u003Cbr>This is an open access article under the CC BY-SA license. |\n| Corresponding Author: |  |  |\n| Rifaa Sadegh\u003Cbr>Scientific Computing, Computer Science and Data Science, Department of Computer Science Faculty of Science and Technology, University of Nouakchott\u003Cbr>Nouakchott, Mauritania\u003Cbr>Email: [rifasadegh@gmail.com](rifasadegh@gmail.com) |  |  |\n\n1. INTRODUCTION  \nMissing data is a pervasive challenge that affects nearly every scientific discipline, from medicine [1] to geology [2], energy [3] and environmental sciences [4] . Rubin [5] defined missing data as unobserved values that could yield critical insights if available. These gaps introduce biases, distort analysis, and reduce the effectiveness of algorithms, ultimately impairing decision-making processes.  \nThe origins of missing data are diverse, arising from incomplete data collection, recording errors, or hardware malfunctions [5] . These gaps skew results and misrepresent the studied population [6], creating a need for robust and scalable solutions to ensure reliable research outcomes. Addressing missing data has proven to be a multifaceted problem, requiring methods that vary depending on the type and complexity of the dataset. Initial approaches, such as listwise deletion, were simple but often discarded valuable information along with the missing data [7] . Over time, more sophisticated imputation techniques emerged, including statistical methods, machine learning algorithms, and deep learning models. Among these, generative adversarial networks (GANs) have gained prominence for their ability to model complex data distributions and address non-linear dependencies effectively. Despite their potential, implementing GANs for data imputation comes  \nwith challenges, including: i) high computational costs due to complex training processes; ii) sensitivity to hyperparameter tuning, which affects model stability; and iii) risk of overfitting, particularly when handling small datasets.  \nThis paper provides a comprehensive review of missing data imputation methods. We analyze traditional statistical approaches, machine learning techniques, and deep learning models, with a particular ","cbCaitacdP0cCuBh","https://ap.wps.com/l/cbCaitacdP0cCuBh","pdf",849838,1,13,"English","en",105,"# Introduction\n## Missing data challenges\n## Scope and structure of the paper\n# Missing data mechanisms and types of variables\n## Missing data categories\n## Variable types and imputation approach classification\n# Comparative analysis of imputation approaches\n## Statistical, machine learning, and deep learning methods\n## GAN-based imputation methods\n# Implications and ethical considerations\n## Results implications in sensitive domains\n# Conclusion and future research directions","[{\"question\":\"What missing data mechanisms are covered in the study?\",\"answer\":\"The survey focuses on MCAR (missing completely at random), MAR (missing at random), and MNAR (missing not at random), and explains how these mechanisms affect imputation choices.\"},{\"question\":\"Which GAN-based imputation models are highlighted?\",\"answer\":\"The paper highlights GAN-based approaches including GAIN, VIGAN, and SolarGAN, emphasizing their adaptability for complex data types.\"},{\"question\":\"What challenges come with using GANs for data imputation?\",\"answer\":\"GAN imputation is associated with high computational demands, sensitivity to hyperparameter tuning, and a risk of overfitting, especially for small datasets.\"}]","A survey of missing data imputation techniques - statistical methods, machine learning models, and GAN-based approaches | PDF",1785938400,33,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"a-survey-of-missing-data-imputation-techniques-statistical-methods-machine-learning-models-and-gan-based-approaches","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-survey-of-missing-data-imputation-techniques-statistical-methods-machine-learning-models-and-gan-based-approaches/127344/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-23","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What missing data mechanisms are covered in the study?","Question",{"text":76,"@type":77},"The survey focuses on MCAR (missing completely at random), MAR (missing at random), and MNAR (missing not at random), and explains how these mechanisms affect imputation choices.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"Which GAN-based imputation models are highlighted?",{"text":81,"@type":77},"The paper highlights GAN-based approaches including GAIN, VIGAN, and SolarGAN, emphasizing their adaptability for complex data types.",{"name":83,"@type":74,"acceptedAnswer":84},"What challenges come with using GANs for data imputation?",{"text":85,"@type":77},"GAN imputation is associated with high computational demands, sensitivity to hyperparameter tuning, and a risk of overfitting, especially for small datasets.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]