[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85101-en":3,"doc-seo-85101-105":29,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":13,"seo_description":14,"update_tm":27,"read_time":28},85101,687197207057,"Sage","https://ap-avatar.wpscdn.com/davatar_29158cc5080c5b710cf443261637dec0",8,"Research & Report","ImputeViz A Visual Analytics Dashboard for Diagnosing Missing Data and Comparing Imputation Methods","Missing data undermines analyses in scientific, social science, and public health studies by introducing bias and shifting accountability to analysts for how missing values are handled. IMPUTEVIZ is an integrated visual analytics dashboard for diagnosing missingness, configuring imputation models, and evaluating results. It unifies MICE, Random Forest, XGBoost, and kNN, plus gKNN for geospatial reasoning, exposing donor provenance. Coordinated heatmaps, diagnostics, and method-comparison views report MAE, RMSE, ∆RMSE, runtime, and variable discrepancies. Case studies show robust strategy selection and improved interpretability for downstream summaries, including non-random missingness considerations (MCAR/MAR/MNAR).","ImputeViz: A Visual Analytics Dashboard for Diagnosing Missing Data and  \nComparing Imputation Methods  \nAitik Dandapat* Stony Brook University  \nLalith Punepalle Raveendrareddy†  \nStony Brook University  \nMithilesh Kumar Singh‡  \nStony Brook University  \nKlaus Mueller § Stony Brook University  \narXiv :2607 .08579v 1 [ cs .HC] 9 Jul 2026  \nABSTRACT  \nMissing data is a persistent obstacle in scientific, social science, and public health research, often biasing analyses and placing accountability on analysts for how they handle missing values. We introduce IMPUTEVIZ, an integrated visual analytics dashboard that supports diagnosing missingness, configuring imputation models, and evaluating results. The system brings together widely used methods—including MICE, Random Forest, XGBoost, and kNN—within an interactive environment that makes missingness patterns explicit. To support geospatial reasoning, we introduce gKNN, a geographically informed kNN variant that blends socio-economic and spatial distances and exposes donor contributions, enabling provenance-based visual accountability by showing which regions drive each estimate. Our primary contribution is a method-agnostic visual analytics environment that makes cross-method comparison a first-class visual task and integrates gKNN alongside standard methods. Coordinated views reveal missingness structure through heatmaps, co-missingness summaries, and distributional diagnostics that help analysts reason about missingness patterns (MCAR/MAR) and cases where missingness may be non-random (MNAR) . Users can compare and tune models and interrogate results via distributional overlays, a Method Comparison Summary reporting MAE, RMSE, ∆RMSE, and runtime for each algorithm on the current target and mask, along with variable-level discrepancy views. Cached per-method results and locked axis scales reduce cognitive overhead from shifting ranges during method switching. These comparisons highlight where methods disagree, which variables are sensitive, and how imputation choices affect downstream summaries. Case studies demonstrate how IMPUTEVIZ helps analysts select effective strategies, surface sensitive variables, and assess model robustness.  \nIndex Terms: Visual analytics, missing data, imputation, spatial epidemiology, choropleth visualization, statistical evaluation.  \n1 INTRODUCTION  \nMissing data remains a persistent obstacle in scientific and publichealth analysis. In county-level public-health surveillance, privacy suppression and sparse reporting are often systematic: missing entries cluster in low-count regions, which can bias downstream trend analysis and resource allocation if imputed poorly. At the sametime, many applied workflows treat imputation as a black-box preprocessing step and provide limited support for cross-method visual comparison or for inspecting why a value was imputed in a particular way – often forcing analysts to compare models sequentially and mentally reconcile shifting scales and parameters across runs.  \nWe present ImputeViz, a visual analytics dashboard that supports the full imputation workflow: diagnosing missingness, selecting and tuning models, comparing outcomes, and inspecting provenance.  \n* e-mail: [adandapat@cs.stonybrook.edu](adandapat@cs.stonybrook.edu)[ ](adandapat@cs.stonybrook.edu)†e-mail: [lpunepallera@cs.stonybrook.edu](lpunepallera@cs.stonybrook.edu)[ ](lpunepallera@cs.stonybrook.edu)‡[e-mail: mkssingh@cs.stonybrook.edu](e-mail: mkssingh@cs.stonybrook.edu)  \n§[e-mail: mueller@cs.stonybrook.edu](e-mail: mueller@cs.stonybrook.edu)  \nThe system integrates common imputers (MICE, Random Forest, XGBoost, linear/kNN baselines) and a geospatially informed method (gKNN) that blends socioeconomic and spatial distance. Rather than promoting one model universally, ImputeViz is designed to make method choice auditable through linked diagnostics, stable-scale comparisons, and quantitative summaries.  \nOur contributions are threefold: (1) a visual ana","cbCaio2Lr3KQTKHd","https://ap.wps.com/l/cbCaio2Lr3KQTKHd","pdf",6614555,1,5,"English","en",105,"# Abstract\n# Introduction\n# Related Work\n# Data","[{\"question\":\"What problems does IMPUTEVIZ address in missing-data analysis?\",\"answer\":\"IMPUTEVIZ targets the bias and interpretability challenges caused by missing values by supporting missingness diagnosis, configurable imputation, and quantitative evaluation. It also helps analysts compare methods and inspect why imputations occur through provenance-aware views.\"},{\"question\":\"Which imputation methods does the dashboard support?\",\"answer\":\"The dashboard integrates widely used methods including MICE, Random Forest, XGBoost, and kNN. It also adds gKNN, a geographically informed kNN variant that blends socio-economic and spatial distances.\"},{\"question\":\"How does IMPUTEVIZ make cross-method comparisons reliable and understandable?\",\"answer\":\"It uses linked diagnostics with stable, locked axis scales to reduce cognitive overhead, and provides a method-comparison summary reporting MAE, RMSE, ∆RMSE, and runtime under consistent holdout-based evaluation. Variable-level discrepancy views help reveal where methods disagree.\"}]",1784201111,13,{"code":4,"msg":30,"data":31},"ok",{"site_id":24,"language":23,"slug":32,"title":13,"keywords":33,"description":14,"schema_data":34,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":27},"imputeviz-a-visual-analytics-dashboard-for-diagnosing-missing-data-and-comparing-imputation-methods","",{"@graph":35,"@context":85},[36,53,68],{"@type":37,"itemListElement":38},"BreadcrumbList",[39,43,47,50],{"item":40,"name":41,"@type":42,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":44,"name":45,"@type":42,"position":46},"https://docshare.wps.com/document/","Document",2,{"item":48,"name":12,"@type":42,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":42,"position":52},"https://docshare.wps.com/document/imputeviz-a-visual-analytics-dashboard-for-diagnosing-missing-data-and-comparing-imputation-methods/85101/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":23,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":40,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-17","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problems does IMPUTEVIZ address in missing-data analysis?","Question",{"text":75,"@type":76},"IMPUTEVIZ targets the bias and interpretability challenges caused by missing values by supporting missingness diagnosis, configurable imputation, and quantitative evaluation. It also helps analysts compare methods and inspect why imputations occur through provenance-aware views.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Which imputation methods does the dashboard support?",{"text":80,"@type":76},"The dashboard integrates widely used methods including MICE, Random Forest, XGBoost, and kNN. It also adds gKNN, a geographically informed kNN variant that blends socio-economic and spatial distances.",{"name":82,"@type":73,"acceptedAnswer":83},"How does IMPUTEVIZ make cross-method comparisons reliable and understandable?",{"text":84,"@type":76},"It uses linked diagnostics with stable, locked axis scales to reduce cognitive overhead, and provides a method-comparison summary reporting MAE, RMSE, ∆RMSE, and runtime under consistent holdout-based evaluation. Variable-level discrepancy views help reveal where methods disagree.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,109,114,119,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":45,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":46,"doc_module":4,"doc_module_name":45,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":45,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":21,"doc_module":4,"doc_module_name":45,"category_name":106,"show_sort_weight":107,"slug":108},"Comic",60,"comic",{"id":110,"doc_module":4,"doc_module_name":45,"category_name":111,"show_sort_weight":112,"slug":113},6,"Technology",50,"technology",{"id":115,"doc_module":4,"doc_module_name":45,"category_name":116,"show_sort_weight":117,"slug":118},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":45,"category_name":12,"show_sort_weight":120,"slug":121},30,"research-report",{"id":123,"doc_module":4,"doc_module_name":45,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":45,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":45,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":45,"category_name":136,"show_sort_weight":21,"slug":137},19,"General","general"]