[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-128040-en":3,"doc-seo-128040-105":30,"detail-sidebar-cat-0-en-105":84},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},128040,5909887256941,"Levi","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Adaptive optimization for prediction with missing data","Training predictive models on datasets containing missing entries commonly relies on an impute-then-predict pipeline. This work reframes prediction with missing data as a two-stage adaptive optimization problem and introduces adaptive linear regression models whose coefficients adapt to the set of observed features. The paper shows equivalences to jointly learned imputation and downstream regression, extends the framework to non-linear models, and demonstrates 2–10% out-of-sample accuracy gains in strongly non-MAR settings.","Adaptive optimization for prediction with missing data  \nDimitris Bertsimas1 · Arthur Delarue2 · Jean Pauphilet3  \nReceived: 1 February 2024 / Revised: 23 January 2025 / Accepted: 24 February 2025 © The Author(s) 2025  \nAbstract  \nWhen training predictive models on data with missing entries, the most widely used and versatile approach is a pipeline technique where we first impute missing entries and then compute predictions. In this paper, we view prediction with missing data as a two-stage adaptive optimization problem and propose a new class of models, adaptive linear regression models, where the regression coefficients adapt to the set of observed features. We show that some adaptive linear regression models are equivalent to learning an imputation rule and a downstream linear regression model simultaneously instead of sequentially. We leverage this joint-impute-then-regress interpretation to generalize our framework to non-linear models. In settings where data is strongly not missing at random, our methods achieve a 2–10% improvement in out-of-sample accuracy.  \nKeywords Missing data · Adaptive optimization  \n1 Introduction  \nReal-world datasets usually combine information from multiple sources, with different measurement units, data encoding, or structure, leading to a myriad of inconsistencies. In particular, they often come with partially observed features. In contrast, most supervised learning models require that all input features are available for every data point.  \nEditor: Scott Sanner.  \n* Arthur Delarue [arthur.delarue@isye.gatech.edu](arthur.delarue@isye.gatech.edu)  \nDimitris Bertsimas  \n[dbertsim@mit.edu](dbertsim@mit.edu)  \nJean Pauphilet  \n[jpauphilet@london.edu](jpauphilet@london.edu)  \n1 Sloan School of Management, Massachusetts Institute of Technology, 77 Massachusetts Ave, Cambridge, MA, USA  \n2 H. Milton Stewart School of Industrial and Systems Engineering, Georgia Institute of Technology,  \n755 Ferst Dr, Atlanta, GA, USA  \n3 London Business School, Regent’s Park, NW1 4SA London, UK  \nMissing data have been studied in the statistical inference literature for several decades (see, for instance, Little & Rubin, 2019) . The typical assumption is that data entries are missing at random (MAR) . Under the MAR assumption, a valid inference methodology is to impute-then-estimate, i.e. , guess the missing values as accurately as possible, then estimate the parameter of interest on the imputed dataset (and even generate confidence intervals, as in Rubin, 1987) . However, most inference guarantees become invalid as soon as the data is not missing at random (NMAR); and except in some special circumstances, the validity of the MAR assumption cannot be tested from the data (Little, 1988 ; Jaeger, 2006) .  \nIn practice, prediction problems with missing data are often treated as if they were inference problems. A typical approach is impute-then-regress, where one first imputes missing values, then trains a model on the imputed dataset (see Emmanuel et al. , 2021, for a review) . These two models (imputation and regression) are often learned in isolation. Because they are designed with inference in mind, imputation methods are sound under the MAR assumption (which makes it possible to guess missing feature values from observed values of the same feature) but may not be reliable when the data is not MAR.  \nOur work contributes to an active stream of research that revisits the findings of the inference literature on missing data for supervised learning tasks. For example, Bertsimas et al. (2024) and Josse et al. (2024) promote, with theoretical and empirical evidence, the use of simple imputation rules (e.g., mean imputation) for prediction, despite the fact that they are not suited for inference. Saar-Tsechansky and Provost (2007) provide early evidence that imputation may be sub-optimal and that training different models based on the set of features available can achieve higher performance. Their approach, which corresp","cbCaiszX6dEXVEIF","https://ap.wps.com/l/cbCaiszX6dEXVEIF","pdf",2586151,1,37,"English","en",105,"# Introduction\n## Missing data and assumptions (MAR vs NMAR)\n## Impute-then-regress and its limitations\n## Adaptive prediction approaches and related work\n# Adaptive linear regression framework","[{\"question\":\"How does performance change when data is strongly not missing at random (NMAR)?\",\"answer\":\"The methods achieve a 2–10% improvement in out-of-sample accuracy in settings where data is strongly not missing at random.\"}]","Adaptive optimization for prediction with missing data | PDF",1785944317,93,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":79,"head_meta":81,"extra_data":83,"updated_unix":28},"adaptive-optimization-for-prediction-with-missing-data","",{"@graph":36,"@context":78},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/adaptive-optimization-for-prediction-with-missing-data/128040/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-24","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72],{"name":73,"@type":74,"acceptedAnswer":75},"How does performance change when data is strongly not missing at random (NMAR)?","Question",{"text":76,"@type":77},"The methods achieve a 2–10% improvement in out-of-sample accuracy in settings where data is strongly not missing at random.","Answer","https://schema.org",{"og:url":52,"og:type":80,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":82,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":85},[86,90,94,98,103,108,113,116,121,124,128],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":87,"show_sort_weight":88,"slug":89},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":91,"show_sort_weight":92,"slug":93},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Exam",70,"exam",{"id":99,"doc_module":4,"doc_module_name":46,"category_name":100,"show_sort_weight":101,"slug":102},5,"Comic",60,"comic",{"id":104,"doc_module":4,"doc_module_name":46,"category_name":105,"show_sort_weight":106,"slug":107},6,"Technology",50,"technology",{"id":109,"doc_module":4,"doc_module_name":46,"category_name":110,"show_sort_weight":111,"slug":112},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":114,"slug":115},30,"research-report",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},9,"Religion & Spirituality",20,"religion-spirituality",{"id":119,"doc_module":4,"doc_module_name":46,"category_name":122,"show_sort_weight":119,"slug":123},"World Cup","world-cup",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":125,"slug":127},10,"Lifestyle","lifestyle",{"id":129,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":99,"slug":131},19,"General","general"]