[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-116863-en":3,"doc-seo-116863-105":30,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},116863,1099513958762,"Logic","https://ap-avatar.wpscdn.com/avatar/1000023916a998db790?x-image-process=image/resize,m_fixed,w_180,h_180&k=1784791008015729253",8,"Research & Report","Missing Values and the Dimensionality of Expected Returns - Missing Value Imputation in Machine Learning Asset Pricing","Combining many cross-sectional return predictors in machine learning often forces missing-value imputation because dropping incomplete stocks can remove most of the sample. The study compares ad hoc mean imputation with maximum likelihood methods, including EM. Despite expectations of divergence, the imputations yield nearly identical inferences for mean returns and Sharpe ratios. The paper attributes invariance to weak cross-sectional correlations and limited variance explained by principal components, then confirms similar behavior in neural network portfolios.","arXiv :2207 . 13071v3 [ stat .ME] 4 May 2023  \nMissing Values and the Dimensionality of Expected Returns  \nAndrew Y. Chen  \nFederal Reserve Board  \nJack McCoy Columbia University  \nMay 2023*  \nAbstract  \nCombining many cross-sectional return predictors (for example, in machine learning) often requires imputing missing values. We compare adhoc mean imputation with several methods including maximum likelihood. Surprisingly, maximum likelihood and ad-hoc methods lead to similar results. This is because predictors are largely independent: Correlations cluster near zero and 10 principal components (PCs) span less than 50% of total variance. Independence implies observed predictors are uninformative about missing predictors, making ad-hoc methods valid. In PC regression tests, 50 PCs are required to capture equal-weighted expected returns (30 PCs value-weighted), regardless of the imputation. We ﬁnd similar invariance in neural network portfolios.  \nJEL Classiﬁcation: G0, G1  \nKeywords: stock market predictability, stock market anomalies, missing values  \n* First submitted to [arXiv.org:](arXiv.org:) July 20, 2022. Emails: [andrew.y.chen@frb.gov](andrew.y.chen@frb.gov) (Chen, corresponding author) [and jmccoy26@gsb.columbia.edu](and jmccoy26@gsb.columbia.edu) (McCoy). This project originated from many conversations with Fabian Winkler. We thank Heiner Beckmeyer, Charlie Clarke, Harry Mamaysky, Markus Pelger, Yinan Su, and an anonymous referee for helpful comments. The views expressed herein are those of the authors and do not necessarily reﬂect the position of the Board of Governors ofthe Federal Reserve or the Federal Reserve System.  \n1 Introduction  \nA growing literature applies methods from machine learning to asset pricing. These studies combine the information in dozens, or even hundreds, of crosssectional stock return predictors. Freyberger et al. (2020) combine 62; Kelly et al.(2023) combine 138; and Han et al. (2022) combine up to 193 predictors. Each of these papers ﬁnds that there are economic gains to expanding the set of predictors beyond the ﬁve used in Fama and French (2015) .1  \nBuried in this literature is the problem of missing values. When learning from many predictors, the standard practice of dropping stocks with missing values is often untenable. For example, applying the standard practice to the 125 mostobserved predictors inthe Chen and Zimmermann (2022) dataset drops 99.8% of stocks. So even though imputing missing values may seem dangerous, machine learning researchers often have no choice but to impute.  \nWe ﬁnd that simply imputing with cross-sectional averages (as in Kozak, Nagel and Santosh (2020); Gu, Kelly and Xiu (2020)) does a surprisingly good job of capturing expected returns. We compare this simple mean imputation with several alternatives in tests that sort stocks on ﬁtted expected returns (a la Lewellen (2015)) . The expected return models include principal components regression, gradient-boosted regression trees, and neural networks. For almost all expected return models, all of the imputations lead to very similar inferences about mean returns and Sharpe ratios. This invariance holds even for imputations that employ maximum likelihood via the expectation-maximization (EM) algorithm.  \nThis invariance comes from the fact that cross-sectional predictors are largely uncorrelated (Green, Hand and Zhang (2013); McLean and Pontiff (2016); Chen and Zimmermann (2022)) . Almost all cross-sectional correlations lie between ¡0.25 and Å0.25, and the ﬁrst 10 principal components span only 50% of total variance. As a result, observed predictors provide little information about the missing predictors, and one might as well simply impute missing values with cross-sectional means. More precisely, our EM imputation boils down to a set  \n1 Lewellen (2015) ﬁnds that linear regressions on 15 predictors, including size, B/M, momentum, proﬁtability, and investment lead to an annualized Sharpe Ratio of around 0.8 (e","cbCaikcwOGnoNQT3","https://ap.wps.com/l/cbCaikcwOGnoNQT3","pdf",691464,1,53,"English","en",105,"# Abstract\n# Introduction\n## Motivation: Why missing values matter\n## Predictors, expected returns, and model comparisons\n## Dimensionality and out-of-sample tests\n## Invariance across imputation and forecasting methods","[{\"question\":\"Why is imputing missing values necessary when using many stock return predictors?\",\"answer\":\"With dozens or hundreds of predictors, dropping stocks with missing values can become impractical and remove nearly the entire sample. Imputation is therefore often required to continue building predictive models and portfolios.\"},{\"question\":\"How do ad hoc mean imputation and maximum likelihood (EM) compare in this study?\",\"answer\":\"Both approaches produce very similar results in tests of mean returns and Sharpe ratios. The invariance persists even when maximum likelihood is implemented via the EM algorithm.\"},{\"question\":\"What explains the invariance in results across different imputations?\",\"answer\":\"The paper argues that cross-sectional predictors are largely uncorrelated and that the first principal components span limited total variance. As a result, observed predictors contain little information about missing predictors, making simple ad hoc imputation effectively valid.\"}]","Missing Values and the Dimensionality of Expected Returns - Missing Value Imputation in Machine Learning Asset Pricing | PDF",1785672124,134,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":28},"missing-values-and-the-dimensionality-of-expected-returns-missing-value-imputation-in-machine-learning-asset-pricing","",{"@graph":36,"@context":86},[37,54,69],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/missing-values-and-the-dimensionality-of-expected-returns-missing-value-imputation-in-machine-learning-asset-pricing/116863/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05","2026-08-02",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"Why is imputing missing values necessary when using many stock return predictors?","Question",{"text":76,"@type":77},"With dozens or hundreds of predictors, dropping stocks with missing values can become impractical and remove nearly the entire sample. Imputation is therefore often required to continue building predictive models and portfolios.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How do ad hoc mean imputation and maximum likelihood (EM) compare in this study?",{"text":81,"@type":77},"Both approaches produce very similar results in tests of mean returns and Sharpe ratios. The invariance persists even when maximum likelihood is implemented via the EM algorithm.",{"name":83,"@type":74,"acceptedAnswer":84},"What explains the invariance in results across different imputations?",{"text":85,"@type":77},"The paper argues that cross-sectional predictors are largely uncorrelated and that the first principal components span limited total variance. As a result, observed predictors contain little information about missing predictors, making simple ad hoc imputation effectively valid.","https://schema.org",{"og:url":52,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,124,129,132,136],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":46,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":46,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":46,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":122,"slug":123},30,"research-report",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":126,"show_sort_weight":127,"slug":128},9,"Religion & Spirituality",20,"religion-spirituality",{"id":127,"doc_module":4,"doc_module_name":46,"category_name":130,"show_sort_weight":127,"slug":131},"World Cup","world-cup",{"id":133,"doc_module":4,"doc_module_name":46,"category_name":134,"show_sort_weight":133,"slug":135},10,"Lifestyle","lifestyle",{"id":137,"doc_module":4,"doc_module_name":46,"category_name":138,"show_sort_weight":107,"slug":139},19,"General","general"]