[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123511-en":3,"doc-seo-123511-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123511,13056703020460,"Valentina","https://ap-avatar.wpscdn.com/avatar/be000253dac470eee5d?_k=1778207105932848923",8,"Research & Report","Double Machine Learning meets Panel Data - Promises, Pitfalls, and Potential Solutions","Estimating causal effects with machine learning (ML) can ease assumptions about functional forms, especially when paired with formal frameworks such as double/debiased machine learning (DML). Most DML theory, however, targets cross-sectional designs, while researchers frequently work with panel data where unobserved heterogeneity is handled by classical techniques. This work adapts DML to panel settings with unobserved heterogeneity and evaluates multiple estimators via simulations, studying cross-fitting validity and heterogeneous impacts on performance.","arXiv :2409 .01266v1 [ econ .EM] 2 Sep 2024  \nDouble Machine Learning meets Panel Data-Promises, Pitfalls,  \nand Potential Solutions  \nJonathan Fuhr and Dominik Papies  \nSchool of Business and Economics, University of T¨ubingen, T¨ubingen, Germany  \nLast edited: September 4, 2024  \nAbstract  \nEstimating causal effect using machine learning (ML) algorithms can help to relax functional form assumptions if used within appropriate frameworks. However, most of these frameworks assume settings with cross-sectional data, whereas researchers often have access to panel data, which in traditional methods helps to deal with unobserved heterogeneity between units. In this paper, we explore how we can adapt double/debiased machine learning (DML) (Chernozhukov et al., 2018) for panel data in the presence of unobserved heterogeneity. This adaptation is challenging because DML’s cross-fitting procedure assumes independent data and the unobserved heterogeneity is not necessarily additively separable in settings with nonlinear observed confounding. We assess the performance of several intuitively appealing estimators in a variety of simulations. While we find violations of the cross-fitting assumptions to be largely inconsequential for the accuracy of the effect estimates, many of the considered methods fail to adequately account for the presence of unobserved heterogeneity. However, we find that using predictive models based on the correlated random effects approach (Mundlak, 1978) within DML leads to accurate coefficient estimates across settings, given a sample size that is large relative to the number of observed confounders. We also show that the influence of the unobserved heterogeneity on the observed confounders plays a significant role for the performance of most alternative methods.  \n1 Introduction  \nAcross multiple quantitative disciplines, recent years have seen an explosion of research on methodologies that try to use machine learning (ML) to help estimate causal effects (e.g., Athey et al. , 2019; Chernozhukov et al., 2018) . These methods aim to relax assumptions in the causal estimation process by using modern ML methods to learn certain properties of the data. Arguably one of the most popular of these methods is the double/debiased machine learning (DML) framework by Chernozhukov et al. (2018) . DML can help relax assumptions about how to adjust for observed confounders by modeling the confounding relationships with flexible ML methods. That is, in the case of a large number of potentially important confounders, or in cases where the functional forms of the confounding influences are unknown, DML uses flexible ML methods to pick the most important confounders and adjust for them flexibly. This framework has seen applications in a variety of disciplines (e.g., Felderer et al., 2023; Gordon et al., 2022; Parpouchi et al., 2021), as well as further developments and extensions to settings beyond the original ones (e.g., Bodory et al., 2022; Chiang et al., 2022; Liu et al., 2021) .  \nAt the same time, many more traditional methods from statistics and econometrics are still the default approaches for credible causal effect estimation. This is mostly because they aim to relax assumptions that are stronger than the estimation assumptions DML can relax. These methods address assumptions about causal identification, e.g., how to estimate causal effects in the presence of unobserved confounding. Examples are panel data methods, difference-in-differences, synthetic control, instrumental variables, and regression discontinuity designs (e.g. , Cunningham, 2021; Huntington-Klein, 2022) . While there has been progress in using some of these methods within the DML framework (see, e.g., Chernozhukov et al., 2024), very little research has explored how to adapt DML to settings with panel data, which will be the focus of our paper.  \nIn many applications, we observe the same units (e.g., individuals, firms, cities, etc.) repeatedly over time. Thi","cbCaidd3jukhfUiG","https://ap.wps.com/l/cbCaidd3jukhfUiG","pdf",2627924,1,47,"English","en",105,"# Introduction\n## Background: DML and causal inference\n## Panel data and unobserved heterogeneity\n## Challenges in adapting DML to panels","[{\"question\":\"Why is adapting DML to panel data with unobserved heterogeneity challenging?\",\"answer\":\"DML’s cross-fitting relies on independent data, which becomes nontrivial when additional dimensions like time are present. Also, unobserved heterogeneity may not be additively separable under nonlinear observed confounding, so standard ways to handle it are unclear within DML.\"},{\"question\":\"How do the authors evaluate different estimators?\",\"answer\":\"They assess several intuitively appealing estimators using simulations across a variety of settings. The evaluation focuses on accuracy and how well each method accounts for unobserved heterogeneity.\"},{\"question\":\"What approach shows the most accurate results in their experiments?\",\"answer\":\"Using predictive models built on the correlated random effects approach (Mundlak, 1978) within DML yields accurate coefficient estimates across settings, particularly when the sample size is large relative to the number of observed confounders.\"}]","Double Machine Learning meets Panel Data - Promises, Pitfalls, and Potential Solutions | PDF",1785816974,118,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"double-machine-learning-meets-panel-data-promises-pitfalls-and-potential-solutions","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/double-machine-learning-meets-panel-data-promises-pitfalls-and-potential-solutions/123511/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is adapting DML to panel data with unobserved heterogeneity challenging?","Question",{"text":75,"@type":76},"DML’s cross-fitting relies on independent data, which becomes nontrivial when additional dimensions like time are present. Also, unobserved heterogeneity may not be additively separable under nonlinear observed confounding, so standard ways to handle it are unclear within DML.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How do the authors evaluate different estimators?",{"text":80,"@type":76},"They assess several intuitively appealing estimators using simulations across a variety of settings. The evaluation focuses on accuracy and how well each method accounts for unobserved heterogeneity.",{"name":82,"@type":73,"acceptedAnswer":83},"What approach shows the most accurate results in their experiments?",{"text":84,"@type":76},"Using predictive models built on the correlated random effects approach (Mundlak, 1978) within DML yields accurate coefficient estimates across settings, particularly when the sample size is large relative to the number of observed confounders.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]