[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126591-en":3,"doc-seo-126591-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},126591,687207020761,"Patrick","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","A Double Machine Learning Approach to Combining Experimental and Observational Data","Experimental and observational studies often lack validity due to assumptions that cannot be directly tested. The document proposes a double machine learning framework that combines both data sources to detect assumption violations and to estimate treatment effects with consistency. It evaluates violations of external validity and ignorability under milder conditions. When exactly one assumption fails, it yields semiparametrically efficient estimators. A no-free-lunch result underscores the need to correctly identify which assumption is violated, and three case studies demonstrate practical relevance.","arXiv :2307 .01449v1 [ stat .ME] 4 Jul 2023  \nA Double Machine Learning Approach to Combining Experimental and Observational Data  \nMarco Morucci* 1 , Vittorio Orlandi*2 , Harsh Parikh*3 ,  \nSudeepa Roy3 , Cynthia Rudin3 , Alexander Volfovsky2  \n*denotes joint ﬁrst authors (in the alphabetical order of their last names)  \n1 Center for Data Science, New York University  \n2 Department of Statistical Science, Duke University  \n3 Department of Computer Science, Duke University  \nAbstract  \nExperimental and observational studies often lack validity due to untestable assumptions. We propose a double machine learning approach to combine experimental and observational studies, allowing practitioners to test for assumption violations and estimate treatment effects consistently. Our framework tests for violations of external validity and ignorability under milder assumptions. When only one assumption is violated, we provide semiparametrically efﬁcient treatment effect estimators. However, our no-free-lunch theorem highlights the necessity of accurately identifying the violated assumption for consistent treatment effect estimation. We demonstrate the applicability of our approach in three real-world case studies, highlighting its relevance for practical settings.  \nKeywords: Data Fusion, Generalizability, External Validity, Observational Study  \n1 Introduction  \nExperiments and observational studies are both indispensable for estimating treatment effects and investigating causal questions of interest. Experiments, or randomized control trials (RCTs), guarantee unbiased estimation of sample average treatment effects via the randomization of treatment across the experimental population. However, they are typically small, expensive, and slow to design and implement. Another concern involves the experimental units themselves, which may be recruited via a non-randomized process and therefore may not be representative of the superpopulation of interest; in such cases we would say that the experiment lacks external validity. If this is the case, then the conclusions of the experiment may not be generalizable to the superpopulation, which is oftentimes of primary interest. On the other hand, observational studies, in which units self-select into treatment, tend to be larger and consist of units that better represent the superpopulation than experimental units. However, analysis of observational studies is plagued by observed and unobserved confounders – covariates that affect both units' propensity to self-select into treatment and their response. While a variety of methods – matching and weighting being two large classes among them – can be used to adjust for confounders, they typically rely on the crucial assumptions that they can accurately model the treatment-covariate or response-covariate relationships and that the available covariates are sufﬁcient for doing so. This latter assumption, which is often referred to as strong or conditional ignorability, is untestable using observational data alone [Rosenbaum and Rubin, 1983] .  \nIn this paper, we propose methods for jointly using experimental and observational data to (i) test for violations of at least one of validity and ignorability and (ii) estimate population average treatment effects even under such violations. The scope of our work is broader than that of previous work, which typically does not attempt the ﬁrst task and only focuses on a subset of the second, where external validity is violated but conditional ignorability isnot. Crucially, our estimators are also valid under violations of certain sets of assumptions that are often required in  \nthe literature, making our methods applicable in settings where other methods may be inappropriate. We provide a detailed discussion of the literature in Section 2 . Broadly, our methods build on the double machine learning literature [Chernozhukov et al., 2018] by estimating the nuisance parameters of interest (e.g., the propensity score) and","cbCaipftuotKO02V","https://ap.wps.com/l/cbCaipftuotKO02V","pdf",812019,1,39,"English","en",105,"# Introduction\n## Validity and ignorability in experiments and observational studies\n## Related work and double machine learning foundation\n# Proposed double machine learning framework\n## Testing assumption violations\n## Estimating treatment effects\n# Theoretical properties\n## Asymptotic normality and unbiasedness\n## Tests for unobserved confounding\n# Empirical applications\n## Project STAR case study\n## Coronary Artery Surgery Study case study\n## Additional real-world case study","[{\"question\":\"Why do experiments and observational studies often face validity problems?\",\"answer\":\"Experiments can suffer from external validity issues when sampled units are not representative of the target superpopulation. Observational studies are affected by confounding from both observed and unobserved factors that influence treatment selection and outcomes.\"},{\"question\":\"What does the proposed double machine learning approach achieve?\",\"answer\":\"It jointly uses experimental and observational data to test for violations of external validity and ignorability and to estimate population average treatment effects even when such violations occur.\"},{\"question\":\"What is the key message of the no-free-lunch theorem in this framework?\",\"answer\":\"Consistent treatment effect estimation requires correctly identifying which assumption is violated; otherwise the estimation can fail despite the use of advanced methods.\"}]","A Double Machine Learning Approach to Combining Experimental and Observational Data | PDF",1785933558,98,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-double-machine-learning-approach-to-combining-experimental-and-observational-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-double-machine-learning-approach-to-combining-experimental-and-observational-data/126591/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-05",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do experiments and observational studies often face validity problems?","Question",{"text":75,"@type":76},"Experiments can suffer from external validity issues when sampled units are not representative of the target superpopulation. Observational studies are affected by confounding from both observed and unobserved factors that influence treatment selection and outcomes.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does the proposed double machine learning approach achieve?",{"text":80,"@type":76},"It jointly uses experimental and observational data to test for violations of external validity and ignorability and to estimate population average treatment effects even when such violations occur.",{"name":82,"@type":73,"acceptedAnswer":83},"What is the key message of the no-free-lunch theorem in this framework?",{"text":84,"@type":76},"Consistent treatment effect estimation requires correctly identifying which assumption is violated; otherwise the estimation can fail despite the use of advanced methods.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]