[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123554-en":3,"doc-seo-123554-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123554,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","GENERIC MACHINE LEARNING INFERENCE ON HETEROGENOUS TREATMENT EFFECTS IN RANDOMIZED EXPERIMENTS","This paper proposes strategies for estimating and making valid inference on key features of heterogeneous treatment effects in randomized experiments. It targets best linear predictors using machine-learning proxies, impact-group sorted average effects, and averages for the most and least affected units. The method works in high-dimensional settings by post-processing ML outputs via repeated data splitting, avoiding overfitting and producing uniformly valid inference through median-based p-values and confidence intervals with adjusted nominal levels.","View metadata, citation and similar [papers at ](papers at core.ac.uk)[core.ac.uk](papers at core.ac.uk) brought to you by CORE  \n[provided by](provided by arXiv.org)[ arXiv.org](provided by arXiv.org) e-Print Archive  \narXiv : 17 12 .04802v4 [ stat .ML] 3 Sep 2019  \nGENERIC MACHINE LEARNING INFERENCE ON HETEROGENOUS TREATMENT EFFECTS IN RANDOMIZED EXPERIMENTS  \nVICTOR CHERNOZHUKOV, MERT DEMIRER, ESTHER DUFLO, AND IV´AN FERN´ANDEZ-VAL  \nAbstract. We propose strategies to estimate and make inference on key features of heterogeneous eﬀects in randomized experiments. These key features include best linear predictors of the eﬀects using machine learning proxies, average eﬀects sorted by impact groups, and average characteristics of most and least impacted units. The approach is valid in high dimensional settings, where the eﬀects are proxied by machine learning methods. We post-process these proxies into the estimates of the key features. Our approach is generic, it can be used in conjunction with penalized methods, deep and shallow neural networks, canonical and new random forests, boosted trees, and ensemble methods. It does not rely on strong assumptions. In particular, we don’t require conditions for consistency of the machine learning methods. Estimation and inference relies on repeated data splitting to avoid overﬁtting and achieve validity. For inference, we take medians of p-values and medians of conﬁdence intervals, resulting from many diﬀerent data splits, and then adjust their nominal level to guarantee uniform validity. This variational inference method is shown to be uniformly valid and quantiﬁes the uncertainty coming from both parameter estimation and data splitting. We illustrate the use of the approach with two randomized experiments in development on the eﬀects of microcredit and nudges to stimulate immunization demand.  \nKey words: Agnostic Inference, Machine Learning, Conﬁdence Intervals, Causal Eﬀects, Variational P-values and Conﬁdence Intervals, Uniformly Valid Inference, Quantiﬁcation of Uncertainty, Sample Splitting, Multiple Splitting, Assumption-Freeness, Microcredit, Immunization Incentives  \nJEL: C18, C21, D14, G21, O16  \n1. Introduction  \nRandomized experiments play an important role in the evaluation of social and economic programs and medical treatments (e.g., Imbens and Rubin (2015); Duﬂo et al. (2007)) . Researchers and policy makers are often interested in features of the impact of the treatment that go beyond the simple average treatment eﬀects. In particular, very often, they want to know whether treatment eﬀect depends on covariates, such as gender, age, etc. It is essential to assess if the impact of the program would generalize to a diﬀerent population with diﬀerent characteristics, and for economists, to better understand the driving mechanism behind the eﬀects of a particular program. In a review of 189 RCT published in top economic journals since 2006, we found that 76  \nDate: September 4, 2019 .  \nWe thank Susan Athey, Moshe Buchinsky, Denis Chetverikov, Siyi Luo, Max Kasy, Susan Murphy, Whitney Newey, and seminar participants at ASSA 2018, Barcelona GSE Summer Forum 2019, NYU, UCLA and Whitney Newey’s Contributions to Econometrics conference for valuable comments. We gratefully acknowledge research support from the National Science Foundation.  \n2 VICTOR CHERNOZHUKOV, MERT DEMIRER, ESTHER DUFLO, AND IV´AN FERN´ANDEZ-VAL  \n(40%) report at least one subgroup analysis, wherein they report treatment eﬀects in subgroups formed by baseline covariates.1  \nOne issue with reporting treatment eﬀects split by subgroups, however, is that there are often a large number of potential sample splits: choosing subgroups ex-post opens the possibility of overﬁtting. To solve this problem, medical journals and the FDA require pre-registering the sub-sample of interest in medical trials in advance. In economics, this approach has gained some traction, with the adoption of pre-analysis plans (which can be ﬁle","cbCailFSogF7xPsC","https://ap.wps.com/l/cbCailFSogF7xPsC","pdf",667064,1,53,"English","en",105,"# Introduction\n## Heterogeneity beyond average treatment effects\n## Subgroup analysis and overfitting risk\n## Pre-registration and pre-analysis plans\n## Using machine learning to explore heterogeneity\n# Proposed generic approach\n## Prediction-to-inference via post-processing\n## Repeated splitting for validity","[{\"question\":\"What heterogeneous treatment-effect features does the paper focus on?\",\"answer\":\"It targets best linear predictors of effects using ML proxies, impact-group sorted average effects, and averages comparing the most vs. least affected units.\"},{\"question\":\"How does the approach ensure valid inference when using machine learning?\",\"answer\":\"It relies on repeated data splitting to avoid overfitting and uses median-based p-values and confidence intervals, then adjusts nominal levels to guarantee uniform validity.\"},{\"question\":\"What kinds of machine learning models can the method be used with?\",\"answer\":\"The framework is generic and can be combined with penalized methods, deep or shallow neural networks, canonical or new random forests, boosted trees, and ensemble methods.\"}]","GENERIC MACHINE LEARNING INFERENCE ON HETEROGENOUS TREATMENT EFFECTS IN RANDOMIZED EXPERIMENTS | PDF",1785817296,134,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"generic-machine-learning-inference-on-heterogenous-treatment-effects-in-randomized-experiments","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/generic-machine-learning-inference-on-heterogenous-treatment-effects-in-randomized-experiments/123554/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What heterogeneous treatment-effect features does the paper focus on?","Question",{"text":75,"@type":76},"It targets best linear predictors of effects using ML proxies, impact-group sorted average effects, and averages comparing the most vs. least affected units.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the approach ensure valid inference when using machine learning?",{"text":80,"@type":76},"It relies on repeated data splitting to avoid overfitting and uses median-based p-values and confidence intervals, then adjusts nominal levels to guarantee uniform validity.",{"name":82,"@type":73,"acceptedAnswer":83},"What kinds of machine learning models can the method be used with?",{"text":84,"@type":76},"The framework is generic and can be combined with penalized methods, deep or shallow neural networks, canonical or new random forests, boosted trees, and ensemble methods.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]