[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119669-en":3,"doc-seo-119669-105":30,"detail-sidebar-cat-0-en-105":95},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119669,7971461741311,"Ophelia","https://ap-avatar.wpscdn.com/avatar/74000253aff267980c6?x-image-process=image/resize,m_fixed,w_180,h_180&k=1779345379180704826",8,"Research & Report","Selective machine learning of doubly robust functionals - Research paper overview","Selective machine learning addresses model selection in semiparametric settings where nuisance parameters may be high-dimensional and the target is a finite-dimensional functional. The work introduces two model selection criteria that reduce bias by using a new pseudo-risk definition tailored to doubly robust estimating functions. A multi-fold cross-validation version is shown to have an oracle property, nearly matching an oracle with knowledge of each candidate model’s pseudo-risk, and a smooth approximation enables valid post-selection inference. The method is applied to selecting semiparametric models for average treatment effect estimation using ensembles of machine learners to handle confounding in observational studies.","View metadata, citation and similar [papers at ](papers at core.ac.uk)[core.ac.uk](papers at core.ac.uk) brought to you by CORE  \n[provided by](provided by arXiv.org)[ arXiv.org](provided by arXiv.org) e-Print Archive  \narXiv : 19 11 .02029v2 [ stat .ME] 29 Apr 2020  \nSelective machine learning of doubly robust functionals  \nYifan Cui, Eric Tchetgen Tchetgen  \nAbstract  \nWhile model selection is a well-studied topic in parametric and nonparametric regression or density estimation, model selection of possibly high-dimensional nuisance parameters in semiparametric problems is far less developed. In this paper, we propose a selective machine learning framework for making inferences about a 􀀌nite-dimensional functional de􀀌ned on a semiparametric model, when the latter admits a doubly robust estimating function. We introduce two model selection criteria for bias reduction of functional of interest, each based on a novel de􀀌nition of pseudo-risk for the functional that embodies this double robustness property and thus may be used to select the candidate model that is nearest to ful􀀌lling this property even when all models are wrong. We establish an oracle property for a multi-fold cross-validation version of the new model selection criteria which states that our empirical criteria perform nearly as well as an oracle with a priori knowledge of the pseudo-risk for each candidate model. We also describe a smooth approximation to the selection criteria which allows for valid post-selection inference. Finally, we apply the approach to model selection of a semiparametric estimator of average treatment e􀀋ect given an ensemble of candidate machine learners to account for confounding inan observational study.  \nkeywords Machine Learning, Doubly Robust, In􀀍uence Function, Model Selection, Average Treatment E􀀋ect  \n1 Introduction  \nModel selection is a well-studied topic in statistics, econometrics and machine learning. In fact, methods for model selection and corresponding theory abound in these disciplines, although primarily in settings of parametric and nonparametric regression and density estimation. Model selection methods are far less developed in settings where one aims to make inferences about a 􀀌nite dimensional, pathwise di􀀋erentiable functional de􀀌ned on a semiparametric model. Model selection for the purpose of estimating such a functional may involve selection of an in􀀌nite dimensional parameter, say a nonparametric regression for the purpose of more accurate estimation of functional in view, which can be considerably more challenging than selecting a regression model strictly for the purpose of prediction. This is because whereas the latter admitsa risk, e.g., mean squared error loss, that can be estimated unbiasedly and therefore can be minimized with small error, the risk of a semiparametric functional will typically not admit an unbiased estimator and therefore may not be minimized without excessive error. This is an important gap in both model selection and semiparametric theory which this paper aims to address.  \nSpeci􀀌cally, we propose a novel approach for model selection of a functional de-􀀌ned on a semiparametric model, in settings where inferences about the targeted functional involves in􀀌nite dimensional nuisance parameters, and the functional ofscienti􀀌c interest admits a doubly robust estimating function. Doubly robust inference (Robins et al. , 1994) has received considerable interest in the past few years across multiple disciplines including statistics, epidemiology and econometrics. An estimator is said to be doubly robust if it remains consistent if one of two nuisance parameters needed for estimation is consistent, even if both are not necessarily consistent. The class of functionals that admit doubly robust estimators is quite rich,  \nand includes estimation of pathwise di􀀋erentiable functionals in missing data problems under missing at random assumptions, and also in more complex settings where missingness pr","cbCairnAcKmoTTcM","https://ap.wps.com/l/cbCairnAcKmoTTcM","pdf",512521,1,54,"English","en",105,"# Introduction\n## Doubly robust inference and motivating challenges\n## Proposed selective model selection framework\n## Pseudo-risk criteria and oracle results\n## Cross-validation and post-selection inference\n## Application to average treatment effect estimation","[{\"question\":\"What problem does the paper target in model selection?\",\"answer\":\"It targets model selection in semiparametric inference when nuisance parameters can be high-dimensional, and standard selection based on prediction-oriented losses may not work for finite-dimensional functionals.\"},{\"question\":\"What is the core idea behind the proposed selection criteria?\",\"answer\":\"Both criteria minimize a cross-validated quadratic pseudo-risk defined to embody double robustness, so the selected candidate model is biased-reduced relative to satisfying that property.\"},{\"question\":\"How does the paper justify the performance of the criteria?\",\"answer\":\"It establishes an oracle property for the multi-fold cross-validation version, showing the empirical criteria perform nearly as well as an oracle that knows the pseudo-risk for each candidate model.\"},{\"question\":\"How is the approach used in practice?\",\"answer\":\"The paper applies it to selecting a semiparametric estimator of the average treatment effect by using an ensemble of candidate machine learners to address confounding in observational studies.\"}]","Selective machine learning of doubly robust functionals - Research paper overview | PDF",1785725605,136,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":90,"head_meta":92,"extra_data":94,"updated_unix":28},"selective-machine-learning-of-doubly-robust-functionals-research-paper-overview","",{"@graph":36,"@context":89},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/selective-machine-learning-of-doubly-robust-functionals-research-paper-overview/119669/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81,85],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper target in model selection?","Question",{"text":75,"@type":76},"It targets model selection in semiparametric inference when nuisance parameters can be high-dimensional, and standard selection based on prediction-oriented losses may not work for finite-dimensional functionals.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What is the core idea behind the proposed selection criteria?",{"text":80,"@type":76},"Both criteria minimize a cross-validated quadratic pseudo-risk defined to embody double robustness, so the selected candidate model is biased-reduced relative to satisfying that property.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the paper justify the performance of the criteria?",{"text":84,"@type":76},"It establishes an oracle property for the multi-fold cross-validation version, showing the empirical criteria perform nearly as well as an oracle that knows the pseudo-risk for each candidate model.",{"name":86,"@type":73,"acceptedAnswer":87},"How is the approach used in practice?",{"text":88,"@type":76},"The paper applies it to selecting a semiparametric estimator of the average treatment effect by using an ensemble of candidate machine learners to address confounding in observational studies.","https://schema.org",{"og:url":52,"og:type":91,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":93,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":96},[97,101,105,109,114,119,124,127,132,135,139],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":106,"show_sort_weight":107,"slug":108},"Exam",70,"exam",{"id":110,"doc_module":4,"doc_module_name":46,"category_name":111,"show_sort_weight":112,"slug":113},5,"Comic",60,"comic",{"id":115,"doc_module":4,"doc_module_name":46,"category_name":116,"show_sort_weight":117,"slug":118},6,"Technology",50,"technology",{"id":120,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":122,"slug":123},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":125,"slug":126},30,"research-report",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":130,"slug":131},9,"Religion & Spirituality",20,"religion-spirituality",{"id":130,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":130,"slug":134},"World Cup","world-cup",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":136,"slug":138},10,"Lifestyle","lifestyle",{"id":140,"doc_module":4,"doc_module_name":46,"category_name":141,"show_sort_weight":110,"slug":142},19,"General","general"]