[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-117739-en":3,"doc-seo-117739-105":30,"detail-sidebar-cat-0-en-105":83},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},117739,1099513958607,"Jiven","https://ap-avatar.wpscdn.com/avatar/100002390cf8733938c?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778829742770036399",8,"Research & Report","Robustness Against Weak or Invalid Instruments - Exploring Nonlinear Treatment Models with Machine Learning","Causal inference in observational studies can fail when instrumental variables (IVs) violate core assumptions, especially when IV strength is weak or validity is questionable. This work introduces two-stage curvature identification (TSCI), combining machine learning for a nonlinear treatment model with a second-stage adjustment for different violation forms. A bias-correction step addresses excess complexity in the ML stage. The resulting TSCI estimator is asymptotically unbiased and Gaussian, even when the ML treatment model is inconsistent, and supports data-driven selection among candidate violation forms, illustrated via education-earnings analysis.","arXiv :2203 . 12808v4 [ stat .ME] 5 Jan 2024  \nRobustness Against Weak or Invalid Instruments: Exploring Nonlinear Treatment Models with Machine Learning  \nZijian Guo, Mengchu Zheng  \nDepartment of Statistics, Rutgers University, USA  \nPeter Bhlmann  \nSeminar for Statistics, ETH Zrich, Switzerland.  \nSummary. We discuss causal inference for observational studies with possibly invalid instrumental variables. We propose a novel methodology called two-stage curvature identification (TSCI) by exploring the nonlinear treatment model with machine learning. The first-stage machine learning enables improving the instrumental variable’s strength and adjusting for different forms of violating the instrumental variable assumptions. The success of TSCI requires the instrumental variable’s effect on treatment to differ from its violation form. A novel bias correction step is implemented to remove bias resulting from the potentially high complexity of machine learning. Our proposed TSCI estimator is shown to be asymptotically unbiased and Gaussian even if the machine learning algorithm does not consistently estimate the treatment model. Furthermore, we design a data-dependent method to choose the best among several candidate violation forms. We apply TSCI to study the effect of education on earnings.  \nKeywords: Generalized IV strength, Confidence interval, Random forests, Boosting, Neural network.  \n1. Introduction  \nObservational studies are major sources for inferring causal effects when randomized experiments are not feasible. But such causal inference from observational studies requires strong assumptions and may be invalid due to the presence of unmeasured confounders. The instrumental variable (IV) regression is a practical and highly popular causal inference approach in the presence of unmeasured confounders. The IVs are required to satisfy three assumptions: conditioning on the baseline covariates, (A1) the IVs are associated with the treatment; (A2) the IVs are not associated with the unmeasured confounders; (A3) the IVs do not directly affect the outcome.  \nDespite the popularity of the IV method, there is a significant concern about whether the used IVs satisfy (A1)-(A3) in practice. Assumption (A1) requires the IV to be strongly associated with the treatment variable, which can be checked with an F-test ina first-stage linear regression model. Inference with the assumption (A1) being violated has been actively investigated under the name of weak IV (Stock et al., 2002; Staiger and Stock, 1997, e.g.) . Assumptions (A2) and (A3) ensure that the IV only affects the outcome through the treatment. If an IV violates either (A2) or (A3), we call it an invalid IV and define its functional form of violating (A2) and (A3) as the violation form, e.g., a linear violation. In the just-identification regime, most empirical analyses  \n2 Guo, Zheng and Bhlmann  \nrely on external knowledge to argue about the validity of (A2) and (A3) . However, thereis a pressing need to develop a robust causal inference method against the proposed IVs violating the classical assumptions.  \n1. 1. Our results and contribution  \nWe aim to devise a robust IV framework that leads to causal identification even if the IVs proposed by domain experts may violate assumptions (A1) to (A3) . Our framework provides a robustness guarantee even when all proposed IVs are invalid, including the most common regime with a single IV that is possibly invalid. It is well known that the treatment effect is not identifiable when there is no constraint on how the IVs violate assumptions (A2) and (A3) . Our key identification assumption is that the violations of (A2) and (A3) arise from simpler forms than the association between the treatment and the IVs; that is, we exclude “special coincidences” such as the IVs violate (A2) and (A3) by linear forms and the conditional mean model of the treatment given IVs is also linear. This identification condition can be evaluated with the general","cbCaibFyyyixiIHX","https://ap.wps.com/l/cbCaibFyyyixiIHX","pdf",1331017,1,62,"English","en",105,"# Introduction\n## Our results and contribution\n## Comparison to existing literature\n## Method overview (TSCI)","[{\"question\":\"Why is bias correction needed, and what theoretical guarantees are provided?\",\"answer\":\"Bias correction removes errors caused by high complexity and potential overfitting from the ML stage. The paper shows the TSCI estimator is asymptotically unbiased and Gaussian under sufficiently large generalized IV strength, and proposes a data-dependent method to choose the best violation form.\"}]","Robustness Against Weak or Invalid Instruments - Exploring Nonlinear Treatment Models with Machine Learning | PDF",1785679289,156,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":78,"head_meta":80,"extra_data":82,"updated_unix":28},"robustness-against-weak-or-invalid-instruments-exploring-nonlinear-treatment-models-with-machine-learning","",{"@graph":36,"@context":77},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/robustness-against-weak-or-invalid-instruments-exploring-nonlinear-treatment-models-with-machine-learning/117739/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-02",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71],{"name":72,"@type":73,"acceptedAnswer":74},"Why is bias correction needed, and what theoretical guarantees are provided?","Question",{"text":75,"@type":76},"Bias correction removes errors caused by high complexity and potential overfitting from the ML stage. The paper shows the TSCI estimator is asymptotically unbiased and Gaussian under sufficiently large generalized IV strength, and proposes a data-dependent method to choose the best violation form.","Answer","https://schema.org",{"og:url":52,"og:type":79,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":81,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":84},[85,89,93,97,102,107,112,115,120,123,127],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":86,"show_sort_weight":87,"slug":88},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":90,"show_sort_weight":91,"slug":92},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Exam",70,"exam",{"id":98,"doc_module":4,"doc_module_name":46,"category_name":99,"show_sort_weight":100,"slug":101},5,"Comic",60,"comic",{"id":103,"doc_module":4,"doc_module_name":46,"category_name":104,"show_sort_weight":105,"slug":106},6,"Technology",50,"technology",{"id":108,"doc_module":4,"doc_module_name":46,"category_name":109,"show_sort_weight":110,"slug":111},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":113,"slug":114},30,"research-report",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},9,"Religion & Spirituality",20,"religion-spirituality",{"id":118,"doc_module":4,"doc_module_name":46,"category_name":121,"show_sort_weight":118,"slug":122},"World Cup","world-cup",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":124,"slug":126},10,"Lifestyle","lifestyle",{"id":128,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":98,"slug":130},19,"General","general"]