[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-125851-en":3,"doc-seo-125851-105":31,"detail-sidebar-cat-0-en-105":93},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},125851,1099523882367,"Hazel","https://ap-avatar.wpscdn.com/davatar_9964176cb1d06d4a9deccf72a44ae3dc",8,"Research & Report","Statistical Performance Guarantee for Subgroup Identification with Generic Machine Learning","Machine learning is widely used to find exceptional responders—individuals most likely to benefit from a treatment—or those harmed by it. A common workflow estimates conditional average treatment effects (CATE) and then ranks individuals by the predicted effects, but CATE estimates can be biased or noisy, and using the same data for both subgroup selection and group evaluation creates a multiple testing issue. This work provides uniform confidence bands for grouped average treatment effects (GATES) by generic ML algorithms, enabling statistically guaranteed subgroup identification under random assignment and random sampling without modeling assumptions, with no resampling burden. Simulations and a late-stage prostate cancer trial demonstrate informative bands and reasonable coverage.","Statistical Performance Guarantee for Subgroup Identification with Generic Machine Learning ∗  \nMichael Lingzhi Li† Kosuke Imai‡  \narXiv :2310 .07973v3 [ stat .ME] 31 Aug 2025  \nSeptember 3, 2025  \nAbstract  \nAcross a wide array of disciplines, many researchers use machine learning (ML) algorithms to identify a subgroup of individuals who are likely to benefit from a treatment the most (“exceptional responders”) or those who are harmed by it. A common approach to this subgroup identification problem consists of two steps. First, researchers estimate the conditional average treatment effect (CATE) using an ML algorithm. Next, they use the estimated CATE to select those individuals who are predicted to be most affected by the treatment, either positively or negatively. Unfortunately, CATE estimates are often biased and noisy. In addition, utilizing the same data to both identify a subgroup and estimate its group average treatment effect results ina multiple testing problem. To address these challenges, we develop uniform confidence bands for estimation of the group average treatment effect sorted by generic ML algorithm (GATES) . Using these uniform confidence bands, researchers can identify, with a statistical guarantee, a subgroup whose GATES exceeds a certain effect size, regardless of how this effect size is chosen. The validity of the proposed methodology depends solely on randomization of treatment and random sampling of units. Importantly, our method does not require modeling assumptions and avoids a computationally intensive resampling procedure. A simulation study shows that the proposed uniform confidence bands are reasonably informative and have an appropriate empirical coverage even when the sample size is as small as 100 . We analyze a clinical trial of late-stage prostate cancer and find a relatively large proportion of exceptional responders.  \nKey Words: causal inference, exceptional responders, heterogeneous treatment effects, treatment prioritization, uniform confidence bands  \n∗ The proposed methodology is implemented through an open-source R package, evalITR, which is freely available for download at the Comprehensive R Archive Network (CRAN; [https://CRAN.R-project.org/package=evalITR](https://CRAN.R-project.org/package=evalITR)). We thank an anonymous reviewer from the Alexander and Diviya Magaro Peer Pre-Review Program at Harvard’s Institute for Quantitative Social Science for helpful comments.  \n†Technology and Operations Management, Harvard Business School, Boston, MA 02163 . Email: [mili@hbs.edu](mili@hbs.edu), URL: [https://michaellz.com](https://michaellz.com)  \n‡Department of Government and Department of Statistics, Harvard University, Cambridge, MA 02138 . Phone: 617–384–6778, Email: Imai@Harvard.Edu, URL: [https://imai.fas.harvard.edu](https://imai.fas.harvard.edu)  \n1 Introduction  \nIn a diverse set of fields, machine learning (ML) algorithms are frequently used to identify a subgroup of individuals who are likely to benefit from a treatment the most (so called “exceptional responders”) or those who are negatively affected by it. In clinical studies, for example, identifying those who experience surprisingly long-lasting positive health outcomes when a majority of patients do not, is useful for understanding why a treatment works and determining which patients should receive it (e.g., Takebe et al. , 2015; Conley et al. , 2020) . Similarly, it is critical to identify asubpopulation of patients who are harmed by a commonly used treatment. In public policy, finding those individuals who are likely to be helped by a treatment is a key methodological step for improving targeting strategies (e.g., Imai and Ratkovic, 2013; Chernozhukov et al. , 2019) .  \nA widely adopted strategy for subgroup identification proceeds in two steps. First, an ML algorithm is used to estimate the conditional average treatment effect (CATE) or its proxy such as biomarker response, conditional on pre-treatment covariates. Sec","cbCaibDX48ZUBKnJ","https://ap.wps.com/l/cbCaibDX48ZUBKnJ","pdf",1020241,5,1,40,"English","en",105,"# Introduction","[{\"question\":\"What problem does the paper address in subgroup identification with machine learning?\",\"answer\":\"It addresses biased and noisy CATE estimates and the multiple testing problem that arises when the same data is used to select a subgroup and estimate that subgroup’s average treatment effect.\"},{\"question\":\"How does the proposed method provide statistical guarantees?\",\"answer\":\"It develops uniform confidence bands for the grouped average treatment effect (GATES) sorted by generic ML algorithms, allowing researchers to identify a subgroup exceeding a chosen effect size with a statistical guarantee.\"},{\"question\":\"What assumptions are required for validity and what does the method avoid?\",\"answer\":\"Validity depends only on randomized treatment assignment and random sampling of units. The method does not require modeling assumptions and avoids computationally intensive resampling.\"}]","Statistical Performance Guarantee for Subgroup Identification with Generic Machine Learning | PDF",1785901583,101,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":88,"head_meta":90,"extra_data":92,"updated_unix":29},"statistical-performance-guarantee-for-subgroup-identification-with-generic-machine-learning","",{"@graph":37,"@context":87},[38,55,70],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":54},"https://docshare.wps.com/document/statistical-performance-guarantee-for-subgroup-identification-with-generic-machine-learning/125851/",4,{"url":53,"name":13,"@type":56,"author":57,"headline":13,"publisher":59,"fileFormat":62,"inLanguage":24,"description":14,"dateModified":63,"datePublished":64,"encodingFormat":62,"isAccessibleForFree":65,"interactionStatistic":66},"DigitalDocument",{"name":9,"@type":58},"Person",{"url":42,"name":60,"@type":61},"DocShare","Organization","application/pdf","2026-08-22","2026-08-05",true,{"@type":67,"interactionType":68,"userInteractionCount":20},"InteractionCounter",{"@type":69},"ViewAction",{"@type":71,"mainEntity":72},"FAQPage",[73,79,83],{"name":74,"@type":75,"acceptedAnswer":76},"What problem does the paper address in subgroup identification with machine learning?","Question",{"text":77,"@type":78},"It addresses biased and noisy CATE estimates and the multiple testing problem that arises when the same data is used to select a subgroup and estimate that subgroup’s average treatment effect.","Answer",{"name":80,"@type":75,"acceptedAnswer":81},"How does the proposed method provide statistical guarantees?",{"text":82,"@type":78},"It develops uniform confidence bands for the grouped average treatment effect (GATES) sorted by generic ML algorithms, allowing researchers to identify a subgroup exceeding a chosen effect size with a statistical guarantee.",{"name":84,"@type":75,"acceptedAnswer":85},"What assumptions are required for validity and what does the method avoid?",{"text":86,"@type":78},"Validity depends only on randomized treatment assignment and random sampling of units. The method does not require modeling assumptions and avoids computationally intensive resampling.","https://schema.org",{"og:url":53,"og:type":89,"og:title":13,"og:site_name":60,"og:description":14},"article",{"robots":91,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":94},[95,99,103,107,111,116,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":96,"show_sort_weight":97,"slug":98},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":100,"show_sort_weight":101,"slug":102},"Literature",80,"literature",{"id":54,"doc_module":4,"doc_module_name":47,"category_name":104,"show_sort_weight":105,"slug":106},"Exam",70,"exam",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":22,"slug":119},7,"Healthcare","healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":20,"slug":138},19,"General","general"]