[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123991-en":3,"doc-seo-123991-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123991,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","A structured comparison of causal machine learning methods to assess heterogeneous treatment effects in spatial data","A structured evaluation of causal machine learning approaches for heterogeneous treatment effects is presented for geographically referenced data. The study addresses a key limitation of standard causal-forest designs: random train/test splits break the spatial structure that links nearby geographic entities. Using a simulated dataset with known average and conditional average treatment effects, the work compares causal-forest model performance under alternative split definitions and introduces a spatial T-learner for unit-level heterogeneous effect estimation. Results show machine learning models outperform ordinary least squares for recovering the true average effect. The preferred model is applied to Valley Metro light rail construction, estimating how pre-treatment transit and pedestrian shares versus auto commuting relate to on-road per-capita CO2 changes across block groups in Maricopa County, Arizona.","Journal of Geographical Systems  \n[https://doi.org/10.1007/s10109-023-00413-0](https://doi.org/10.1007/s10109-023-00413-0)  \nORIGINAL ARTICLE  \nA structured comparison of causal machine learning methods to assess heterogeneous treatment effects in spatial data  \nKevin Credit1 · Matthew Lehnert2  \nReceived: 30 August 2022 / Accepted: 27 April 2023 © The Author(s) 2023  \nAbstract  \nThe development of the “causal” forest by Wager and Athey (J Am Stat Assoc 113(523): 1228–1242, 2018) represents a significant advance in the area of explanatory/causal machine learning. However, this approach has not yet been widely applied to geographically referenced data, which present some unique issues: the random split of the test and training sets in the typical causal forest design fractures the spatial fabric of geographic data. To help solve this issue, we use a simulated dataset with known properties for average treatment effects and conditional average treatment effects to compare the performance of CF models across different definitions of the test/train split. We also develop a new “spatial” T-learner that can be implemented using predictive methods like random forest to provide estimates of heterogeneous treatment effects across all units. Our results show that all of the machine learning models outperform traditional ordinary least squares regression at identifying the true average treatment effect, but are not significantly different from one another. We then apply the preferred causal forest model in the context of analysing the treatment effect of the construction of the Valley Metro light rail (tram) system on on-road CO2 emissions per capita at the block group level in Maricopa County, Arizona, and find that the neighbourhoods most likely to benefit from treatment are those with higher pre-treatment proportions of transit and pedestrian commuting and lower proportions of auto commuting.  \nKeywords Causal forest · Heterogeneous treatment effects · Machine learning · Causal inference · Spatial · CO2 emissions · Transit  \nJEL Classification C21 · C52 · C54 · C63  \n* Kevin Credit [kevin.credit@mu.ie](kevin.credit@mu.ie)  \n1 National Centre for Geocomputation, Maynooth University, Maynooth, Co. Kildare, Ireland  \n2 Satelytics, Perrysburg, Ohio, USA  \n1 3  \n1 Introduction  \nInitially, the use of machine learning techniques such as random forest and neural networks focused mostly on prediction tasks because of the “black box,” nonlinear nature of the relationship between input variables and outputs. But given the high predictive performance of these models compared to traditional statistical methods (like linear regression) (Strittmatter 2019 ; Farbmacher et al. 2021 ; Hagenauer et al. 2019 ; Yoshida and Seya 2021 ; Credit 2022), there has been considerable interest in developing new approaches for using these models to answer explanatory—and even causal—research questions in the social sciences. The development of the “causal” forest by Wager and Athey (2018) represents a significant advance in this area. By applying the fundamental logic of the principles of causal inference—“the science” of Rubin (2005)—to the powerful random forest model format, the causal forest not only provides an estimate of the average treatment effect (ATE) across the entire dataset, but also unit-level conditional average treatment effects (CATE) that allow researchers to draw conclusions about the effectiveness of treatment across various subpopulations (Athey and Imbens 2016) . Providing estimates of heterogeneous treatment effects (HTE) for all treated and untreated units are particularly valuable in the context of local urban policy decision-making and could allow planners to assess the viability of infrastructure investments or policies in candidate areas.  \nSeveral studies have used causal forests to analyse HTE for the effect of various interventions on outcomes such as student achievement, agricultural yields, crime, and corporate investment (Athey an","cbCaiaVMe9v7prya","https://ap.wps.com/l/cbCaiaVMe9v7prya","pdf",1570268,1,28,"English","en",105,"# Introduction\n## From prediction-focused ML to causal questions\n## Causal forests and heterogeneous treatment effects\n## Challenges for geographically referenced data","[{\"question\":\"Why do standard causal forest train/test splits cause problems for spatial data?\",\"answer\":\"Because random splitting fractures the spatial fabric of geographic observations, ignoring relationships driven by distance that should be reflected during modeling.\"},{\"question\":\"How does the study compare causal ML methods for heterogeneous treatment effects?\",\"answer\":\"It uses a simulated dataset with known ATE and CATE to evaluate causal-forest variants under different test/train split definitions, and compares them to ordinary least squares.\"},{\"question\":\"How is the proposed approach used in the real-world application?\",\"answer\":\"The preferred causal forest model estimates how the Valley Metro light rail construction affects on-road CO2 emissions per capita, identifying neighborhoods likely to benefit based on pre-treatment transit and pedestrian versus auto commuting shares.\"}]","A structured comparison of causal machine learning methods to assess heterogeneous treatment effects in spatial data | PDF",1785819697,71,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"a-structured-comparison-of-causal-machine-learning-methods-to-assess-heterogeneous-treatment-effects-in-spatial-data","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/a-structured-comparison-of-causal-machine-learning-methods-to-assess-heterogeneous-treatment-effects-in-spatial-data/123991/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why do standard causal forest train/test splits cause problems for spatial data?","Question",{"text":75,"@type":76},"Because random splitting fractures the spatial fabric of geographic observations, ignoring relationships driven by distance that should be reflected during modeling.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the study compare causal ML methods for heterogeneous treatment effects?",{"text":80,"@type":76},"It uses a simulated dataset with known ATE and CATE to evaluate causal-forest variants under different test/train split definitions, and compares them to ordinary least squares.",{"name":82,"@type":73,"acceptedAnswer":83},"How is the proposed approach used in the real-world application?",{"text":84,"@type":76},"The preferred causal forest model estimates how the Valley Metro light rail construction affects on-road CO2 emissions per capita, identifying neighborhoods likely to benefit based on pre-treatment transit and pedestrian versus auto commuting shares.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]