[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123996-en":3,"doc-seo-123996-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123996,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","Variable importance measure for spatial machine learning models with application to air pollution exposure prediction","Exposure assessment underpins air pollution cohort studies, where exposures must be predicted for individuals at locations lacking measurements. Beyond producing accurate predictions to reduce exposure measurement error, interpreting the mechanisms learned by complex machine learning models requires robust notions of variable importance, which are often not unified and are further obscured by spatial correlation. This work addresses the problem using leave-one-out variable importance for models with separable mean and covariance components, evaluated on two datasets: PM2.5 sulfur sub-species and ultrafine particles from Seattle traffic-related pollution. The method yields interpretable, comparable measures and distinguishes model mechanisms even when out-of-sample predictive accuracies are similar.","arXiv :2406 .01982v1 [ stat .AP] 4 Jun 2024  \nVariable importance measure for spatial machine learning models with application to air pollution exposure prediction  \nSi Cheng 1† Magali N. Blanco2 Lianne Sheppard 1 ,2 Ali Shojaie 1⋆  \nAdam Szpiro 1⋆  \n1 Department of Biostatistics, University of Washington  \n2 Department of Environmental & Occupational Health Sciences, University of Washington  \nAbstract  \nExposure assessment is fundamental to air pollution cohort studies. The objective is to predict air pollution exposures for study subjects at locations without data in order to optimize our ability to learn about health effects of air pollution. In addition to generating accurate predictions to minimize exposure measurement error, understanding the mechanism captured by the model is another crucial aspect that may not always be straightforward due to the complex nature of machine learning methods, as well as the lack of unifying notions of variable importance. This is further complicated in air pollution modeling by the presence of spatial correlation. We tackle these challenges in two datasets: sulfur (S) from regulatory United States national PM2.5 sub-species data and ultrafine particles (UFP) from a new Seattle-area traffic-related air pollution dataset. Our key contribution is a leave-one-out approach for variable importance that leads to interpretable and comparable measures for a broad class of models with separable mean and covariance components. We illustrate our approach with several spatial machine learning models, and it clearly highlights the difference in model mechanisms, even for those producing similar predictions. We leverage insights from this variable importance measure to assess the relative utilities of two exposure models for S and UFP that have similar out-of-sample prediction accuracies but appear to draw on different types of spatial information to make predictions.  \nKeywords: air pollution, machine learning, spatial modeling, variable importance  \n⋆ indicates co-senior authors  \n† contact: [si.cheng@aya.yale.edu](si.cheng@aya.yale.edu) ; Hans Rosling Center for Population Health, Box 351617, Seattle, WA 98195  \n1 Introduction  \nSpatial prediction models are versatile tools that provide deeper understanding of social or natural mechanisms and guide decision making in practice. Examples include crime analysis in sociology (Chainey et al. , 2008; Zhao and Tang, 2017; Yi et al. , 2018), nature disaster forecasting (Aggarwalet al. , 1975; Arnaud et al. , 2002; Parker et al. , 2017; Bui et al. , 2018; Karimzadeh et al. , 2019), and exposure assessment in public health (Monn, 2001; Kibria et al. , 2002; Kim et al. , 2009; Diasand Tchepel, 2018; Xu et al. , 2022) . The flexibility of machine learning (ML) models make them useful in prediction tasks with potentially complicated underlying mechanisms, but approaches to handling spatial structures in such models are relatively limited, compared to the abundance of ML methods, despite their practical importance.  \nKanevski (2009); Li et al. (2011); Du et al. (2020) provided reviews and discussions on the application of ML models in spatial settings. Some approaches incorporate spatial information into the features that are used in vanilla ML models (e.g. Kovacevic et al. , 2009; Cracknell and Reading, 2014; Hengl et al., 2015 , 2018), which are straightforward to implement but do not provide explicit information on spatial heterogeneity and/or correlation; some combine ML methods and spatial smoothing into two-step models (e.g. Bergen et al. , 2013; Liu et al. , 2018; Chen et al. , 2019; Blanco, 2021), which are flexible but may not partition the heterogeneity attributable to the mean and covariance components in an optimal way; and joint spatial-ML modeling (e.g. Datta et al. , 2016; Wai et al. , 2020; Saha et al. , 2021; Georganos et al. , 2021), which are better-suited for spatial prediction, but may lead to more intensive computation and/or less clear theo","cbCaid3xcS0oS5wS","https://ap.wps.com/l/cbCaid3xcS0oS5wS","pdf",7535635,1,45,"English","en",105,"# Abstract\n# Introduction\n## Spatial prediction models and applications\n## Challenges in model interpretation and variable importance\n## Permutation-based and other variable importance approaches\n## Difficulties of generalizing variable importance to spatial settings","[{\"question\":\"What problem does the paper address in air pollution exposure prediction?\",\"answer\":\"It addresses how to predict air pollution exposures at locations without direct data while also explaining what the spatial machine learning model is learning. It focuses on defining variable importance in the presence of spatial correlation.\"},{\"question\":\"Why is variable importance difficult in spatial machine learning models?\",\"answer\":\"Common variable importance approaches may become unreliable when features are correlated, which is common for spatial variables derived from geographic information system data. Differences among model classes also lead to diverse and often incomparable variable importance measures.\"},{\"question\":\"What key contribution does the paper propose?\",\"answer\":\"It introduces a leave-one-out approach for variable importance that produces interpretable and comparable measures for a broad class of spatial models. The method is designed for models with separable mean and covariance components.\"}]","Variable importance measure for spatial machine learning models with application to air pollution exposure prediction | PDF",1785819729,113,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"variable-importance-measure-for-spatial-machine-learning-models-with-application-to-air-pollution-exposure-prediction","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/variable-importance-measure-for-spatial-machine-learning-models-with-application-to-air-pollution-exposure-prediction/123996/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in air pollution exposure prediction?","Question",{"text":75,"@type":76},"It addresses how to predict air pollution exposures at locations without direct data while also explaining what the spatial machine learning model is learning. It focuses on defining variable importance in the presence of spatial correlation.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"Why is variable importance difficult in spatial machine learning models?",{"text":80,"@type":76},"Common variable importance approaches may become unreliable when features are correlated, which is common for spatial variables derived from geographic information system data. Differences among model classes also lead to diverse and often incomparable variable importance measures.",{"name":82,"@type":73,"acceptedAnswer":83},"What key contribution does the paper propose?",{"text":84,"@type":76},"It introduces a leave-one-out approach for variable importance that produces interpretable and comparable measures for a broad class of spatial models. The method is designed for models with separable mean and covariance components.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]