[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-123987-en":3,"doc-seo-123987-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},123987,8796095462418,"Noah","https://ap-avatar.wpscdn.com/avatar/80000253c1241d02b47?x-image-process=image/resize,m_fixed,w_180,h_180&k=1778826106357471780",8,"Research & Report","On the use of adversarial validation for quantifying dissimilarity in geospatial machine learning prediction","Recent geospatial machine learning research shows that cross-validation (CV) outcomes depend strongly on how dissimilar the sample data are from the prediction locations. This paper introduces a feature-space method to quantify that dissimilarity on a 0–100% scale using adversarial validation via a binary classifier. Experiments on synthetic and real datasets with increasing dissimilarities demonstrate full-range quantification and clarify CV behavior under varying mismatch levels.","arXiv :2404 . 12575v1 [ cs .LG] 19 Apr 2024  \nOn the use of adversarial validation for quantifying dissimilarity in geospatial machine learning prediction  \nYanwen Wanga,* , Mahdi Khodadadzadeha , and Ra´ul Zurita-Millaa  \na Faculty of Geo-Information Science and Earth Observation (ITC), University of Twente, 7522NH Enschede, the Netherlands  \nARTICLE HISTORY  \nCompiled April 22, 2024  \nABSTRACT  \nRecent geospatial machine learning studies have shown that the results of model  \nevaluation via cross-validation (CV) are strongly affected by the dissimilarity be  \ntween the sample data and the prediction locations. In this paper, we propose a  \nmethod to quantify such a dissimilarity in the interval 0 to 100%, and from the  \nperspective of the data feature space. The proposed method is based on adversarial  \nvalidation, which is an approach that can check whether sample data and prediction  \nlocations can be separated with a binary classifier. To study the effectiveness and  \ngenerality of our method, we tested it on a series of experiments based on both syn  \nthetic and real datasets and with gradually increasing dissimilarities. Results show  \nthat the proposed method can successfully quantify dissimilarity across the entire  \nrange of values. Next to this, we studied how dissimilarity affects CV evaluations  \nby comparing the results of random CV and of two spatial CV methods, namely  \nblock and spatial+ CV. Our results showed that CV evaluations follow similar pat  \nterns in all datasets and predictions: when dissimilarity is low (usually lower than  \n30%), random CV provides the most accurate evaluation results. As dissimilarity  \nincreases, spatial CV methods, especially spatial+ CV, become more and more ac  \ncurate and even outperforming random CV. When dissimilarity is high (>=90%),  \nno CV method provides accurate evaluations. These results show the importance of  \nconsidering feature space dissimilarity when working with geospatial machine learn  \ning predictions, and can help researchers and practitioners to select more suitable  \nCV methods for evaluating their predictions.  \nKEYWORDS  \nMachine learning; geospatial regression; model evaluation; cross-validation; feature  \nspace.  \n1. Introduction  \nMachine learning (ML) is widely used in geospatial prediction to estimate unknown values at specific prediction locations (Hengl et al. 2018; Aguilar et al. 2018; Usmanet al. 2023) . These predictions are often done to create spatially continuous products, for example, mineral (Khodadadzadeh and Gloaguen 2019), health risk (Garcia-Marti et al. 2018), or phenological (Zurita-Milla, Laurent, and van Gijsel 2015) maps. In these and many other applications, predictions come from ML regression models trained on limited sample data, where the number of samples is typically much smaller than the  \nnumber of prediction locations. This imbalance between samples and prediction locations, is mostly due to practical limitations such as accessibility (Lamichhane, Kumar, and Wilson 2019) or sampling costs (Hengl et al. 2015) . For similar reasons, collecting additional data for an independent evaluation of geospatial ML prediction is rarely feasible (Valavi et al. 2019) . To address these operational constraints, the evaluation of geospatial ML models is mainly conducted by splitting the available sample data into training and validation subsets (de Bruin et al. 2022; Wang, Khodadadzadeh, and Zurita-Milla 2023) . Random k-fold cross-validation (RDM-CV) stands out as the most popular evaluation method (Chen et al. 2018; Nesha et al. 2020; Guo et al. 2022) . Asthe name indicates, RMD-CV randomly splits the sample data into k equal-size folds, and then, it iteratively uses one of them as a validation subset and the remaining ones as a training subset. When sample data are collected by a probability sampling strategy, such as simple random sampling (Brus, Kempen, and Heuvelink 2011; Wanget al. 2012) and regular sampling (Lagacherie et al. 20","cbCaij4Ujphtjj4V","https://ap.wps.com/l/cbCaij4Ujphtjj4V","pdf",27571593,1,18,"English","en",105,"# Introduction\n## Motivation for geospatial ML evaluation and CV\n## Random k-fold cross-validation and its limitations\n## Feature-space dissimilarity and spatial prediction settings","[{\"question\":\"What problem does the paper address in geospatial machine learning evaluation?\",\"answer\":\"It addresses that CV results for geospatial predictions are strongly affected by the dissimilarity between training sample data and the prediction locations.\"},{\"question\":\"How does the proposed method quantify dissimilarity?\",\"answer\":\"It uses adversarial validation, training a binary classifier to test whether sample data and prediction locations can be separated in the data feature space, producing a 0–100% dissimilarity measure.\"},{\"question\":\"How does dissimilarity level affect which CV method works best?\",\"answer\":\"When dissimilarity is low (typically below 30%), random CV is most accurate; as dissimilarity increases, spatial CV methods—especially spatial+ CV—become more accurate; at very high dissimilarity (\\u003e=90%), no CV method yields accurate evaluations.\"}]","On the use of adversarial validation for quantifying dissimilarity in geospatial machine learning prediction | PDF",1785819664,45,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"on-the-use-of-adversarial-validation-for-quantifying-dissimilarity-in-geospatial-machine-learning-prediction","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/on-the-use-of-adversarial-validation-for-quantifying-dissimilarity-in-geospatial-machine-learning-prediction/123987/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-04",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What problem does the paper address in geospatial machine learning evaluation?","Question",{"text":75,"@type":76},"It addresses that CV results for geospatial predictions are strongly affected by the dissimilarity between training sample data and the prediction locations.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How does the proposed method quantify dissimilarity?",{"text":80,"@type":76},"It uses adversarial validation, training a binary classifier to test whether sample data and prediction locations can be separated in the data feature space, producing a 0–100% dissimilarity measure.",{"name":82,"@type":73,"acceptedAnswer":83},"How does dissimilarity level affect which CV method works best?",{"text":84,"@type":76},"When dissimilarity is low (typically below 30%), random CV is most accurate; as dissimilarity increases, spatial CV methods—especially spatial+ CV—become more accurate; at very high dissimilarity (>=90%), no CV method yields accurate evaluations.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]