[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-126357-en":3,"doc-seo-126357-105":31,"detail-sidebar-cat-0-en-105":92},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":28,"seo_description":14,"update_tm":29,"read_time":30},126357,962085564381,"Clementine","https://ap-avatar.wpscdn.com/davatar_6f874abed73319feea01a86fa6f0fab8",8,"Research & Report","Downscaling soil moisture to sub-km resolutions with simple machine learning ensembles - Paper Supplement","Paper supplement detailing model components and data preparation for downscaling soil moisture predictions to sub-kilometer resolutions using simple machine learning ensembles. It specifies a probabilistic layer that maps posterior multivariate normals to fixed priors via KL divergence to regularize training, compares time-padded versus temporally accurate datasets, and describes feature selection grounded in links between NDVI, LST, evapotranspiration, soil texture, and topography. It also provides resampling and normalization/scaling factors, including variable-specific spatial and temporal resolutions.","Paper Supplement  \nContents  \n1 Model Architecture Supplement 1  \n2 Dataset Supplement 2  \n2.1 Feature Selection ........................................ 2  \n2.2 Scaling .............................................. 3  \n2.2.1 Precipitation Data ................................... 3  \n2.3 Composition .......................................... 3  \n2.4 OK Datasets .......................................... 4  \n3 Spatial Predictions 7  \n4 Cross Validation 9  \n5 SHAP 9  \n6 Ensemble Advantage 10  \n7 Domain Preference 10  \n8 Code and Data availability 12  \n1 Model Architecture Supplement  \nProbabilistic Layer  \nThe hidden probabilistic layer in the Prob model serves to learn and map a posterior multivariate normal distribution onto a prior multivariate normal distribution. The prior multivariate distribution has fixed standard deviations of 0.5 with learnable mean values. The posterior multivariate normal distribution has both learnable mean and standard deviation.  \nDuring the training process, the layer tries to learn the prior distribution for all values fed into it. It then tries to learn the distribution for the posterior given the inputs. These two values are then compared via the Kullback-Leibler (KL) divergence to gauge similarity. This is done to ensure the posterior distribution is not overfitting a specific sample and is penalized for deviating too far from the prior distribution. The posterior distribution is then condensed to two values and passed forward in the network to the Independent Normal layer.  \nTemporal Resolution  \nThe padded model doesn’t just out perform the daily model on the dataset as a whole. The padded model also outperformed the daily model on a site by site basis for each metric except for R on the daily validation dataset (Fig. 1b) . This is not unexpected as the daily model had greater LST variations in it’s dataset than the padded model. The daily training model exhibits slight biased against low SWC readings. This is visible in the heatmaps of Figure 1a. At near zero in-situ SWC measurement readings the daily model has a strong cluster of predictions around 0 .1 m3 /m3 . The mechanism for this is unknown, however, it seems apparent that training on additional samples helped the model  \nidentify lower SWC trends.  \nDaily Model Padded Model  \nPred icted (m³/m³ )  \nPadded Dataset Daily Dataset  \n100  \n50  \n0  \n100  \n50  \n0  \n(b)  \nFigure 1: a) Predictions for a model trained on a time-padded dataset which contains much more samples (658,000) to learn from and a model trained on a temporally accurate dataset (372,000) . Both models predict on the validation sets for each dataset. b) Head to head for these models on sites in each dataset. If a model outperforms the other in a metric the bar increases by one.  \n2 Dataset Supplement  \n2.1 Feature Selection  \nThe variables selected SMAP, ND VI, LST, Precipitation, Sand and Clay content, pH, Evapotranspiration, and Topography/Elevation. are linked to SWC through multiple mechanisms.  \nNDVI, LST, and ET  \nVegetation Index (NDVI), and Evapotranspiration (ET) Land surface temperature (LST) has a very strong coupling with SWC. As LST increases, more energy is available for SWC to harness in order to evaporate and leave the soil. This relationship is well established and exploited to benefit in DisPATCH algorithms. NDVI corresponds to plant greenness and plant cover over an area. Because plants require water for healthy efficient production, NDVI has been correlated to SWC on multiple occasions[1][2] . Evapotranspiration (ET) is also included as a variable as it is directly associated with SWC.  \nSoil Texture  \nSWC is directly influenced by the physical properties of the soil, such as texture and composition. Porosity and grain size directly influence the cohesive and adhesive properties of water which permit capillary rise. The greater the surface area by volume, the easier it is for water to adhere to mineral surfaces and resist extracting forces such as the","cbCaib8PaaLz8jKl","https://ap.wps.com/l/cbCaib8PaaLz8jKl","pdf",3821799,4,1,12,"English","en",105,"# 1 Model Architecture Supplement\n## Probabilistic Layer\n## Temporal Resolution\n# 2 Dataset Supplement\n## 2.1 Feature Selection\n## 2.2 Scaling\n# 3 Spatial Predictions\n# 4 Cross Validation\n# 5 SHAP\n# 6 Ensemble Advantage\n# 7 Domain Preference\n# 8 Code and Data availability","[{\"question\":\"What role does the probabilistic layer play in the model?\",\"answer\":\"It learns mappings between posterior and prior multivariate normal distributions and regularizes training by penalizing divergence from the prior using KL divergence.\"},{\"question\":\"How does the time-padded dataset affect performance compared with the temporally accurate dataset?\",\"answer\":\"The padded model improves results site-by-site for most metrics, with an exception for the R metric on the daily validation dataset, attributed to different LST variation patterns between datasets.\"},{\"question\":\"Which factors are selected as predictors for soil water content (SWC)?\",\"answer\":\"Selected variables include SMAP, NDVI, LST, precipitation, sand/clay content, pH, evapotranspiration, and topography/elevation, justified by multiple physical and ecohydrological mechanisms linking them to SWC.\"}]","Downscaling soil moisture to sub-km resolutions with simple machine learning ensembles - Paper Supplement | PDF",1785904648,30,{"code":4,"msg":32,"data":33},"ok",{"site_id":25,"language":24,"slug":34,"title":13,"keywords":35,"description":14,"schema_data":36,"social_meta":87,"head_meta":89,"extra_data":91,"updated_unix":29},"downscaling-soil-moisture-to-sub-km-resolutions-with-simple-machine-learning-ensembles-paper-supplement","",{"@graph":37,"@context":86},[38,54,69],{"@type":39,"itemListElement":40},"BreadcrumbList",[41,45,49,52],{"item":42,"name":43,"@type":44,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":46,"name":47,"@type":44,"position":48},"https://docshare.wps.com/document/","Document",2,{"item":50,"name":12,"@type":44,"position":51},"https://docshare.wps.com/document/research-report/",3,{"item":53,"name":13,"@type":44,"position":20},"https://docshare.wps.com/document/downscaling-soil-moisture-to-sub-km-resolutions-with-simple-machine-learning-ensembles-paper-supplement/126357/",{"url":53,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":24,"description":14,"dateModified":62,"datePublished":63,"encodingFormat":61,"isAccessibleForFree":64,"interactionStatistic":65},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":42,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-21","2026-08-05",true,{"@type":66,"interactionType":67,"userInteractionCount":20},"InteractionCounter",{"@type":68},"ViewAction",{"@type":70,"mainEntity":71},"FAQPage",[72,78,82],{"name":73,"@type":74,"acceptedAnswer":75},"What role does the probabilistic layer play in the model?","Question",{"text":76,"@type":77},"It learns mappings between posterior and prior multivariate normal distributions and regularizes training by penalizing divergence from the prior using KL divergence.","Answer",{"name":79,"@type":74,"acceptedAnswer":80},"How does the time-padded dataset affect performance compared with the temporally accurate dataset?",{"text":81,"@type":77},"The padded model improves results site-by-site for most metrics, with an exception for the R metric on the daily validation dataset, attributed to different LST variation patterns between datasets.",{"name":83,"@type":74,"acceptedAnswer":84},"Which factors are selected as predictors for soil water content (SWC)?",{"text":85,"@type":77},"Selected variables include SMAP, NDVI, LST, precipitation, sand/clay content, pH, evapotranspiration, and topography/elevation, justified by multiple physical and ecohydrological mechanisms linking them to SWC.","https://schema.org",{"og:url":53,"og:type":88,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":90,"canonical":53},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":93},[94,98,102,106,111,116,121,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":47,"category_name":95,"show_sort_weight":96,"slug":97},"Story & Novel",90,"story-novel",{"id":48,"doc_module":4,"doc_module_name":47,"category_name":99,"show_sort_weight":100,"slug":101},"Literature",80,"literature",{"id":20,"doc_module":4,"doc_module_name":47,"category_name":103,"show_sort_weight":104,"slug":105},"Exam",70,"exam",{"id":107,"doc_module":4,"doc_module_name":47,"category_name":108,"show_sort_weight":109,"slug":110},5,"Comic",60,"comic",{"id":112,"doc_module":4,"doc_module_name":47,"category_name":113,"show_sort_weight":114,"slug":115},6,"Technology",50,"technology",{"id":117,"doc_module":4,"doc_module_name":47,"category_name":118,"show_sort_weight":119,"slug":120},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":47,"category_name":12,"show_sort_weight":30,"slug":122},"research-report",{"id":124,"doc_module":4,"doc_module_name":47,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":47,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":47,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":47,"category_name":137,"show_sort_weight":107,"slug":138},19,"General","general"]