[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119053-en":3,"doc-seo-119053-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119053,4398048950312,"Violet","https://ap-avatar.wpscdn.com/avatar/400002538284de19e3c?_k=1778320343897328908",8,"Research & Report","Metamorphic testing of machine learning and conceptual hydrologic models","Predicting hydrologic-system responses to altered driving forces beyond historical patterns is essential for estimating climate-change impacts and management effects, yet reliability is hard to assess because such predictions cannot be directly tested against newly observed outcomes. Metamorphic testing evaluates models by defining input transformations with known qualitative expected responses and checking model consistency. The study extends this idea through a multi-model framework and sensitivity analysis tied to training and calibration choices, enabling quantitative comparison across model structures. Applied to CAMELS-calibrated conceptual and machine learning hydrologic models, results show ML models outperform during calibration and validation, while response magnitudes and even response signs can depend on training data, and quantitative outputs can differ despite passing metamorphic tests.","Hydrol. Earth Syst. Sci., 28, 2505–2529, 2024 [https://doi.org/10.5194/hess-28-2505-2024](https://doi.org/10.5194/hess-28-2505-2024)[ ](https://doi.org/10.5194/hess-28-2505-2024)© Author(s) 2024 . This work is distributed under the Creative Commons Attribution 4 .0 License.  \nMetamorphic testing of machine learning and conceptual hydrologic models  \nPeter Reichert 1,􀀕 ;􀀔 , Kai Ma2,3 ;􀀔 , Marvin Höge1 , Fabrizio Fenicia 1 , Marco Baity-Jesi 1 , Dapeng Feng4 , and Chaopeng Shen4  \n1Eawag: Swiss Federal Institute of Aquatic Science and Technology, Dübendorf, Switzerland  \n2Institute of International Rivers and Eco-Security, Yunnan University, Kunming, China  \n3Yunnan Key Laboratory of International Rivers and Transboundary Eco-security, Yunnan University, Kunming, China  \n4 Civil and Environmental Engineering, Pennsylvania State University, University Park, State College, PA, USA  \n􀀕 retired  \n􀀔 These authors contributed equally to this work.  \nCorrespondence: Peter Reichert (peter.reichert@emeriti.eawag.ch)  \nReceived: 10 July 2023 – Discussion started: 26 July 2023  \nRevised: 15 April 2024 – Accepted: 20 April 2024 – Published: 13 June 2024  \nAbstract. Predicting the response of hydrologic systems to modiﬁed driving forces beyond patterns that have occurred in the past is of high importance for estimating climate change impacts or the effect of management measures. This kind of prediction requires a model, but the impossibility of testing such predictions against observed data makes it difﬁcult to estimate their reliability. Metamorphic testing offers a methodology for assessing models beyond validation with real data. It consists of deﬁning input changes for which the expected responses are assumed to be known, at least qualitatively, and testing model behavior for consistency with these expectations. To increase the gain of information and reduce the subjectivity of this approach, we extend this methodology to a multi-model approach and include a sensitivity analysis of the predictions to training or calibration options. This allows us to quantitatively analyze differences in predictions between different model structures and calibration options in addition to the qualitative test of the expectations. In our case study, we apply this approach to selected conceptual and machine learning hydrological models calibrated for basins from the CAMELS data set. Our results conﬁrm the superiority of the machine learning models over the conceptual hydrologic models regarding the quality of ﬁt during calibration and validation periods. However, we also ﬁnd that the response of machine learning models to modiﬁed inputs can deviate from the expectations and the magnitude, and even the sign of the response can depend on the training  \ndata. In addition, even in cases in which all models passed the metamorphic test, there are cases in which the quantitative response is different for different model structures. This demonstrates the importance of this kind of testing beyond and in addition to the usual calibration–validation analysis to identify potential problems and stimulate the development of improved models.  \n1 Introduction  \nThe availability of hydrologic and meteorological data and catchment attributes for a large number of catchments in the USA (Newman et al., 2015 ; Addor et al., 2017) has greatly stimulated hydrologic research in the past few years (Kratzert et al., 2018 ; Shen, 2018 ; Kratzert et al., 2019a, b ; Razavi, 2021 ; Ng et al., 2023 ; Feng et al., 2020) . In particular, it has been shown that the training of machine learning models jointly with hydrologic data from a large number of catchments leads to an extraordinary performance of these models, even for the prediction of the output of catchments that had not been used for training (Kratzert et al., 2018, 2019a, b ; Feng et al., 2020, 2021) . Arguably, this breakthrough was made possible by the combination of two elements:  \n1. using machine learning models, in particu","cbCain5OSfVMjquU","https://ap.wps.com/l/cbCain5OSfVMjquU","pdf",5718370,1,25,"English","en",105,"# Abstract\n# 1 Introduction","[{\"question\":\"Why is model reliability difficult to assess for hydrologic predictions under modified driving forces?\",\"answer\":\"Because predictions for altered future driving forces cannot be directly validated against observed data for those specific conditions. This makes reliability estimation challenging with standard validation approaches.\"},{\"question\":\"What does metamorphic testing do in this context?\",\"answer\":\"It defines input changes where the expected responses are known at least qualitatively, then tests whether the model outputs remain consistent with those expectations.\"},{\"question\":\"How does the study extend metamorphic testing to improve information gain?\",\"answer\":\"It adds a multi-model approach and includes sensitivity analysis to training or calibration options, allowing quantitative comparison of differences across model structures in addition to qualitative expectation checks.\"}]","Metamorphic testing of machine learning and conceptual hydrologic models | PDF",1785722101,63,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"metamorphic-testing-of-machine-learning-and-conceptual-hydrologic-models","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/metamorphic-testing-of-machine-learning-and-conceptual-hydrologic-models/119053/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is model reliability difficult to assess for hydrologic predictions under modified driving forces?","Question",{"text":75,"@type":76},"Because predictions for altered future driving forces cannot be directly validated against observed data for those specific conditions. This makes reliability estimation challenging with standard validation approaches.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What does metamorphic testing do in this context?",{"text":80,"@type":76},"It defines input changes where the expected responses are known at least qualitatively, then tests whether the model outputs remain consistent with those expectations.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the study extend metamorphic testing to improve information gain?",{"text":84,"@type":76},"It adds a multi-model approach and includes sensitivity analysis to training or calibration options, allowing quantitative comparison of differences across model structures in addition to qualitative expectation checks.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]