[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-119927-en":3,"doc-seo-119927-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":4,"is_deleted":4,"is_public":20,"is_downloadable":20,"audit_status":20,"page_count":21,"language":22,"language_code":23,"site_id":24,"html_lang":23,"table_of_contents":25,"faqs":26,"seo_title":27,"seo_description":14,"update_tm":28,"read_time":29},119927,4810365810221,"Aurora","https://ap-avatar.wpscdn.com/davatar_155a257f0dc6eb9ab79c44ca47cae57d",8,"Research & Report","Improving river water quality prediction with hybrid machine learning and temporal analysis","River systems deliver essential ecosystem services but many are degraded by diffuse and point-source pollutant inputs, making water quality evaluation critical for remediation and management. Machine learning predictive models can enhance monitoring networks, yet reduced or redundant training information can weaken performance and cause overtraining. This study examines historical Santiago River (Mexico) data to select variable, representative subsets, using clustering and time series analysis to train ANFIS, ANN, and SVM models with improved accuracy.","Ecological Informatics 82 (2024) 102655  \nContents lists available at ScienceDirect  \nEcological Informatics  \njournal [homepage: www.elsevier.com/locate/ecolinf](homepage: www.elsevier.com/locate/ecolinf)  \n| Improving river water quality prediction with hybrid machine learning and temporal analysis\u003Cbr>Alberto Fern´andez del Castillo a, Marycarmen Verduzco Garibaya, Diego Díaz-V´azquez a, Carlos Yebra-Montesb, Lee E. Brown c, Andrew Johnson c, Alejandro Garcia-Gonzalez d, *, Misael Sebasti´an Gradilla-Hern´andez a, *\u003Cbr>a Tecnologico de Monterrey, Escuela de Ingenieria y Ciencias, Laboratorio de Sostenibilidad y Cambio Clim´atico, Av. General Ramon Corona 2514, Nuevo M´exico, CP 45138 Zapopan, Jalisco, Mexico\u003Cbr>b ENES-Le´on, Universidad Nacional Aut´onoma de M´exico, Blvd. UNAM 2011, Predio el Saucillo y El Po-trero, CP, 37684 Le´on, Guanajuato, Mexico c School of Geography and water@leeds, University of Leeds, Leeds LS2 9JT, UK\u003Cbr>d Tecnologico de Monterrey, Escuela de Medicina y Ciencias de la Salud, Av. General Ramon Corona 2514, Nuevo Mexico, CP, 45138 Zapopan, Jalisco, Mexico |  |  |\n| --- | --- | --- |\n| A R T I C L E I N F O |  | A B S T R A C T |\n| Keywords:\u003Cbr>Water Quality Index\u003Cbr>Highly polluted river\u003Cbr>Time series analysis Cluster analysis Monitoring network Data Science |  | River systems provide multiple ecosystem services to society globally, but these are already degraded or threatened in many areas of the world due to water quality issues linked to diffuse and point-source pollutant inputs. Water quality evaluation is essential to develop remediation and management strategies. Computational tools such as machine learning based predictive models have been developed to improve monitoring network capabilities. The model’s performance is reduced when datasets composed of reductant information are used for training, on the other hand, the selection of most representative and variable water quality scenarios could result in higher precision. This study analyzed historical water quality behavior in the Santiago River, Mexico, to identify the most variable and representative data available to train machine learning models (Adaptive Neuro Fuzzy Inference System – ANFIS, Artificial Neural Network – ANN, and Support Vector Machine-SVM). Thirteen monitoring sites were clustered according to their water quality variability from 2009 to 2022. Subsequently, a Time Series Analysis (TSA) was used to select the most representative monitoring station from each cluster. Data for 6/13 monitoring sites were retained for the Best Training Subset (BTS) used to train restricted models that performed with similar (ANN and SMV) or higher (ANFIS) prediction accuracy (in terms of RMSE, MAE, MSE and R2) for both training and testing. This study provides evidence of water quality data containing redundant information that is not useful to improve machine learning model performance, in turn leading to overtraining. Combined analytical approaches can maximize the representativeness and variability of data selected for machine learning applications, leading to improved prediction. |\n\n1. Introduction  \nAnthropogenic activities including population growth, waste generation, changes in land use and climate change have driven major changes in river water quality worldwide (Li et al., 2022). Degradation of water quality leads to significant issues for biodiversity, agricultural crop growth and water consumption. The cost of degraded watersheds for water supply utilities alone has been estimated as >$5 billion annually (McDonald et al., 2016). Water quality monitoring programs (WQMP) aim to accurately assess the type and extension of water source pollution (Duan et al., 2016), through the constant and long-term  \nmeasurement of biological, physical, and chemical parameters. However, non-specialists can find it challenging to evaluate, interpret and synthesize the emergent complex datasets (Gitau et al., 2016). Thus, water quality indices (WQIs) ","cbCaiiwP8NUdqFJZ","https://ap.wps.com/l/cbCaiiwP8NUdqFJZ","pdf",5273214,1,12,"English","en",105,"# Introduction\n## Water quality monitoring and indices\n## Computational predictive modeling and limitations\n## Study approach and objectives","[{\"question\":\"Why is water quality prediction important for river ecosystem services?\",\"answer\":\"River systems provide ecosystem services, but water quality degradation from pollutant inputs threatens biodiversity, agriculture, and water consumption. Accurate evaluation supports remediation and management decisions.\"},{\"question\":\"What problem does the study address in machine learning model training data?\",\"answer\":\"Model performance can drop when training uses reductant or less representative information, while selecting redundant scenarios may lead to overtraining and reduced generalization.\"},{\"question\":\"How does the study select representative monitoring stations for model training?\",\"answer\":\"Thirteen monitoring sites are clustered by water quality variability (2009–2022). Time series analysis then selects a representative station from each cluster, and a Best Training Subset is retained for restricted model training.\"}]","Improving river water quality prediction with hybrid machine learning and temporal analysis | PDF",1785727036,30,{"code":4,"msg":31,"data":32},"ok",{"site_id":24,"language":23,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"improving-river-water-quality-prediction-with-hybrid-machine-learning-and-temporal-analysis","",{"@graph":36,"@context":85},[37,54,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,48,51],{"item":41,"name":42,"@type":43,"position":20},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":47},"https://docshare.wps.com/document/","Document",2,{"item":49,"name":12,"@type":43,"position":50},"https://docshare.wps.com/document/research-report/",3,{"item":52,"name":13,"@type":43,"position":53},"https://docshare.wps.com/document/improving-river-water-quality-prediction-with-hybrid-machine-learning-and-temporal-analysis/119927/",4,{"url":52,"name":13,"@type":55,"author":56,"headline":13,"publisher":58,"fileFormat":61,"inLanguage":23,"description":14,"dateModified":62,"datePublished":62,"encodingFormat":61,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":57},"Person",{"url":41,"name":59,"@type":60},"DocShare","Organization","application/pdf","2026-08-03",true,{"@type":65,"interactionType":66,"userInteractionCount":4},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"Why is water quality prediction important for river ecosystem services?","Question",{"text":75,"@type":76},"River systems provide ecosystem services, but water quality degradation from pollutant inputs threatens biodiversity, agriculture, and water consumption. Accurate evaluation supports remediation and management decisions.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"What problem does the study address in machine learning model training data?",{"text":80,"@type":76},"Model performance can drop when training uses reductant or less representative information, while selecting redundant scenarios may lead to overtraining and reduced generalization.",{"name":82,"@type":73,"acceptedAnswer":83},"How does the study select representative monitoring stations for model training?",{"text":84,"@type":76},"Thirteen monitoring sites are clustered by water quality variability (2009–2022). Time series analysis then selects a representative station from each cluster, and a Best Training Subset is retained for restricted model training.","https://schema.org",{"og:url":52,"og:type":87,"og:title":13,"og:site_name":59,"og:description":14},"article",{"robots":89,"canonical":52},"index,follow",{"doc_id":7,"site_id":24},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,122,127,130,134],{"id":20,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":47,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":53,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":29,"slug":121},"research-report",{"id":123,"doc_module":4,"doc_module_name":46,"category_name":124,"show_sort_weight":125,"slug":126},9,"Religion & Spirituality",20,"religion-spirituality",{"id":125,"doc_module":4,"doc_module_name":46,"category_name":128,"show_sort_weight":125,"slug":129},"World Cup","world-cup",{"id":131,"doc_module":4,"doc_module_name":46,"category_name":132,"show_sort_weight":131,"slug":133},10,"Lifestyle","lifestyle",{"id":135,"doc_module":4,"doc_module_name":46,"category_name":136,"show_sort_weight":106,"slug":137},19,"General","general"]