Environmental monitoring and evidence-based decision-making in water resources require robust water-quality (WQ) modeling frameworks. Dissolved oxygen (DO) and pH are key indicators addressed in this study using a data-driven approach. The South Platte River Basin, United States, is selected as the case study, with data obtained from the United States Geological Survey (USGS) and preprocessed prior to modeling. This study proposes a dynamically optimized weighted ensemble deep learning (DL-EDL) framework, in which ensemble weights are adaptively determined through nonlinear programming rather than fixed or heuristically assigned schemes. The proposed model achieves R² values of 0.72 for DO and 0.82 for pH, indicating satisfactory predictive performance for complex environmental time-series data. In addition, uncertainty analysis based on bootstrap and Monte Carlo uncertainty analyses shows that the DL-EDL model yields lower prediction variability, with reduced standard deviations, 0.035 for DO and 0.004 for pH in bootstrap analysis, confirming the stability and reliability of the model outputs. Overall, the proposed framework provides a practical tool for water-quality monitoring and decision support in water resource management.
Recent years have seen significant advancements in the field of water resources management (WRM). It is notable that river flow regimes have undergone abrupt shifts, which have led to an increase in turbulence. This study presents a new approach to ensemble machine learning (EML) that utilizes a weighted-based machine learning (ML) framework to create an optimized ensemble model. The study concentrates on the Kashkan River basin in Lorestan Province, Iran. The primary dataset for this study was historical river discharge data obtained from the Iran Water Resources Company (IWRC). Between 2017 and 2018, 2,068 monthly river Debi measurements were examined for analysis. This study utilizes time-series modeling by determining the effect lags and subjecting them to ML and EML models to simulate the stream flow. To enhance predictive accuracy, three ensemble models were developed: Weighted Ensemble Machine Learning (WEML), Linear-Programmed Ensemble Machine Learning (LPEML), and Linear-Programmed Weighted Ensemble Machine Learning (LPWEML). According to the evaluation results, the WEML and LPEML models demonstrated the lowest computational errors, achieving R² values of 0.939 and 0.934, respectively. The efficiency of WEML and LPEML models can be seen through validation approaches. However, the LPWEML performed poorly compared to other MLs. According to these findings, the proposed methodologies are effective in increasing the accuracy of streamflow prediction.
The survival of ecosystems, agriculture, industry, and drinking water supply depends on rivers, which are among the most vital freshwater resources. The focus of this study is on long-term river modeling. This study examines how machine learning (ML) models can predict river discharge and incorporates explainable artificial intelligence (XAI) principles to enhance model transparency and reliability. Bootstrap Tree (BT), Decision Tree (DT), Histogram Gradient Boosting (HGB), Extreme Gradient Boosting (XGB), and Categorical Boosting (CB) are among the models that were examined. A comprehensive set of performance metrics was used to assess these models and they were verified through rigorous validation. The study area was the South Platte River in the United States, which is a hydrological system that is increasingly affected by discharge variability due to climate change. The results demonstrate CB, XGB, and HGB’s superior predictive accuracy and minimal errors. The top-performing models demonstrated high performance indices with coefficients of determination (R²) of 0.939, 0.852, and 0.817, respectively. Long-term modeling indicates a continuous decline in streamflow, with projections suggesting potential drying within the next two decades if no intervention occurs. These findings underscore the need for urgent action to address water resource challenges and are closely aligned with the United Nations Sustainable Development Goals (SDGs), especially SDG 6, SDG 13, and SDG 15.
Global warming and population growth have significantly intensified the challenges in securing drinking water supplies. This study investigates transient instabilities of streamflow using ensemble machine learning (EML) and machine learning (ML) methodologies on the South Platte river in the United States. The United States Geological Survey’s online database was utilized to obtain the primary dataset. Several technical approaches were employed for preprocessing the initial dataset: cleaning outlier data, clean missing data, and 10 fold cross-validation. Nonlinear programming, genetic algorithms, least squares, linear programming, gradient descent, particle swarm optimization, Nelder Mead, and simulated annealing were employed algorithms to develop eight-weighted EML models. The results showed that the ensemble learning approach and the aggregation of weak learners by mentioned algorithms have been significantly successful. Particularly, the nonlinear programming-EML (NLP-EML) outperformed others, achieving the highest prediction accuracy with an R2 coefficient equal to 0.97. The probability density function showed that NLP-EML was the most reliable model. Overall, the findings highlight the superior performance and reliability of EML approaches in hydrological modeling, offering practical guidance to experts on the creation of robust ensemble models for improved prediction accuracy.
This study focuses on enhancing long-term, multi-step forecasting of dissolved oxygen (DO), a key indicator of river water quality. We introduce a novel hybrid method, Hidden Pattern Feature Extraction-Statistical Mode Decomposition (HPFE-SMD), integrated with explainable ensemble learning models, namely Random Forest (RF) and Extra Trees Regressor (ETR), both in standalone and hybrid configurations (HPFE-RF and HPFE-ETR). The models were trained and evaluated using monthly DO data spanning 1974-2023 from two sites within the Mississippi River Basin, across forecasting horizons of 1, 3, 9, and 15 months. The hybrid models consistently outperformed their standalone counterparts. For instance, at a 15-month horizon for Site 1, the HPFE-ETR model reduced the Mean Absolute Error (MAE) by 98.1 % compared to standalone ETR. In comparison with TVF-EMDbased models, HPFE-SMD achieved a 10.8 % and 4.3 % reduction in Mean Absolute Percentage Error (MAPE) for RF and ETR, respectively, at the 9-month horizon. Overall, HPFE-RF and HPFE-ETR achieved high predictive performance with RMSE values below 0.25 mg/L and R2 values exceeding 0.99, even for long-term forecasts. SHAP (SHapley Additive exPlanations) analysis revealed that key statistical features, such as vibration amplitude (RMS), energy, skewness, kurtosis, and crest factor, played a dominant role in model predictions. Additionally, the proposed method demonstrated strong generalizability by accurately forecasting other water quality parameters, including total nitrogen, pH, total dissolved solids, and sodium adsorption ratio. These results highlight the added value of the HPFE-SMD approach over traditional decomposition or standalone ML models, showcasing its potential for integration into advanced water quality monitoring and management systems.
Rivers provide irreplaceable resources for human life, and the problem of water scarcity has attracted serious attention worldwide. In this study, Kashkan River located in Loristan Province of Iran was studied using data obtained from the database of Iran Water Resources Company (IWRC). Three distinct machine learning (ML) models – Regression Tree (RT), Random Search Regression Tree (RSRT), and Bayesian Optimization Regression Tree (BORT) – were utilized to enhance water resource management practices. The primary model used was RT, a method that uses Bayesian optimization and stochastic search algorithms to provide an accurate estimate of the maximum flow within a river. The two hybrid models, RSRT and BORT, were introduced to improve the model performance. Through a comprehensive comparison and analysis of the results generated by these models, valuable insights were gained. Among the three models, the RSRT model demonstrated superior performance and accuracy metrics in streamflow (SF) modeling, closely aligning its results with a DR line of 1, indicating an optimal fit. The BORT and RT models also achieved excellent results, with their performance being on par with that of the top-performing RSRT model.
The environmental protection water quality is a critical subject in dominance of the sustainable water resources management. Accordingly, the essential indicators of water quality were considered, which are not only appropriate indicators for water quality, but they can also serve as crucial indicators for the health of water environment and ecosystems. Therefore, this study uses well-known ensemble machine learning methodologies to investigate and predict the fluctuations of water quality parameters. Optimization procedures used to assemble machine learnings were non-linear programming (NLP), genetic, gradient descent, and least square algorithms, linear programming, particle swarm optimization, Nelder-Mead optimization, and simulating annealing optimization. Using optimization procedures, the basic MLs were assembled and eight new ensemble machine learning were developed. The studied area was the South Platte River basin, USA. The primary dataset was obtained through the online database of the United States Geological Survey, which contained sampling information on river water related to 2023-2024. Then, using clean missing and outlier data preprocessing techniques, the dataset was modified. Finally, using the 10-fold cross-validation technique, the primary data was validated. The results showed that NLP significantly improved the accuracy and performance of models, achieving the best performance with R2 of 0.9836 and 0.9031 across DO and pH modelings. The modeling results indicated that the pH parameter fluctuated within the safe range. While the DO seems that tolerated in unsafe domain for aquatic ecosystems. The findings of this research could help a wide range of decision-makers.
As global population increases, water resources are under increased pressure, while traditional monitoring approaches struggle with equipment and economical limitations. To address this challenge, this study uses artificial intelligence (AI) to enhance water quality (WQ) assessment, a crucial component of water resources management (WRM). Specifically, it integrates single-layer feedforward networks (SLFN) with particle swarm optimization (PSO) and genetic algorithm (GA) to enhance prediction accuracy and efficiency. The Colorado River, USA, was selected as the study area, focusing on two key indicators, dissolved oxygen (DO) and electrical conductivity (EC). The hybrid PSO-SLFN and GA-SLFN significantly outperformed traditional method. PS-SLFN(EC) achieved the highest accuracy with R2 = 0.865 and RMSE = 0.119, while GA-SLFN(DO) was most effective for DO prediction (R2 = 0.512, RMSE = 0.515). Uncertainty analysis confirmed the reliability of the models. The PS-SLFN(EC) showed the lowest mean error (0.0043) and standard deviation (0.12); GA-SLFN(DO) demonstrated consistent results with a mean error of 0.153 and standard deviation of 0.253. These outcomes highlight the prominence of hybrid AI models in achieving robust and data-efficient WQ modeling without any need for costly tools or equipments.
Many factors impact water quality (WQ), such as climate change and population growth. Thus, the present work aims to propose an accurate and potent solution for the WQ instabilities challenge in the South Platte River in United States. The data driven model based on the machine learning model tuned with Kalman filter (KF) was considered to reduce input data noise. The least absolute shrinkage and selection operator (LASSO) algorithm were used to analyze the importance of features and select the best inputs. The US Geological Survey (USGS) archive provided the primary database related to 2023-2024, with over 38,000 samples. The random forest (RF) was combined with KF and LASSO to reduce noise and analyze the importance of features due to the high number of samples. Artificial neural network (ANN), linear regression (LR), and support vector machine (SVM) were developed to compare the accuracy of the proposed model. The proposed model had the highest coefficient of determination values, which were between 0.95 and 0.99. Modeling the indicators revealed that some WQ variations could negatively affect aquatic ecosystems.
Today, humanity has managed to overcome many challenges related to water. One of the most significant challenges concerning surface water resources is their preservation and maintenance. The investigation of the variations of Miqan Lake, which is located in the Markazi province of Iran, is the focus of this study. Given its vicinity to the city of Arak and considering the consecutive years of droughts in Iran, this study examines the fluctuations in Miqan Lake water level. We utilized four artificial intelligence models to study the trend of lake changes. These models include an essential machine learning model, namely a single-layer feed-forward neural network, which is combined with three evolutionary algorithms: particle swarm optimization, genetic algorithm, and imperialist competitive algorithm. Subsequently, in addition to the SLFFN model, three new hybrid evolutionary models were developed to address the modeling of changes and fluctuations related to water in the Miqan lake. This research's initial data corresponds to monthly samples collected from Miqan Lake for over 15 years. To examine the mutual effects of modeling and potential errors, besides regression receiver operation characteristic (RROC) analysis, a probability density function (PDF) analysis was also conducted. The results showed that PSELM had the lowest error in estimating lake water level fluctuations, obtaining the best performance in the evaluation indicators: RMSE, MAPE, and SI of 0.59, 0.028, and 0.018, respectively. The best match in actual and estimated values also belonged to ICELM. In the analysis of RROC, ICELM was superior, and the results of PDF analysis confirmed the ICELM superiority. Ultimately, these results indicate a noticeable reduction in the water level of Miqan Lake, leading to various risks such as pollution of freshwater resources and the onslaught of dust storms in the city of Arak.
Today, humanity faces the complex phenomenon of global development. Problems such as limited resources, financial constraints, time constraints, and the involved nature of issues surrounding water resources management pose challenges to effectively monitor water resources issue parameter management (WRIPM). Nevertheless, the importance of this issue cannot be ignored. The focus of this study is on WRIPM issues, and new machine learning models are utilized to investigate and estimate fluctuations in WRIPM parameters. The initial dataset utilized in this study was sourced from the US Geological Survey, specifically sampling stations information from the South Platte River in the United States. Two new models were developed: the Ensemble Bagged Machine (EBM) and the Stochastic Weighted Ensemble Bagged Machine (SWEBM), which further optimized the characteristics of the EBM model using Bayesian Optimization (BN). These models were employed to simulate Dissolved Oxygen (DO), Electrical Conductivity (EC), Power of hydrogen (pH), and river flow rate (Debi) parameters. Additionally, the research employed various scenarios for evaluation and validation. Uncertainties were calculated using the Wilson analysis. The results demonstrated the superiority of the SWEBM in all modeling aspects, yielding the best indices, namely, R2, MAE, and RMSE, at 0.984, 0.0259, and 0.0394, respectively. Performing an receiver operating characteristic analysis revealed that the area over the RROC curve value for the superior model was 36.16, demonstrating excellent performance in modeling the pH parameter. The WM analysis method confirmed that the SWEBM model was the most trustworthy, delivering the greatest precision and best modeling for the pH parameter (WOUB=0.00683, MOPE=0.00138, and STD=0.1665) even though it was slightly overestimated. The probability density function analysis revealed that the SWEBM was the best model. While other models had results that were relatively close. This research represents valuable findings that can be shared with stakeholders, including experts, scholars, and organizations responsible for overseeing water quality administration.
This study presents a new method based on three types of deep learning-based models (DLM) for estimation of water parameters. The DLM models were recurrent neural networks (RNN), long short-term memory (LSTM), and bidirectional long short-term memory (BiLSTM). The study areas were the Colorado River basin in the United States and the Mighan Wetland in Iran. The electrical conductivity (EC), dissolved oxygen (DO), total dissolved solids (TDS), chloride ions (Cl), and river flow rate (debi) were simulated by the DLM models. The Wilson score (WS) uncertainty analysis results for Colorado modelling showed that LSTMdebi, RNNDO, and RNNEC were the best models in simulating due to having the lowest errors (Mean ei equal to 0.36, -1.50, and -0.59), respectively. Finally, the highest value of the R2 index, 0.998, was achieved by the LSTM model in modelling the debi parameter, and 0.996 in EC modelling, in the Mighan Wetland.
Recently, due to global climate change and population growth, environmental protection has become more interested. Water is the main critical issue because it is the most significant environmental resource. Therefore, this study introduces a novel approach to examine, modeling, and addressing the monitoring of water quality (WQ) critical scenario related to unexpected extreme variations of crucial indicators (UEVCI). Therefore, this research integrates ensemble machine learning (EML) techniques with Non-linear programming (NLP) and Simulated annealing algorithm (SAA) to develop an optimal weighted ensemble models. New development models were nonlinear-programmed ensemble machine learning (NLEML) and simulated annealing ensemble machine learning (SAEML). Besides, we developed least-squared boosted regression tree (LsBRT), artificial neural network (ANN), and multiple linear regression (MLR) models individually to compare the performance of new ensemble models. The South Platte River Basin in Colorado, USA was the study region. The initial dataset was extracted through the United States Geologic Survey (USGS) from 2023 to 2024. Preprocessing approaches such as cleaning missing data (CMD), cleaning outlier data (COD), and k-fold cross validation (KFCV) with k = 5 were used to prepare the dataset. The final dataset was utilized to examine variations of essential parameters that affect water health and quality, including the power of hydrogen (pH) and dissolved oxygen (DO). The results showed that the NLEML provided the most accurate results in estimating fluctuation of pH parameter with an R2 coefficient of 0.85. Also, the NLEML estimated the variance of the DO parameter with an R2 equal to of 0.79, resulting in an outperforming simulation.
Humanity is witnessing scientific advances and since one of the vital concerns has been accessing water for drinking and agriculture, scientific innovations have been used to solve water challenges. Recently, Machine Learning (ML) is one of the most promising developments, therefore, for utilization of this innovation in simulating the Water Quality-Quantity Assessment (WQA) issues, the author has developed the Extreme Learning Machine (ELM) as the stand-alone model and its combination with evolutionary algorithms (EA) to optimize the modeling of the WQA parameters. So, the Particle Swarm Optimization (PSO), Genetic Algorithm (GA), and Imperialist Competitive Algorithm (ICA) are implemented to improve the performance of WQA parameters prediction. Then, the three novel models are developed for this purpose simultaneously with the stand-alone ML model. The novel models are GAELM, PSELM, and ICELM. The study area was the Colorado River basin in the USA, which was the input dataset extracted from the US Geological Survey. The considered parameters were Power of hydrogen (pH), river flow or Debi, Electrical Conductivity (EC), and Dissolved Oxygen (DO). To over- and under-fitting ML models, the input dataset was detrended and randomized by K-fold cross-validation. Furthermore, this study used seven evaluation methods for determining the models' performances. Based on the evaluation metrics, the ICELM model was the best in the EC, DO, and pH modeling. Also, the ELM model was superior in Debi prediction. Additionally, the discrepancy percent charts for all models in simulations were drawn. Moreover, the error percent plots also were drawn for all modeling. Finally, the Wilson method analyzed the prediction errors and related uncertainties.
Along with the global population growth, the human need for safe drinking water sources has increased. With global warming, the water challenge is perhaps the most crucial challenge for the world community. At the same time, scientific methods are one of the best tools to help humanity. Considering that in many natural phenomena, it is possible to describe them based on complex relationships, it is almost impossible to solve them analytically and mathematically. Therefore, it is necessary to use methods with the ability, accuracy, and high speed to justify nonlinear relationships. One of these methods is Artificial Intelligence (AI). This research used the Extreme Learning Machine (ELM) model and Genetic Algorithm (GA) to create a new hybrid model Genetic Extreme Learning Machine (GAELM). AI and hybrid models were used to simulate and predict the water quality parameters changes. The study area in this work was the Colorado River Basin in the United States. The desired qualitative parameters were Electrical Conductivity (EC) and Dissolved Oxygen (DO). Finally, using seven approaches, the models' performance was compared. The results showed that the best simulation related to the GAELM hybrid model in the EC parameter modeling with indices RMSE and R2 equal to 0.1304, and 0.8619, respectively. Also, the ELM model was ranked in second place in accuracy. Based on the uncertainty analysis (UA-WSM) results, the GAELM(EC) model was the most accurate, with the minimum average prediction error equal to 0.01.
Qualitative analysis of water resources is one of the most widely used topics in water resources research today. Researchers use various analysis methods of water parameters to achieve the desired goals in this field. This research uses artificial intelligence (AI), learning machine (LM), data mining, and mathematical techniques to simulate water behavior and estimate its parametric changes. The proposed model used in this study was a Self-adaptive Extreme learning machine (SAELM) to estimate hydrogeological parameters of the Meghan wetland located in Markazi province in Iran. In addition, SAELM simulation results were compared to Least square support vector machine (LSSVM), Multiple linear regression (MLR), and Adaptive Neuro-fuzzy inference system (ANFIS) models. The simulated parameters were Electrical Conductivity (EC), Total Dissolved Solids (TDS), Groundwater Level (GWL), and salinity. This information was related to sampling for 175 months in the study area. Finally, after simulation operation, four models were introduced as superior models. Mentioned exceptional models were SAELM in GWL modeling, SAELM in modeling the EC, MLR in salinity simulation, and LSSVM in the simulation of TDS parameters. Moreover, by five approaches, the models' performance was evaluated. Suggested strategies were performance evaluation by statistical indicators, Wilson score method uncertainty analysis (WSMUA), response & correlation plots, discrepancy ratio charts, and distribution error diagrams. Based on statistical indicators, the SAELM(GWL) model was the most accurate model with RMSE, MAPE, and R-2 indices equal to 0.1496, 0.0043, and 0.9933, respectively. The ANFIS model had the worst results in simulation.
Abstract Along with the global population growth, the human need for safe drinking water sources has increased. With global warming, the water challenge is perhaps the most crucial challenge for the world community. At the same time, scientific methods are one of the best tools to help humanity. Considering that in many natural phenomena, it is possible to describe them based on complex relationships, it is almost impossible to solve them analytically and mathematically. Therefore, it is necessary to use methods with the ability, accuracy, and high speed to justify nonlinear relationships. One of these methods is Artificial Intelligence (AI). This research used the Extreme Learning Machine (ELM) model and Genetic Algorithm (GA) to create a new hybrid model Genetic Extreme Learning Machine (GAELM). AI and hybrid models were used to simulate and predict the water quality parameter changes. The study area in this work was the Colorado River Basin in the United States. The desired qualitative parameters were Electrical Conductivity (EC) and Dissolved Oxygen (DO). Finally, using seven approaches, the models' performance was compared. The results showed that the best simulation related to the GAELM hybrid model in the EC parameter modeling with indices RMSE and R2 equal to 0.1304, and 0.8619, respectively. Also, the ELM model was ranked in second place in accuracy. Based on the uncertainty analysis (UA-WSM) results, the GAELM(EC) model was the most accurate, with the minimum average prediction error equal to 0.01.
Today, various methods have been developed to extract drinking water resources, which scientists use to simulate the quantitative and qualitative water resources parameters. Due to Iran's geographical and climatic characteristics, this region is located on the drought belt in Asia. In this research, some Artificial Intelligence (AI) and mathematical models have been used for groundwater level prediction. The AI models used for this research are Extreme Learning Machine (ELM), Least Square Support Vector Machine (LSSVM), Adaptive Neuro-Fuzzy Inference System (ANFIS), and Multiple Linear Regression (MLR) model. In this study, simultaneously, these models were used to simulate and estimate groundwater level (GWL). The database used in the simulation is the data related to the Total Dissolved Solids (TDS), Electrical Conductivity (EC), Salinity (S), and Time (t) parameters. The results showed that ELM was more accurate than other methods. In Uncertainty Wilson Score Method (UWSM) analysis, ELM had an Underestimation performance and was determined as the more precise model.
In recent years, as a result of climate change as well as rainfall reduction in arid and semi‐arid regions, modelling qualitative and quantitative parameters belonging to aquifers has become crucially important. In Iran, as aquifers are treated as the most commonly used drinking water resources, modelling their qualitative and quantitative parameters is enormously important. In this paper, for the first time, values of salinity, total dissolved solids (TDS), groundwater level (GWL) and electrical conductivity (EC) of the Arak Plain, located in Markazi Province, Iran, are simulated by means of four modern artificial intelligence models including extreme learning machine (ELM), wavelet extreme learning machine (WELM), online sequential extreme learning machine (OSELM) and wavelet online sequential extreme learning machine (WOSELM) as well as the MODFLOW software for a 15‐year period monthly. To develop the hybrid artificial intelligence models, the wavelet is employed. First, the effective lags in estimating the qualitative and quantitative parameters of the groundwater are identified using the autocorrelation function (ACF) and the partial autocorrelation function (PACF) analysis. After that, four different models are developed by the selected input combinations and also the ACF and the PACF in the form of different lags for each of ELM, WAELM, OSELM and WOSELM methods. Then, the superior models in simulating the groundwater qualitative and qualitative parameters are detected by conducting a sensitivity analysis. To forecast the electrical conductivity (EC) by the best WOSELM model, the values of the Nash–Sutcliffe efficiency coefficient (NSC), Mean Absolute Error (MAE) and the scatter index (SI) are obtained to be 0.991, 18.005 and 4.28E‐03, respectively. In addition, the most effective lags in estimating these parameters are introduced. Subsequently, the results found by the MODFLOW model are compared with those of the artificial intelligence models and it is concluded that the latter are more accurate. For instance, the scatter index and Nash–Sutcliffe efficiency coefficient values calculated by WOSELM for TDS, respectively, are 5.34E‐03 and 0.991. Finally, an uncertainty analysis is conducted to evaluate the performance of different numerical models. For example, MODFLOW has an underestimated performance in simulating the salinity parameter.