
Distributed hydrologic models have been widely used for their functional diversity and rationality in theory. However, calibration of distributed models is computationally expensive since it requires a large number of model runs, even if an efficient multi-objective optimization algorithm is employed. To alleviate the burden of computation, we develop a two-stage surrogate-based model calibration framework by coupling standard backpropagation neural network with AdaBoost (ADBP) to efficiently calibrate the parameters of a large and complex distributed hydrologic model, that is, the Variable Infiltration Capacity (VIC) model. The first-stage model (SM-I) picks out the parameter sets whose simulated outputs are in the crucial range and the second-stage model (SM-II) estimates the values of outputs accurately with the parameter sets picked out by SM-I. The proposed surrogate model is tested in the Lanjiang River basin and the Xiangjiang River basin for parameter calibration. With the similar calibrated parameter sets for VIC, the surrogate model gains up to 20 times speedup compared with the original VIC model in the basin with a large area and complex physical conditions.
Water level is an important guide for water resource management and wetland ecosystems, defining one of the most basic processes in hydrology. This research seeks to investigate the possibility of complementing numerical modeling with a Machine Learning (ML) model to forecast daily water levels in the southern Everglades in Florida, USA. An exact analytical solution to water level may not be possible, but using the computational methods afforded by ML, the traditional numerical techniques may be enhanced to generate more robust, scalable predictions. Five locations were chosen for application of the Time-Delayed Neural Network (TDNN) and Long-Short Term Memory Recurrent Neural Network (LSTM-RNN) ML models, which were built to estimate water level with 1, 2, 3, 7 and 10 day forecasts using a simulation step of 1 day. The results showed that rainfall forecasts from weather models could improve water-level forecasts if the accuracy and performance of the weather models can be improved. The ML models presented here improve water-level predictions from a historical hydrologic model for a 24 hour forecast horizon.
The Grande de San Miguel catchment in El Salvador generates floods almost every year and produces significant damage to transport infrastructure, houses, and agricultural land, as well as causing people to evacuate. To anticipate flood conditions, an early warning system (EWS) has been in place since 2001 in collaboration with local inhabitants, civil protection, and the National Hydrometeorological Service, but a rainfall-runoff forecast model has not been implemented in the catchment to predict these flood events. In this study, a Multilayer Perceptron Artificial Neural Network (MLP-ANN) model is proposed as an operational flood forecast model for this catchment to address the short-term flood forecast. The model results show excellent performance for lead time no larger than 4 hours and good results for lead time between 5 and 9 hours, but an unsatisfactory result for a forecast lead time of 12 hours since there are some high-flow events that the model underestimates.
A data preprocessing technique, known as Wavelet Analysis from the field of signal processing, was used to enhance an artificial neural network (ANN) model by addressing its primary drawbacks, notably nonstationarity and noise. Wavelet transforms are well-suited for handling time series contaminated by nonstationary traits and can extract relevant information at various time series frequencies. To improve the model's performance, the ANN calibration weighting scheme assigns more weight to representative subseries and reduces the impact of noisy data on the rest. Noise filtering techniques (soft/hard single threshold) where applied to manipulate the subseries, further enhancing the model's predictive skill. This approach resulted in the creation of a Noise Filter with Wavelet Analysis in Artificial Neural Network (NOWANN) scheme. To identify the best model, a model-based approach was used for input variable selection, k-fold cross-validation to determine the optimal training split, early stopping to prevent overfitting, and an exhaustive search in the variable space (hidden nodes versus decomposition levels) to find the most suitable ANN structure. Other data-driven methods (DDM) testing the effectiveness of wavelets as data preprocessors were briefly discussed, including least median squares, model tree (M5P), pace regression, and linear regressions. For wavelet coefficients time-invariance, the a-trous wavelet transform was used. The results indicate that the NOWANN schemes outperformed all previous models, confirming the effectiveness of this approach.
In recent years, there has been a surge of interest in machine learning (ML) and artificial intelligence (AI) due to the effectiveness of deep learning algorithms and the increasing availability of large data sets. This chapter provides a brief overview of the applications of AI and ML techniques in hydroinformatics, a field that deals with advanced information technology, data analytics, and modeling for aquatic environment management. Data-driven models are becoming more common in water management as they can reveal hidden patterns in data and offer improved accuracy in certain situations. This chapter highlights the importance of spatiotemporal data analysis, pattern recognition, and optimization approaches in water resources management under uncertainty. It does not offer a comprehensive review of all methods but rather focuses on selected ML techniques widely used in water-related problems. Additionally, the chapter discusses the challenges associated with using ML models, such as black-box criticisms, and the potential of hybrid models that combine the strengths of ML and physically based process models for more robust solutions in hydroinformatics.
Large-scale river basins in the southern part of China suffer frequently from flooding problems, and storage areas are often used for minimizing downstream flooding risk, which in turn causes damage to economic activities within these areas. Such damages need to be minimized, even though this is often conflicting with the main objective of reducing downstream flooding risk. This problem is especially highlighted in the middle part of the Huai River basin. Previous studies focused on solving this multi-objective optimization problem using generic algorithms coupled with hydrodynamic models. This study further develops probability analysis using a robust optimization and similarity-based selection method for investigating the possibility to do real-time operations. The results show that, in most cases, both robust strategies and a similarity selection method can effectively reduce flood risk downstream, and, at the same time, reduce storage area damage. This suggests that the strategies from similarity selection are more effective in reducing flooding risk but cause more storage area damage and vice versa.
For a long time, the classical problem of identifying the optimal modeling structure and/or parameters followed the calibration-validation norm, originating from the iconic split-sample scheme by Vit Klemeš. A common feature of such approaches is their dependence on the length and representativeness of the available data. This introduces several questions since the inferred parameters are selected according to a specific subset (or subsets) of historical data, while the rest of data is used for validation. In this vein, we propose a conceptually simple approach driven by the well-known stochastic simulation paradigm, which builds upon the idea of calibrating models using alternative, yet probabilistically consistent, synthetic data. Decoupling this way, the available data now become the basis to generate stochastic inputs, as well as for model validation and parameter uncertainty assessment. This allows for embedding the stochasticity of real-world drivers (rainfall, evapotranspiration) and responses (runoff) and thus their hydrological uncertainty. Furthermore, it results to stable and robust models , as calibration is performed using long enough time series that reproduce important properties that are associated with the changing climate (e.g., long-term persistence), which are generally hidden in the short historical samples. Identifying this way, the derived parameters are optimal not only for the historical data set, but for any alternative plausible realization of the modeled processes.
The increasing rise of computational power has led to significant advances in the application of machine learning (ML) techniques in hydrological simulation. These methods are specifically developed to discover constitutive relations by implementing an unbiased implicit approach to capture unforeseen patterns in massive data sets. This study applied multiple ML approaches for daily rainfall-runoff simulation across a mixed urban-rural drainage system, the Northeast Cape Fear River Basin in North Carolina (NC), USA. Multiple ML algorithms such as Support Vector Machines (SVM), Bayesian Lasso, and Random Forest (RF) models, along with the Sacramento Soil Moisture Accounting (SAC-SMA) rainfall-runoff model, were applied to predict sequential daily streamflow records based on a set of collected data from climate and streamflow gauging stations. Analysis suggests that the effects of input data on model performance, the error associated with forcing data, the amount of training data, and the correlation among different attributes of data series have strong influences on ML computation. Compared to the SAC-SMA, the Bayesian Lasso model was able to simulate the temporal dependencies among the observations and thereby was capable of accurately modeling the multivariate sequences of complex daily rainfall-runoff records across a mixed urban-rural catchment. The results provided an algorithmically informed simulation on the dynamics of daily streamflow simulation that may apply to other complex catchments and climate settings.
An effective pattern recognition technique based on incorporating an unsupervised clustering algorithm into a spatial random sampling toolbox (SRS-GDA; Wang & Xuan, 2020) is used to identify and classify extreme rainfall patterns in England, Wales, and Scotland in different sizes over the last century. The spatial features of the rainfall patterns, such as their geographic location, size, shape, and orientation, can be automatically extracted from the clustering analysis. The top three of the dominating patterns are identified and the temporal variation of their attributes are also investigated. The enhanced toolbox presents great potential in auto-labeling clusters to support deep learning of complex environmental spatial-temporal features over large data sets, demonstrated by an example of convolution neural network, which is able to pick up the labeled rainfall patterns with high accuracy.
Water quality in rivers is influenced by natural factors and human activities that interact in complex and nonlinear ways, which make water quality modeling a challenging task. The concept of complex networks (CN), a recent development in network theory, seems to provide new avenues to unravel the connections and dynamics of water quality phenomenon, including clandestine teleconnections. This study explores the spatial patterns of water quality using CN concepts at both catchment scale and larger national scale. Three major water quality parameters–dissolved oxygen, permanganate index (CODMn), and ammonia nitrogen (NH 3 -N)–measured weekly for 12 years at 91 monitoring stations across China, are analyzed. The results show that the degree centrality and clustering coefficient values for water quality indicators is DO > NH 3 -N > CODMn at both basin scale and national scale. The findings improve understanding of water quality dynamics and suggest new methods for environment system analysis and watershed management.
Earth and Space Science Open Archive This preprint has been submitted to and is under consideration at AGU Books. ESSOAr is a venue for early communication or feedback before peer review. Data may be preliminary.Learn more about preprints preprintOpen AccessYou are viewing an older version [v1]Go to new versionThree-dimensional clustering in the characterization of spatiotemporal drought dynamics: cluster size filter and drought indicator threshold optimizationAuthorsVitaliDiaziDGerald AugustoCorzo PereziDHennyVan LaneniDDimitriSolomatineiDSee all authors Vitali DiaziDCorresponding Author• Submitting AuthorIHE Delft Institute for Water EducationDelft University of TechnologyiDhttps://orcid.org/0000-0002-5502-4099view email addressThe email was not providedcopy email addressGerald Augusto Corzo PereziDUNESCO-IHE Institute for Water EducationiDhttps://orcid.org/0000-0002-2773-7817view email addressThe email was not providedcopy email addressHenny Van LaneniDWageningen UniversityiDhttps://orcid.org/0000-0001-9226-3921view email addressThe email was not providedcopy email addressDimitri SolomatineiDIHE Delft Institute for Water EducationDelft University of TechnologyiDhttps://orcid.org/0000-0003-2031-9871view email addressThe email was not providedcopy email address
Human behavior and decision making are dynamically influenced by digital media, making digital news a vast data source of events and points of view. Meanwhile, artificial intelligence now has advanced capacity for natural language processing (NLP) through tools such as sentiment analysis and topic identification. This research uses machine learning algorithms to prove the hypothesis that it is possible to find correlations between news information and water resource problems. The 207 water bodies in the Magdalena River Basin, Colombia, were analyzed alongside 19,490 news articles published between 2016 and 2020 in 42 online newspapers. A platform for the visualization of spatiotemporal information was developed as a proof of concept. There has been a noticeable increase in digital news about the Magdalena River Basin in recent years and some correlation with extreme events, but the comparison between the sentiments and measured physical variables was inconclusive. This is likely due to the difficulty in filtering the topic from obtained news, as well as in precisely identifying the spatial location and temporal range of events. In the future, NLP techniques to analyze news articles could be used to provide additional information for decision making at the basin level.
Earth and Space Science Open Archive This preprint has been submitted to and is under consideration at AGU Books. ESSOAr is a venue for early communication or feedback before peer review. Data may be preliminary.Learn more about preprints preprintOpen AccessYou are viewing the latest version by default [v2]Fuzzy-Committees of Conceptual distributed ModelsAuthorsMostafaFarragiDMostafaFarragiDGeraldCorzo PereziDDimitriSolomatineiDSee all authors Mostafa FarragiDCorresponding Author• Submitting AuthorGFZ German Research Centre for GeosciencesiDhttps://orcid.org/0000-0002-1673-0126view email addressThe email was not providedcopy email addressMostafa FarragiDGFZ German Research Centre for GeosciencesiDhttps://orcid.org/0000-0002-1673-0126view email addressThe email was not providedcopy email addressGerald Corzo PereziDIHE DelftiDhttps://orcid.org/0000-0002-2773-7817view email addressThe email was not providedcopy email addressDimitri SolomatineiDDelft University of TechnologyiDhttps://orcid.org/0000-0003-2031-9871view email addressThe email was not providedcopy email address
Finding a balance between conflicting interests in multipurpose reservoirs is an important challenge for decision makers. This study assesses the use of different computational tools to obtain optimal reservoir operations at the Hatillo dam in the Dominican Republic. A multi-objective optimization approach is applied to models that simulate reservoir operations and three different machine learning (ML) models are employed to learn the real operation of the system. A general model is proposed to simulate daily reservoir operations (2009–2019), integrating water balances, physical constraints of the dam components, and the ML models, the latter defining daily controlled discharges. In the optimization process, the ML parameters are the decision variables, while the objectives evaluated are irrigation, hydropower generation, and flood control. The results are compared with the actual operation of the reservoir. The flood control objective was found to have a wide room for improvement over the real operation of the reservoir, and several of the solutions were found to improve the real operation for the three proposed objectives. The multilayer perceptron models tended to generate the best results for this case study and the nondominated sorting generic algorithm (NSGA II) optimizer generated the best optimization results.
Extreme runoff modeling is hindered by the lack of sufficient and relevant ground information and the low reliability of physically based models. This chapter proposes combining precipitation Remote Sensing (RS) products, Machine Learning (ML) modeling, and hydrometeorological knowledge to improve extreme runoff modeling. The approach applied to improve the representation of precipitation is the object-based Connected Component Analysis (CCA), a method that enables classifying and associating precipitation with extreme runoff events. Random Forest (RF) is employed as an ML model. Two and a half years of near-real-time hourly RS precipitation from the PERSIANN-CCS and IMERG-early run databases and runoff at the outlet of a basin in the tropical Andes of Ecuador were used. The developed models show the ability to simulate extreme runoff for the cases of long-duration precipitation events regardless of the spatial extent, obtaining Nash-Sutcliffe efficiencies above 0.72. On the contrary, there was an unacceptable model performance for a combination of short-duration and spatially extensive precipitation events. The strengths/weaknesses of the developed ML models are attributed to the ability/difficulty to represent complex precipitation-runoff responses.
Submarine mass movement generates different depositional units that include creep, slides, slumps, debris flows, etc., which can be collectively termed the mass transport complexes or Mass Transport Deposits (MTDs). Their interpretation is crucial, as such deposits during translation over the instable slope may lead to several catastrophic submarine events, e.g., landslides, tsunamis, or avalanches, and hence possess precursory threats for subsea installations. This chapter shows how to design a workflow and to compute a new attribute, the MTD cube meta-attribute, through an artificial neural network. This has illuminated the structural architecture of the MTD from 3-D seismic data in the Karewa prospect of offshore Taranaki Basin, New Zealand.
The study of seismic attributes helps to elucidate the subsurface geologic body, but no single attribute will always correspond to a particular structure or feature. However, a hybrid-attribute can be designed by amalgamating a set of attributes into a single attribute which, in turn, can delimit a particular geologic feature with greater certainty. This chapter describes these hybrid attributes, termed meta-attributes, which have evolved recently to augment the interpretation of seismic data, especially for 3D seismic data.
Earth observation data are a vital resource for studying long-term changes, but the large data volumes can be challenging to analyze. Time-series analysis in particular is hampered by the typical thin-time-slice file organization. We examine several potential solutions inspired in large part by the data-parallel methods that have arisen with cloud computing. These solutions include various combinations of data reorganization, spatial indexing, distributed storage, and pre-computation that we term Analytics Optimized Data Stores (AODS) . We find that even simple solutions (such as a data cube) produce more than an order of magnitude improvement; the best provide two to three orders of magnitude improvement. The most performant solutions have tradeoffs in terms of generality or storage footprint, but may nonetheless be useful components in data analytics frameworks where performance is critical.