Recently, a hybrid framework combining machine learning (ML) and process-based equations, termed differentiable modeling, has shown comparable accuracy to pure ML models while offering enhanced interpretability and spatial generalizability. However, it remained unclear how well hybrid models generalize to extreme floods outside of the range of training data, and whether optimizing models for extreme events jeopardizes spatial generalizability and the physical significance of internal variables. Here we evaluated multiple versions of a differentiable model (delta HBV1.0 and delta HBV1.1p) for predicting unseen extreme events, and benchmarked them against a widely-applied long short-term memory (LSTM) network on the CAMELS data set. We found that both delta HBV and LSTM models performed well, with delta HBV1.1p outperforming LSTM for events with a return period of 5 years or more. This advantage was more pronounced as the return period increased (0.06 higher median Nash-Sutcliffe efficiency and lower peak flow errors for 80% of the 50-year or rarer events). Loss function choice had a larger impact on delta HBV1.1p than on LSTM, and we showed the proper loss led to delta HBV models that further surpassed LSTM in different performance aspects. Furthermore, allowing more dynamic parameters improved the extreme metrics, had no negative impact on spatial generalization, and exerted a minimal influence on the untrained variables. We hypothesize that delta HBV's mass balance and first-order exchange terms help constrain and inform its responses to mitigate the underestimation of peaks compared to LSTM. We conclude that adopting interpretable structural priors can improve generalizability to unseen cases and thus increase model reliability to better inform stakeholder preparedness.
Understanding streamflow dynamics in watersheds affected by human activity and climate variability is important for sustainable water and environmental resource management. This study evaluates the vulnerability of Alabama watersheds to anthropogenic and climatic changes using an integrated framework combining GIS, remote sensing, hydrological modeling, and machine learning (ML). Three Soil and Water Assessment Tool (SWAT) models, differing in spatial resolution and soil inputs, were developed to simulate streamflow under baseline and land-use/land cover (LULC) scenarios from 1990 to 2023. The model, built with consistent 100 × 100 m rasters and fine-resolution SSURGO (Soil Survey Geographic Database) soil data, achieved the best calibration and was selected for detailed analysis. Streamflow trends were assessed over two periods (1993–2009 and 2010–2023) to help isolate climate variability (from LULC effects), while LULC changes were evaluated using 1992, 2011, and 2021 maps. A Long Short-Term Memory (LSTM) model further enhanced simulation accuracy by integrating partially calibrated SWAT outputs. Watershed vulnerability was ranked using a multi-criteria framework. Two watersheds were classified as highly vulnerable, nine as moderately vulnerable, and three as having low vulnerability. Basin-level contrasts revealed moderate climate impacts in the Tombigbee Basin, greater climate sensitivity in the Black Warrior Basin, and LULC-dominated impacts in the Alabama Basin. Overall, LULC change exerted stronger and more spatially variable effects on streamflow than climate variability. This study introduces a transferable SWAT–ML vulnerability ranking framework to guide watershed and environmental management in data-scarce, human-modified regions.
ABSTRACT This technical note describes recent efforts to integrate machine learning (ML) models, specifically long short‐term memory (LSTM) networks and differentiable parameter learning conceptual hydrological models (δ conceptual models), into the next‐generation water resources modeling framework (Nextgen) to enhance future versions of the U.S. National Water Model (NWM). We address three specific methodology gaps of this new modeling framework: (1) assess model performance across many ungauged catchments, (2) diagnostic‐based model selection, and (3) regionalization based on catchment attributes. We demonstrate that an LSTM trained on CAMELS catchments can make large‐scale predictions with Nextgen across the New England region and match the average flow duration curve observed by stream gauges for streamflow with low exceedance probability (high flows), but diverges from the mean in high exceedance probability (low flows). We demonstrate improvements in peak flow predictions when using δ conceptual model, but results also suggest that performance increases may come at a cost of accurately representing hydrologic states within the conceptual model. We propose a novel approach using ML to predict the most performant mosaic modeling approach and demonstrate improved distributions of efficiency scores throughout the large sample of basins. Our findings advocate for the future development of ML capabilities within Nextgen for advancing operational hydrological modeling.
We calculated metrics of climate change, land use-land cover change, and hydrologic nonstationarity in 671 catchments across the Contiguous United States (CONUS) that are known not to have relatively little urbanization and anthropogenic land cover. Climate change is correlated with hydrologic nonstationarity in these basins. Land use-land cover change has no correlation with hydrologic nonstationarity in these basins. We present kriging maps over CONUS showing climate change, land use-land cover change, and hydrologic nonstationarity.
Detecting river centrelines and estimating river water surface widths are valuable for measuring planform geometry. Extracting the river centreline and water surface width estimation from satellite images enables a better understanding of the river network dynamics and can be useful in predicting future changes. This study introduces a novel approach to detect river centrelines and water surface widths that leverages the powerful feature extraction capabilities of DeepLabV3, a state-of-the-art semantic segmentation model and integrates them with the geometric analysis strengths of the Medial Axis Transform (MAT). Our MAT approach identifies the midpoints of the river to detect the centreline. It calculates the distance from the centreline to the edge of the river and estimates the water surface width of the river. The effectiveness of this approach was validated through case studies of the Sipsey, Coosa, Tennessee and Mississippi Rivers. This approach utilizes Sentinel-1A Synthetic Aperture Radar (SAR) imagery, enabling data acquisition independent of the weather conditions. We validated this approach using the National Hydrography Dataset Plus (NHDPlus) and in situ water surface width measurements from the HYDRoacoustic dataset in support of the Surface Water Oceanographic Topography (HYDRoSWOT) and cross-section width from the United States Geological Survey (USGS). The MAT approach accurately extracted centrelines and estimated water surface widths for the straight and meandering river sections. Quantitatively, the fine-tuned DeepLabV3 model achieved a 0.933 F1-score for water mask extraction, while the resulting centreline RMSEs against NHDPlus ranged from 0.55 m (Sipsey River) to 3.10 m (Mississippi River), and water surface width estimations generally varied by 2-15% from in-situ measurements. The accuracy of the method is high for straight and meandering rivers; however, errors, primarily caused by complex river morphology, increased in the Mississippi River because its braided channel system challenged the ability of MAT to define a consistent centreline and water surface width. However, it exhibited reduced accuracy and significant spatial deviations when applied to complex braided sections of the Mississippi River. This approach will significantly advance the field of planform geometry measurement by providing researchers with reliable, scalable and practical methodologies that can be used to develop robust and efficient tools.
River hydrodynamics are influenced by numerous factors that traditional models often fail to fully capture. Simulating complex hydrographs can benefit from parsimonious upscaling models, such as fractional derivative equations, that reduce the need to account for all variables. While fractional Saint-Venant equations (SVEs) have been mathematically explored, they lack clear physical interpretation and have not been applied in practical scenarios. This study introduces novel fractional-order Saint-Venant equations (FSVEs) for simulating river flow dynamics, addressing limitations in conventional modeling. Three models—constant, tempered, and variable time-fractional SVEs (CtFSVE, TtFSVE, and VtFSVE)—are developed to capture peak attenuation and tailing more effectively. Numerical experiments indicate that lower time-fractional derivative values enhance retention, producing a lower peak, delayed peak arrival, and pronounced late-time tailing. TtFSVE models transient tailing in hydrographs, VtFSVE captures transient evolution where inflow and outflow differ, and CtFSVE balances accuracy and simplicity with a single added parameter for various hydrographs. In the simulation of real-world hydrograph data, the fractional SVEs show high predictive accuracy. However, they should be regarded as effective proxy models that require parameter calibration and are not yet fully ‘plug-and-play’ predictive models. Comparative analysis with the Long Short-Term Memory (LSTM) machine learning model and distributed domain coupling model (DDCM) shows CtFSVE’s superior performance in capturing complex flow dynamics with minimal data, while field validation demonstrates its accuracy over traditional SVE, underscoring its practicality for complex river networks. The fractional engine shows promise as an effective tool for upscaling surface flow without the prohibitive burden of mapping detailed system heterogeneity.
Abstract Rapid and accurate maps of floods across large domains, with high temporal resolution capturing event peaks, have applications for flood forecasting and resilience, damage assessment, and parametric insurance. Satellite imagery produces incomplete observations spatially and temporally, and hydrodynamic models require tradeoffs between computational efficiency and accuracy. We address these challenges with a novel flood model which predicts surface water area from the U.S. National Water Model using a convolutional neural network (NWM‐CNN). We trained NWM‐CNN on 780 flood events, at a 250 m resolution with an RMSE of 4.58% on held out validation geographies. We demonstrate NWM‐CNN across California during the 2023 atmospheric rivers, comparing predictions against Sentinel‐1 mapped flood observations. We compared historical predictions from 1979 to 2023 to flood damage reports in Sacramento County, California. Results show that NWM‐CNN captures inundation extent better than the Height Above Nearest Drainage (HAND) approach (25%–36% RMSE, respectively).
Accurate representation of the turbulent exchange of carbon, water, and heat between the land surface and the atmosphere is critical for modelling global energy, water, and carbon cycles in both future climate projections and weather forecasts. Evaluation of models' ability to do this is performed in a wide range of simulation environments, often without explicit consideration of the degree of observational constraint or uncertainty and typically without quantification of benchmark performance expectations. We describe a Model Intercomparison Project (MIP) that attempts to resolve these shortcomings, comparing the surface turbulent heat flux predictions of around 20 different land models provided with in situ meteorological forcing evaluated with measured surface fluxes using quality-controlled data from 170 eddy-covariance-based flux tower sites. Predictions from seven out-of-sample empirical models are used to quantify the information available to land models in their forcing data and so the potential for land model performance improvement. Sites with unusual behaviour, complicated processes, poor data quality, or uncommon flux magnitude are more difficult to predict for both mechanistic and empirical models, providing a means of fairer assessment of land model performance. When examining observational uncertainty, model performance does not appear to improve in low-turbulence periods or with energy-balance-corrected flux tower data, and indeed some results raise questions about whether the energy balance correction process itself is appropriate. In all cases the results are broadly consistent, with simple out-of-sample empirical models, including linear regression, comfortably outperforming mechanistic land models. In all but two cases, latent heat flux and net ecosystem exchange of CO2 are better predicted by land models than sensible heat flux, despite it seeming to have fewer physical controlling processes. Land models that are implemented in Earth system models also appear to perform notably better than stand-alone ecosystem (including demographic) models, at least in terms of the fluxes examined here. The approach we outline enables isolation of the locations and conditions under which model developers can know that a land model can improve, allowing information pathways and discrete parameterisations in models to be identified and targeted for future model development.
We explore deep learning for the Next Generation Water Resources Modeling Framework (NextGen). We present results from Random forest-based multi-model ensembles, Long Short-Term Memory (LSTM) and differentiable parameter learning hydrological models (δ conceptual models) and attribute sensitivity for ungauged basins. This poster was presented at the CIROH Developers Conference May 29 – June 1, 2024, at the University of Utah in Salt Lake City.
It has been proposed that conservation laws might not be beneficial for accurate hydrological modeling due to errors in input (precipitation) and target (streamflow) data, and this might explain why deep learning models (which are not based on enforcing closure) can out-perform catchment-scale conceptual and process-based models at predicting streamflow. We test this hypothesis using physics-informed machine learning and find that: (1) enforcing closure in the rainfall-runoff mass balance does appear to harm the overall skill of hydrological models, (2) deep learning models learn to account for spatiotemporally variable biases in data, however (3) this “closure” effect accounts for only a small fraction of the difference in predictive skill between deep learning and conceptual models.
Long short-term memory (LSTM) models have been shown to be efficient for rainfall-runoff modeling, and to a lesser extent, for groundwater depth forecasting. In this study, LSTMs were applied to quantify the spatiotemporal evolution of surface and subsurface hydrographs in Alabama in the Southeastern United States, where water sustainability has not been fully quantified across spatiotemporal scales. First, the surface water LSTM model with extensive dynamic (precipitation and other weather variables) and static (basin characteristics) inputs predicted the main characteristics of streamflow for six years at 19 gauged basins in Alabama. The model tended to underestimate extremely high streamflow but adding drainage density as an input feature slightly improved the predictions of extreme events. Second, to predict the groundwater depth evolution, a groundwater LSTM (GW-LSTM) model was proposed and applied using static inputs capturing the aquifers' hydrogeological properties and dynamic inputs of meteorological information. Three precipitation scenarios were also explored to evaluate the groundwater hydrograph evolution in the next two decades. The GW-LSTM model predicted the general trend of daily groundwater depth fluctuations (at 21 wells distributed across Alabama from 1990 to 2021) including most extremely high groundwater levels, and recovered groundwater depth for locations withheld from model training and validation. This study, therefore, extended the application of LSTMs in quantifying the spatiotemporal evolution of surface water and groundwater, two manifestations of a single integrated resource.
We used Natural Language Processing (NLP) to assess topic diversity in all research articles (∼75,000) from eighteen water science and hydrology journals published between 1991 and 2019. We found that individual water science and hydrology research articles are becoming increasingly interdisciplinary in the sense that, on average, the number of equally-common topics represented in individual articles is increasing. This is true even though the body of water science and hydrology literature as a whole is not becoming more topically diverse. These findings suggest that the National Research Council’s (1991) recommendation to increase multidisciplinarity of hydrological research has been followed. Topics with the largest increases in popularity were Climate Change Impacts, Water Policy & Planning, and Pollutant Removal. Topics with the largest decreases in popularity were Stochastic Models and Numerical Models. At a journal level, Water Resources Research, Journal of Hydrology, and Hydrological Processes are the three most topically diverse journals in the discipline. We also identified topics that are becoming increasingly isolated, and which could potentially benefit from integrating more with the wider hydrology discipline.
We used Natural Language Processing (NLP) to assess topic diversity in all research articles (similar to 75,000) from eighteen water science and hydrology journals published between 1991 and 2019. We found that individual water science and hydrology research articles are becoming increasingly diverse in the sense that, on average, the number of topics represented in individual articles is increasing, which may be a sign of increasing interdisciplinarity. This is true even though the body of water science and hydrology literature as a whole is not becoming more topically diverse. Topics with the largest increases in popularity were Climate Change Impacts, Water Policy & Planning, and Pollutant Removal. Topics with the largest decreases in popularity were Stochastic Models and Numerical Models. At a journal level, Water Resources Research, Journal of Hydrology, and Hydrological Processes are the three most topically diverse journals among the corpus that we studied.
Hydrologic exchange flows (HEFs) of river-aquifer systems are known to affect water flow, but the quantitative influences of lateral HEFs and the riparian zone's hydraulic conductivity (K) distribution on stream fluxes remain obscure under varying hydrologic conditions. To fill this knowledge gap, this study proposed a physical based, distributed domain, coupled (open channel and groundwater) flow model (DDCM) to quantify the effect of lateral HEFs on hydrograph characteristics, including especially the peak discharge and tailing decay which are practically important. Numerical experiments showed that (1) the interaction between the lateral HEFs and river hydrodynamics reduced the peak discharge and flood flow rates, (2) a heterogeneous K field of the riparian zone generated multi-rate HEFs (which then changed flow response) represented by the hydrographs with various declining rates (varying from exponential to power-law), significantly expanding the flow process, and (3) the probability density function of K also affected the tailing and peak of the hydrograph. A preliminary test showed that the DDCM captured the overall pattern of hydrographs observed from a catchment in the Wadi Ahin West, Oman. This study, therefore, provided a model-based quantification of the mechanisms and factors of the lateral HEFs affecting the hydrograph pattern in flood events, and further applications are needed to test the applicability of the DDCM in capturing real-world hydrographs affected by the lateral HEFs.
Abstract. The most accurate rainfall-runoff predictions are currently based on deep learning. There is a concern among hydrologists that data-driven models based on deep learning may not be reliable in extrapolation or for predicting extreme events. This study tests that hypothesis using Long Short-Term Memory networks (LSTMs) and an LSTM variant that is architecturally constrained to conserve mass. The LSTM (and the mass-conserving LSTM variant) remained relatively accurate in predicting extreme (high return-period) events compared to both a conceptual model (the Sacramento Model) and a process-based model (US National Water Model), even when extreme events were not included in the training period. Adding mass balance constraints to the data-driven model (LSTM) reduced model skill during extreme events.