We propose an approximate Bayesian method for quantifying the total uncertainty in inverse partial differential equation (PDE) solutions obtained with machine learning surrogate models, including operator learning models. The proposed method accounts for uncertainty in the observations, PDE, and surrogate models. First, we use the surrogate model to formulate a minimization problem in the reduced space for the maximum a posteriori (MAP) inverse solution. Then, we randomize the MAP objective function and obtain samples of the posterior distribution by minimizing different realizations of the objective function. We test the proposed framework by comparing it with the iterative ensemble smoother and deep ensembling methods for a nonlinear diffusion equation with an unknown space-dependent diffusion coefficient. Among other applications, this equation describes the flow of groundwater in an unconfined aquifer. Depending on the training dataset and ensemble sizes, the proposed method provides similar or more descriptive posteriors of the parameters and states than the iterative ensemble smoother method. Deep ensembling underestimates uncertainty and provides less-informative posteriors than the other two methods. Our results show that, despite inherent uncertainty, surrogate models can be used for parameter and state estimation as an alternative to the inverse methods relying on (more accurate) numerical PDE solvers.
We evaluate the performance of the Karhunen-Lo & egrave;ve Deep Neural Network (KL-DNN) framework for surrogate modeling and approximate Bayesian parameter estimation in partial differential equation models. In the surrogate model, the Karhunen-Lo & egrave;ve (KL) expansions are used for the dimensionality reduction of the number of unknown parameters and variables, and a deep neural network is employed to relate the reduced space of parameters to that of the state variables. The KL-DNN surrogate model is used to formulate a maximum-a-posteriori-like least-squares problem, which is randomized to draw samples of the posterior distribution of the parameters. We test the proposed framework for a hypothetical unconfined aquifer via comparison with the forward MODFLOW and inverse PEST++ iterative ensemble smoother (IES) solutions as well as the state-of-theart Fourier neural operator (FNO) and deep operator networks (DeepONets) operator learning surrogate models. Our results show that the KL-DNN surrogate model outperforms FNO and DeepONet for forward predictions. For solving inverse problems, the randomized algorithm provides the same or more accurate Bayesian predictions of the parameters than IES as evidenced by the higher log predictive probability of both the estimated parameter field and the forecast hydraulic head. The posterior mean obtained from the randomized algorithm is closer to the reference parameter field than that obtained with FNO as the maximum a posteriori estimate.
The National Weather Service (NWS) Office of Water Prediction (OWP), in conjunction with the National Center for Atmospheric Research and the NWS National Centers for Environmental Prediction (NCEP) implemented version 2.1 of the National Water Model (NWM) into operations in April of 2021. As with the initial version implemented in 2016, NWM v2.1 is an hourly cycling analysis and forecast system that provides streamflow guidance for millions of river reaches and other hydrologic information on high-resolution grids. The NWM provides complementary hydrologic guidance at current NWS river forecast locations and significantly expands guidance coverage and water budget information in underserved locations. It produces a full range of hydrologic fields, which can be leveraged by a broad cross section of stakeholders ranging from the emergency responder and water resource communities, to transportation, energy, recreation and agriculture interests, to other water-oriented applications in the government, academic and private sectors. Version 2.1 of the NWM represents the fifth major version upgrade and more than doubles simulation skill with respect to hourly streamflow correlation, Nash Sutcliffe Efficiency, and bias reduction, over its original inception in 2016. This paper will discuss the driving factors underpinning the creation of the NWM, provide a brief overview of the model configuration and performance, and discuss future efforts to improve NWM components and services.
In the face of escalating instances of inland and flash flooding spurred by intense rainfall and hurricanes, the accurate prediction of rapid streamflow variations has become imperative. Traditional data assimilation methods face challenges during extreme rainfall events due to numerous sources of error, including structural and parametric model uncertainties, forcing biases, and noisy observations. This study introduces a cutting-edge hybrid ensemble and optimal interpolation data assimilation scheme tailored to precisely and efficiently estimate streamflow during such critical events. Our hybrid scheme uses an ensemble-based framework, integrating the flow-dependent background streamflow covariance with a climatological error covariance derived from historical model simulations. The dynamic interplay (weight) between the static background covariance and the evolving ensemble is adaptively computed both spatially and temporally. By coupling the National Water Model (NWM) configuration of the WRF-Hydro modeling system with the Data Assimilation Research Testbed (DART), we evaluate the performance of our hybrid prediction system using two impactful case studies: (1) West Virginia's flash flooding event in June 2016 and (2) Florida's inland flooding during Hurricane Ian in September 2022. Our findings reveal that the hybrid scheme substantially outperforms its ensemble counterpart, delivering enhanced streamflow estimates for both low and high flow scenarios, with an improvement of up to 50 %. This heightened accuracy is attributed to the climatological background covariance, mitigating bias and augmenting ensemble variability. The adaptive nature of the hybrid algorithm ensures reliability, even with a very small time-varying ensemble. Moreover, this innovative hybrid data assimilation system propels streamflow forecasts up to 18 h in advance of flood peaks, marking a substantial advancement in flood prediction capabilities.
We propose a reduced-order deep-learning surrogate model for dynamic systems described by time-dependent partial differential equations. This method employs space-time Karhunen-Lo & egrave;ve expansions (KLEs) of the state variables and space-dependent KLEs of space-varying parameters to identify the reduced (latent) dimensions. Subsequently, a deep neural network (DNN) is used to map the parameter latent space to the state variable latent space. An approximate Bayesian method is developed for uncertainty quantification (UQ) in the proposed KL-DNN surrogate model. The KL-DNN method is tested for the linear advection- diffusion and nonlinear diffusion equations, and the Bayesian approach for UQ is compared with the deep ensembling (DE) approach, commonly used for quantifying uncertainty in DNN models. It was found that the approximate Bayesian method provides a more informative distribution of the PDE solutions in terms of the coverage of the reference PDE solutions (the percentage of nodes where the reference solution is within the confidence interval predicted by the UQ methods) and log predictive probability. The DE method is found to underestimate uncertainty and introduce bias. For the nonlinear diffusion equation, we compare the KL-DNN method with the Fourier Neural Operator (FNO) method and find that KL-DNN is 10% more accurate and needs less training time than the FNO method.
COPYRIGHT © 2023 Siddique, Sharma and McCreight. This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms. Editorial: Hydrological modeling, analyses, and predictions: Opportunities and challenges
In situ gauge networks are often used in hydrologic model calibration, but these networks are limited or nonexistent in many regions. The upcoming Surface Water Ocean Topography (SWOT) mission promises to fill this observation gap by providing discharge estimates for rivers wider than 100 m. SWOT observation utility for model parameter selection in regions devoid of in situ gauges is assessed using proxy SWOT discharge estimates derived from an observing system simulation experiment and Monte Carlo methods. The sensitivity of the parameter selection to measurement error and observation frequency is also evaluated. Single‐ and multi‐point parameter selection are performed for ten sub‐basins within the Susitna and upper Tanana river basins in Alaska. SWOT is expected to observe Alaskan river points 4–7 times per 21‐day repeat cycle with 120‐km swath coverage. For an expected SWOT measurement error of 35%, parameter estimation is successful for 50% (90%) of sub‐basins using single‐ (multi‐) point parameter selection. Decreasing observation frequency to simulate lower latitudes resulted in success for only 10% of midlatitude and tropical sub‐basins for single‐point selection, whereas multi‐point selection was successful in 80% (60%) of midlatitude (tropical) sub‐basins. Single‐point parameter selection is more sensitive to measurement error than multi‐point parameter selection. The results strongly support the use of multi‐point over single‐point parameter selection, yielding robust results nearly independent of observation frequency. Most importantly, this study suggests SWOT can be used to successfully select hydrologic model parameters in basins without an in situ gauge network.
Streamflow timing errors (in the units of time) are rarely explicitly evaluated but are useful for model evaluation and development. Wavelet-based approaches have been shown to reliably quantify timing errors in streamflow simulations but have not been applied in a systematic way that is suitable for model evaluation. This paper provides a step-by-step methodology that objectively identifies events, and then estimates timing errors for those events, in a way that can be applied to large-sample, high-resolution predictions. Step 1 applies the wavelet transform to the observations and uses statistical significance to identify observed events. Step 2 utilizes the cross-wavelet transform to calculate the timing errors for the events identified in step 1; this includes the diagnostic of model event hits, and timing errors are only assessed for hits. The methodology is illustrated using real and simulated stream discharge data from several locations to highlight key method features. The method groups event timing errors by dominant timescales, which can be used to identify the potential processes contributing to the timing errors and the associated model development needs. For instance, timing errors that are associated with the diurnal melt cycle are identified. The method is also useful for documenting and evaluating model performance in terms of defined standards. This is illustrated by showing the version-over-version performance of the National Water Model (NWM) in terms of timing errors.
Abstract. Predicting major floods during extreme rainfall events remains an important challenge. Rapid changes in flows over short time-scales combined with multiple sources of model error makes it difficult to accurately simulate intense floods. This study presents a general data assimilation framework that aims to improve flood predictions in channel routing models. Hurricane Florence, which caused catastrophic flooding and damages in the Carolinas in September 2018, is used as a case study. The National Water Model (NWM) configuration of the WRF-Hydro modeling framework is interfaced with the Data Assimilation Research Testbed (DART) to produce ensemble streamflow forecasts and analyses. Hourly streamflow observations from 107 United States Geological Survey (USGS) gauges are assimilated for a period of one month. The data assimilation (DA) system developed in this paper explores two novel contributions: (1) Along-The-Stream (ATS) covariance localization and (2) spatially and temporally varying adaptive covariance inflation. ATS localization aims to mitigate not only spurious correlations, due to limited ensemble size, but also physically incorrect correlations between unconnected and indirectly connected state variables in the river network. We demonstrate that ATS localization provides improved information propagation during the model update. Adaptive prior inflation is used to tackle errors in the prior, including large model biases. Analysis errors incurred during the update are addressed using posterior inflation. Results show that ATS localization is a crucial ingredient of our hydrologic DA system, providing at least 40% more accurate (RMSE) streamflow estimates than regular, Euclidean distance-based localization. Assessment of hydrographs indicates that adaptive inflation is extremely useful and perhaps indispensable for improving the forecast skill during flooding events with significant model errors. We argue that adaptive prior inflation is able to serve as a vigorous bias correction scheme which varies both spatially and temporally. Major improvements over the model's severely underestimated streamflow estimates are suggested along Pee Dee River in South Carolina and many other locations in the domain, where inflation is able to avoid filter divergence and thereby assimilate significantly more observations.
Accurate representation of channel properties is important for forecasting in hydrologic models, as it affects the height, celerity, and attenuation of flood waves. Yet, considerable uncertainty in the parameterization of channel geometry and hydraulic roughness (Manning's n) exists within the NOAA National Water Model (NWM), due largely to data scarcity; only ∼2800 out of the 2.7×106 river reach segments in the NWM have measured channel properties. In this study, we seek to improve channel representativeness by updating channel geometry and roughness parameters using a large, previously unpublished hydraulic geometry (HyG) dataset of approximately 48 000 gauges. We begin with a Sobol' sensitivity analysis of channel geometry parameters for 12 small, semi-natural basins across the continental U.S., which reveals an outsized sensitivity of simulated flow to Manning's n relative to channel geometry parameters. We then develop and evaluate a set of regression-based regionalizations of channel parameters estimated using the HyG dataset. Finally, we compare the model output generated from updated channel parameter sets to observations and the current NWM v2.1 parameterization. We find that while the NWM land surface model holds the most influence over flow, given its control over total volume, the updated channel parameterization leads to improvements in simulated streamflow performance relative to observed flows, with a statistically significant mean R2 increase from 0.479 to 0.494 across approximately 7400 gauge locations. HyG-based channel geometry and roughness provide a substantial overall improvement in channel representation over the default parameterization, updating the previous set value for most reaches of Manning's n=0.060 to a new range between 0.006 and 0.537 (median 0.077). This research provides a more representative, observationally based channel parameter dataset for the NWM routing module and new insight into the influence of the routing module within the overall modeling framework.
Earth and Space Science Open Archive This preprint has been submitted to and is under consideration at Water Resources Research. ESSOAr is a venue for early communication or feedback before peer review. Data may be preliminary.Learn more about preprints preprintOpen AccessYou are viewing an older version [v1]Go to new versionHydrologic Model Parameter Estimation in Ungauged Basins using Simulated SWOT Discharge ObservationsAuthorsNicholas JElmeriDJames LMcCreightiDChristopher R.HainSee all authors Nicholas J ElmeriDCorresponding Author• Submitting AuthorNASA Postdoctoral ProgramiDhttps://orcid.org/0000-0002-6343-6210view email addressThe email was not providedcopy email addressJames L McCreightiDNational Center for Atmospheric Research (UCAR)iDhttps://orcid.org/0000-0001-6018-425Xview email addressThe email was not providedcopy email addressChristopher R. HainMarshall Space Flight Center, NASAview email addressThe email was not providedcopy email address
Predicting major floods during extreme rainfall events remains an important challenge. Rapid changes in flows over short timescales, combined with multiple sources of model error, makes it difficult to accurately simulate intense floods. This study presents a general data assimilation framework that aims to improve flood predictions in channel routing models. Hurricane Florence, which caused catastrophic flooding and damages in the Carolinas in September 2018, is used as a case study. The National Water Model (NWM) configuration of the WRF-Hydro modeling framework is interfaced with the Data Assimilation Research Testbed (DART) to produce ensemble streamflow forecasts and analyses. Instantaneous streamflow observations from 107 United States Geological Survey (USGS) gauges are assimilated for a period of 1 month. The data assimilation (DA) system developed in this paper explores two novel contributions, namely (1) along-the-stream (ATS) covariance localization and (2) spatially and temporally varying adaptive covariance inflation. ATS localization aims to mitigate not only spurious correlations, due to limited ensemble size, but also physically incorrect correlations between unconnected and indirectly connected state variables in the river network. We demonstrate that ATS localization provides improved information propagation during the model update. Adaptive prior inflation is used to tackle errors in the prior, including large model biases which often occur in flooding situations. Analysis errors incurred during the update are addressed using posterior inflation. Results show that ATS localization is a crucial ingredient of our hydrologic DA system, providing at least 40 % more accurate (root mean square error) streamflow estimates than regular, Euclidean distance-based localization. An assessment of hydrographs indicates that adaptive inflation is extremely useful and perhaps indispensable for improving the forecast skill during flooding events with significant model errors. We argue that adaptive prior inflation is able to serve as a vigorous bias correction scheme which varies both spatially and temporally. Major improvements over the model's severely underestimated streamflow estimates are suggested along the Pee Dee River in South Carolina, and many other locations in the domain, where inflation is able to avoid filter divergence and, thereby, assimilate significantly more observations.
The Data Assimilation Research Testbed (DART) is a community facility for ensemble data assimilation developed and maintained by the National Center for Atmospheric Research (NCAR). DART provides ensemble data assimilation capabilities for NCAR community earth system models and many other prediction models. It is straightforward to add interfaces for new models and new observations to DART. DART provides traditional ensemble data assimilation algorithms that implicitly assume Gaussianity and linearity. Traditional algorithms can still work when these assumptions are violated. However, it is possible to greatly improve results by extending ensemble algorithms to explicitly account for aspects of nonlinearity and non-Gaussianity. Two new algorithms have been added to DART. 1). Anamorphosis transforms variables to make the assimilation problem more linear and Gaussian before transforming posterior estimates back to the original model variables; 2). The marginal correction rank histogram filter (MCRHF) directly represents arbitrary non-Gaussian distributions. These methods are particularly valuable for data assimilation for bounded quantities like tracers or streamflow. DART is being applied to a number of novel applications. Examples in the poster include 1). An eddy-resolving global ocean ensemble reanalysis with the POP ocean model and an ensemble optimal interpolation; 2). The WRF-Hydro/DART system now includes a multi-parametric ensemble, anamorphosis, and spatially-correlated noise for the forcing fields. 3). Results from the Carbon Monitoring System over Mountains using CLM5 to assimilate remotely-sensed observations (LAI, biomass, and SIF) for a field site in Colorado; 4). Assimilation of MODIS snow cover fraction and daily GRACE total water storage data and its impact on soil moisture using the DART/NOAH-MP system. 5). An ensemble atmospheric reanalysis using the CAM general circulation model.
The Data Assimilation Research Testbed (DART) has been coupled with the community WRF-Hydro modeling system with the intent of providing efficient and flexible support for assimilating a wide range of streamflow and soil moisture observations and delivering an ensemble of model states useful for quantifying streamflow uncertainties. The coupled framework, named Hydro-DART, is used to study and assess the flooding consequences of Hurricane Florence over the Carolinas during August-September 2018 period.Several extensions to earlier versions of Hydro-DART have been explored. These include: (1) a multi-configuration ensemble in which different ensemble members are run with different physical parameters (e.g., Manning's roughness and channel geometry) in order to create additional ensemble variability, (2) a variable transform, anamorphosis, which is introduced such that bounded quantities (e.g., streamflow) are transformed to a Gaussian space prior to the Kalman update as a way to avoid non-physical state updates, (3) a spatially-correlated noise, which is introduced to represent uncertainty of input forcings (e.g., overland and subsurface fluxes) in a physically meaningful way, and (4) an along-the-stream localization, which considers precipitation correlation length scale, rather than physical proximity. Hourly streamflow gauge data, from the flood-affected area, is used to test the impact of these extensions on the overall prediction accuracy. Analyses and hindcasts are compared to those based on the nudging assimilation currently employed in the National Water Model (NWM) operations. Standard streamflow forecast metrics are also supplemented by a wavelet-based event timing error metric.
Accurately estimating the snow water equivalent (SWE) that is stored in the worlds mountains remains a challenging and important unsolved problem. The SWE reconstruction approach, where the remotely sensed seasonal depletion of fractional snow-covered area (fSCA) is used with a snow model to build up the snowpack in reverse, has been used for decades to help tackle this problem retrospectively. Despite some success, this deterministic approach ignores uncertainties in the snow model, the meteorological forcing, and the remotely sensed fSCA. A trade-off has also existed between the desired temporal and spatial resolution of the satellite-retrieved fSCA depletion. Recently, ensemble-based data assimilation techniques that can account for the uncertainties inherent in the reconstruction exercise have allowed for probabilistic snow reanalyses. In addition, new higher resolution optical satellite constellations such as Sentinel-2 and the PlanetScope cubesats have been launched into polar orbit, potentially eliminating the aforementioned trade-off. We combine these two developments, namely ensemble-based data assimilation and the emerging remotely sensed data streams, to see if snow reanalyses can be improved at the hillslope (100 m) scale in complex terrain. As a first step, we develop accurate high-resolution binary snow-cover maps using a terrestrial automatic camera system installed on a mountaintop near Ny-Ålesund (Svalbard, Norway). These maps are used to validate fSCA retrieved from various satellite sensors (MODIS, Sentinel-2 MSI, and Landsat 8 OLI) using algorithms ranging from simple thresholding of the normalized difference snow index to spectral unmixing. Through the validation, we demonstrate that the spectral unmixing technique can obtain unbiased fSCA retrievals at the hillslope scale. Next, we move to the Mammoth Lakes basin in the Californian Sierra Nevada, USA, where we have access to independent validation data retrieved from several Airborne Snow Observatory (ASO) and Airborne Visible/Infrared Imaging Spectrometer (AVIRIS) flights. Using these airborne retrievals as a reference, we show that fSCA can be retrieved at the hillslope scale with reasonable accuracy at an unprecedented near daily revisit period using a combination of the Landsat, Sentinel-2 MSI, RapidEye, and PlanetScope satellite constellations. In a series of data assimilation experiments we show how the combination of these constellations can lead to significant improvements in hillslope scale snow reanalyses as gauged by various evaluation metrics. Furthermore, it is suggested that an iterative ensemble smoother data assimilation scheme can provide more robust SWE estimates than other smoothers that have previously been proposed for snow reanalysis. We briefly conclude with thoughts as to the current impediments to conducting a global hillslope scale snow reanalysis and propose avenues for further research, such as how snow reanalyses can help in the prediction exercise.
The Surface Water Ocean Topography (SWOT) mission, launching next year, will provide high-spatial resolution measurements of terrestrial surface water, including global rivers with widths greater than 50-100 m. SWOT measurements are naturally suited for stream hydrology, and many previous studies have worked to quantify the impact of SWOT observations on the modeling of channel flow. This work highlights the application of SWOT for WRF-Hydro modeling in Alaska for data assimilation and model calibration to support ongoing National Oceanic and Atmospheric Administration (NOAA) National Water Model development. Results demonstrate the effectiveness of using SWOT discharge estimates to calibrate WRF-Hydro in ungauged basins, and quantifies the impact of SWOT data assimilation on WRF-Hydro performance.