Hydrologic science lacks a comprehensive theory of stormflow generation, preventing the development of a general hydrologic model. Studies show that models focusing on dominant local processes often outperform general models that rely on parameter tuning, leading to higher confidence solutions. For continental-scale hydrologic and hydraulic prediction, regional mosaics of models may outperform a single-model approach. However, variations in model inputs, programming languages, solvers, and discretizations hinder interoperability and comparisons. To address these challenges, we developed the Next Generation Water Resources Modeling Framework (NextGen): a model-agnostic, standards-based architecture for model interoperability and evaluation. Two standards enable the Framework: (1) the Basic Model Interface (BMI) version 2.0, for model control, coupling, and querying; and (2) the Open Geospatial Consortium WaterML 2.0 part 3 Hydrologic Features (HY_Features) conceptual data model to describe the "hydrofabric" of surface water hydrologic and hydraulic features. In the NextGen Framework, models retain their unique solution methods while becoming interoperable through BMI variable exchange tied to a common hydrofabric. The Framework enables scientific evaluation of water prediction models that simulate diverse hydrologic and hydraulic processes. Its design supports models written in multiple programming languages and runs on laptops, cloud and distributed memory supercomputers.
Existing process-based models for simulating coastal foredune evolution largely use the same analytical approach for estimating wind-induced surface shear stress distributions over spatially variable topography. Originally developed for smooth, low-sloping hills, these analytical models face significant limitations when the topography of interest exhibits large height-to-length ratios and/or steep, localized features. In this work, we utilize computational fluid dynamics (CFD) to examine the error trends of a commonly used analytical shear stress model for a series of idealized two-dimensional dune profiles. It is observed that the prediction error of the analytical model increases compared to the CFD simulations for increasing height-to-length ratio and localized slope values. Furthermore, we explore two data-driven methodologies for generating alternative shear stress prediction models, namely, symbolic regression and linear, projection-based, non-intrusive reduced-order modeling. These alternative modeling strategies demonstrate reduced overall error but still suffer in their generalizability to broader sets of dune profiles outside of the training data. Finally, the impact of these improvements on aeolian sediment transport fluxes is examined to demonstrate that even modest improvements to the shear stress prediction can have significant impacts on dune evolution simulations over engineering-relevant timescales.
Training object detection algorithms to operate in complex geo-environments remains a significant challenge, necessitating large and diverse datasets (i.e., unique backgrounds and conditions) that are not always readily available. Physically generating requisite data can also be both cost and time prohibitive depending on the object(s) and area(s) of interest -- especially in the case of multi-spectral and hyper-spectral imagery. Thus, there is increasing interest in the use of synthetic data to supplement existing physical datasets. To this end, the US Army Engineer Research and Development Center (ERDC) continues to develop a computational test-bed with a tool suite called the VESPA or, the Virtual Environmental Simulation for Physics-based Analysis, to support synthetic multi-spectral and hyper-spectral EO/IR imagery generation. The VESPA consists of integrated (1) scene generation tools, (2) multi-fidelity models for simulating heat and mass transfer and atmospheric energy propagation in geo-environments and climates worldwide that are optimized for high performance computing (3) data interrogation utilities, and (4) component-level sensor models capable of producing AI/ML ready near- and far-field imagery that is comparable to that produced by real sensors. This study presents an overview of the VESPA, new advances/capabilities, and results from a recent detailed validation and verification study.
Estimation of nearshore bathymetry is important for accurate prediction of nearshore wave conditions. However, direct data collection is expensive and time-consuming while accurate airborne lidar-based survey is limited by breaking waves and decreased light penetration affected by water turbidity. Instead, tower-based platforms or Unmanned Aircraft System (UAS) can provide indirect video-based observations. The video-based time-series imagery provides wave celerity information and time-averaged (timex) or variance enhanced (var) images identify persistent regions of wave breaking. In this work, we propose a rapid and improved bathymetry estimation method that takes advantage of image-derived wave celerity and a first-order bathymetry estimate from Parameter Beach Tool (PBT), software that fits parameterized sandbar and slope forms to the timex or var images. Two different sources of the data, PBT and wave celerity, are combined or blended optimally based on their assumed accuracy in a statistical framework. The PBT-derived bathymetry serves as "prior" coarse-scale background information and then is updated and corrected with the imagery-derived wave data through the dispersion relationship, which results in a better bathymetry estimate that is consistent with imagery-based wave data. To illustrate the accuracy of our proposed method, imagery data sets collected in 2017 at the US Army EDRC's Field Research Facility in Duck, NC under different weather and wave height conditions are tested. Estimated bathymetry profiles are remarkably close to the direct survey data. The computational time for the estimation from PBT-based bathymetry and imagery-derived wave celerity is only about five minutes on a free Google Cloud node with one CPU core. These promising results indicate the feasibility of reliable real-time bathymetry imaging during a single flight of UAS.
Numerous applications in biology, statistics, science, and engineering require generating samples from high-dimensional probability distributions. In recent years, the Hamiltonian Monte Carlo (HMC) method has emerged as a state-of-the-art Markov chain Monte Carlo technique, exploiting the shape of such high-dimensional target distributions to efficiently generate samples. Despite its impressive empirical success and increasing popularity, its wide-scale adoption remains limited due to the high computational cost of gradient calculation. Moreover, applying this method is impossible when the gradient of the posterior cannot be computed (for example, with black-box simulators). To overcome these challenges, we propose a novel two-stage Hamiltonian Monte Carlo algorithm with a surrogate model. In this multi-fidelity algorithm, the acceptance probability is computed in the first stage via a standard HMC proposal using an inexpensive differentiable surrogate model, and if the proposal is accepted, the posterior is evaluated in the second stage using the high-fidelity (HF) numerical solver. Splitting the standard HMC algorithm into these two stages allows for approximating the gradient of the posterior efficiently, while producing accurate posterior samples by using HF numerical solvers in the second stage. We demonstrate the effectiveness of this algorithm for a range of problems, including linear and nonlinear Bayesian inverse problems with in-silico data and experimental data. The proposed algorithm is shown to seamlessly integrate with various low-fidelity and HF models, priors, and datasets. Remarkably, our proposed method outperforms the traditional HMC algorithm in both computational and statistical efficiency by several orders of magnitude, all while retaining or improving the accuracy in computed posterior statistics.
The increasing deployment of AI in critical sectors necessitates advancements in explainable AI (XAI) to ensure transparency and trustworthiness of AI decisions. This paper introduces a novel methodology that leverages the Virtual Environmental Simulation for Physics-based Analysis (VESPA) framework in conjunction with Randomized Input Sampling for Explanation (RISE) to provide enhanced explainability for AI models, particularly in complex simulated environments. VESPA, known for its high-fidelity, physics-based simulations across diverse conditions, generates a vast dataset encompassing various sensor configurations, environmental factors, and material responses. This dataset serves as the foundation for applying RISE, a model-agnostic approach that generates pixel-level importance maps by probing the AI model with masked versions of the input images. Through this integration, we offer a systematic way to visualize and understand the influence of different environmental elements on AI decisions. Our approach not only sheds light on the "black box" of AI decision-making processes but also provides a scalable framework for evaluating AI models' robustness and reliability under a wide array of simulated scenarios.
Nearshore bathymetry changes on scales of hours to months in ways that strongly impact coastal processes. However, even at the best-monitored sites, surveys are typically not conducted with sufficient frequency to capture important changes such as sandbar migration. As a result, nearshore models often rely on outdated bathymetric boundary conditions, which may introduce significant errors. In this study, we investigate ensemble optimal interpolation (EnOI) as a method to update survey-derived bathymetry with altimetric measurements that are spatially sparse but have high temporal availability. We present the results of two synthetic examples and two field data experiments that demonstrate the ability of the method to accurately track morphological change between surveys. The method reduces the RMSE relative to a static bathymetry (corresponding to the day before the first assimilation step) by 23% to 68%. When compared with an estimate linearly interpolated between survey-derived bathymetries, the EnOI analysis reduces the RMSE by 19% to 47% in three out of the four experiments.
Autonomous vehicles (AVs) employ a wide range of sensing modalities including LiDAR, radar, RGB cameras, and more recently infrared (IR) sensors. IR sensors are becoming an increasingly common component of AVs’ sensor packages to provide redundancy and enhanced capabilities in conditions that are adverse for other types of sensors. For example, while RGB cameras are sensitive to lighting conditions and LiDAR performance is degraded in inclement weather such as rain, IR sensors are unaffected by lighting conditions and can contribute additional meaningful information in inclement weather. The US Army Corps of Engineers, Engineer Research and Development Center (ERDC) has developed the ERDC Computational Test Bed (CTB) to provide a suite of tools that can be used to support virtual development and testing of AVs. The CTB includes physics-based vehicle-terrain interaction, sensor and environment modeling, geo-environmental thermal modeling, software-inthe- loop capabilities, and virtual environment generation. Thermal modeling capabilities within the CTB utilize decades of near-surface phenomenology and autonomy research. Recent additions have been made to support large-domains commonly required for autonomous vehicle operations. These additions provide high-fidelity, physics-based thermal transfer and IR sensor models for creating high-quality synthetic imagery simulating IR sensors mounted on AVs. Highly parallelized thermal and IR sensor models for large-domain AV operations will be presented in this paper.
Process-Based Modeling (PBM) and Machine Learning (ML) are often perceived as distinct paradigms in the geosciences. Here we present differentiable geoscientific modeling as a powerful pathway toward dissolving the perceived barrier between them and ushering in a paradigm shift. For decades, PBM offered benefits in interpretability and physical consistency but struggled to efficiently leverage large datasets. ML methods, especially deep networks, presented strong predictive skills yet lacked the ability to answer specific scientific questions. While various methods have been proposed for ML-physics integration, an important underlying theme -- differentiable modeling -- is not sufficiently recognized. Here we outline the concepts, applicability, and significance of differentiable geoscientific modeling (DG). "Differentiable" refers to accurately and efficiently calculating gradients with respect to model variables, critically enabling the learning of high-dimensional unknown relationships. DG refers to a range of methods connecting varying amounts of prior knowledge to neural networks and training them together, capturing a different scope than physics-guided machine learning and emphasizing first principles. Preliminary evidence suggests DG offers better interpretability and causality than ML, improved generalizability and extrapolation capability, and strong potential for knowledge discovery, while approaching the performance of purely data-driven ML. DG models require less training data while scaling favorably in performance and efficiency with increasing amounts of data. With DG, geoscientists may be better able to frame and investigate questions, test hypotheses, and discover unrecognized linkages.
The goal of this study is to leverage emerging machine learning (ML) techniques to develop a framework for the global reconstruction of system variables from potentially scarce and noisy observations and to explore the epistemic uncertainty of these models. This work demonstrates the utility of exploiting the stochasticity of dropout and batch normalization schemes to infer uncertainty estimates of super-resolved field reconstruction from sparse sensor measurements. A Voronoi tessellation strategy is used to obtain a structured-grid representation from sensor observations, thus enabling the use of fully convolutional neural networks (FCNN) for global field estimation. An ensemble-based approach is developed using Monte-Carlo batch normalization (MCBN) and Monte-Carlo dropout (MCD) methods in order to perform approximate Bayesian inference over the neural network parameters, which facilitates the estimation of the epistemic uncertainty of predicted field values. We demonstrate these capabilities through numerical experiments that include sea-surface temperature, soil moisture, and incompressible near-surface flows over a wide range of parameterized flow configurations.
Wave runup observations are key data for understanding coastal response to storms. Lidar scanners are capable of collecting swash elevation data at high spatial and temporal resolution in a range of environmental conditions. Efforts to develop automated algorithms that effectively separate returns off of the beach and the sea surface are complicated by environmental noise, thus requiring time-intensive data quality control or manual digitization. In this study, a fully convolutional neural network (FCNN) was trained and validated on 966 30-min lidar linescan time series of the beach and swash zone and tested on an additional 99 30-min linescan time series to improve the automated classification of lidar returns off of the beach and water, facilitating the extraction of a depth-defined wave runup time series. Lidar returns were classified as beach or water, and beach points at each cross-shore location were interpolated through time, creating a time-varying beach elevation surface that was used to calculate instantaneous swash depths. Runup was defined as the most landward position of 3-cm water depth through time. Overall, the runup time series determined using the manually and machine learning (ML)-digitized beach-water interface agreed well (0.02-m root-mean-square difference (RMSD) 3-cm contour location and 0.06-m RMSD 2% runup exceedance elevation R2%). The trained model was found to be robust to noise and moderate data gaps and applicable in a range of wave conditions. Results demonstrate the potential of the ML model to replace manual data processing steps and significantly reduce the time and effort required to extract the instantaneous runup location from lidar linescan time series.
Abstract Groundwater modeling is widely relied upon by environmental scientists and engineers to advanced understanding, make predictions, and design solutions to water resource problems of importance to society. Groundwater models are tools used to approximate subsurface behavior, including the movement of water, the chemical composition of the phases present, and the temperature distribution. As a model is a simplification of a real-world system, approximations and uncertainties are inherent to the modeling process. Due to this, special consideration must be given to the role of uncertainty quantification, as essentially all groundwater systems are stochastic in nature.
Immediate estimation of nearshore bathymetry is crucial for accurate prediction of nearshore wave conditions and coastal flooding events. However, direct bathymetry data collection is expensive and time-consuming, while accurate airborne lidar-based survey is limited by breaking waves and decreased light penetration affected by water turbidity. Several recent efforts have been made to apply interpolation and inverse modeling approaches to indirect remote sensed observations along with sparse direct survey data points. Example indirect observations include video-based observations such as time-series snapshots and time-averaged (Timex) images across the surf zone taken from tower-based platforms and Unmanned Aircraft Systems (UASs), while stationary LiDAR tower and UAS flights with infrared camera capability or imagery-based structure-from-motion (SfM) algorithms have been used to provide beach topographic data. In this work, we present three bathymetry estimation tools for real-time nearshore characterization using different types of information.
Our world has been facing a number of serious challenges in the past few decades, such as an increased rate of energy and water consumption, climate change, and natural hazards, that have led to complex human-resource dynamics, requiring hydrologists and geoscientists to consider interdisciplinary approaches. The multi-scale and multi-physics nature of these problems have created high-dimensional equations, or processes whose governing equations are not known accurately, making traditional approaches that are based on modeling and solving the physical processes directly, computationally very demanding or inaccurate. Furthermore, many processes in hydrology and geoscience involve solving inverse problems, due to parameters that are either difficult or impossible to measure directly. Inverse problems typically require running of forward solvers many times, aggravating already existing computational issues. This chapter discusses how data-driven approaches can be used to address these challenges. In particular, reduced order modeling (ROM) can be used to capture the most prominent features present in high-dimensional data efficiently, creating fast forward and inverse solvers of physical processes. Further, computational speed-up and increase in the accuracy is achievable by combining ROMs with deep learning techniques
Short-term probabilistic forecasts of the trajectory of the COVID-19 pandemic in the United States have served as a visible and important communication channel between the scientific modeling community and both the general public and decision-makers. Forecasting models provide specific, quantitative, and evaluable predictions that inform short-term decisions such as healthcare staffing needs, school closures, and allocation of medical supplies. Starting in April 2020, the US COVID-19 Forecast Hub ( https://covid19forecasthub.org/ ) collected, disseminated, and synthesized tens of millions of specific predictions from more than 90 different academic, industry, and independent research groups. A multimodel ensemble forecast that combined predictions from dozens of groups every week provided the most consistently accurate probabilistic forecasts of incident deaths due to COVID-19 at the state and national level from April 2020 through October 2021. The performance of 27 individual models that submitted complete forecasts of COVID-19 deaths consistently throughout this year showed high variability in forecast skill across time, geospatial units, and forecast horizons. Two-thirds of the models evaluated showed better accuracy than a naïve baseline model. Forecast accuracy degraded as models made predictions further into the future, with probabilistic error at a 20-wk horizon three to five times larger than when predicting at a 1-wk horizon. This project underscores the role that collaboration and active coordination between governmental public-health agencies, academic modeling teams, and industry partners can play in developing modern modeling capabilities to support local, state, and federal response to outbreaks.
Physical systems governed by advection-dominated partial differential equations (PDEs) are found in applications ranging from engineering design to weather forecasting. They are known to pose severe challenges to both projection-based and non-intrusive reduced order modeling, especially when linear subspace approximations are used. In this work, we develop an advection-aware (AA) autoencoder network that can address some of these limitations by learning efficient, physics-informed, nonlinear embeddings of the high-fidelity system snapshots. A fully non-intrusive reduced order model is developed by mapping the high-fidelity snapshots to a latent space defined by an AA autoencoder, followed by learning the latent space dynamics using a long-short-term memory (LSTM) network. This framework is also extended to parametric problems by explicitly incorporating parameter information into both the high-fidelity snapshots and the encoded latent space. Numerical results obtained with parametric linear and nonlinear advection problems indicate that the proposed framework can reproduce the dominant flow features even for unseen parameter values.
The United State Army Corp of Engineers (USACE) Engineering Research and Development Center (ERDC) has developed a suite of computational tools called the Computational Test Bed (CTB) for advanced high-fidelity physics-based autonomous vehicle sensor and environment simulations. These tools provide insights into onboard navigation, image processing, sensor fusion techniques, and rapid data generation for artificial intelligence and machine learning techniques across the full spectrum (visible, NIR, MWIR, and LWIR) and for various sensor modalities (LiDAR, EO, radar). This paper presents ERDC's CTB that allows the community to design, develop, test, and evaluate the entire autonomy space from machine learning algorithm development using augmented synthetic data to large-scale autonomous system testing.
Estimation of riverbed profiles, also known as bathymetry, plays a vital role in many applications, such as safe and efficient inland navigation, prediction of bank erosion, land subsidence, and flood risk management. The high cost and complex logistics of direct bathymetry surveys, i.e, depth imaging, have encouraged the use of indirect measurements such as surface flow velocities. However, estimating high-resolution bathymetry from indirect measurements is an inverse problem that can be computationally challenging. Here, we propose a reduced-order model (ROM) based approach that utilizes a variational autoencoder (VAE), a type of deep neural network with a narrow layer in the middle, to compress bathymetry and flow velocity information and accelerate bathymetry inverse problems from flow velocity measurements. In our application, the shallow -water equations (SWE) with appropriate boundary conditions (BCs), e.g., the discharge and/or the free surface elevation, constitute the forward problem, to predict flow velocity. Then, ROMs of the SWEs are constructed on a nonlinear manifold of low dimensionality through a variational encoder and the bathymetry inversion problem is derived on the low-dimensional latent space in a Hierarchical Bayesian setting. The reformulation allows variational inference with a small number (e.g., O(100) of ROM runs and efficient uncertainty quantification. We have tested our inversion approach on a one-mile reach of the Savannah River, GA, USA. Once the neural network is trained (offline stage), the proposed technique can perform the inversion operation orders of magnitude faster than traditional inversion methods that are commonly based on linear projections, such as principal component analysis (PCA), or the principal component geostatistical approach (PCGA). Furthermore, tests show that the algorithm can estimate the bathymetry with good accuracy even with sparse flow velocity measurements.
Timely observations of nearshore water depths are important for a variety of coastal research and management topics, yet this information is expensive to collect using in situ survey methods. Remote methods to estimate bathymetry from imagery include using either ratios of multi-spectral reflectance bands or inversions from wave processes. Multi-spectral methods work best in waters with low turbidity, and wave-speed-based methods work best when wave breaking is minimal. In this work, we build on the wave-based inversion approaches, by exploring the use of a fully convolutional neural network (FCNN) to infer nearshore bathymetry from imagery of the sea surface and local wave statistics. We apply transfer learning to adapt a CNN originally trained on synthetic imagery generated from a Boussinesq numerical wave model to utilize tower-based imagery collected in Duck, North Carolina, at the U.S. Army Engineer Research and Development Center’s Field Research Facility. We train the model on sea-surface imagery, wave conditions, and associated surveyed bathymetry using three years of observations, including times with significant wave breaking in the surf zone. This is the first time, to the authors’ knowledge, an FCNN has been successfully applied to infer bathymetry from surf-zone sea-surface imagery. Model results from a separate one-year test period generally show good agreement with survey-derived bathymetry (0.37 m root-mean-squared error, with a max depth of 6.7 m) under diverse wave conditions with wave heights up to 3.5 m. Bathymetry results quantify nearshore bathymetric evolution including bar migration and transitions between single- and double-barred morphologies. We observe that bathymetry estimates are most accurate when time-averaged input images feature visible wave breaking and/or individual images display wave crests. An investigation of activation maps, which show neuron activity on a layer-by-layer basis, suggests that the model is responsive to visible coherent wave structures in the input images.