Estimating constituent loads in streams and rivers is a crucial but challenging task due to low-frequency sampling in most watersheds. While predictive modeling can augment sparsely sampled water quality data, it can be challenging due to the complex and multifaceted interactions between several sub-watershed eco-hydrological processes. Traditional water quality prediction models, typically calibrated for individual sites, struggle to fully capture these interactions. This study introduces XGBest, a machine learning-based tool, that integrates hydrological data, land cover, and physical watershed attributes at a regional scale to predict daily concentrations of Total Nitrogen (TN), Total Phosphorus (TP), and Total Suspended Solids (TSS). XGBest leverages 29 environmental variables, including daily and antecedent discharge, temporal features, and landscape characteristics, to comprehensively evaluate water quality dynamics across a large hydrologic region. To explore the robustness of the developed tool, XGBest was validated using observed water quality data in three different hydrologic regions in the eastern United States, encompassing 499 water quality monitoring sites characterized by diverse hydro-climatic conditions and watershed attributes. This study also employed the legacy United States Geological Survey (USGS) tools - LOADEST and WRTDS as benchmarks to evaluate the performance of XGBest in these regions.The results demonstrated that XGBest outperformed LOADEST and WRTDS in predictive accuracy and revealed critical insights into the spatial and temporal variability of nutrient and sediment loads. In addition, SHapley Additive exPlanations (SHAP) values highlighted the importance of integrating static and dynamic watershed attributes, such as land cover, antecedent discharge, and seasonality, in capturing the complex concentration-discharge (C-Q) relationships. This study positions XGBest as a robust and scalable water quality prediction tool that bridges the gap between hydrology and broader environmental management. By combining multiple environmental factors into a unified predictive framework, XGBest enhances our understanding of water quality and supports more effective environmental monitoring and management strategies.
This study enhances hydrological modeling in ungauged watersheds by employing physical similarity and machine learning-based clustering for regionalizing the Soil and Water Assessment Tool (SWAT) model parameters at the HUC12 (hydrological unit code) watershed scale within a HUC02 basin. Eleven features, including environmental, topographical, soil, and hydrological properties, were utilized to identify physical similarities for watershed clustering. Machine learning techniques, including random forest and hierarchical clustering, were employed to transfer calibrated parameters from gauged to ungauged watersheds. Validation of parameter transfer over gauged SWAT model projects showed that 88% of the projects achieved calibrated status (KGE ≥ 0.5; PBIAS ≤ 25%). Additional validation using MODIS satellite evapotranspiration measurements confirmed the robustness of the approach. Results indicated that the proposed approach successfully captures physical similarities, and effectively captures flow patterns. Overall, the study highlights the potential of physical similarity-based clustering and machine learning techniques for improving hydrological modeling in ungauged watersheds.
In environmental data analysis, source apportionment can be an important approach to extract useful information that might otherwise be hidden within the data. The United States Environmental Protection Agency (EPA) has developed an open-source python package, the Environmental Source Apportionment Toolkit (ESAT), which enables source apportionment modeling and error estimation workflows. ESAT is intended to replace Positive Matrix Factorization v5 (PMF5) that has substantial data size limitations. ESAT is currently in alpha testing with development plans for enhanced functionality and support of large datasets, High-performance Computing (HPC) execution through a command line interface (CLI), and a standalone desktop graphical user interface (GUI). The alpha product of ESAT is publicly available and offers a complete application programming interface (API) to replicate the workflows and functionality of PMF5, with examples provided through Jupyter Notebooks. The ESAT computing module currently contains two non-negative matrix factorization (NMF) algorithms for model training, with the module designed for other algorithms to be easily added. The two algorithms currently available are the least-squares NMF (LS-NMF) and weighted-semi NMF (WS-NMF). Each algorithm offers different benefits depending on project or data requirements. The ESAT python codebase has been optimized to run in a highly parallelized manner, with most of the numerical computations implemented in Rust, a low-level language comparable in performance to C. ESAT replicates the model error estimation methods of PMF5, namely bootstrap, displacement, and a hybrid method. To facilitate experimentation and testing, ESAT contains a synthetic dataset generator and model simulator that can evaluate how well ESAT can recreate synthetic factors and contributions. Continuous development of new features are tested and added to the python package on a regular basis. One such feature is the addition of an uncertainty perturbation workflow, which will run a collection of models while slightly perturbing the uncertainty matrix, and then evaluating the impact on the solution profiles and contributions. The alpha version of the ESAT python package is available for installation from pypi at https://pypi.org/project/esat/. Further testing and development of the alpha version will proceed to a full release in late 2025. The development of a GUI desktop application is currently planned to begin after the ESAT full release.
Subalpine lakes are valuable resources that are at increasing environmental risk. Aquatic ecosystem models are useful tools for understanding dynamics of lakes, however, there are few examples of these models being applied to subalpine lakes, which may differ from temperate lakes in dimensions such as physical setting and ecology. Here we apply the aquatic ecosystem model AQUATOX to the Loch, a well-studied lake in Rocky Mountain National Park in Colorado, USA, to assess the applicability of this model to a subalpine lake setting, and to identify modeling gaps. We found that AQUATOX could represent phytoplankton dynamics during the ice-off period. Attempting to calibrate the model during the ice-on period using the same structure and parameters as the ice-off period underestimated winter chlorophyll a concentrations. Additionally, the model was used to simulate a nutrient bioassay experiment - nitrate was well simulated; P was overestimated but consistent with the observed pattern. These results support the use of current models for subalpine lakes when data are sufficient. Areas identified for future model development include better models for boundary conditions, improved light data, modeling of mixotrophy, and better representation of ice formation and under-ice stratification.
This forecasting approach may be useful for water managers and associated public health managers to predict near-term future high-risk cyanobacterial harmful algal blooms (cyanoHAB) occurrence. Freshwater cyanoHABs may grow to excessive concentrations and cause human, animal, and environmental health concerns in lakes and reservoirs. Knowledge of the timing and location of cyanoHAB events is important for water quality management of recreational and drinking water systems. No quantitative tool exists to forecast cyanoHABs across broad geographic scales and at regular intervals. Publicly available satellite monitoring has proven effective in detecting cyanobacteria biomass near-real time within the United States. Weekly cyanobacteria abundance was quantified from the Ocean and Land Colour Instrument (OLCI) onboard the Sentinel-3 satellite as the response variable. An Integrated Nested Laplace Approximation (INLA) hierarchical Bayesian spatiotemporal model was applied to forecast World Health Organization (WHO) recreation Alert Level 1 exceedance >12 mu g L-1 chlorophyll-a with cyanobacteria dominance for 2192 satellite resolved lakes in the United States across nine climate zones. The INLA model was compared against support vector classifier and random forest machine learning models; and Dense Neural Network, Long Short-Term Memory (LSTM), Recurrent Neural Network (RNN), and Gneural Network (GNU) neural network models. Predictors were limited to data sources relevant to cyanobacterial growth, readily available on a weekly basis, and at the national scale for operational forecasting. Relevant predictors included water surface temperature, precipitation, and lake geomorphology. Overall, the INLA model outperformed the machine learning and neural network models with prediction accuracy of 90% with 88% sensitivity, 91% specificity, and 49% precision as demonstrated by training the model with data from 2017 through 2020 and independently assessing predictions with data from the 2021 calendar year. The probability of true positive responses was greater than false positive responses and the probability of true negative responses was less than false negative responses. This indicated the model correctly assigned lower probabilities of events when they didn't exceed the WHO Alert Level 1 threshold and assigned higher probabilities when events did exceed the threshold. The INLA model was robust to missing data and unbalanced sampling between waterbodies.
The first phase of a national scale Soil and Water Assessment Tool (SWAT) model calibration effort at the HUC12 (Hydrologic Unit Code 12) watershed scale was demonstrated over the Mid-Atlantic Region (R02), consisting of 3036 HUC12 subbasins. An R-programming based tool was developed for streamflow calibration including parallel processing for SWAT-CUP (SWAT- Calibration and Uncertainty Programs) to streamline the computational burden of calibration. Successful calibration of streamflow for 415 gages (KGE ≥0.5, Kling-Gupta efficiency; PBIAS ≤15%, Percent Bias) out of 553 selected monitoring gages was achieved in this study, yielding calibration parameter values for 2106 HUC12 subbasins. Additionally, 67 more gages were calibrated with relaxed PBIAS criteria of 25%, yielding calibration parameter values for an additional 150 HUC12 subbasins. This first phase of calibration across R02 increases the reliability, uniformity, and replicability of SWAT-related hydrological studies. Moreover, the study presents a comprehensive approach for efficiently optimizing large-scale multi-site calibration.
Reservoirs are dominant features of the modern hydrologic landscape and provide vital services. However, the unique morphology of reservoirs can create suitable conditions for excessive algae growth and associated cyanobacteria blooms in shallow in-flow reservoir locations by providing warm water environments with relatively high nutrient inputs, deposition, and nutrient storage. Cyanobacteria harmful algal blooms (cyanoHAB) are costly water management issues and bloom recurrence is associated with economic costs and negative impacts to human, animal, and environmental health. As cyanoHAB occurrence varies substantially within different regions of a water body, understanding in-lake cyanoHAB spatial dynamics is essential to guide reservoir monitoring and mitigate potential public exposure to cyanotoxins. Cloud-based computational processing power and high temporal frequency of satellites enables advanced pixel-based spatial analysis of cyanoHAB frequency and quantitative assessment of reservoir headwater in-flows compared to near-dam surface waters of individual reservoirs. Additionally, extensive spatial coverage of satellite imagery allows for evaluation of spatial trends across many dozens of reservoir sites. Surface water cyanobacteria concentrations for sixty reservoirs in the southern U.S. were estimated using 300 m resolution European Space Agency (ESA) Ocean and Land Colour Instrument (OLCI) satellite sensor for a five year period (May 2016-April 2021). Of the reservoirs studied, spatial analysis of OLCI data revealed 98% had more frequent cyanoHAB occurrence above the concentration of >100,000 cells/mL in headwaters compared to near-dam surface waters (P < 0.001). Headwaters exhibited greater seasonal variability with more frequent and higher magnitude cyanoHABs occurring mid-summer to fall. Examination of reservoirs identified extremely high concentration cyanobacteria events (>1,000,000 cells/mL) occurring in 70% of headwater locations while only 30% of near-dam locations exceeded this threshold. Wilcoxon signed-rank tests of cyanoHAB magnitudes using paired-observations (dates with observations in both a reservoir's headwater and near-dam locations) confirmed significantly higher concentrations in headwater versus near-dam locations (p < 0.001).
We developed statistical models to generate runoff time-series at National Hydrography Dataset Plus Version 2 (NHDPlusV2) catchment scale for the Continental United States (CONUS). The models use Normalized Difference Vegetation Index (NDVI) based Curve Number (CN) to generate initial runoff time-series which then is corrected using statistical models to improve accuracy. We used the North American Land Data Assimilation System 2 (NLDAS-2) catchment scale runoff time-series as the reference data for model training and validation. We used 17 years of 16-day, 250-m resolution NDVI data as a proxy for hydrologic conditions during a representative year to calculate 23 NDVI based-CN (NDVI-CN) values for each of 2.65 million NHDPlusV2 catchments for the Contiguous U.S. To maximize predictive accuracy while avoiding optimistically biased model validation results, we developed a spatio-temporal cross-validation framework for estimating, selecting, and validating the statistical correction models. We found that in many of the physiographic sections comprising CONUS, even simple linear regression models were highly effective at correcting NDVI-CN runoff to achieve Nash-Sutcliffe Efficiency values above 0.5. However, all models showed poor performance in physiographic sections that experience significant snow accumulation.
The Piscine Stream Community Estimation System (PiSCES) provides users with a hypothesized fish community for any stream reach in the conterminous United States using information obtained from Nature Serve, the US Geological Survey (USGS), StreamCat, and the Peterson Field Guide to Freshwater Fishes of North America for over 1000 native and non-native freshwater fish species. PiSCES can filter HUC8-based fish assemblages based on species-specific occurrence models; create a community abundance/biomass distribution by relating relative abundance to mean body weight of each species; and allow users to query its database to see ancillary characteristics of each species (e.g., habitat preferences and maximum size). Future efforts will aim to improve the accuracy of the species distribution database and refine/augment increase the occurrence models. The PiSCES tool is accessible at the EPA's Quantitative Environmental Domain (QED) website at https://qed.epacdx.net/pisces/
Gridded precipitation datasets are becoming a convenient substitute for gauge measurements in hydrological modeling; however, these data have not been fully evaluated across a range of conditions. We compared four gridded datasets (Daily Surface Weather and Climatological Summaries [DAYMET], North American Land Data Assimilation System [NLDAS], Global Land Data Assimilation System [GLDAS], and Parameter-elevation Regressions on Independent Slopes Model [PRISM]) as precipitation data sources and evaluated how they affected hydrologic model performance when compared with a gauged dataset, Global Historical Climatology Network-Daily (GHCN-D). Analyses were performed for the Delaware Watershed at Perry Lake in eastern Kansas. Precipitation indices for DAYMET and PRISM precipitation closely matched GHCN-D, whereas NLDAS and GLDAS showed weaker correlations. We also used these precipitation data as input to the Soil and Water Assessment Tool (SWAT) model that confirmed similar trends in streamflow simulation. For stations with complete data, GHCN-D based SWAT-simulated streamflow variability better than gridded precipitation data. During low flow periods we found PRISM performed better, whereas both DAYMET and NLDAS performed better in high flow years. Our results demonstrate that combining gridded precipitation sources with gauge-based measurements can improve hydrologic model performance, especially for extreme events.
There is always a challenge of obtaining the “best” data to inform environmental models. Here we present different types of available precipitation datasets while detailing temporal and spatial resolution, potential errors in the dataset, and optimal performance scenarios. Our goal is to inform modelers of the various types, resolutions, and sources of precipitation data available for environmental modeling. Precipitation is the main driver in the hydrological cycle and modelers use this information to understand water quality and water availability. Environmental models use observed precipitation information for modeling past or current conditions, while simulated data are used to predict future conditions as well as recreate historic conditions. Several precipitation datasets and data generation methods such as National Climatic Data Center (NCDC) rain gauges, National and Global Land Data Assimilation (NLDAS, GLDAS), Next Generation Weather Radar, and Stochastic Weather Generators are described, giving their strengths and weaknesses. USEPA’s Hydrologic Micro Services (HMS) project has developed a collection of interoperable water quantity and quality modeling components that leverage existing internet-based data sources and sensors via a web service. The precipitation component of USEPA’s Hydrologic Micro Service (HMS) will provide the information and data from multiple sources through the web service for modeling purposes.
United States Environmental Protection Agency (EPA) has developed a collection of microservices called Hydrologic Micro Services (HMS) for building hydrologic and water quality modeling workflows. HMS components are available as RESTful web services as well as desktop libraries. An HMS component may have multiple implementations addressing varying levels of underlying physical process details and assumptions. HMS components can be used in desktop and web-based workflows. A workflow can call into a specific implementation of an HMS component depending upon the details suitable for the problem statement being addressed by the workflow. Building a workflow from HMS components enables modelers to address hydrologic and water quality problem statements more precisely, in contrast to the current state of modeling where using existing models forces modelers into a potentially sub-optimal workflow. Model selection to address a problem statement has several drawbacks: the selected model may not have the appropriate level of complexity, the model may not address all parts of the problem statement without making less desirable assumptions, or the model may have more features and requirements than necessary. HMS components include data provisioning and simulation algorithms for water quantity and quality modeling. Workflows built using HMS components can in turn be used as components in larger workflows. For example, precipitation data provisioning components can download data from various data sources such as NLDAS, GLDAS, DAYMET, NCDC, PRISM, and WGEN. A simple workflow was developed as an HMS component to compare precipitation data from different sources. Comparison is performed using multiple rainfall statistics.