Estimating constituent loads in streams and rivers is a crucial but challenging task due to low-frequency sampling in most watersheds. While predictive modeling can augment sparsely sampled water quality data, it can be challenging due to the complex and multifaceted interactions between several sub-watershed eco-hydrological processes. Traditional water quality prediction models, typically calibrated for individual sites, struggle to fully capture these interactions. This study introduces XGBest, a machine learning-based tool, that integrates hydrological data, land cover, and physical watershed attributes at a regional scale to predict daily concentrations of Total Nitrogen (TN), Total Phosphorus (TP), and Total Suspended Solids (TSS). XGBest leverages 29 environmental variables, including daily and antecedent discharge, temporal features, and landscape characteristics, to comprehensively evaluate water quality dynamics across a large hydrologic region. To explore the robustness of the developed tool, XGBest was validated using observed water quality data in three different hydrologic regions in the eastern United States, encompassing 499 water quality monitoring sites characterized by diverse hydro-climatic conditions and watershed attributes. This study also employed the legacy United States Geological Survey (USGS) tools - LOADEST and WRTDS as benchmarks to evaluate the performance of XGBest in these regions.The results demonstrated that XGBest outperformed LOADEST and WRTDS in predictive accuracy and revealed critical insights into the spatial and temporal variability of nutrient and sediment loads. In addition, SHapley Additive exPlanations (SHAP) values highlighted the importance of integrating static and dynamic watershed attributes, such as land cover, antecedent discharge, and seasonality, in capturing the complex concentration-discharge (C-Q) relationships. This study positions XGBest as a robust and scalable water quality prediction tool that bridges the gap between hydrology and broader environmental management. By combining multiple environmental factors into a unified predictive framework, XGBest enhances our understanding of water quality and supports more effective environmental monitoring and management strategies.
We present a methodology for creating a roof rainwater harvesting (RWH) contaminant sensing-recording-grading (SRG) system comprising hardware and software components like low-cost sensors, a solar-powered data logger, a publicly available Arduino Integrated Development Environment (IDE) software, and smart mobile and web applications for data retrieval. We illustrate the prototype SRG system designed for monitoring basic roof-RWH quality parameters (i.e., electrical conductivity (μS/cm), temperature (°C), and depth (mm)) with a temporal frequency of 15 min from February to August 2024 in a rain barrel receiving rooftop runoff from a U.S. Environmental Protection Agency (EPA) building located in Georgia, USA. We established data validation protocols and verified the performance of the sensors by using an alternative set of sensors. We performed minimal data filling and comparable data cleaning for intermittent data gaps, which were partly attributed to extreme weather conditions or hardware or software issues. Analysis of the cleaned data set showed a robust performance of the tested sensors comparable to the validation sensor, with strong Pearson correlation coefficients between the two sensors' conductivity (0.99) and temperature measurements (1.00) and similar data spreads and mean values. The clean data analysis also showed that the RWH conductivity ranged from 7 to 116 μS/cm, the temperature ranged from 5 to 29 °C, and the depth ranged from 29 to 838 mm, from February to August 2024.
This study enhances hydrological modeling in ungauged watersheds by employing physical similarity and machine learning-based clustering for regionalizing the Soil and Water Assessment Tool (SWAT) model parameters at the HUC12 (hydrological unit code) watershed scale within a HUC02 basin. Eleven features, including environmental, topographical, soil, and hydrological properties, were utilized to identify physical similarities for watershed clustering. Machine learning techniques, including random forest and hierarchical clustering, were employed to transfer calibrated parameters from gauged to ungauged watersheds. Validation of parameter transfer over gauged SWAT model projects showed that 88% of the projects achieved calibrated status (KGE ≥ 0.5; PBIAS ≤ 25%). Additional validation using MODIS satellite evapotranspiration measurements confirmed the robustness of the approach. Results indicated that the proposed approach successfully captures physical similarities, and effectively captures flow patterns. Overall, the study highlights the potential of physical similarity-based clustering and machine learning techniques for improving hydrological modeling in ungauged watersheds.
In environmental data analysis, source apportionment can be an important approach to extract useful information that might otherwise be hidden within the data. The United States Environmental Protection Agency (EPA) has developed an open-source python package, the Environmental Source Apportionment Toolkit (ESAT), which enables source apportionment modeling and error estimation workflows. ESAT is intended to replace Positive Matrix Factorization v5 (PMF5) that has substantial data size limitations. ESAT is currently in alpha testing with development plans for enhanced functionality and support of large datasets, High-performance Computing (HPC) execution through a command line interface (CLI), and a standalone desktop graphical user interface (GUI). The alpha product of ESAT is publicly available and offers a complete application programming interface (API) to replicate the workflows and functionality of PMF5, with examples provided through Jupyter Notebooks. The ESAT computing module currently contains two non-negative matrix factorization (NMF) algorithms for model training, with the module designed for other algorithms to be easily added. The two algorithms currently available are the least-squares NMF (LS-NMF) and weighted-semi NMF (WS-NMF). Each algorithm offers different benefits depending on project or data requirements. The ESAT python codebase has been optimized to run in a highly parallelized manner, with most of the numerical computations implemented in Rust, a low-level language comparable in performance to C. ESAT replicates the model error estimation methods of PMF5, namely bootstrap, displacement, and a hybrid method. To facilitate experimentation and testing, ESAT contains a synthetic dataset generator and model simulator that can evaluate how well ESAT can recreate synthetic factors and contributions. Continuous development of new features are tested and added to the python package on a regular basis. One such feature is the addition of an uncertainty perturbation workflow, which will run a collection of models while slightly perturbing the uncertainty matrix, and then evaluating the impact on the solution profiles and contributions. The alpha version of the ESAT python package is available for installation from pypi at https://pypi.org/project/esat/. Further testing and development of the alpha version will proceed to a full release in late 2025. The development of a GUI desktop application is currently planned to begin after the ESAT full release.
Subalpine lakes are valuable resources that are at increasing environmental risk. Aquatic ecosystem models are useful tools for understanding dynamics of lakes, however, there are few examples of these models being applied to subalpine lakes, which may differ from temperate lakes in dimensions such as physical setting and ecology. Here we apply the aquatic ecosystem model AQUATOX to the Loch, a well-studied lake in Rocky Mountain National Park in Colorado, USA, to assess the applicability of this model to a subalpine lake setting, and to identify modeling gaps. We found that AQUATOX could represent phytoplankton dynamics during the ice-off period. Attempting to calibrate the model during the ice-on period using the same structure and parameters as the ice-off period underestimated winter chlorophyll a concentrations. Additionally, the model was used to simulate a nutrient bioassay experiment - nitrate was well simulated; P was overestimated but consistent with the observed pattern. These results support the use of current models for subalpine lakes when data are sufficient. Areas identified for future model development include better models for boundary conditions, improved light data, modeling of mixotrophy, and better representation of ice formation and under-ice stratification.
The first phase of a national scale Soil and Water Assessment Tool (SWAT) model calibration effort at the HUC12 (Hydrologic Unit Code 12) watershed scale was demonstrated over the Mid-Atlantic Region (R02), consisting of 3036 HUC12 subbasins. An R-programming based tool was developed for streamflow calibration including parallel processing for SWAT-CUP (SWAT- Calibration and Uncertainty Programs) to streamline the computational burden of calibration. Successful calibration of streamflow for 415 gages (KGE ≥0.5, Kling-Gupta efficiency; PBIAS ≤15%, Percent Bias) out of 553 selected monitoring gages was achieved in this study, yielding calibration parameter values for 2106 HUC12 subbasins. Additionally, 67 more gages were calibrated with relaxed PBIAS criteria of 25%, yielding calibration parameter values for an additional 150 HUC12 subbasins. This first phase of calibration across R02 increases the reliability, uniformity, and replicability of SWAT-related hydrological studies. Moreover, the study presents a comprehensive approach for efficiently optimizing large-scale multi-site calibration.
Reservoirs are dominant features of the modern hydrologic landscape and provide vital services. However, the unique morphology of reservoirs can create suitable conditions for excessive algae growth and associated cyanobacteria blooms in shallow in-flow reservoir locations by providing warm water environments with relatively high nutrient inputs, deposition, and nutrient storage. Cyanobacteria harmful algal blooms (cyanoHAB) are costly water management issues and bloom recurrence is associated with economic costs and negative impacts to human, animal, and environmental health. As cyanoHAB occurrence varies substantially within different regions of a water body, understanding in-lake cyanoHAB spatial dynamics is essential to guide reservoir monitoring and mitigate potential public exposure to cyanotoxins. Cloud-based computational processing power and high temporal frequency of satellites enables advanced pixel-based spatial analysis of cyanoHAB frequency and quantitative assessment of reservoir headwater in-flows compared to near-dam surface waters of individual reservoirs. Additionally, extensive spatial coverage of satellite imagery allows for evaluation of spatial trends across many dozens of reservoir sites. Surface water cyanobacteria concentrations for sixty reservoirs in the southern U.S. were estimated using 300 m resolution European Space Agency (ESA) Ocean and Land Colour Instrument (OLCI) satellite sensor for a five year period (May 2016-April 2021). Of the reservoirs studied, spatial analysis of OLCI data revealed 98% had more frequent cyanoHAB occurrence above the concentration of >100,000 cells/mL in headwaters compared to near-dam surface waters (P < 0.001). Headwaters exhibited greater seasonal variability with more frequent and higher magnitude cyanoHABs occurring mid-summer to fall. Examination of reservoirs identified extremely high concentration cyanobacteria events (>1,000,000 cells/mL) occurring in 70% of headwater locations while only 30% of near-dam locations exceeded this threshold. Wilcoxon signed-rank tests of cyanoHAB magnitudes using paired-observations (dates with observations in both a reservoir's headwater and near-dam locations) confirmed significantly higher concentrations in headwater versus near-dam locations (p < 0.001).
We developed statistical models to generate runoff time-series at National Hydrography Dataset Plus Version 2 (NHDPlusV2) catchment scale for the Continental United States (CONUS). The models use Normalized Difference Vegetation Index (NDVI) based Curve Number (CN) to generate initial runoff time-series which then is corrected using statistical models to improve accuracy. We used the North American Land Data Assimilation System 2 (NLDAS-2) catchment scale runoff time-series as the reference data for model training and validation. We used 17 years of 16-day, 250-m resolution NDVI data as a proxy for hydrologic conditions during a representative year to calculate 23 NDVI based-CN (NDVI-CN) values for each of 2.65 million NHDPlusV2 catchments for the Contiguous U.S. To maximize predictive accuracy while avoiding optimistically biased model validation results, we developed a spatio-temporal cross-validation framework for estimating, selecting, and validating the statistical correction models. We found that in many of the physiographic sections comprising CONUS, even simple linear regression models were highly effective at correcting NDVI-CN runoff to achieve Nash-Sutcliffe Efficiency values above 0.5. However, all models showed poor performance in physiographic sections that experience significant snow accumulation.
Input data acquisition and preprocessing is time-consuming and difficult to handle and can have major implications on environmental modeling results. US EPA's Hydrological Micro Services Precipitation Comparison and Analysis Tool (HMS-PCAT) provides a publicly available tool to accomplish this critical task. We present HMS-PCAT's software design and its use in gathering, preprocessing, and evaluating precipitation data through web services. This tool simplifies catchment and point-based data retrieval by automating temporal and spatial aggregations. In a demonstration of the tool, four gridded precipitation datasets (NLDAS, GLDAS, DAYMET, PRISM) and one set of gauge data (NCEI) were retrieved for 17 regions in the United States and evaluated on 1) how well each dataset captured extreme events and 2) how datasets varied by region. HMS-PCAT facilitates data visualizations, comparisons, and statistics by showing the variability between datasets and allows users to explore the data when selecting precipitation datasets for an environmental modeling application.
The Piscine Stream Community Estimation System (PiSCES) provides users with a hypothesized fish community for any stream reach in the conterminous United States using information obtained from Nature Serve, the US Geological Survey (USGS), StreamCat, and the Peterson Field Guide to Freshwater Fishes of North America for over 1000 native and non-native freshwater fish species. PiSCES can filter HUC8-based fish assemblages based on species-specific occurrence models; create a community abundance/biomass distribution by relating relative abundance to mean body weight of each species; and allow users to query its database to see ancillary characteristics of each species (e.g., habitat preferences and maximum size). Future efforts will aim to improve the accuracy of the species distribution database and refine/augment increase the occurrence models. The PiSCES tool is accessible at the EPA's Quantitative Environmental Domain (QED) website at https://qed.epacdx.net/pisces/
Gridded precipitation datasets are becoming a convenient substitute for gauge measurements in hydrological modeling; however, these data have not been fully evaluated across a range of conditions. We compared four gridded datasets (Daily Surface Weather and Climatological Summaries [DAYMET], North American Land Data Assimilation System [NLDAS], Global Land Data Assimilation System [GLDAS], and Parameter-elevation Regressions on Independent Slopes Model [PRISM]) as precipitation data sources and evaluated how they affected hydrologic model performance when compared with a gauged dataset, Global Historical Climatology Network-Daily (GHCN-D). Analyses were performed for the Delaware Watershed at Perry Lake in eastern Kansas. Precipitation indices for DAYMET and PRISM precipitation closely matched GHCN-D, whereas NLDAS and GLDAS showed weaker correlations. We also used these precipitation data as input to the Soil and Water Assessment Tool (SWAT) model that confirmed similar trends in streamflow simulation. For stations with complete data, GHCN-D based SWAT-simulated streamflow variability better than gridded precipitation data. During low flow periods we found PRISM performed better, whereas both DAYMET and NLDAS performed better in high flow years. Our results demonstrate that combining gridded precipitation sources with gauge-based measurements can improve hydrologic model performance, especially for extreme events.
Many watershed models simulate overland and instream microbial fate and transport, but few provide loading rates on land surfaces and point sources to the waterbody network. This paper describes the underlying equations for microbial loading rates associated with 1) land-applied manure on undeveloped areas from domestic animals; 2) direct shedding (excretion) on undeveloped lands by domestic animals and wildlife; 3) urban or engineered areas; and 4) point sources that directly discharge to streams from septic systems and shedding by domestic animals. A microbial source module, which houses these formulations, is part of a workflow containing multiple models and databases that form a loosely configured modeling infrastructure which supports watershed-scale microbial source-to-receptor modeling by focusing on animal- and human-impacted catchments. A hypothetical application - accessing, retrieving, and using real-world data - demonstrates how the infrastructure can automate many of the manual steps associated with a standard watershed assessment, culminating in calibrated flow and microbial densities at the watershed's pour point.
We employ Monte Carlo simulation and sensitivity analysis techniques to describe the population dynamics of pesticide exposure to a honey bee colony using the VarroaPop + Pesticide model. Simulations are performed of hive population trajectories with and without pesticide exposure to determine the effects of weather, queen strength, foraging activity, colony resources, and Varroa populations on colony growth and survival. The daily resolution of the model allows us to conditionally identify sensitivity metrics. Simulations indicate queen strength and forager lifespan are consistent, critical inputs for colony dynamics in both the control and exposed conditions. Adult contact toxicity, application rate and nectar load become critical parameters for colony dynamics within exposed simulations. Daily sensitivity analysis also reveals that the relative importance of these parameters fluctuates throughout the simulation period according to the status of other inputs.
Cyanobacterial harmful algal blooms (cyanoHAB) cause human and ecological health problems in lakes worldwide. The timely distribution of satellite-derived cyanoHAB data is necessary for adaptive water quality management and for targeted deployment of water quality monitoring resources. Software platforms that permit timely, useful, and cost-effective delivery of information from satellites are required to help managers respond to cyanoHABs. The Cyanobacteria Assessment Network (CyAN) mobile device application (app) uses data from the European Space Agency Copernicus Sentinel-3 satellite Ocean and Land Colour Instrument (OLCI) in near realtime to make initial water quality assessments and quickly alert managers to potential problems and emerging threats related to cyanobacteria. App functionality and satellite data were validated with 25 state health advisories issued in 2017. The CyAN app provides water quality managers with a user-friendly platform that reduces the complexities associated with accessing satellite data to allow fast, efficient, initial assessments across lakes.
Users of Integrated Environmental Modeling (IEM) systems are responsible for defining individual chemicals and their properties, a process that is time-consuming at best and overwhelming at worst, especially for new chemicals with new structures. A software tool is needed to allow users to define a chemical structure, predict transformation products within an environmental setting, and calculate relevant physicochemical properties. Independent software provides relevant chemical and environmental descriptors to parameterize IEM systems that support fate/transport of organics by integrating cheminformatic applications and software technologies. These 1) encode process science using SMART reaction strings, an extension of SMILES notation; 2) generate transformation products based on functional group analysis (nitroaromatics, azo aromatics, halogenated alkanes), environmental conditions (aerobic or anaerobic), and reaction processes (reduction, hydrolysis, photolysis, biodegradation); 3) generate molecular descriptors (partition coefficients, electron affinities) through calculators; 4) collect environmental descriptors (pH, Fe(II), dissolved organic carbon, soil organic carbon content) from a user or via web-based databases (National Water Quality Database); and 5) retrieve and analyze generated data (quantitative structure activity relationships) in structure-based databases. The results are a web-based tool where the user identifies the organic chemical by structure, common or IUPAC name, or CASID; selects reaction conditions and media; provides environmental descriptors from site-specific data; and chooses a specific transformation process from a reaction library to produce transformation pathways and products, with their physicochemical properties, for IEM consumption.
Eight software applications are compared for their performance in estimating the octanol-water partition coefficient (Kow), melting point, vapor pressure and water solubility for a dataset of polychlorinated biphenyls, polybrominated diphenyl ethers, polychlorinated dibenzodioxins, and polycyclic aromatic hydrocarbons. The predicted property values are compared against a curated dataset of measured property values compiled from the scientific literature with careful consideration given to the analytical methods used for property measurements of these hydrophobic chemicals. The variability in the predicted values from different calculators generally increases for higher values of Kow and melting point and for lower values of water solubility and vapor pressure. For each property, no individual calculator outperforms the others for all four of the chemical classes included in the analysis. Because calculator performance varies based on chemical class and property value, the geometric mean and the median of the calculated values from multiple calculators that use different estimation algorithms are recommended as more reliable estimates of the property value than the value from any single calculator.
Microbial fate and transport in watersheds should include a microbial source apportionment analysis that estimates the importance of each source, relative to each other and in combination, by capturing their impacts spatially and temporally under various scenarios. A loosely configured software infrastructure was used in microbial source-to-receptor modeling by focusing on animal- and human-impacted mixed-use watersheds. Components include data collection software, a microbial source module that determines loading rates from different sources, a watershed model, an inverse model for calibrating flows and microbial densities, tabular and graphical viewers, software to convert output to different formats, and a model for calculating risk from pathogen exposure. The system automates, as much as possible, the manual process of accessing and retrieving data and completes input data files of the models. The workflow considers land-applied manure from domestic animals on undeveloped areas; direct shedding (excretion) on undeveloped lands by domestic animals and wildlife; pastureland, cropland, forest, and urban or engineered areas; sources that directly release to streams from leaking septic systems; and shedding by domestic animals directly to streams. The infrastructure also considers point sources from regulated discharges. An application is presented on a real-world watershed and helps answer questions such as: What are the major microbial sources? What practices contribute to contamination at the receptor location? What land-use types influence contamination at the receptor location? and Under what conditions do these sources manifest themselves? This research aims to improve our understanding of processes related to pathogen and indicator dynamics in mixed-use watershed systems.