Hydrologic modelling plays a vital role in water resources management but often falls short of achieving the positive change modelers envision. In this position paper we argue that a key contributing factor is the lack of trust and shared understanding among modelers, decision-makers, and the public. Models need to be trusted first-a social challenge as well as a technical one. Through three case studies-involving groundwater development, a national-scale runoff model, and a declining lake ecosystem-we analyze the interactions between technical modelling, stakeholder engagement, and policy outcomes, drawing on principles from both hydrologic and social sciences. We recommend that hydrologic modelers foster transparency, balance model authority with flexibility, and tailor stakeholder engagement to overall project needs. Implementing these recommendations will enhance the legitimacy of hydrologic models, increasing the likelihood of achieving positive, sustainable change in water resource systems.
Global hydrological models are essential for managing water resources and predicting hydrological events. However, the local-scale usability of global models challenges big-data management, communication, adoption, and validation. Validation is the biggest challenge bercause of the need for large-scale data management and model calibration, which requires extensive and often inaccessible observed data. This study assesses the GEOGLOWS-ECMWF Global Hydrologic Model, revealing systematic biases that impact its accuracy. We propose a bias-correction methodology using flow duration curves to align non-exceedance probabilities of simulated and observed streamflow, significantly improving the GEOGLOWS model. Unfortunately, this approach does not inherently improve simulations in ungauged locations. The methodology not only enhances the GEOGLOWS model's accuracy but also stands as a versatile solution applicable across various hydrological models. This bias correction approach provides a tool for improving hydrological predictions and gives users the confidence to use global models for local water resource management and decision-making processes.
Flooding is a global problem that impacts people, communities, and governments every year. A better understanding of flooding in an area can enable an improved emergency response before a flood hits. Flood maps are a crucial tool to translate what, for most, is an abstract streamflow into a more understandable and actionable representation of who and what is at risk. Satellite-based flood maps are a useful tool that has potential global applications. We developed methods to determine areas that are suitable for generating satellite-based synthetic flood maps. For our processes, we used Forecasting Inundation Extents using REOF analysis (FIER), a data-driven method of synthesizing flood maps by correlating extracted spatial and temporal patterns from satellite imagery with historical hydrological variables. To overcome the limitation of only using places where gauges are installed, we used large-scale hydrological models, namely the National Water Model (NWM) and the GEOGLOWS Streamflow Model, to provide simulated retrospective streamflow data to train our model. We evaluated locations where both optical and radar imagery would be suitable for creating these models. The procedures we developed and the results that we obtained are potentially transferable to many satellite data sources and methods of model generation.
From 1998 to 2024, we collected field samples at 45 selected lakes in Yellowstone National Park during the months of April through October. We estimated inflows, outflows, and Secchi depths for most lakes. We analyzed the samples for total phosphorous and chlorophyll-a. We used these data to classify the lake trophic states using the Carlson TSI (CTSI), Vollenweider (VW), and Larsen–Mercier (LM) models to assess how trophic states evolved over this 26-year period. This longitudinal dataset is unique because of its extensive 26-year time span gathered from difficult-to-access locations. We found that the data depended on lake size, lake elevation, and the month when data were collected. Most of the lakes exhibit mesotrophic conditions, with variations depending on the trophic state model used. The CTSI distribution shows median values typically between 40 and 55, while the VW and LM index distributions present a somewhat similar pattern but with fewer lakes categorized due to data requirements. We visualized temporal patterns using heatmaps and analyzed trends using the Mann–Kendall test to identify trends and if they were statistically significant. We found only four lakes with statistically significantly increasing trends and two with decreasing trends. Because of the difference in the months when data were collected, the increasing trends in three of the lakes are less certain. We found that, except for four lakes, the trophic states of Yellowstone lakes were maintained or improved over this ~20-year period. Only the trophic state of Nymph Lake clearly deteriorated. The remaining lakes had stable trophic states, with three having weak evidence of worsening conditions. This long-term dataset, which we publish for others’ use, provides an opportunity to better understand eutrophication processes and water quality dynamics in Yellowstone, providing critical information for park management and conservation efforts.
Water-related disasters, including floods and droughts, require innovative forecasting tools for effective risk management and decision-making. The National Water Level Forecast (NWLF) application addresses this need by integrating global hydrological simulations from the GEOGLOWS ECMWF Global Hydrological Model with locally observed data. Developed using the Tethys platform, NWLF employs a pre-computed backend architecture to deliver reliable water level forecasts, ensuring operational stability and an intuitive user interface. Central to its method is the Discharge-to-Water Level Transformation (DWLT), which uses monthly duration curves to align simulated and observed data through non-exceedance probability, significantly improving forecast accuracy. Case studies across Brazil, Colombia, Ecuador, and Peru demonstrate NWLF's scalability and adaptability in addressing diverse regional hydrological challenges. Enhancements in NWLF architecture, such as pre-computation, enable faster analyses. Opportunities to improve the system include collaborative threshold refinements and enhanced stakeholder engagement, NWLF has the potential to bridge the gap between global models and local applications. This study presents the NWLF application and highlights its potential for advancing hydrological forecasting, contributing to Sustainable Development Goals (SDGs) related to water resource management and disaster risk reduction while paving the way for continuous improvement and broader adoption.
Reproducible environmental modelling often relies on spatial datasets as inputs, typically manually subset for specific areas. Yet, models can benefit from a data distribution approach facilitated by online repositories, and automating processes to foster reproducibility. This study introduces a method leveraging diverse state-scale spatial datasets to create cohesive packages for GIS-based environmental modelling. These datasets were generated and shared via GeoServer and THREDDS Data Server connected to HydroShare, contrasting with conventional distribution methods. Using the Regional Hydro-Ecologic Simulation System (RHESSys) across three U.S. catchment-scale watersheds, we demonstrate minimal errors in spatial inputs and model streamflow outputs compared to traditional approaches. This spatial data-sharing method facilitates consistent model creation, fostering reproducibility. Its broader impact allows scientists to tailor the method to various use cases, such as exploring different scales beyond state-scale or applying it to other online repositories using existing data distribution systems, eliminating the need to develop their own.
Assessing groundwater storage changes is important in semiarid regions such as the Volta Basin in sub-Saharan Africa but can be challenging due to data scarcity. This study presents a multi-source approach combining NASA's Gravity Recovery and Climate Experiment (GRACE) mission data, Climate Hazards group Infrared Precipitation with Stations (CHIRPS) precipitation data, the new Global Land Data Assimilation System (GLDAS) v2.2 CLSM groundwater storage anomaly dataset, traditional non-groundwater terrestrial water storage components from GLDAS v2.1, and Copernicus satellite data for surface water dynamics to analyze groundwater storage changes and recharge rates in the Volta Basin. Our study reveals: (1) A small change in groundwater storage from 2002 to 2012, followed by a significant increase from 2012 to 2022; (2) A total groundwater increase of approximately 30 cubic kilometers or about 10 cm of liquid water equivalent over the entire study period; (3) Water storage in the basin is dominated by fluctuations in Lake Volta, which accounts for ∼50 % of terrestrial water storage; (4) Groundwater recharge rates estimated using the Water Table Fluctuation Method on GRACE- and GLDAS-derived data, aligned with values from monitoring wells and previous studies; (5) The GLDAS v2.2 dataset exhibits similar seasonal trends with periodic peaks and troughs, yet the amplitude of anomalies from GLDAS are larger compared to the GRACE dataset; and (6) Groundwater recharge shows a weak correlation with extreme precipitation events, suggesting more water percolates rather than evaporates during such events. However, other factors, including land use changes and agricultural practices, may also impact groundwater storage and recharge.
We estimate long-term groundwater storage loss in California's Central Valley (CV) using a novel data imputation method that combines in situ data with Earth Observations to generate temporally and spatially interpolated groundwater elevations. We combine these data with storage coefficient maps to produce time series of groundwater volume changes which compare well with previously published groundwater storage change estimates for the valley. We also compare our results to groundwater storage changes we calculated using Gravity Recovery and Climate Experiment (GRACE) mission data and show that the two storage estimates are well correlated, but the GRACE volume estimates are lower due the well-known "leakage" effect. While other researchers have accounted for leakage by scaling the GRACE results using various factors and assumptions, our method demonstrates a direct method for calibrating GRACE estimated groundwater change, which can then be applied to future GRACE results in the CV with confidence.
We introduce an open-source web-based Application Programming Interface (API) developed within a representational state transfer (REST) architecture framework that provides access to the operational streamflow forecasts from the U.S. National Water Model (NWM). We built this API within the Google Cloud infrastructure, taking advantage of Google's API Gateway, BigQuery, and the Google Cloud Run architecture. We ran a data transformation from netCDF to tabular formation then ingested data into BigQuery using Apache Beam parallel processing technology. The API activates functions deployed to Cloud Run that executes SQL queries within the BigQuery data warehouse. The API gives users granular control to specify queries based on forecast type, reference datetime, stream segment specifics, and forecast ensemble members. This API greatly simplifies access and use of current and historic NWM forecasts, an otherwise arduous task due to the large volume of data and unwieldy storage methods. Retrieved forecast data can be used for many critical water management applications providing actionable intelligence for, e.g., flood management, safety, recreation, agriculture, and related uses.
Download This Paper Open PDF in Browser Add Paper to My Library Share: Permalink Using these links will ensure access to this page indefinitely Copy URL Copy DOI
Rice is a vital cereal crop providing essential food for Nepal. To mitigate climate risks and ensure food security, especially with climate change and increasing extreme weather events, it is imperative and a prerequisite to have timely and reliable rice yield prediction. While traditional vegetation indices (VIs) and environmental variables are commonly used in rice yield prediction, the potential of solar-induced chlorophyll fluorescence (SIF), despite its greater sensitivity, has been explored in limited studies. Furthermore, integrating SIF with other remote sensing and environmental data for regional crop yield modeling still needs to be explored. This study aims to address these gaps by incorporating SIF to enhance rice yield prediction in Nepal, offering a novel approach to improving model accuracy. This study uses multi-source data and machine learning algorithms to study 22 major rice-growing districts in the Terai region of Nepal. Random forest (RF), support vector machine (SVR), gradient boosting regressor (GBR), and least absolute shrinkage and selection operator (LASSO) algorithms were employed to predict crop yield using various combinations of input features. Machine learning models were trained and tested over four different time windows within the growing period of rice, using 15 years of rice yield data from 2003 to 2017. Our analysis reveals that GBR out-performed other machine learning algorithms when using input features from July to September, achieving root mean square error (RMSE) and the ratio of RMSE (RRMSE) to observed mean values of 250 kg/ha and 8.6%, respectively. The findings also reveal that incorporating SIF alongside traditional remote sensing and environmental variables significantly enhances the model's prediction accuracy. These findings are valuable for ensuring food security, optimizing agricultural resource management, and supporting decision-making for farmers, policymakers, and supply chain stakeholders.
Download This Paper Open PDF in Browser Add Paper to My Library Share: Permalink Using these links will ensure access to this page indefinitely Copy URL Copy DOI
Recent decades have witnessed a massive increase in the volume and quality of hydrologic data available to aid water resources decision makers, managers, and scientists. This has been accompanied by exponential growth in both desktop and cloud computing, as well as data storage capabilities. As a result, there are abundant opportunities to drastically change how water data is collected, managed, disseminated, and analyzed – which should ultimately have significant positive impacts on water science, engineering, and management. We are at the cusp of a new era in water data science which brings with it many exciting technological and scientific challenges and opportunities. Many of these challenges are alleviated and opportunities are multiplied when hydrology is viewed as a “team sport” rather than as an individual activity. These factors formed the motivation for the development of the HydroShare open-source software and the hydroshare.org operational system. This retrospective paper reviews a decade of HydroShare development and operation by presenting the general architecture, functionality, key contributions of the project to earth science and cyberinfrastructure research, current usage metrics, and future directions.
Surface water is a vital component of the Earth’s water cycle and characterizing its dynamics is essential for understanding and managing our water resources. Satellite-based remote sensing has been used to monitor surface water dynamics, but cloud cover can obscure surface observations, particularly during flood events, hindering water identification. The fusion of optical and synthetic aperture radar (SAR) data leverages the advantages of both sensors to provide accurate surface water maps while increasing the temporal density of unobstructed observations for monitoring surface water spatial dynamics. This paper presents a method for generating dense time series of surface water observations using optical–SAR sensor fusion and gap filling. We applied this method to data from the Copernicus Sentinel-1 and Landsat 8 satellite data from 2019 over six regions spanning different ecological and climatological conditions. We validated the resulting surface water maps using an independent, hand-labeled dataset and found an overall accuracy of 0.9025, with an accuracy range of 0.8656–0.9212 between the different regions. The validation showed an overall false alarm ratio (FAR) of 0.0631, a probability of detection (POD) of 0.8394, and a critical success index (CSI) of 0.8073, indicating that the method generally performs well at identifying water areas. However, it slightly underpredicts water areas with more false negatives. We found that fusing optical and SAR data for surface water mapping increased, on average, the number of observations for the regions and months validated in 2019 from 11.46 for optical and 55.35 for SAR to 64.90 using both, a 466% and 17% increase, respectively. The results show that the method can effectively fill in gaps in optical data caused by cloud cover and produce a dense time series of surface water maps. The method has the potential to improve the monitoring of surface water dynamics and support sustainable water management.
Recent years have witnessed a significant increase in the availability and number of geographic simulation models across various domains, leading to challenges in evaluating their relative value. Traditional model evaluations typically compare simulation results with measured data or other models. This report presents the application of the newly “Model Academic Influence Index (MAI)" method which focuses on evaluating a model's academic contributions. It offers both annual and lifetime index, and reflects the model's major application areas covered. The report evaluates the MAI of 205 models and 22 methods in 2022 from trusted digital repositories and emphasizes the importance of open-source models, providing URLs and licenses. Recognizing the complexity and importance of this task, we invite ongoing discussion and feedback from the modeling community. This report aims to support more informed decision-making in academia and the public and promote the development of a more open and scientific modeling profession and community.
Obtaining and managing groundwater data is difficult as it is common for time series datasets representing groundwater levels at wells to have large gaps of missing data. To address this issue, many methods have been developed to infill or impute the missing data. We present a method for improving data imputation through an iterative refinement model (IRM) machine learning framework that works on any aquifer dataset where each well has a complete record that can be a mixture of measured and input values. This approach corrects the imputed values by using both in situ observations and imputed values from nearby wells. We relied on the idea that similar wells that experience a similar environment (e.g., climate and pumping patterns) exhibit similar changes in groundwater levels. Based on this idea, we revisited the data from every well in the aquifer and “re-imputed” the missing values (i.e., values that had been previously imputed) using both in situ and imputed data from similar, nearby wells. We repeated this process for a predetermined number of iterations—updating the well values synchronously. Using IRM in conjuncture with satellite-based imputation provided better imputation and generated data that could provide valuable insight into aquifer behavior, even when limited or no data were available at individual wells. We applied our method to the Beryl-Enterprise aquifer in Utah, where many wells had large data gaps. We found patterns related to agricultural drawdown and long-term drying, as well as potential evidence for multiple previously unknown aquifers.
Multidimensional, georeferenced data are used extensively in hydrology, meteorology, and water science and engineering. These data are produced, shared, and used by diverse organizations globally. Conventions have been developed to standardize the metadata and format of these datasets to ensure compatibility with current and future software and web services. However, the most common conventions are complex and difficult to implement correctly, resulting in datasets that are unusable for many applications due to a lack of compliance with the conventions. We have developed a method and software module for programmatically assigning metadata and guiding the dataset creation, validating, and cleaning process, so that convention-compliant datasets can be consistently and repeatably created by people with a limited knowledge of file formats and data standards. These datasets can then be used in any application that supports the particular standard. Specifically, this paper examines the process of building multidimensional, georeferenced netCDF datasets that are compliant with the NetCDF Climate and Forecast Conventions. We present a new free and open-source Python package called cfbuild that helps to automate the process of building or updating datasets, making them sufficiently compliant with the Climate and Forecast Conventions and the Attribute Conventions for Data Discovery so that they can be reliably served using a THREDDS Data Server and shared via OPeNDAP.
Artificial intelligence applications to environmental modelling, Modelling, Simulation and optimisation of environmental processes, in particular for water management,
The Group on Earth Observations (GEO) Global Water Sustainability (GEOGloWS) hydrologic model provides global river discharge hindcasts and daily forecasts at approximately one million subbasins worldwide. The model is meant to sustainably provide discharge data during emergency situations and to underdeveloped countries which do not have sufficient local capacity. The primary model error is biased flow magnitudes which reduce the usefulness of the results. We applied a revised implementation of the SABER bias correction method to correct GEOGloWS model results. SABER uses a combination of watershed clustering with machine learning, geospatial analysis, and statistics to generalize bias patterns in gauged basins so they can also be applied to ungauged basins. We validated the bias corrected data created using the improved SABER method at 12,965 gauges globally and showed that this method reduced the mean error at 90% of gauges. We present an analysis of the improved SABER method using several metrics including mean error, root mean squared error, and Kling Gupta Efficiency. We found that the GEOGloWS model is usually biased high but our results indicate a reduction in the bias of the GEOGloWS model worldwide. We evaluate the varied performance of the bias correction procedure and significance of the improvements which vary based on stream order the watershed classifications derived in our analysis. We provide guidance on the use of bias corrected global data to local scale applications and discuss implications for the GEOGloWS model in the future.
Alva Couch合作论文数Computer Science Department
Tufts University12