Groundwater is a significant part of the global water cycle and an essential water source for domestic, industrial and agricultural use. In England, groundwater supplies ∼30% of public water, exceeding 75% in the most densely populated and water-stressed Thames and Southern regions. We analyse groundwater level trends of 2902 wells across England with 9–189 year records. We show that about half of the stations experience long-term trends or sudden changes. Long-term trends are very spatially heterogeneous. We investigate the links between these trends and various potential drivers and other proxy variables, and find that population density is the primary explanatory factor for increasing trends in densely populated urban areas, and irrigation intensity is the dominant factor for decreasing trends in intensive irrigated areas. Our results demonstrate that the spatial variability of groundwater trends across England is characterised by widespread anthropogenic footprints.
The integration of artificial intelligence into Earth system models (ESMs) has revolutionized the simulation and prediction of complex environmental dynamics. However, this shift introduces substantial challenges for reproducibility, a cornerstone of scientific progress. In particular, artificial intelligence-infused hybrid ESMs face amplified issues of numerical instability, procedural opacity and asymmetric access to computational resources. If left unaddressed, these challenges risk turning hybrid ESMs into opaque and weakly verifiable systems, reducing model traceability, weakening cumulative knowledge building and narrowing the evidential basis for climate risk assessment and policy guidance. This Perspective argues that reproducibility should be reframed to reflect the epistemological and operational realities of hybrid ESMs. We propose an integrated roadmap that couples a theory of reproducibility assessment with practical pathways for implementation in modelling practices. Within this context, we introduce Reproducibility in hybrid Earth system models (RHEM) as a reference guideline for governing transparent, trustworthy, and reproducible hybrid ESMs. Building on this foundation, the pathways operationalize the framework’s criteria into actionable measures that embed transparency and trustworthiness in the modelling process itself. By linking conceptual structure with operational guidance, reproducibility is repositioned from a post hoc requirement to a structural property of hybrid ESMs and established as a foundational principle for Earth system science in the artificial intelligence era. AI integration in Earth system models enhances prediction and modelling capabilities but also amplifies challenges for reproducibility. This Perspective introduces a framework for assessing reproducibility and provides practical ways to strengthen reproducibility in hybrid Earth system models.
In hydrology, deep learning (DL) models have already achieved remarkable breakthroughs in predicting streamflow. These DL models are fed with meteorological time-series and static catchment attributes across large samples of catchments, and predict streamflow remarkably well in both gauged and ungauged situations. In recent years, some studies have transferred the idea of constructing multi-basin/station DL models – particularly Long-Short Term Memory (LSTM) neural networks – to large-sample groundwater level modelling to explore their potential for temporal and spatial extrapolation. To the best of authors’ knowledge, existing multi-station LSTM applications are limited to three, covering 76 climate-sensitive stations in Northern France, 108 nationwide stations in Germany, and 1,800 coastal stations across nine countries/regions. Notably, spatial generalisation was investigated solely in the German study, which suggests that the model utilised static features primarily as ‘unique identifiers’ to memorise local behaviour rather than deriving the generalisable hydrological insights required for spatial extrapolation. Given the limited number of studies and the potentially biased datasets, the generalisation ability of multi-station DL models for groundwater level modelling is still under exploration.A newly released large-sample groundwater dataset by the Environment Agency of England, comprising more than 200,000 daily and 200 million sub-daily sampling observations for over 3,400 wells, offers a unique opportunity to test the generalisation ability of multi-station DL models in time and space, and whether these models can yield process-relevant insights on groundwater dynamic mechanisms. In this study, we want to investigate the following questions:1) How well can multi-station DL models simulate the groundwater variability across England?2) Which input features does the DL model use to make its predictions (especially in places where it does well)?
Groundwater is a central component of the Earth system. However, our understanding of how it is dynamically interlinked with the atmosphere, hydrosphere, cryosphere, biosphere, geosphere, and anthroposphere remains limited. In the pursuit of understanding groundwater dynamics across diverse global settings, we present GROW (the global-scale integrated GROundWater package). This analysis-ready, quality-controlled dataset combines depth to groundwater and level time series from 55 countries, 91% from North America, India, Europe, and Australia, with associated Earth system variables. The dataset contains >200,000 time series with either daily, monthly, or yearly temporal resolution, accompanied by 36 time series or static attributes of meteorological, hydrological, geophysical, vegetation, and anthropogenic variables (e.g., precipitation, drainage density, rock type, NDVI, land use). 34 data flags regarding well features (e.g., coordinates and country), as well as time series characteristics (e.g., gap fraction or autocorrelation), facilitate quick data filtering. GROW provides a foundation for understanding large-scale groundwater processes in space and time, as well as for calibrating and evaluating models that simulate groundwater dynamics within the Earth system.
In highly seasonal regimes hydrologic models generally achieve high scores on common performance metrics such as the Nash-Sutcliffe Efficiency (NSE) and the Kling-Gupta Efficiency (KGE). However, variance in streamflow time series is composed of seasonal, interannual, and irregular variance, and the NSE and KGE do not differentiate between these components. Differences in performance on these three components have not been evaluated across a broad spectrum of hydrologic models and regions. We analyse open-access simulations from 18 regional and global hydrologic models. We find that these models consistently achieve the highest NSE and KGE in highly seasonal catchments where they are worse at simulating interannual variability, compared to less seasonal catchments. Simulated year-to-year changes in ecologically relevant hydrologic signatures are less accurate in highly seasonal catchments, and the NSE of the interannual variance component is usually lower. This suggests that these hydrologic models may struggle to predict long-term responses to climate change, especially in highly seasonal tropical, alpine, and polar regions, which are some of the most vulnerable to climate change. We encourage hydrologic modellers to explicitly evaluate skill at simulating interannual variability, rather than relying only on aggregate measures such as the NSE and KGE.
Reliable quantification of global water-cycle components, such as river flow and land evapotranspiration, remains a major challenge. Here we refine estimates of global water partitioning by combining outputs from multiple Earth system models with river flow observations from 50 large basins, applying the emergent constraint approach. Between 1980 and 2014, global river flow was (39.1 ± 5.4) × 103 km3 yr−1, with a river flow-to-precipitation ratio of 0.35 ± 0.03, both lower than previous estimates. Land evapotranspiration reached (73.4 ± 6.2) × 103 km3 yr−1. Under climate change, we project global river flow to rise by 7.8 ± 5.5 mm per year per degree of warming. This estimate, refined through the emergent constraint method, is 9.3
Flood loss models are increasingly used in the (re)insurance sector to inform a range of financial decisions, and more broadly in research and policy analysis to understand present-day and future flood risk trends. These models simulate the interactions between flood hazard, vulnerability and exposure over large spatial domains, requiring a range of input information and modelling assumptions. Due to this high level of complexity, evaluating the impact of uncertain input data and assumptions on modelling results, and therefore the overall model “acceptability”, remains a very complex process. In this paper, we advocate for the use of global sensitivity analysis (GSA), a generic technique to analyse the propagation of multiple uncertainties through mathematical models, to improve the sensitivity testing of flood loss models and the identification of their key sources of uncertainty. We discuss key challenges in the application of GSA to large-scale flood loss models, propose pragmatic strategies to overcome these challenges, and showcase the type of insights that can be obtained by GSA through two proof-of-principle applications to a commercial model, JBA Risk Management's flood loss model, for the transboundary Rhine River basin in Europe, and Queensland in Australia.
We are in a simultaneous state of exuberance and starvation of Earth system data. Model ensembles of increasing complexity provide petabytes of output, while remote sensing products offer terabytes of new data every day. On the other hand, we still lack data on some key processes that are more challenging to observe, like groundwater recharge, or only from particular regions of the world (often regions already heavily impacted by anthropogenic change). This leaves us with highly imbalanced datasets. Our ability to produce and collect mountains of data contrasts with our progress in improving scientific process understanding. How can we harness model simulations and data alike to enhance our knowledge and test scientific hypotheses about process relationships despite data gaps and poorly known biases in modelled and observational datasets? Our talk discusses methods to approach this problem while being agnostic to the data source (model simulations or observations). We present a new approach to interrogate a given dataset and identify correlational and possibly causal relationships between its variables. We test the method on an ensemble of complex global hydrological model simulations and observations from the ISIMIP experiments, and demonstrate its usefulness and limitations. We show that our approach can provide powerful insights into dominant process controls while scaling with large amounts of data.
We increasingly rely on complex models to understand the functioning and fate of our planet. But what hydrological theory underlies these models and thus the conclusions we draw from them? How do old and rapidly increasing new observations help to advance both theory and models? In this contribution, we discuss the importance of functional relationships for (large-scale) hydrology. We define functional relationships as relationships between two or more variables that characterize the functioning of hydrological systems, such as relationships between forcing and response variables (e.g. precipitation and runoff). Functional relationships are not only a central part of hydrological theory, but they also inform how we contextualize and make measurements, and they help us to build, constrain, and evaluate models. To illustrate their value, we first provide an overview of some relationships in large-scale hydrology. We then show how such relationships can be used to evaluate global water models. We conclude by discussing what our current state of knowledge can tell us about what we should explore next, in particular the need for mechanistic explanations of empirical relationships and the potential of linking multiple hydrological fluxes within a unified framework.
Hydrological models differ in the way how hydrological processes are implemented. A rigorous comparison of different hydrological model structures is needed to disentangle the link between similarities and differences in process representations and simulated hydrological processes, states and fluxes. A major challenge in model comparison is to identify effects of individual processes. To move a step in this direction, we developed controlled experiments and compared three hydrological models (HBV, mHM, SWAT+) in nine German catchments (400-3000 km²) along an elevation gradient. We aim at presenting a framework for a consistent comparison of process representations in model structures consisting of three steps: (1) A model comparison protocol was developed for a detailed comparison of process representations in model structures. Consistency was achieved by using the same input data for all models. By grouping the processes in a standardized way, differences and similarities between the models were identified. (2) To investigate the dominant model components, a daily parameter sensitivity analysis was carried out for the three models with different hydrological variables as target variables (e.g. actual evapotranspiration, soil moisture, snow and discharge). The dominant model parameters and associated processes vary more between the models than between the catchments. This also applies to the temporal variability of the parameter sensitivity. (3) The model performance was analysed for a set of different performance criteria. The optimal parameter values differ greatly depending on which performance criteria were selected. This is in particular true for soil and evapotranspiration parameters. Typical patterns can be derived between catchments of different landscapes. The joint analysis of these three methodological steps demonstrates the benefit of a detailed process analysis in model structures for a better understanding of suitable process representations. Therefore, it shows the potentials for improving model structures.
ABSTRACTCoastal groundwater is a vital resource for coastal communities around the globe, and submarine groundwater discharge (SGD) delivers nutrients to coastal marine ecosystems. Climatic changes and anthropogenic actions alter coastal hydrology, causing seawater intrusion (SWI) globally. However, the selection of SWI and SGD study sites may be highly biased, limiting our process knowledge. Here, we analyse hydroenvironmental characteristics of coastal basins studied in 1298 publications on SGD and SWI to understand these potential biases. We find that studies are biased towards basins with gross domestic product per capita below (SWI) and above (SGD) the median of all global coastal basins. Urban coastal basins are strongly overrepresented compared to rural coastal basins, limiting our progress in understanding undisturbed natural processes. Despite the connection between anthropogenic activity and coastal groundwater issues, and the consequential overrepresentation of urban basins in coastal groundwater studies, perceptual (or conceptual) models of coastal groundwater rarely include anthropogenic influences aside from pumping (e.g., subsidence, land use change). Taking a holistic view on coastal groundwater flows, we have developed an editable perceptual model illustrating the current understanding, including both natural and anthropogenic drivers. As SGD and SWI in new areas of the globe are studied, we advocate for researchers to utilise and further edit this perceptual model to openly communicate our process understanding and study assumptions.
Groundwater, Earth’s largest nonfrozen freshwater reservoir, is vital for water supply security. Groundwater models help to manage complex domestic, agricultural, and industrial water demands while preserving ecosystem health under climate change. The community-driven groundwater model portal (GroMoPo) hosts groundwater model metadata to analyse biases and distribution of groundwater models. Over 450 models are currently featured on GroMoPo, with most models from high-GDP countries at local-to-regional scales. The GroMoPo initiative addresses current knowledge gaps and facilitates future collaboration and data sharing.
Global hydrological models are valuable tools to predict flood hazard across data-scarce regions and future climate scenarios. Their ability to create spatially coherent projections means their results are broadly used for scientific analysis and policy planning. However, the complexity of the models, coupled with the high volume of data they generate, poses significant challenges in evaluating the process representation contained within the models. Existing analysis show, how a model transfers input into output varies strongly between global water models in a long-term analysis. Yet, flood event prediction needs to take place at daily or higher temporal resolution. Are global hydrological models able to accurately represent flood generation? And do they accurately combine different flood-generating processes, such as extreme rainfall, snowmelt, or wet antecedent conditions, into extreme flows?In this analysis, we compare simulations from five global hydrological models. The models are part of the global water sector within the third simulation round of the Inter-Sectoral Impact Model Intercomparison Project (ISIMIP3a). In ISIMIP, all models are run with the same forcing data, on a daily resolution from 1901 to 2019. We extract and compare runoff time series across the 67400 land cells. For each cell, a threshold-based flood event extraction allows calculation of flood duration, magnitude, number of extreme events, etc. Additionally, we use the extracted events to compare model inputs, such as precipitation, or model fluxes, such as snowmelt, that contribute to high-flow generation.Five models (CWatM, H08, LPJmL, ORCHIDEE, WaterGAP2), with four input variables and fluxes (precipitation, runoff, soil moisture, and snowmelt) at daily resolution over 67400 land cells results in 58 billion data points to analyse. Extracting this process-based statistical information from the model data reduces the dimensionality and scope of the high-resolution data to a form where comparison between models is possible. How do high flow statistics compare between models? Does the same extreme rainfall result in extreme flow across all models? What role does snowmelt and soil moisture play in runoff generation between models? These questions support an evaluation of flood events within global models through process-based model intercomparison.
It is projected that the likelyhood and duration of extreme soil moisture (SM) droughts will increase in Germany under future warming scenarios. Annual precipitation changes are small under climate change in Germany with increases in winter and decreasing precipitation in summer for some parts of Germany. Generally, the climate ensemble spread in the future precipitation signal is large. Furthermore, impacts of SM droughts depend largely on the soil volume evaluated. We identified a gradient of stronger soil drying in shallow SM compared to deeper SM under global warming, leading to different effects on shallow-rooted vegetation compared to deep-rooted vegetation (agriculture versus forestry). In addition, spatial characteristics such as soil properties can strongly influence the dynamics of SM and thus shape the response of SM drought to changing meteorological conditions. In this work we evaluate the impact of the considered soil depth and spatial features on simulated changes in SM droughts in Germany. We compare this influence to the uncertainty in meteorological changes. We use a large climate ensemble based on Euro-Cordex regional climate model simulations, which were bias-adjusted and spatially disaggregated to run the mesoscale hydrological model (mHM) (mhm-ufz.org) with a high spatial resolution of 1.2x1.2km. This work aims to expand the picture of climate change impacts on SM droughts in Germany. The results can contribute to an improved definition of sector-specific drought indicators that will support national efforts to ensure climate change resilient water management.
Topography affects the distribution and movement of water on Earth, yet new insights about topographic controls continue to surprise us and exciting puzzles remain. Here we combine literature review and data synthesis to explore the influence of topography on the global terrestrial water cycle, from the atmosphere down to the groundwater. Above the land surface, topography induces gradients and contrasts in water and energy availability. Long-term precipitation usually increases with elevation in the mid-latitudes, while it peaks at low- to mid-elevations in the tropics. Potential evaporation tends to decrease with elevation in all climate zones. At the land surface, topography is expressed in snow distribution, vegetation zonation, geomorphic landforms, the critical zone, and drainage networks. Evaporation and vegetation activity are often highest at low- to mid-elevations where neither temperature, nor energy availability, nor water availability—often modulated by lateral moisture redistribution—impose strong limitations. Below the land surface, topography drives the movement of groundwater from local to continental scales. In many steep upland regions, groundwater systems are well connected to streams and provide ample baseflow, and streams often start losing water in foothills where bedrock transitions into highly permeable sediment. We conclude by presenting organizing principles, discussing the implications of climate change and human activity, and identifying data needs and knowledge gaps. A defining feature resulting from topography is the presence of gradients and contrasts, whose interactions explain many of the patterns we observe in nature and how they might change in the future.
While measured streamflow is commonly used for hydrological model evaluation and calibration, an increasing amount of data on additional hydrological variables is available. These data have the potential to improve process consistency in hydrological modeling and consequently for predictions under change, as well as in data-scarce or ungauged regions. Here, we show how these hydrological data beyond streamflow are currently used for model evaluation and calibration. We consider storage and flux variables, namely snow, soil moisture, groundwater level, terrestrial water storage, evapotranspiration, and altimetric water level. We aim at summarizing the state-of-the-art and providing guidance for the use of additional hydrological variables for model evaluation and calibration. Based on a review of the current literature, we summarize observation methods and uncertainties of currently available data sets, challenges regarding their implementation, and benefits for model consistency. The focus is on catchment modeling studies with study areas ranging from a few km 2 to ~500,000 km 2 . We discuss challenges for implementing alternative variables that are related to differences in the spatio-temporal resolution of observations and models, as well as to variable-specific features, for example, discrepancy between observed and simulated variables. We further discuss advancements required to deal with uncertainties of the hydrological data and to integrate multiple, potentially inconsistent datasets. The increased model consistency and improvement shown by most reviewed studies regarding the additional variables often come at the cost of a slight decrease in streamflow model performance.
Global water models increasingly allow us to explore the terrestrial water cycle in earth-sized digital laboratories to support science and guide policy. However, these models are still subject to considerable uncertainties that mainly originate from three sources: (1) imbalances in data quality and availability across geographical regions and between hydrologic variables, (2) poorly quantified human influence on the water cycle, and (3) difficulties in tailoring process representations to regionally diverse hydrologic systems. New, more accurate, and larger datasets, as well as better accumulated and even improved knowledge, will help to reduce these uncertainties and eventually lead to model advancement. In this review, we explore sources of uncertainty critical to global water models and define actions to reduce them where possible, therefore providing a guide for global model advancement. Following this path will increase the robustness of model outputs, which is urgently needed to tackle key scientific and societal challenges.
Rainfall-runoff models are often calibrated by defining feasible parameter ranges and constraining them with streamflow data, and occasionally other hydrological variables. Traditionally, global prior ranges have served as a baseline, containing a wide range of parameter values suitable for various catchment types. However, there might be more information available to reduce a priori parameter uncertainty in a structured way. This study addresses this gap by defining plausible prior parameter ranges based on the distribution of identifiable parameters and their relationship to catchment characteristics. Using a version of the conceptual Probability Distributed Moisture (PDM) model, the study focuses on a large sample of catchments in the United States, covering diverse climatic, land cover, geological, and landscape types. Thus, investigating the effects of these physical and climatic properties on parameter prior ranges. The combination of automatic grouping and catchment attributes resulted in significant reductions in parameter space, with high average predictive accuracies for traditional efficiency measures. Surprisingly, we find distinct and spatially coherent regions within the US where specific prior parameter ranges maintain high levels of performance. More than 75% of the catchments show NSE values above 0.6 and KGE values above 0.7. Our results suggest that regionalizing prior parameter ranges can significantly reduce parameter uncertainty. These findings have significant implications for the prediction of hydrological responses in ungauged catchments.