To improve hydrological uncertainty estimation, recent studies have explored machine learning (ML)-based post-processing approaches that enable both enhanced predictive performance and hydrologically informed probabilistic streamflow predictions. Among these, random forests (RF) and their probabilistic extension, quantile random forests (QRF), are increasingly used for their balance between interpretability and performance. However, the application of QRF in regional post-processing settings remains unexplored. In this study, we develop a hydrologically informed QRF post-processor trained in a multi-site setting and compare its performance against a locally (at-site) trained QRF using probabilistic evaluation metrics. The QRF framework leverages simulations and state variables from the GR6J process-based hydrological model, along with readily available catchment descriptors, to predict daily streamflow uncertainty. Our results show that the regional QRF approach is beneficial for hydrological uncertainty estimation, particularly in catchments where local information is insufficient. The findings highlight that multi-site learning enables effective information transfer across hydrologically similar catchments and is especially advantageous for high-flow events. However, the selection of appropriate catchment descriptors is critical to achieving these benefits.
Abstract. Reliable transferability of hydrological model parameters across time and space remains a challenge for large‑scale water resources assessment. In this study, we investigate whether a differentiable hybrid framework can identify robust and physically coherent parameter sets for annual streamflow modeling across a large‑sample dataset of 3,044 catchments from eight countries. To focus on temporal and spatial transferability analysis, we work at the annual time scale using what we consider to be the simplest possible model: an annual anomaly model of climate elasticities, coupled with the Turc–Mezentsev formulation for the long-term streamflow mean (MQ). A dense neural network is trained in an end‑to‑end fashion to map catchment descriptors to the four model parameters, with gradients propagated through the entire modeling chain. We evaluate the framework using three cross‑validation settings inspired by Klemeš (1986): temporal, spatial, and combined temporal–spatial cross-validation. As a benchmark, we compare the hybrid model against local, catchment‑by‑catchment linear regressions under temporal cross-validation. Our results show that, for temporal transferability, our parameter learning approach outperforms local calibration, yielding higher Nash-Sutcliffe efficiency (NSE) values while producing elasticity coefficients that remain within plausible physical ranges, despite lacking explicit parameter constraints. By contrast, spatial transferability reveals a marked limitation: the anomaly component extrapolates well spatially, but regionalizing MQ from descriptors proves difficult, with MQ errors dominating the loss of performance in spatial and spatiotemporal cross-validation. Experiments with random descriptors further show that our parameter learning uses attributes mainly as catchment identifiers in temporal cross-validation but relies on their physical content to sustain spatial transfer, particularly for MQ. Overall, the study demonstrates that simple differentiable hybrid annual models can learn robust and interpretable anomaly parameters, while highlighting MQ regionalization as the main remaining bottleneck for spatially transferable annual streamflow predictions.
Relating variations in annual streamflow to a climate anomaly, commonly referred to as streamflow elasticity to climate, is central for a rapid assessment of the impact of climate change on water resources. This elasticity is classically estimated via a multiple linear regression between anomalies in streamflow and climate variables. However, this approach does not explicitly account for the fact that elasticity depends on aridity as suggested by “Budyko-type” water balance formulas. Using a large dataset of 4,122 catchments from four continents, we first verify empirically the link between elasticity and aridity. Then, we propose a method to constrain elasticity coefficients with derivatives from a “Budyko-type” water balance formula, that allows introducing an explicit dependency between elasticity and aridity. We show that adding this dependency produces a regionalized elasticity formula with physically-realistic elasticity coefficients.
In bucket-type hydrological models (also known as conceptual models), the paths of the water in a catchment are represented in a simplified way. In general, there is one way for water to enter – via precipitation – and two ways to leave – streamflow and evaporation. However, as the simple water balance equation P=Q+E this concept is based on is often not fulfilled, many bucket-type models include ‘sweep parameters’, parameters that represent an additional way for water to enter or leave the catchment. Sweep parameters come as correction factors that are used to align inputs and outputs, but also in more sophisticated ways, such as representations of groundwater inflows or outflows. The X2 parameter (Intercatchment Groundwater Flow parameter) in the GR4J model is a well-known example of a sweep parameter.Compared to a model in which the water balance is enforced, a model that includes a sweep parameter is usually more successful in simulating streamflow volumes: Too much or too little water can be compensated thanks to the sweep parameter, while otherwise the only options for compensation are via evaporation or large simulated storage volumes.Because including a sweep parameter improves model performance, sweep parameters are often seen as ‘cheat parameters’. This accusation is understandable, since sweep parameters can also compensate for incorrect input data. Still, there are many reasons why the use of sweep parameters should not be frowned upon. Many catchments are not closed systems along their topographic borders and sweep parameters are one way of representing this knowledge. In addition, we should avoid compensating for incorrect input data or additional water gains or losses via evaporation, a flux that is generally not included in model calibration – and that could be considered cheating as well. If a mismatch in the basic water balance can be represented via a sweep parameter, this is arguably a reasonable and transparent way to do so.To investigate the effects of sweep parameters, we tested the model performance and model robustness towards variations in precipitation input data for hydrological models with and without a sweep parameter. Using a large-sample approach for more than 500 catchments in France, we could not find any evidence that model robustness is affected by the use of a sweep parameter. Furthermore, we clearly illustrate that models benefit from using a sweep parameter. Based on these results, we argue that it is justifiable to decide to sweep, but also stress that the way and effect of the sweeping should be communicated transparently and interpreted with caution.
One of the most basic questions asked of hydrologists is the quantification of catchment response to climatic variations, i.e., the variations around the average annual flow given the climatic anomaly of a particular year. This paper presents an analysis based on 4122 catchments from four continents, where we investigate how annual streamflow variability depends on climate variables – rainfall and potential evaporation – and on the synchronicity between precipitation and potential evaporation. We use catchment data to verify the existence of this link and show that, in all countries and under the main climates represented, anomalies in this synchronicity are the second most important factor to explain annual streamflow anomalies, after precipitation, but before potential evaporation. Introducing the synchronicity between precipitation and potential evaporation as an independent variable improves the prediction of annual streamflow variability with an average additional explained variance of 6 % globally.
The hydrological analysis of high elevation catchments is particularly difficult for two reasons: . first, precipitation measurements are scarce at higher elevations, . second, even when there are precipitation measurements, the collected amounts are strongly biased due to the well-known effect of wind on snowflakes. Several formulations have been proposed to correct this wind-dependent underestimation of solid precipitation amounts. They all depend on at least one parameter, which must be calibrated for the specific location. At a few locations in the world, a double-fenced shielded raingage can be used to provide a reference precipitation amount, and the parameter of the correction can be determined experimentally. But at most locations, we have no real way to parameterize the adjustment relationship. We use here a newly released dataset comprising 30 years of data for 11 stations located at high elevation in Armenia, where the precipitation gage network is strongly impacted by snow undercatch. Using ground snow surveys jointly with a degree-day based snow accumulation and melt model, we show that we can propose an adapted parameterization of the correction formula.
Hydrologists are requested to quantify the response of catchments with respect to climatic variability or climatic changes: for this, they need to be able to assess the climate elasticity of streamflow. Here, we present a large sample study, based on 4122 catchments from four continents, investigating to which extent the climate elasticity of streamflow depends on aridity, i.e. the ratio of the long-term average values of potential evaporation to precipitation. After examining the example of the “Budyko-type” water balance formulas – which embed the dependency between elasticity and aridity – we use catchment data to verify empirically the existence of this link and we discuss the possibilities to impose the dependency to aridity in elasticity in order to obtain more physically-consistent elasticity coefficients.
Groundwater sustains human water use globally, as it provides about a quarter and half of the total water withdrawn for irrigation and domestic purposes, respectively. Intense groundwater pumping has an impact both under- and above-ground, by lowering the water table and reducing streamflow in surrounding rivers. However, groundwater abstraction is often neglected in hydrological models because of the large uncertainties involved. These modelling uncertainties arise from the lack of data to constrain natural processes (including groundwater recharge and discharge, intercatchment groundwater flow) and anthropogenic processes (abstraction rates and their spatiotemporal patterns). Therefore, there is a need to represent groundwater abstraction in hydrological models, to consider its uncertainties, and to determine the appropriate level of complexity in process representation given data availability.This study examines the uncertainties in groundwater abstraction for streamflow predictions over a sample of catchments in France. To this end, we use a parsimonious lumped hydrological model at the daily time step (GR6J), which represents groundwater storage through an exponential store (Michel et al, 2003). Groundwater abstraction is modelled by taking water from this exponential reservoir. We account for the uncertainties in both the water withdrawal input data and the hydrological model parameters (that describe the natural processes). We adopt annual abstraction data from the French national dataset (BNPE), that we temporally disaggregate using different assumptions. Regarding the model parameters, we select an ensemble of parameter sets that produce simulations that are consistent with the observations (streamflow, groundwater levels). Our results reveal that, beyond streamflow observations, piezometric data help to reduce the uncertainty in the parameters such as the capacity of the exponential store. Overall, our study shows the importance of accounting for groundwater abstraction and its uncertainties for streamflow predictions.Michel, C., Perrin, C. & Andréassian, V., 2003. The exponential store: a correct formulation for rainfall-runoff modelling. Hydrological Sciences Journal, 48(1): 109-124, https://dx.doi.org/10.1623/hysj.48.1.109.43484
When evaluating the calibration of hydrological models, in some discharge-only calibration cases, we sometimes find that the selected parameter sets succeed in discharge estimation, while producing unreasonable simulations of other fluxes. In this paper, we hypothesize that the realism of the unmeasurable intercatchment groundwater flow (IGF) may be improved by introducing a constraint on the actual evaporation (AE). This paper evaluates three multi-objective calibration cases in which an extra constraint on AE is added to the classical objective function, which only focus on discharge efficiency. The results show that for the multi-objective cases (1) the discharge efficiency is not significantly compromised; (2) the spatial efficiency (SPAEF) metric with the reference AE data on the annual scale improves considerably; (3) an agreement with each other for the partition of annual water balance components, especially AE and IGF, is reached. These results support the improved realism of the AE and thus the IGF estimation.
The use of machine learning (ML) methods in rainfall-runoff modelling has apparently led to better prediction, but there are some concerns about the interpretability of these models. The emergence of hybrid modelling, which couples the data driven approach with the classical physics-based conceptual approach, has shown promise in enhancing both interpretability and accuracy. ML models and conceptual models each come with their own modelling practices and habits. To develop a hybrid approach, it is necessary to consider them. While some of the steps in these modelling chains are similar (for instance the selection of the right metric during the calibration or learning step), others are more specifics, such as the optimization of the hyper-parameters of ML models. Furthermore, the hybrid approach comes with specific methodological challenges that emerge when coupling the two different types of models. For instance, depending on the choice made by the modeller, the parameters of the conceptual model are either trained with the ML model parameters or calibrated separately by a non-ML method. There is a need to better understand the variety of hybrid approaches and to estimate the impact of their methodological choices. This work is based on a literature review and on large-sample modelling experiments with hybridizations of two classical models running at different time steps: the monthly GR2M model and the daily GR4J model.
Streamflow forecasting is useful for various purposes, from ensuring the safety of populations during floods to managing hydraulic structures. The aim of this work is to combine two hydrological modelling approaches widely used in streamflow forecasting in order to define their benefits and limits in a probabilistic framework: the multi-model approach (which accounts for structural and parametric model uncertainty) and the semidistributed approach (which considers explicitly the spatial variability of precipitation and hydrological processes). The study focuses on 12 tributaries of the Rhone River, which were modelled using 39 hydrological model configurations. Tests were carried out at an hourly time step for lead times ranging from 1 h to 120 h, considering ensemble meteorological forecasts. The results show that explicitly considering uncertainty with a probabilistic super-ensemble (meteorological ensemble chained to a multi-model approach) improves the quality of streamflow forecasts. On the other hand, there is no clear benefit from a semi-distributed approach compared with a lumped framework. This paper also explored the structure and size of the super-ensemble, showing that it is possible to reduce its complexity through model selection or combination methods without impairing predictive performance. This study provides valuable insights into the strengths and limitations of a super-ensemble approach and how to limit its complexity, contributing to the ongoing efforts to improve streamflow forecasting for operational purposes.
Abstract. One of the most basic questions asked to hydrologists is that of the quantification of catchment response to climatic variations, i.e. that of the variations around the average annual flow given the climatic anomaly of a given year. This paper presents a large sample analysis based on 4122 catchments from four continents, where we investigate how annual streamflow variability depends on climate variables – rainfall and potential evaporation – and on the season when precipitation occurs, i.e. on the synchronicity between precipitation and potential evaporation. We use catchment data to verify the existence of this link, and show that, in all countries and under the main climates represented, synchronicity anomalies come as the second most important factor to explain annual streamflow anomalies: after precipitation but before potential evaporation. Introducing the synchronicity between precipitation and potential evaporation as an independent variable improves the prediction of annual streamflow variability significantly.
Estimation of future streamflows is generally done using rainfall-runoff models to generate streamflow projections based on future climate inputs. Unfortunately, the performance of these models degrades significantly when predicting values outside of their calibration range, which undermines the credibility of projected scenarios. This abstract presents a method to analyze and improve the equations constituting a rainfall-runoff model structure in the context of climate change scenario modelling demonstrated with an application to the GR2M model and 201 catchments in South-East Australia. The method, termed "Data Assimilation Informed model Structure Improvement" (DAISI), enhances a rainfall-runoff model by combining data assimilation with polynomial updates of the state equations. The method is generic and modular, and consistently improves model performance across various metrics, including KGE, NSE on log-transformed flow, and flow duration curve bias. The updated model exhibits higher elasticity of runoff to rainfall, indicating potential significance for climate change simulations. The DAISI diagnostic identifies a reduced number of update configurations in the GR2M structure, with distinct regional patterns in three sub-regions (Western Victoria, central region, and Northern New South Wales). We suggest potential improvements for DAISI, such as incorporating additional observed variables like actual evapotranspiration to better constrain internal model fluxes.
Surface hydrological models usually define their modelling units using topographic catchment boundaries, which are then connected by the river network. Inter-catchment Groundwater Flows (IGFs) are water fluxes that do not respect these topographic boundaries. They can significantly influence river discharge. Therefore, surface hydrological models usually estimate IGFs indirectly, by adjusting the water balance across the topographic catchment. As they cannot be measured directly, the realism of these simulated fluxes can be questioned.Here, we investigate how a model calibration strategy could help to improve the physical realism of simulated IGFs. We propose a multi-objective calibration strategy, where we optimise the model simulation on two fluxes: river discharge and actual evapotranspiration using MODIS satellite estimates. Indeed, we hypothesize that better IGFs could be estimated if the water balance is more constrained by evaporation. We explore different objective functions to identify the most efficient way to use satellite data by looking at the model robustness in time and space.The Seine catchment is characterised by a complex, multi-layered aquifer system where the river loses water in some places and gains water in others. We evaluate the ability of the GRSD model, a semi-distributed hydrological model that implements the lumped GR5J in each subcatchment, to consistently describe this system thanks to this calibration strategy. The influence of four upstream dams is also considered in the modelling, as they have a significant impact on the hydrology. In particular, this work could help to understand the extent to which low flows are maintained naturally by groundwater or artificially by these dams.This work is partly funded by the ANR (CIPRHES project) and by the European Union’s HORIZON Research and Innovation Actions Programme under Grant Agreement No. 101059372 (STARS4Water project).
The transferability of hydrological models over contrasting climate conditions, also identified as model robustness, has been the subject of much research in recent decades. The occasional lack of robustness identified in such models is not only an operational challenge – since it affects the confidence that can be placed in projections of climate change impact – it also hints at possible deficiencies in the structures of these models. This paper presents a large-scale application of the robustness assessment test (RAT) for three hydrological models with different levels of complexity: GR6J, HYPE and MIKE SHE. The dataset comprises 352 catchments located in Denmark, France and Sweden. Our aim is to evaluate how robustness varies over the dataset and between models and whether the lack of robustness can be linked to some hydrological and/or climate characteristics of the catchments (thus providing a clue as to where to focus model improvement efforts). We show that, although the tested models are very different, they encounter similar robustness issues over the dataset. However, models do not necessarily lack robustness in the same catchments and are not sensitive to the same hydrological characteristics. This work highlights the applicability of the RAT regardless of model type and its ability to provide a detailed diagnostic evaluation of model robustness issues.
Over the last decade, large-sample approaches, i.e., based on large catchment sets, have become increasingly popular in hydrological studies. Efforts were made to assemble and disseminate national catchment datasets. This article aims to make a contribution to the construction of a large international database of catchments by proposing the CAMELS-FR dataset, a contribution to the CAMELS (Catchment Attributes and MEteorology for Large-sample Studies) initiative. The first version presented here gathers hydroclimatic data and physical attributes for a set of 654 catchments in France. These catchments cover a wide spectrum of hydroclimatic conditions (from oceanic to continental, mountainous, or Mediterranean conditions) and are considered to have limited human influence. Data include time series of daily streamflow (with at least 30 years over the 1970-2021 period, also aggregated to monthly and yearly time steps) and of 11 catchment-scale daily climate variables (including precipitation, potential evaporation, and air temperature), as well as a total of 255 catchment attributes organized into 10 classes (e.g., geology, soil, land cover). River flow time series were quality-checked. Along with the database itself, two graphical tools are proposed, namely dynamic graphs to visualize time series and graphical fact sheets to summarize the main catchment characteristics. Care was taken to provide as many metadata as possible to help users interpret their results based on this dataset. We intend to update the database regularly to include new available data and account for end users' feedback. CAMELS-FR is available at https://doi.org/10.57745/WH7FJR.
Human activities perturb the large-scale water cycle by withdrawing large amounts of freshwater for agriculture, manufacturing, energy production and drinking water supply, and by operating dams/reservoirs. The risk that human water demand exceeds freshwater availability widely threatens human water security and ecosystem health, in particular in the face of climate change. Therefore, national-scale hydrological models need to integrate representations of human activities to anticipate and address water scarcity and to support the design of adaptation strategies beyond the local scale. However, the lack of detailed observational datasets of human influence at a national scale hinders the development and evaluation of integrated modelling approaches. This study focuses on processing a national observational dataset of human influence for hydrological modelling at the catchment scale in France, where climate change is expected to reduce water resources and increase water demand notably in the sector of irrigation. We collect data of water withdrawal, water release, reservoir operations from a large range of sources. These include national-scale datasets that are typically available at a coarse (annual) temporal resolution only and that are known to have large uncertainties, such as the French national database of quantitative water withdrawals. Covering a large spatial domain and attempting to account for uncertainties, our resulting dataset is a first step toward the development of robust integrated human-water system models at a national scale.