To improve hydrological uncertainty estimation, recent studies have explored machine learning (ML)-based post-processing approaches that enable both enhanced predictive performance and hydrologically informed probabilistic streamflow predictions. Among these, random forests (RF) and their probabilistic extension, quantile random forests (QRF), are increasingly used for their balance between interpretability and performance. However, the application of QRF in regional post-processing settings remains unexplored. In this study, we develop a hydrologically informed QRF post-processor trained in a multi-site setting and compare its performance against a locally (at-site) trained QRF using probabilistic evaluation metrics. The QRF framework leverages simulations and state variables from the GR6J process-based hydrological model, along with readily available catchment descriptors, to predict daily streamflow uncertainty. Our results show that the regional QRF approach is beneficial for hydrological uncertainty estimation, particularly in catchments where local information is insufficient. The findings highlight that multi-site learning enables effective information transfer across hydrologically similar catchments and is especially advantageous for high-flow events. However, the selection of appropriate catchment descriptors is critical to achieving these benefits.
Abstract. Reliable transferability of hydrological model parameters across time and space remains a challenge for large‑scale water resources assessment. In this study, we investigate whether a differentiable hybrid framework can identify robust and physically coherent parameter sets for annual streamflow modeling across a large‑sample dataset of 3,044 catchments from eight countries. To focus on temporal and spatial transferability analysis, we work at the annual time scale using what we consider to be the simplest possible model: an annual anomaly model of climate elasticities, coupled with the Turc–Mezentsev formulation for the long-term streamflow mean (MQ). A dense neural network is trained in an end‑to‑end fashion to map catchment descriptors to the four model parameters, with gradients propagated through the entire modeling chain. We evaluate the framework using three cross‑validation settings inspired by Klemeš (1986): temporal, spatial, and combined temporal–spatial cross-validation. As a benchmark, we compare the hybrid model against local, catchment‑by‑catchment linear regressions under temporal cross-validation. Our results show that, for temporal transferability, our parameter learning approach outperforms local calibration, yielding higher Nash-Sutcliffe efficiency (NSE) values while producing elasticity coefficients that remain within plausible physical ranges, despite lacking explicit parameter constraints. By contrast, spatial transferability reveals a marked limitation: the anomaly component extrapolates well spatially, but regionalizing MQ from descriptors proves difficult, with MQ errors dominating the loss of performance in spatial and spatiotemporal cross-validation. Experiments with random descriptors further show that our parameter learning uses attributes mainly as catchment identifiers in temporal cross-validation but relies on their physical content to sustain spatial transfer, particularly for MQ. Overall, the study demonstrates that simple differentiable hybrid annual models can learn robust and interpretable anomaly parameters, while highlighting MQ regionalization as the main remaining bottleneck for spatially transferable annual streamflow predictions.
Improving droughts forecasting - whether meteorological, agricultural, or hydrological - is a major challenge for the protection of natural ecosystems and for many economic sectors, including agriculture, energy production, drinking water supply, navigation, and tourism. To provide public water managers with robust low-flow forecasting tools in a context of climate change, the French Office for Biodiversity (OFB) and the Water and Biodiversity Direction (DEB) have supported, since 2011, an initiative aimed at developing a national operational low-flow forecasting platform. This platform, known as PREMHYCE, is the result of a long-term scientific and technical collaboration between INRAE, Météo-France, the University of Lorraine, BRGM, and EDF (Tilmant et al., 2023).PREMHYCE relies on five hydrological models and ensembles of meteorological scenarios to produce probabilistic streamflow forecasts, enabling the estimation of risks of falling below low-flow thresholds (typically vigilance, alert, reinforced alert, or crisis levels). Forecast lead times range from a few days to several weeks, depending on management objectives and catchments considered. The platform provides daily streamflow forecasts at more than 1,300 gauging stations across the French hydrographic network, with lead times of up to 90 days. These forecasts are made available to more than fifty operational services across mainland France and Réunion Island. They are used to anticipate low-flow periods within local and national decision-making bodies.In recent years, the PREMHYCE platform has evolved and been upgraded as part of a research project (ANR CIPRHES, 2021–2025), including developments in meteorological forecasting, hydrological modelling, uncertainty quantification, and improvements of the user interface in close collaboration with end users.This communication aims to present the PREMHYCE forecasting chain, its main functionalities, its range of applications, and its recent developments. Key words: low-flow forecasting, water management, hydrological modelling Reference: Tilmant, F., Bourgin, F., François, D., Le Lay, M., Perrin, C., Rousset, F., Vergnes, J.-P., Willemet, J.-M., Magand, C., and Morel, M. (2023). - PREMHYCE, une plateforme nationale pour la prévision des étiages. Sciences Eaux & Territoires. 42, 17–21, https://doi.org/10.20870/Revue-SET.2023.42.7297. Acknowledgements:This work was financially supported by the French National Research Agency (ANR) (grant ANR-20-CE04-0009) within the CIPRHES project, by the French Office for Biodiversity (OFB) and by the Water and Biodiversity Direction (DEB, at the Ministry for ecology).
The use of machine learning (ML) methods in rainfall-runoff modelling has apparently led to better prediction, but there are some concerns about the interpretability of these models. The emergence of hybrid modelling, which couples the data driven approach with the classical physics-based conceptual approach, has shown promise in enhancing both interpretability and accuracy. ML models and conceptual models each come with their own modelling practices and habits. To develop a hybrid approach, it is necessary to consider them. While some of the steps in these modelling chains are similar (for instance the selection of the right metric during the calibration or learning step), others are more specifics, such as the optimization of the hyper-parameters of ML models. Furthermore, the hybrid approach comes with specific methodological challenges that emerge when coupling the two different types of models. For instance, depending on the choice made by the modeller, the parameters of the conceptual model are either trained with the ML model parameters or calibrated separately by a non-ML method. There is a need to better understand the variety of hybrid approaches and to estimate the impact of their methodological choices. This work is based on a literature review and on large-sample modelling experiments with hybridizations of two classical models running at different time steps: the monthly GR2M model and the daily GR4J model.
An hourly hydrological forecasting model (GRP), used for flood forecasting in France, has been enhanced by developing a semi-distributed version (GRPS). This new model addresses some limitations of the original lumped approach by integrating flow observations from upstream stations to improve downstream flood predictions. A comparison of GRP and GRPS was conducted on a large set of flood events in nested catchments in France. Results indicate that GRPS slightly outperforms GRP up to the lead time of the last upstream flow observations reaching the downstream station, though performance gains vary widely across events. Event classification identified detrimental interactions between data assimilation methods and river routing model errors. A sensitivity test suggested that enhancing both the propagation model and the assimilation scheme could reduce poor performance at short lead times for some events. The study also demonstrated that upstream forecast quality significantly impacts downstream forecast accuracy. Overall, the findings confirm the benefits of incorporating multiple flow observations into a semi-distributed model and suggest several avenues for further improvement. Un mod & egrave;le de pr & eacute;vision hydrologique horaire (GRP), utilis & eacute; pour pr & eacute;dire les crues en France, a & eacute;t & eacute; adapt & eacute; en une version semi-distribu & eacute;e (GRPS), pour r & eacute;pondre & agrave; certaines limites de l'approche originale en assimilant les observations de d & eacute;bit des stations en amont pour am & eacute;liorer les pr & eacute;visions en aval. Une comparaison entre GRP et GRPS a & eacute;t & eacute; effectu & eacute;e sur un grand & eacute;chantillon d'& eacute;v & eacute;nements de crue dans des bassins versants embo & icirc;t & eacute;s fran & ccedil;ais. Les r & eacute;sultats indiquent que GRPS surpasse en moyenne l & eacute;g & egrave;rement GRP jusqu'au moment o & ugrave; les derni & egrave;res observations de d & eacute;bit en amont atteignent la station en aval, malgr & eacute; une grande disparit & eacute; d'un & eacute;v & eacute;nement & agrave; l'autre. Une classification des & eacute;v & eacute;nements a permis d'identifier des interactions n & eacute;fastes entre les m & eacute;thodes d'assimilation de donn & eacute;es et les erreurs du mod & egrave;le de routage. Un test de sensibilit & eacute; a sugg & eacute;r & eacute; que l'am & eacute;lioration du mod & egrave;le de propagation et du sch & eacute;ma d'assimilation pourrait r & eacute;duire les mauvaises performances & agrave; court terme pour certains & eacute;v & eacute;nements. L'& eacute;tude a & eacute;galement d & eacute;montr & eacute; que la qualit & eacute; des pr & eacute;visions en amont a un impact significatif sur la pr & eacute;cision des pr & eacute;visions en aval. Cette & eacute;tude confirme l'int & eacute;r & ecirc;t de l'assimilation de multiples observations de d & eacute;bit dans un mod & egrave;le semi-distribu & eacute; et propose quelques pistes pour son am & eacute;lioration.
The evaluation of streamflow predictions forms an essential part of most hydrological modelling studies published in the literature. The evaluation process typically involves the computation of some evaluation metrics, but it can also involve the preliminary processing of the predictions as well as the subsequent processing of the computed metrics. In order for published hydrological studies to be reproducible, these steps need to be carefully documented by the authors. The availability of a single tool performing all of these tasks would simplify not only the documentation by the authors but also the reproducibility by the readers. However, this requires such a tool to be polyglot (i.e. usable in a variety of programming languages) and openly accessible so that it can be used by everyone in the hydrological community. To this end, we developed a new tool named evalhyd that offers metrics and functionalities for the evaluation of deterministic and probabilistic streamflow predictions. It is open source, and it can be used in Python, in R, in C++, or as a command line tool. This article describes the tool and illustrates its functionalities using Global Flood Awareness System (GloFAS) reforecasts over France as an example data set.
We compared the flood forecasts issued by a model used by operational services in France (GRP) and by a model developed to improve the simulation of floods resulting from intense rainfall (GR5H_RI). We selected 10,652 flood events from 19 years of hourly data available for 229 French catchments. The models were combined with a state-updating procedure to produce forecasts at 3, 6, 12 and 24 h lead times. Results indicate that the GR5H_RI model performs better on average than the GRP model at all lead times, particularly for forecasting flash floods (rise time < 12 h), which occur mainly in summer and early autumn. The use of the last observed streamflow to update initial conditions does not compensate for GRP's structural errors in the case of fast catchment response to intense rainfall. The new structure therefore opens valuable operational perspectives.
The hydrological models used for flood forecasting purposes are usually first calibrated in simulation mode and then applied for flood forecasting in combination with some data assimilation procedure allowing to update and correct the model in real-time. Such a two-step procedure may not be the best choice to train model parameters for forecasting purposes. An alternative approach is to calibrate the hydrological model separately for each target lead time in the presence of the updating procedure. However, it is unclear whether this approach can provide the most efficient and the most robust forecasts. We compared in this paper three approaches to calibrate the parameters of a flood forecasting model: calibration in simulation mode (i.e. without data assimilation), lead-time-dependent calibration (i.e. with data assimilation), and a procedure combining both approaches. An hourly flood forecasting model (a parsimonious hydrological model combined with a state updating assimilation procedure) was used to produce forecasts for three lead times on 687 catchments and 40,411 flood events in metropolitan France. The performance of the model calibrated with the three approaches was evaluated according to catchment response times and flood rise times. An analysis of parameter robustness was also carried out. The results showed that lead-time-dependent calibration improves performance for catchments characterised by slow temporal dynamics. However, for catchments with fast temporal dynamics, it leads to degraded performance at short lead times and reduces parameter robustness. We found that the combined calibration approach is the best compromise between parameter robustness and performance at all lead times and for all catchment and flood types.
<p>When they are used for operational forecasting, hydrological models are almost always combined with some kind of updating procedures. Then a question arises: should the model parameters be calibrated with or without the updating procedures? Calibrating with the updating procedures often improves forecast efficiency, but it can also lead to parameter inconsistency and ultimately to a drop in performance in some cases.</p> <p>In this study, we evaluate the pros and cons of making the parameters of a flood forecasting model vary with lead times. We investigate the dependencies of the model parameters to the lead times and determine where and when this procedure significantly improves forecast quality. A modified version of the GR5H hydrological model is used on 229 French catchments where 10,652 events were selected. The model is run at the hourly time step and combined with a simple updating procedure to produce forecasts at four lead times. The model parameters were estimated from a large screening of the parameter space (3 million runs for each catchment). Results show that the parameters related to fast catchment processes are the most dependant on lead times, indicating the need for more specific parameter estimation methods when modelling catchments prone to flash floods.</p>
Machine learning models have recently gained popularity in hydrological modelling at the catchment scale, fuelled by the increasing availability of large-sample data sets and the increasing accessibility of deep learning frameworks, computing environments, and open-source tools. In particular, several large-sample studies at daily and monthly time scales across the globe showed successful applications of the LSTM architecture as a regional model learning of the hydrological behaviour at the catchment scale. Yet, a deeper understanding of how machine learning models close the water balance and how they deal with inter-catchment groundwater flows is needed to move towards better process understanding. We investigate the performance and behaviour of the LSTM architecture at a monthly time step on a large sample French data set coined CHAMEAU – following the CAMELS initiative. To provide additional information to the learning step of the LSTM, we use the parameter sets and fluxes from the conceptual GR2M model that has a dedicated formulation to deal with inter-catchment groundwater flows. We see this study as a contribution towards the development of hybrid hydrological models.
De nombreuses activités humaines peuvent être fortement impactées par les pénuries d'eau (production d’eau potable, irrigation, production électrique, navigation fluviale, loisirs, etc.). Il est donc nécessaire d'anticiper les périodes d'étiage afin d'améliorer la gestion de l'eau pour mieux répondre aux besoins, tout en préservant le fonctionnement des écosystèmes aquatiques. Ceci est renforcé par la perspective d'étiages futurs plus sévères dans le contexte du changement climatique en cours. Pour répondre à ces enjeux, cinq institutions françaises ont développé un outil opérationnel de prévision des bas débits, PREMHYCE. Issu d’un projet de recherche initié en 2011 et soutenu par l’Office français de la biodiversité et la Direction de l’eau et de la biodiversité, il est testé depuis 2018 en temps réel sur plus de mille bassins versants en France métropolitaine et à La Réunion. PREMHYCE comprend cinq modèles hydrologiques qui sont implémentés sur des bassins versants jaugés et assimilent les dernières observations de débits en temps réel. Les prévisions de débits sont émises quotidiennement en considérant des ensembles de scénarios de pluies et de températures (issus d’un modèle météorologique jusqu’à une échéance de 15 jours et d’archives climatiques jusqu’à une échéance de 90 jours). Les résultats de la plateforme sont communiqués aux gestionnaires opérationnels partenaires, à l’aide de supports de communication adaptés, pour les aider dans la prise de décision. Une interface web de visualisation donne une vue d’ensemble en temps réel de la situation observée et prévue.