There are several types of generalization ability that we may wish our models to be capable of. All but the most basic of these require the representation to be suitably interpretable so that it can provide meaningful support for scenario analysis, scientific reasoning, and decision making under system non-stationarity and model transfer. However interpretability of a model can only be meaningfully understood in the context of the ‘language’ used for its construction. In this regard it is important to recognize that, while machine-learning-based (MLB) models tend to prioritize accuracy and precision (in service of predictive performance) and physics-based (PB) models tend to emphasize physical/geo-scientific interpretability (in service of understanding), their learned representations are actually based in related but somewhat different languages, levels of linguistic abstraction, and grammatical rules.Importantly, these differences are not fundamentally necessary. It is my opinion that the future of geo-scientific ML need not compromise accuracy and precision to achieve improved understanding. Instead, we must develop “telescopic” hierarchical representations that prioritize “learning from data” at their fundamental levels, while simultaneously enabling “geo-scientific abstraction” so that higher-level interpretable and understandable representations can be extracted by directed compression. Ultimately, the geosciences will benefit from a specific kind of “interpretable generative modeling” that can learn how to construct causal and/or understandable representations of the underlying physical data generating processes from data, and that can facilitate the kind of hierarchical, multi-level abstraction processes alluded to above.
Soil moisture is a key variable for a range of hydrological and ecological processes, yet capturing its small-scale variability and preferential flow phenomena remains challenging. Recent advancements in deep learning have demonstrated potential in predicting hydrological variables, but conventional data-driven models often struggle to represent small-scale variability effectively. In this study, we integrate Long-Short Term Memory (LSTMs) networks and Gaussian Mixture Models (GMMs) to simulate soil moisture dynamics while explicitly quantifying its associated variability. Unlike deterministic approaches, our probabilistic framework accounts for nonlinear relationships between inputs and outputs while modeling the inherent small-scale variability in soil moisture. We apply this methodology to a comprehensive in-situ soil moisture dataset from the Attert experimental basin, where the experimental design incorporates three replicated soil moisture profiles at each location and depth within a 5-meter radius. These replications are fundamental to our probabilistic framework: they provide direct, co-located observations of the natural spread in soil moisture under identical boundary conditions, allowing the model to learn the statistical structure of small-scale variability. This design enables disentangling sensor noise from genuine spatial heterogeneity and provides an empirical basis for training models that capture both temporal dynamics and local-scale variability.. Our results demonstrate that the proposed model reproduces soil moisture dynamics across multiple depths and scales, achieving an average Kling-Gupta Efficiency (KGE) of 0.52, Rank Correlation of 0.72, and Root Mean Squared Error of 0.036 m3m-3, while also capturing the key aspects of small-scale variability and sensor uncertainty. Furthermore, the modeled distributions offer new insights into the spatiotemporal structure of soil moisture and underscore the value of probabilistic modeling in hydrological approaches. By explicitly incorporating small-scale variability into the modeling process, our approach enhances both the interpretability and reliability of soil moisture predictions. While LSTMs effectively capture temporal dynamics, our findings underscore the necessity of incorporating variability quantification to improve model accuracy and generalization. This study highlights the potential of probabilistic deep learning frameworks in hydrological modeling and supports their broader application for improved soil moisture estimation and variability assessment.
ABSTRACT Artificial intelligence (AI) is transforming hydrological and Earth system science. High‐capacity neural networks can extract more information from data than traditional models, leading to major gains in predictive skill. However, this predictive success does not necessarily imply physically credible or uniquely identifiable process representations. In open, data‐limited and epistemically uncertain systems, multiple model structures and parameterisations can reproduce observed behaviour while implying different process explanations—a condition broadly described as equifinality. High‐capacity AI modelling amplifies such challenge into a neural equifinality regime, where over‐parameterised neural‐network architectures expand the space of plausible internal representations that remain consistent with available observations. This creates predictive mirages, where improved aggregate performance may obscure unphysical compensation for data error, model structure or evaluation design. AI can therefore act as a double‐edged tool for model inference: used purely as a predictive engine, its hyper‐flexibility can conceal structural ambiguity and promote overconfidence; used within explicit uncertainty and physical‐consistency frameworks, it may help expose where observations are insufficient to constrain internal process pathways. We call for evaluating AI‐enabled models against dual criteria of performative and structural acceptability, and for constraining state‐flux behaviours using independent evidence. The challenge ahead is to leverage AI's dual predictive and diagnostic roles to make predictive mirages and structural ambiguity more scientifically visible, interpretable and manageable in hydrology and Earth system modelling.
Reliable uncertainty estimation is essential for decision making, evaluating model performance, and defining the limits of what can be inferred from data. While uncertainty estimation typically requires specifying prior assumptions about a distributional form, we introduce an approach to learn the structure of uncertainty directly from data. Specifically, we introduce a variational long short-term memory network (vLSTM) that uses variational inference to enable flexible, non-parametric probabilistic predictions. The vLSTM is assessed against deep learning baseline models for probabilistic rainfall-runoff prediction. We discuss training dynamics of probabilistic models, including concerns of overfitting, and compare predictive strategies that emphasize coverage versus point-wise accuracy. Results demonstrate that the vLSTM achieves state-of-the-art performance when evaluated using log-likelihood, while offering a distinct approach to uncertainty estimation that lets uncertainty patterns emerge instead of prescribing them. In our case study, the learned predictive distributions closely resemble that of the current baseline approach, which prescribes a mixture of asymmetric Laplacian distributions. This finding validates our approach, but also points to its fundamental strength: our variational approach to learning uncertainty structure has the potential to provide a more fundamental understanding of predictive uncertainty in arbitrary types of dynamic models and applications across scientific disciplines, enabling progress especially in fields where a priori assumptions seem hard to justify. In general, the vLSTM serves as a valuable approach for exploring uncertainty structures before transitioning to more computationally efficient models once the emerging patterns of uncertainty are better understood.
Global Climate Models (GCMs), as physically-based models (PBMs), are robust tools for climate-adaptive planning, yet when time or expertise is limited, stochastic modeling can offer a simpler alternative. Few studies have systematically compared stochastic models with PBMs for forecasting precipitation and temperature. This study evaluates ARMA, ARIMA, NSTF, and an improved NSTF variant (INSTF), while testing the effects of Yeo–Johnson (YJ) and Inverse Hyperbolic Sine (IHS) transformations using limited observational data from Lashkenar Village, Iran (2007–2023), and PBM outputs for 2024–2054. The INSTF model introduces quasi-dynamic (induced) noise via Iterative Fourier Series (IFS) to mimic annual cycles and enforce regime-aware noise selection, replacing the pure noise structure used in NSTF. During the simulation period (2019–2054), the standard uncertainty of the mean—a desirable performance metric—for the profound models was nearly zero for ARMA (2,3) and 0.03 for INSTF in annual precipitation modeling, while for annual temperature modeling, it was − 0.08 for both NSTF and INSTF. The YJ transformation generally outperformed IHS, although in some cases, results were better without any transformation. The De Martonne aridity index, which initially indicated a semi-arid climate, declined post-2024 across all models, indicating intensifying dryness. While stochastic models aligned well with PBMs trends in climate change (CC) and can therefore be recommended for CC forecasting, they remain less reliable for real-time precipitation and temperature event forecasting, particularly in design-phase applications.
The Mass-Conserving Perceptron (MCP) establishes a modeling paradigm in which conceptual hydrologic models can be reformulated as physically constrained, conceptually interpretable neural networks. Here, we develop a snow-water MCP network framework and evaluate it across 513 CAMELS-US basins. We first recast a coupled two-state SOIL-MCP and SNOWMCP conceptual model as a mass-conserving neural network and show that the hydrologic-model and neural-network formulations achieve comparable predictive performance. We then examine cross-node state-information sharing within two-state HYDROMCP architectures and evaluate broader single-layer networks constructed from three types of interpretable MCP units with one to five states. Across CONUS, the median KGEss increases from 0.82 for one-state networks to 0.89 for two-state networks and 0.90 for five-state networks, suggesting diminishing aggregate gains beyond two states. Basin-specific MCP and LSTM selection yields the same median KGEss of 0.90, while the selected MCP networks use fewer parameters on average. Complementary AIC- and KGE-based selection identifies compact, basin-specific directed-graph representations that balance predictive accuracy and model complexity. These analyses provide an empirical basis for identifying the numbers, types, and interactions of states needed for hydrologic representation. Future studies should test joint training against multiple hydrologic responses, such as streamflow, snow water equivalent, and groundwater storage.
Modern deep learning can be used not only to improve predictions, but also to uncover interpretable equations that connect observable properties to the parameters of physically based geoscientific models.
Machine learning models can achieve high predictive accuracy in hydrological applications but often lack physical interpretability. The Mass-Conserving Perceptron (MCP) provides a physics-aware artificial intelligence (AI) framework that enforces conservation principles while allowing hydrological process relationships to be learned from data. In this study, we investigate how progressively embedding physically meaningful representations of hydrological processes within a single MCP storage unit improves predictive skill and interpretability in rainfall-runoff modeling. Starting from a minimal MCP formulation, we sequentially introduce bounded soil storage, state-dependent conductivity, variable porosity, infiltration capacity, surface ponding, vertical drainage, and nonlinear water-table dynamics. The resulting hierarchy of process-aware MCP models is evaluated across 15 catchments spanning five hydroclimatic regions of the continental United States using daily streamflow prediction as the target. Results show that progressively augmenting the internal physical structure of the MCP unit generally improves predictive performance. The influence of these process representations is strongly hydroclimate dependent: vertical drainage substantially improves model skill in arid and snow-dominated basins but reduces performance in rainfall-dominated regions, while surface ponding has comparatively small effects. The best-performing MCP configurations approach the predictive skill of a Long Short-Term Memory benchmark while maintaining explicit physical interpretability. These results demonstrate that embedding hydrological process constraints within AI architectures provides a promising pathway toward interpretable and process-aware rainfall-runoff modeling.
Urban drainage models (UDMs) require observation data from sensors for calibration before managing urban floods. However, there is often the case that sensors are not available when model calibration is required, and hence sensor placing and model calibration need to be simultaneously considered in many engineering practices. While many methods are available for either sensor deployment or model calibration with given observation data, there is a lack of approach to consider these two objectives simultaneously. To address this gap, this paper proposes a method that simultaneously enables sensor placement and model calibration using Bayesian decision theory. First, a graph partition strategy is used to ensure overall uniformity in sensor distribution. Subsequently, Bayesian Experimental Design is employed to identify optimal sensor locations by maximizing expected data worth, measured by the relative entropy between prior and posterior probabilities. The UDM is sequentially calibrated using an ensemble smoother algorithm with observation data collected from these strategically placed sensors. Effectiveness and robustness of the method are tested using two real-world UDMs of different scales under various rainfall and parameter scenarios. Comparisons with two empirical approaches, where sensors are evenly deployed or installed in flood hotspots, show that the proposed method provides more accurate water level and flood predictions with less uncertainties across various scenarios. The proposed method is anticipated to be promising in engineering applications as sensor placing and model calibration are often simultaneously needed especially for many new UDMs or existing UDMs without sensors.
Regionalization is an issue that hydrologists have been working on for decades. It is used, for example, when we transfer parameters from one calibrated model to another, or when we identify similarities between gauged to ungauged catchments. However, there is still no unified method that can successfully transfer parameters and identify similarities between different regions while accounting for differences in meteorological forcing, catchment attributes, and hydrological responses. Machine learning (ML) has shown promising results in the generalization of its results at temporal and spatial scales for streamflow prediction. This suggests that ML models have learned useful regionalization relationships that we could extract. This study explores how the HydroLSTM representation, a modification of traditional Long Short-Term Memory, can learn meaningful relationships between meteorological forcing and catchment attributes. One promising feature of the HydroLSTM representation is that the learned patterns can generate different hydrological responses across the US. These findings indicate that we can learn more about regionalization by studying ML models.
Development of environmental models generally requires available data to be split into “development” and “evaluation” subsets. How this is done can significantly affect a model's outputs and performance. However, data splitting is generally done in a subjective, ad-hoc manner, with little justification, raising questions regarding the reliability of the findings of many modelling studies. To address this issue, we present and demonstrate the value of an R-package along with high-level guidelines for implementing many state-of-the-art data splitting methods in order to develop the model in a considered, defensible, consistent, repeatable and transparent fashion, thereby improving the generalizability of the resulting models. Results from two rainfall-runoff case studies show that models with high generalization ability can be achieved even when the available data contain rare, extreme events. Additionally, data splitting methods can be used to explicitly quantify the parameter uncertainty associated with data splitting and the resulting bounds on model predictions.
Finding similarities between model parameters across different catchments has proved to be challenging. Existing approaches struggle due to catchment heterogeneity and non-linear dynamics. In particular, attempts to correlate catchment attributes with hydrological responses have failed due to interdependencies among variables and consequent equifinality. Machine Learning (ML), particularly the Long Short-Term Memory (LSTM) approach, has demonstrated strong predictive and spatial regionalization performance. However, understanding the nature of the regionalization relationships remains difficult. This study proposes a novel approach to partially decouple learning the representation of (a) catchment dynamics by using the HydroLSTM architecture and (b) spatial regionalization relationships by using a Random Forest (RF) clustering approach to learn the relationships between the catchment attributes and dynamics. This coupled approach, called Regional HydroLSTM, learns a representation of "potential streamflow" using a single cell-state, while the output gate corrects it to correspond to the temporal context of the current hydrologic regime. RF clusters mediate the relationship between catchment attributes and dynamics, allowing identification of spatially consistent hydrological regions, thereby providing insight into the factors driving spatial and temporal hydrological variability. Results suggest that by combining complementary architectures, we can enhance the interpretability of regional machine learning models in hydrology, offering a new perspective on the "catchment classification" problem. We conclude that an improved understanding of the underlying nature of hydrologic systems can be achieved by careful design of ML architectures to target the specific things we are seeking to learn from the data.
Machine learning (ML) is increasingly considered the solution to environmental problems where limited or no physico‐chemical process understanding exists. But in supporting high‐stakes decisions, where the ability to explain possible solutions is key to their acceptability and legitimacy, ML can fall short. Here, we develop a method, rooted in formal sensitivity analysis , to uncover the primary drivers behind ML predictions. Unlike many methods for explainable artificial intelligence (XAI), this method (a) accounts for complex multi‐variate distributional properties of data, common in environmental systems, (b) offers a global assessment of the input‐output response surface formed by ML, rather than focusing solely on local regions around existing data points, and (c) is scalable and data‐size independent, ensuring computational efficiency with large data sets. We apply this method to a suite of ML models predicting various water quality variables in a pilot‐scale experimental pit lake. A critical finding is that subtle alterations in the design of some ML models (such as variations in random seed, functional class, hyperparameters, or data splitting) can lead to different interpretations of how outputs depend on inputs. Further, models from different ML families (decision trees, connectionists, or kernels) may focus on different aspects of the information provided by data, despite displaying similar predictive power. Overall, our results underscore the need to assess the explanatory robustness of ML models and advocate for using model ensembles to gain deeper insights into system drivers and improve prediction reliability.
Hydrology is experiencing a shift from process‐based toward deep learning (DL) models. Entity‐aware (EA) DL models with static features (predominantly physiographic proxies) merged to dynamic forcing features show significant performance improvements. However, recent studies challenge the notion that combining dynamic forcings with static attributes make such models entity aware, suggesting that static features are not effectively leveraged for generalization. We examine entity awareness using state‐of‐the‐art Long‐Short Term Memory (LSTM) networks and the CAMELS‐US data set. We compare EA models provided with physiographic static features to ablated variants not provided with static inputs. Findings suggest that the superior performance of EA models is primarily driven by information provided by meteorological data, with limited contributions from physiographic static features, particularly when tested out‐of‐sample. These results challenge previously held assumptions regarding how physiographic proxies contribute to generalization ability in EA Models, highlighting the need for new approaches for robust generalization in DL models.
Due largely to challenges associated with physical interpretability of machine learning (ML) methods, and because model interpretability is key to credibility in management applications, many scientists and practitioners are hesitant to discard traditional physical-conceptual modeling approaches despite their poorer predictive performance. Here, we examine how to develop parsimonious minimally-optimal representations that can facilitate better insight regarding system functioning. The term “minimally-optimal” indicates that the desired outcome can be achieved with the smallest possible effort and resources, while “parsimony” is widely held to support understanding. Accordingly, we suggest that ML-based modeling should use computational units that are inherently physically-interpretable, and explore how generic network architectures comprised of Mass-Conserving-Perceptron can be used to model dynamical systems in a physically-interpretable manner. In the context of spatially-lumped catchment-scale modeling, we find that both physical interpretability and good predictive performance can be achieved using a “distributed-state” network with context-dependent gating and “information-sharing” across nodes. The distributed-state mechanism ensures a sufficient number of temporally-evolving properties of system storage while information-sharing ensures proper synchronization of such properties. The results indicate that MCP-based ML models with only a few layers (up to two) and relativity few physical flow pathways (up to three) can play a significant role in ML-based streamflow modeling.
While many modern studies are dedicated to ML-based large-sample hydrologic modeling, these efforts have not necessarily translated into predictive improvements that are grounded in enhanced physical-conceptual understanding. Here, we report on a CONUS-wide large-sample study (spanning diverse hydro-geo-climatic conditions) using ML-augmented physically-interpretable catchment-scale models of varying complexity based in the Mass-Conserving Perceptron (MCP). Results were evaluated using attribute masks such as snow regime, forest cover, and climate zone. Our results indicate the importance of selecting model architectures of appropriate model complexity based on how process dominance varies with hydrological regime. Benchmark comparisons show that physically-interpretable mass-conserving MCP-based models can achieve performance comparable to data-based models based in the Long Short-Term Memory network (LSTM) architecture. Overall, this study highlights the potential of a theory-informed, physically grounded approach to large-sample hydrology, with emphasis on mechanistic understanding and the development of parsimonious and interpretable model architectures, thereby laying the foundation for future models of everywhere that architecturally encode information about spatially- and temporally-varying process dominance.
Interpreting complex datasets remains a major challenge for scientists, particularly due to high dimensionality and collinearity among variables. We introduce a novel application of Kolmogorov-Arnold Networks (KANs) to enhance interpretability and parsimony beyond what traditional correlation analyses offer. We present two interpretable, color-coded visualization tools: the Pairwise KAN Matrix (PKAN) and the Multivariate KAN Contribution Matrix (MKAN). PKAN characterizes nonlinear associations between pairs of variables, while MKAN serves as a nonlinear feature-ranking tool that quantifies the relative contributions of inputs in predicting a target variable. These tools support pre-processing (e.g., feature selection, redundancy analysis) and post-processing (e.g., model explanation, physical insights) in model development workflows. Through experimental comparisons, we demonstrate that PKAN and MKAN yield more robust and informative results than Pearson Correlation and Mutual Information. By capturing the strength and functional forms of relationships, these matrices facilitate the discovery of hidden physical patterns and promote domain-informed model development.
Accurate, high-resolution spatiotemporal estimates of rainfall intensity (RI) are essential for effective prevention and control of urban flooding. Traditional methods are often costly and provide inadequate coverage. Meanwhile, the associated uncertainty of RI estimates is often overlooked, potentially leading to poor decisions in management of urban floods. Here we examine the potential of video imagery recorded by surveillance cameras in urban areas, for providing real-time estimates of RI. Specifically, we propose the use of Bayesian deep learning (DL) to estimate RI and its related uncertainty in a real-time manner. Our DL approach combines the strengths of convolutional neural network (CNN) and long short-term memory (LSTM) to construct a suitable model for this task, and uses variational inference to quantify the uncertainty of the CNN-LSTM model. The proposed approach is tested using video imagery captured under various light-intensity conditions, including daytime, nighttime, early morning, and early evening, and evaluated via experiments using both random samples and independent rainfall events. Further, we show that the temporal processing provided by the LSTM network is important to achieve good performance of RI estimation. The results indicate the strong potential for leveraging urban camera networks to obtain high-precision spatiotemporal RI estimates at low cost, which can be extremely valuable for implementing effective flood control measures and planning emergency responses in the urban areas.
Access to accurate spatio-temporal groundwater level data is crucial for sustainable water management in Chile. Despite this importance, a lack of unified, quality-controlled datasets have hindered large-scale groundwater studies. Our objective was to establish a comprehensive, reliable nationwide groundwater dataset. We curated over 120,000 records from 640 wells, spanning 1970-2021, provided by the General Water Resources Directorate. One notable enhancement to our dataset is the incorporation of elevation data. This addition allows for a more comprehensive estimation of groundwater elevation. Rigorous data quality analysis was executed through a classification scheme applied to raw groundwater level records. This resource is invaluable for researchers, decision-makers, and stakeholders, offering insights into groundwater trends to support informed, sustainable water management. Our study bridges a crucial gap by providing a dependable dataset for expansive studies, aiding water management strategies in Chile.
Explainable Artificial Intelligence (XAI) offers the promise of being able to provide additional insight into complex hydrological problems. As the “new kid on the block”, these methods are embraced enthusiastically and often viewed as offering something radically new and different. However, upon closer inspection, many XAI approaches are very similar to more “traditional” methods of “interrogating” existing models, such as sensitivity or break-even analysis. In fact, the approach of developing data-driven models to obtain a better understanding of hydrological processes to inform the development of more physics-based models is as old as hydrology itself. Consequently, rather than being considered a new approach, XAI should be viewed as part of a long-standing tradition, and XAI methods part of an ever-expanding hydrological modelling toolkit, rather than a silver bullet. Critically, there needs to be shift from focusing on how to best eXplain what AI models have learnt (i.e., the X component of XAI) to developing models that are able to capture relationships that are contained within the data in a robust and reliable fashion (i.e., the AI component of XAI), as there is little value in explaining AI-derived relationships if these do not reflect underlying hydrological processes. However, this is often not the case due to a focus on maximising the predictive ability of AI models “at all costs”, not uncommonly resulting in large models that often have thousands or even millions of parameters that are not well defined. Consequently, these models generally do not capture underlying hydrological processes in a robust and reliable fashion. Finally, there is also a need to stop thinking about XAI as a purely technical approach, but a socio-technical approach that views XAI as a process that can assist with solving problems that are situated within broader social and political contexts.