
ABSTRACT Water losses in Water Distribution Networks remain a persistent challenge, with global impacts exceeding $39 billion annually. While Artificial Intelligence (AI) offers promising leak detection solutions, assessing operational readiness and reproducibility is difficult due to literature fragmented across sensing technologies and spatial scales. This systematic review of 86 studies (2015–2026) structures the field through a hierarchical hydroinformatics framework: network-level monitoring, district metered area analysis, and In-situ localization. Each level relies on distinct sensing modalities conditioning feature engineering and model selection. Crucially, the analysis reveals that the primary bottleneck preventing real-world deployment is not algorithmic sophistication, but the prevalence of isolated computational approaches that ignore physical and operational constraints. To bridge the gap between simulated validation and operational utility, this review argues for conceptualizing AI as a sensor-aware, adaptive hierarchical mechanism. Furthermore, it exposes the urgent need for a paradigm shift in evaluation standards by moving beyond traditional metrics to explicitly incorporate operational False Alarm Rates, computational feasibility, and integration with operational systems (SCADA/GIS). By aligning algorithmic evaluation with infrastructural realities, this critical assessment establishes how AI-based approaches must evolve to become a viable tool for water loss reduction.
ABSTRACT With the acceleration of urbanization, urban water consumption prediction has become increasingly critical for refined water resource management. However, existing forecasting methods face three key challenges: insufficient systematic integration of multi-factor influences on water consumption, inadequate exploration of complex temporal structures in water use sequences, and limited capability in capturing local fluctuations in water consumption. To address these gaps, this paper proposes a hybrid forecasting model that integrates Convolutional Neural Networks (CNN) with Patch-based Time Series Transformer (PatchTST), incorporating feature selection and multi-scale decomposition. Specifically, Pearson Correlation Coefficient and Maximal Information Coefficient are combined to screen influential variables, and the water consumption time series is decomposed into trend, seasonal, and residual components. The CNN is embedded into the PatchTST architecture to enhance the extraction of short-term fluctuation features through its local perception capability. Using water consumption data from a residential community in a northern Chinese city, comparative and ablation experiments are conducted. Experimental results demonstrate that the proposed model achieves optimal performance across all metrics, with reductions of 48.5, 61.7, and 56.0% in MSE, MAE, and MAPE, respectively, providing effective technical support for urban water resource management.
ABSTRACT A novel methodology for the optimal aggregation of water demands and leakages in the modelling of intermittent water distribution networks (WDNs) is presented in this paper. The methodology uses the K-means clustering algorithm for performing aggregation based on the spatial data of the network nodes. The optimal aggregation level is evaluated considering both the hydraulic accuracy and the required computational effort of the model. The methodology is applied to a real intermittent WDN located in southern Italy. For the application, a calibrated hydraulic model of the network was used to describe the water demands of users equipped with private tanks, and leakages associated with nodes of the water distribution network. Results show that a suitable balance between accuracy and computational effort of the model can be achieved for levels of aggregation in the range of 0.5–1.0 tanks per km of network. The proposed methodology offers a practical approach for simplifying the modelling of intermittent water distribution networks, without compromising simulation accuracy. The potential for transfer of the methodology and of the obtained results to other networks with different sizes, layouts, and intermittent supply schemes requires further research.
ABSTRACT The detection of water pipe leaks using machine learning to identify ground-based acoustic signals has long been a hot topic in the field of water supply safety. However, achieving reliable leak detection in complex real scenarios remains a significant challenge. Based on an acoustic signal dataset collected from an actual pipeline network, this study proposes a sliding window traversal method with intra-window normalization (SWT-IWN) for data preprocessing, which effectively enhances the separability of acoustic features between leakage and non-leakage pipe segments. On this basis, a dataset fused with spatial position information of sampling points is constructed. Furthermore, a leakage identification model named CST-Net is designed, which uses a convolutional neural network (CNN) to extract time-frequency features and a Swin-Transformer to model the positional sequence correlation of sampling points, thereby realizing collaborative representation of the two feature types. Experimental results show that CST-Net accurately identifies leakage pipe segments of 2.5 m in length, with an identification accuracy of 92.35%. The model also demonstrates strong identification performance and stability under diverse sampling conditions, well meeting the requirements of practical engineering applications.
ABSTRACT Cascade reservoirs reshape phytoplankton interaction networks through enhanced environmental filtering, with partial downstream recovery along physicochemical and habitat gradients. Phytoplankton are key primary producers in river ecosystems, yet their responses to cascade reservoir development remain insufficiently understood. In this study, phytoplankton and water samples were collected from 2021 to 2023 along the Zhimenda–Zhutuo reach in the upper Yangtze River to investigate spatial variations in community structure and assembly mechanisms. Phytoplankton communities exhibited pronounced spatial heterogeneity, with cell abundance, species richness, and diversity following the pattern S3 (downstream recovery section) > S2 (cascade reservoir–impacted section) > S1 (natural river section). Community assembly mechanisms also differed among river sections: stochastic processes dominated in S1 and S3, whereas deterministic processes prevailed in S2. Co-occurrence network analysis revealed a relatively simple community structure in S1, while stronger intra- and interspecific interactions were observed in S2. In S3, community stability gradually increased with distance from the final reservoir but remained slightly lower than that in S1. Partial least squares path modeling further indicated that flow velocity in S2 played a key role in regulating phytoplankton community stability. Overall, these findings provide new insights into how cascade reservoir development reshapes phytoplankton community structure, assembly processes, and stability in large river systems.
Urban surface-flow modelling is highly sensitive to the representativeness, resolution, and conditioning of geospatial inputs, especially in heterogeneous catchments containing engineered barriers and drainage pathways. This study presents an automated, auditable workflow that transforms routinely available elevation, land-use, infrastructure, and design-rainfall datasets into Hydrological Analysis-Ready Data (H-ARD) for high-resolution urban applications. The workflow comprises five stages - Acquisition, Preparation, Enrichment, Processing, and Validation - and produces versioned outputs, audit artefacts, and machine-readable run logs. It was applied to four Norwegian 1 m LiDAR DEM tiles (Bergen, Oslo, Stavanger, and Hamar), each subdivided into urban and rural zones, to quantify how conditioning alters drainage connectivity, stream-network topology, land-use thematic detail, runoff-coefficient fields, and the representativeness of nearest-station versus gridded IDF rainfall assignments. Results show that conditioning materially changes inferred hydrological connectivity in anthropogenically modified areas, refines land-surface parameter fields, and reveals location-dependent disagreement between precipitation assignment approaches. The contribution is methodological: a reproducible workflow and a standardised set of input-impact diagnostics for evaluating conditioned versus raw geospatial products. Hydrological model outputs are not evaluated in this study; any claim of improved discharge, water level, inundation, or flood-risk performance requires separate validation against observations.
Understanding and modeling uplift pressures in hydraulic structures is crucial for structural safety, design optimization, and retrofitting. Despite its importance, research on uplift generated by high-velocity unidirectional flow over offset cracks or joints remains limited. This study establishes accurate uplift models for this particular problem using optimized explainable machine learning techniques, complemented by computational fluid dynamics methods. The models, developed using 558 laboratory experiments, demonstrate high predictive accuracy for both calibration and validation sets. For example, during validation, the models exhibit a mean coefficient of determination (R-2) of 0.99, a root mean square error of 0.02, and a mean absolute error of 0.01. The dominant influencing factors for uplift are the gap width-offset height ratio and the relative offset height, which exhibit negative and positive correlations with uplift, respectively. The proposed methodology is also applied to a prototype flood tunnel, yielding satisfactory predictions, with R-2 = 0.99 and mean error = 6.5%. This study provides an enhanced uplift modeling approach that ultimately contributes to the resilience and sustainability of hydraulic structures.
Flood-related disasters have become increasingly frequent in recent years, as evidenced by severe events in Brazil, Spain, and Germany. The development of impact-based flood early warning systems (IBFWS) has been an essential tool for minimizing human and economic losses. Recent advances in data-driven models show strong potential for improving flood forecasting capabilities. More promising approach that has emerged is the use of physics-enhanced machine learning models. These models incorporate physical and hydrological concepts into data-driven frameworks, which enhance their interpretability and robustness. This paper proposes a physics-enhanced Long Short-Term Memory (LSTM) model to incorporate inter-station lag times into the model's feature selection and temporal configuration, improving flood forecasts. The framework is applied to a flood-prone urban basin using high-resolution (10-minute) rainfall and streamflow data, assessing both overall forecast skill and the accuracy of flood events, particularly the peak magnitude and timing errors. Results demonstrate that the physics-enhanced configuration consistently increases prediction accuracy by reducing redundancy among inputs. Moreover, it maintains the physical coherence of the hydrological processes, supporting the transition from black-box to grey-box modeling. The resulting architecture remained computationally efficient, highlighting the potential of physics-enhanced neural networks for operational and impact-based flood forecasting.
Inland water-quality monitoring systems increasingly generate large volumes of environmental data, creating opportunities for advanced analytical methods to identify subtle pollution signals that may not be captured by traditional threshold-based monitoring approaches. This study proposes an interpretable unsupervised machine-learning framework for detecting anomalies in inland water-quality monitoring data (consisting of more than 17,000 observations across 23 states, curated by the Central Pollution Control Board (2021), from which a processed subset was used for model evaluation) using multiple complementary detection models and explainable AI techniques. The framework uses the four detection models, which are Isolation Forest, One-Class Support Vector Machine, Elliptic Envelope, and Autoencoder and is optimised for ecological feature engineering (dissolved oxygen deficit, BOD/DO ratio, coliform load index) to increase the sensitivity to complex pollution stressors. SHAP-based explanations, t-SNE projections and statistical comparisons were utilised to ensure interpretability and strong validation of anomalies. The results show that about 7.88 and 8% of observations were anomalous, with peri-urban tanks in Karnataka and Uttar Pradesh being identified to have hotspots with characteristics of oxygen depletion and microbial contamination, especially during post-monsoon seasons. The ensemble was able to identify all domain-specific threshold violations with the best performance of Isolation Forest and Autoencoder (F1 > 0.70).
Accurate and robust soil moisture estimation remains a major challenge in precision agriculture, particularly under increasing climate change impacts. Reliable assessment of net soil moisture using precipitation and evapotranspiration (ET) derived from multisensor remote sensing is therefore essential. However, most existing approaches rely on single-sensor products or complex data assimilation frameworks, limiting their operational applicability. This study addresses this gap by proposing a simplified interoperable framework, integrating Climate Hazards Group Infrared Precipitation with Station Data (CHIRPS) precipitation, Moderate Resolution Imaging Spectroradiometer (MODIS) ET, and Soil Moisture Active Passive (SMAP) soil moisture data. Differences between CHIRPS precipitation and MODIS ET datasets were computed at country, state and subdivision scales and validated against SMAP observations. The correlations between the derived differences and SMAP soil moisture yielded R-2 values of 0.66, 0.69, and 0.62, at country, state and subdivision levels, respectively. Furthermore, CHIRPS, MODIS, and SMAP datasets at the subdivision level were used to train a Support Vector Regressor model for the period 2021-2023, at 8-day temporal resolution, with estimations performed for June 2023. The linear kernel SVR achieved superior performance with R-2 values of 0.81 and 0.792 during training and testing at the subdivision level. Validation with ground observations gave an R-2 of 0.7123, confirming robust sensor interoperability. Overall, the study demonstrates efficient integration of CHIPRS, MODIS, and SMAP datasets.
Taiwan is located in the western North Pacific, where typhoon-induced heavy rainfall frequently challenges disaster prevention and decision-making. To improve short-term (1-6 h) rainfall forecasting during typhoon landfall, this study developed a deep learning framework for eastern Taiwan using historical typhoon records and ground-based meteorological observations. Six models were compared: Transformer Encoder model (TransEnc), LSTM, GRU, attention-enhanced LSTM (MSA-LSTM), attention-enhanced GRU (MSA-GRU), and an application-oriented hybrid Encode-Decode-Attention-Recurrent (EDAR) framework. At t + 1, all models broadly captured the main rainfall evolution. As forecast horizon increased, RMSE and MAE generally rose, whereas NSE and correlation declined, indicating recursive error accumulation. MSA-GRU showed the strongest overall multi-step performance, with the highest average NSE and correlation and significantly lower RMSE than TransEnc, LSTM, GRU, and MSA-LSTM (p < 0.05). EDAR achieved the lowest mean MAE and remained competitive in the supplementary Keelung analysis. Hualien served as the primary development site, and a supplementary analysis was also conducted at Keelung. Overall, attention-enhanced recurrent models provided more stable multi-step forecasts, with potential for real-time typhoon heavy-rainfall early warning.
Combined sewer overflows (CSOs) remain a major challenge for urban water management, with risks increasing under more frequent, intense rainfall. This study developed a random forest (RF) classification framework to forecast CSO occurrence in real time at two Quebec City outfalls (U051 and U12B). The model was trained on 5-min rainfall data from nearby rain gauges and water level data measurements at each outfall. Predictors were engineered to represent both short-term impact and antecedent conditions. The RF models achieved strong event-detection performance, with F1-scores of 0.92 (U051) and 0.91 (U12B) for the CSO occurrence class, with corresponding CSOclass precision of 0.94 and 0.91 and recall of 0.91 and 0.92, respectively. Feature importance results showed that lagged water level and recent rainfall intensity/accumulation metrics were the most influential predictors, with differences in dominant drivers between the two outfalls. Compared with traditional hydrodynamic modelling, the proposed RF approach provides a computationally lightweight alternative that does not require detailed network geometry or extensive calibration, making it easier to deploy and update in operational settings. Compared with deep learning approaches, RF offers faster training, robust performance with moderate datasets, and transparent interpretability through feature importance, enabling practical decision support for real-time CSO warning.
This study presents a scenario-based framework for rainfall-runoff modelling that evaluates classical machine learning, tree-based ensembles, and deep learning architectures across distinct flood types in a montane basin. Using 15 years of hourly hydrometeorological and sensor-network observations from the Upper Vydra Basin (2008-2023), we assessed eight models and an equal-weight ensemble to link predictive skill to flood-generating processes. Model performance varied across six flood typologies. The extended LSTM achieved the highest individual accuracy for long-duration and multi-peak events, whereas Random Forest and XGBoost were most effective for short-duration floods under contrasting antecedent wetness. Transformer models showed systematic overprediction, and support vector regression performed weakest. An equal-weight ensemble combining eight ML/DL architectures provided the most robust overall performance, with the highest prediction accuracy (NSE = 0.955) and reduced error variance. SHAP analysis highlighted the dominant influence of precipitation and snowmelt and the value of distributed water-level observations for representing catchment storage dynamics. These findings show that predictive skill depends strongly on the match between model architecture and hydrometeorological context. Scenario-based evaluation combined with ensemble integration offers a practical pathway toward reliable and operationally efficient flood forecasting in montane environments. HIGHLIGHTS center dot Scenario-based ML/DL framework evaluates six flood types in a montane basin. center dot Sensor network and ERA5-Land data enabled peakflow modelling in data-sparse basins. center dot SHAP confirms precipitation, snowmelt and tributary levels as key predictors. center dot xLSTM excels in long floods; tree-based models in short events. center dot Ensembles reduce variance and perform best for snow-influenced floods.
Main-channel migration reflects the stability and adjustment of braided rivers, yet its intra-annual spatiotemporal variation remains poorly quantified in the braided reach of the Lower Yellow River (LYR). Using Sentinel-2 imagery from 2018 to 2022, this study developed a method based on the Euclidean Distance Transform to extract the main-channel centerline and quantify its migration width in a multi-channel river. Results confirmed that main-channel migration was highly episodic and mainly concentrated during the flood season, with additional adjustment in some reaches after the flood period. Spatially, the most active migration was concentrated from Mayugou to Guanzhuangyu and from Heishi to Youfangzhai, where the channel was relatively wide and multi-threaded. Finally, an empirical model was established to predict reach-scale migration width from mean discharge, incoming sediment coefficient, and interval duration. Calibration with 2018-2021 data and validation with 2022 data yielded relative root mean square errors of 23% and 36%, respectively. This study provides a remote-sensing-based framework for quantifying short-term planform adjustment of braided rivers and improves a basis for river regulation and channel-stability assessment in the LYR.
Graphical abstract illustrating how mesh topology affects flow redistribution and, consequently, water levels in an idealised urban layout.Urban flooding involves complex flow redistribution and energy dissipation mechanisms that are difficult to represent with depth-averaged models. This study examines how mesh discretisation influences their numerical representation using four steady-flow experiments representing idealised urban layouts. A mesh-sensitivity analysis is performed with a finite-volume shallow-water model by varying mesh type, resolution, and orientation. Results show that water-depth predictions and discharge partitioning in the streets depend on how energy dissipation is numerically represented. In aligned layouts, models usually fail to correctly represent the flow dynamics, resulting in strong mesh sensitivity and inaccurate flow redistribution unless appropriate mesh types and resolutions are used. In contrast, rotated layouts are dominated by inertia-driven flow redistribution, making the influence of mesh characteristics more limited. A hydraulic-power framework is introduced to quantify head losses in urban layouts, particularly at crossroads, and to highlight unresolved dissipation mechanisms. Finally, the efficiency of a resistance-based closure using a calibrated urban Strickler coefficient is evaluated. While it can improve global water-depth predictions at a lower computational cost, the calibrated parameter is highly case-dependent, limiting its physical meaning and transferability. Overall, this work provides guidance for mesh design and clarifies the limits of resistance calibration in 2D urban flood modelling.HIGHLIGHTSSystematic assessment of mesh type, resolution, and orientation for urban flood modelling. Mesh-induced flow deviation and artificial losses quantified at intersections. Hydraulic-power analysis used to assess regular and singular energy dissipation. Strickler calibration evaluated to compensate unresolved singular losses.
This technical note introduces an efficient workflow for the large-scale computation of near-bed flow velocities under wave action, based on linear airy wave theory. As a case study, rasterized datasets of bathymetry, significant wave height, and wave period for the German Bight are used to estimate the maximum orbital velocity at the seabed. This is implemented within a modular hydroinformatic framework, combining geographic raster data (pre-)processing using pluggable models and interactive visualization. It employs just-in-time compilation via the python library Numba. This allows for execution across all cores of the Central Processing Unit (CPU) or, if present, a Graphics Processing Unit (GPU) using the Compute Unified Device Architecture (CUDA) architecture. This design facilitates large parallel computation across thousands of CUDA cores, enabling faster processing of extensive raster domains for iterative model refinement. Performance benchmarks reveal a more than fourfold reduction in runtime and 70% lower memory consumption compared to a standard CPU implementation. The findings underscore how classical hydrodynamic formulations can be effectively integrated into scalable, open, and reproducible workflows for rapid spatial evaluation of marine processes. The presented method offers a practical foundation for area-wide analyses of seabed dynamics, offshore engineering design, and morphodynamic sensitivity studies within modern hydroinformatics.
The graphical summary adopts a left to right flowchart structure, systematically demonstrating the research logic and core contributions of this study. On the left is the "Research Background and Problems" module, highlighting three core elements: "high-intensity human activities (such as reservoir scheduling)", "water and sediment flux in the lower Yellow River", and "multi time scale evolution characteristics", pointing to the research goal of "aiming to reveal". The middle section is the "Multi Method Analysis Framework" module, which lists three methods in sequence: "Regression and Interpolation", "M-K Test and Seasonal Decomposition", and "Wavelet Analysis", and indicates "Applied to 2016-2022 data". On the right is the "Core Empirical Discoveries" module, which presents three key findings: a step like increase in water and sediment flux in 2018, a continuous and amplitude enhanced annual cycle (12 months), and a stronger nonlinear response of sediment transport compared to runoff. On the far right is the "Research Implications" module, which points out that this study provides empirical evidence for the impact of reservoir operation and provides methodological references for the management of regulated rivers. The arrows run through each module, clearly presenting a complete research narrative of "background method discovery revelation".
Reliable discharge estimation is fundamental to the design and operation of hydraulic structures like weirs, which are critical for irrigation, flood control, and water treatment. Slit weirs offer an efficient solution for precise flow management with minimal energy loss, yet accurately predicting their discharge coefficient (C-d) remains a persistent challenge, particularly for geometrically complex designs such as triangular slit weirs. This paper addresses this challenge by, first, providing a comprehensive characterization of triangular slit weir discharge behavior and proposing a new, highly accurate empirical equation for their C-d. Second, and critically, this study explores the power of advanced computational intelligence to enhance C-d and discharge prediction for both rectangular and triangular slit weirs. It study systematically compares the performance of Artificial Neural Networks (ANN), Adaptive Neuro-Fuzzy Inference Systems (ANFIS), Support Vector Machines (SVM), and traditional regression models. These findings reveal that the ANN model achieves outstanding performance (Training R-2 = 0.98, Testing R-2 = 0.88), surpassing ANFIS (Testing R-2 = 0.86), SVM (Testing R-2 = 0.85), and regression models (Testing R-2 = 0.82).
Accurately forecasting natural water systems is a complex task due to their interconnected structure, where both spatial and temporal dependencies play a critical role. In this work, we applied spatio-temporal graph neural networks of varying complexity to forecast the flow of rivers and the total releases of reservoirs in the Upper Colorado River Basin. Since prolonged droughts driven by climate change can reduce water levels in hydrological systems to critical thresholds, it is essential to forecast to mitigate their negative consequences. The models were trained using five years of historical time series data from directly connected sensor points within a river basin. We evaluated six models and compared their forecasting performance using mean squared error, overall, in boxplots. The graph convolutional recurrent network model performed the best compared to the other five models in the case study, which indicates that the graph convolution with the Chebyshev polynomial has the best forecast accuracy in water system forecasting.