Fecal coliform bacteria are key microbial indicators of water quality and public health risk, yet conventional monitoring methods relying on site-specific sampling and laboratory analysis are spatially limited and lack real-time responsiveness. This study develops a novel framework for large-scale estimation of fecal coliform concentrations by integrating Sentinel-2 multispectral imagery with a convolutional neural network (CNN) model across South Korea's four major river basins: Han, Nakdong, Geum, and Yeongsan. To enhance spectral sensitivity to microbial pollution, backscattering albedo (uT) was derived from reflectance data and incorporated alongside first- and second-order spectral differentiation. A ResNet-18-based CNN model was trained using Sentinel-2 reflectance, backscattering albedo, first and second differentiation combined dataset and in-situ fecal coliform data from 2017 to 2022, achieving performance with R2 values of 0.922 and 0.566 for training and validation datasets, respectively. The model captured low-concentration patterns more consistently, while prediction errors increased in moderate-to-high concentration ranges. Spatial prediction maps revealed contamination hotspots in urban and agricultural zones, particularly downstream of tributary confluences and near regulated flow structures. Model uncertainty was quantified using Maximum Likelihood Estimation, and SHapley Additive exPlanations (SHAP) analysis identified near-infrared and shortwave-infrared backscattering bands as the most influential features, providing a transparent interpretation of the model's behavior. This remote sensing-based approach enables robust, scalable, and explainable estimation of fecal coliform distributions over broad geographic areas, surpassing the spatial and temporal limitations of traditional field-based monitoring. This study presents the large-scale fecal coliform estimation model that directly incorporates backscattering albedo and spectral derivatives, offering a novel remote sensing-based solution that captures microbial contamination dynamics with unprecedented spatial precision and model interpretability.
Numerous gridded precipitation (P) datasets have been developed to address a variety of needs and challenges. However, selecting the most suitable and reliable dataset remains difficult for users. We conducted the most comprehensive global evaluation to date of gridded (sub-)daily P datasets using hydrological modeling. A total of 24 datasets - derived from satellite, (re)analysis, gauge sources, or combinations thereof - were assessed. To evaluate their performance, we calibrated the conceptual hydrological model HBV against observed daily streamflow for 18 428 catchments (each <10000km(2)) worldwide, using each P dataset as input. The Kling-Gupta Efficiency (KGE) was used as performance metric, with the calibration score serving as proxy for P dataset performance. Overall, Multi-Source Weighted-Ensemble Precipitation (MSWEP) V2.8 demonstrated the best performance (median KGE of 0.78), highlighting the value of merging P estimates from diverse data sources and applying daily gauge corrections. Among the purely satellite-based P datasets, the soil moisture- and microwave-based Global Precipitation Mission plus Soil Moisture to RAIN (GPM + SM2RAIN) dataset performed best (median KGE of 0.64). The Global Data Assimilation System (GDAS) analysis ranked highest among the (re)analyses (median KGE of 0.72), slightly outperforming the widely used European Centre for Medium-range Weather Forecasts ReAnalysis 5 (ERA5; median KGE of 0.71). Performance varied across K & ouml;ppen-Geiger climate zones, with the highest scores in polar (E) regions (median KGE of 0.76 across datasets) and the lowest in arid (B) regions (median KGE of 0.53 across datasets). Spatial correlation analysis between catchment attributes and KGE scores identified aridity index, potential evaporation, and P occurrence as the strongest predictors of performance. Our assessment revealed significant regional differences in dataset performance and error characteristics, emphasizing the importance of careful dataset selection for water resource management, hazard assessment, agricultural planning, and environmental monitoring.
Addressing the releases of industrial contaminants to rivers requires rapid assessment tools for emergency response and scenario screening. Although physics-based models such as the Environmental Fluid Dynamics Code can resolve hydrodynamic and contaminant transport processes in detail, their computational costs limit their application to large-scale scenarios and near-real-time application. In this study, we developed an integrated framework combining graph-based surrogate modeling and augmented reality visualization for the rapid simulation of river hydrodynamics and contaminant transport. A Graph Convolutional Network–Long Short-Term Memory framework was used to capture spatial and temporal dependencies through two surrogate modules. The hydrodynamic module employed depth-based classification to improve consistency under dry-channel conditions, achieving a Nash–Sutcliffe efficiency (NSE) of 0.9871 for water depth and 0.9865 for flow. For contaminant transport, a hybrid surrogate model composed of a four-class classifier and three concentration-range-specific regressors was developed to address highly uneven concentration distributions. The integrated contaminant surrogate model achieved NSE = 0.9138 across 243 accident scenarios. Spatiotemporal map comparisons and a point-based time series evaluation showed that the proposed framework accurately reproduced the main transport patterns and temporal concentration dynamics. The predicted results were further implemented in an augmented reality platform to interactively visualize the evolution of the contaminant plume. The proposed framework enables rapid scenario evaluation with substantially reduced computational demand while retaining close agreement with transport behavior obtained using the Environmental Fluid Dynamics Code.
Cyanobacterial blooms threaten aquatic ecosystems and human health, and their occurrence has been intensified by hydrological and weather variations associated with climate change. This study proposes an optimal weir operation framework by coupling Long Short-Term Memory (LSTM) models with a Soft Actor-Critic (SAC)-based deep reinforcement learning (DRL) agent for managing cyanobacterial blooms. The LSTM models simulated hydrological, nutrient, and cyanobacteria-related state variables that were input to the DRL environment to update continuous weir operation. The SAC-based DRL system was trained to determine the optimal weir overflow strategy to reduce cyanobacteria compared to the baseline LSTM simulation within operational storage constraints. Additional sensitivity analyses indicated that average operation metrics were relatively stable under cyanobacteria prediction perturbations and across SAC training seeds, while cyanobacterial reduction performance and peak-event responses remained sensitive to these uncertainties. On the validation set, the LSTM models achieved R2 values ranging from 0.75 to 0.89 for hydrological and nutrient variables and 0.76 for cyanobacteria regression, while the cyanobacteria occurrence classification yielded an accuracy of 92.22%. The DRL-based optimal weir operation reduced cyanobacteria concentrations by up to 13.01% and 8.39% during training and validation, respectively. Hydro-thermal scenarios revealed that higher inflow can partially mitigate thermal stress on cyanobacteria; however, the tested inflow range had limited capacity to reduce cyanobacterial concentrations under rising water temperature conditions. These results demonstrated the potential ability of the proposed LSTM-DRL framework to manage cyanobacterial blooms, though site-specific calibration and further uncertainty evaluation are required before operational application.
Harmful algal blooms (HABs) negatively impact aquatic ecosystems and humans, requiring constant management and forecasting of chlorophyll-a (Chl-a), which can be used to measure HABs. While various studies have predicted the lead time of Chl-a, they did not consider hydrological influences such as the operation of the weirs. This study is innovative in its use of retention times between river weirs as predictive variables that incorporate the dynamics of water flow and storage influenced by weir operations. To account for weir operation, the data were reconstructed by calculating the retention time between weirs and reflecting the retention time as a lead time. Four machine learning architectures, random forest (RF), extreme gradient boosting (XGB), deep neural network (DNN), and convolutional neural network (CNN), were used for prediction. The CNN model successfully predicted downstream Chl-a with a Nash-Sutcliffe efficiency (NSE) of 0.7198, a root mean square error (RMSE) of 18.4410 mg center dot m-3, and a mean absolute error (MAE) of 13.1202 mg center dot m-3. Additionally, the CNN model demonstrated predictive power even for the data not included in the training and validation. The analysis also integrated the Shapley additive explanations (SHAP) model to interpret the influence of input features on model predictions, offering insights into how upstream conditions impact downstream water quality. Successful prediction of downstream Chl-a through upstream flow features and the identification of features important for prediction will facilitate the operation of weirs in the Nakdong and Geum Rivers.
Monitoring total suspended solids (TSS) is critical for understanding water quality and managing pollution in river ecosystems. However, traditional methods face challenges in achieving real-time estimates in resource-constrained environments. This study aims to develop an optimized framework for convolutional neural network (CNN) to estimate TSS concentrations using Sentinel-2 multispectral data, with a focus on lightweight architecture and quantization techniques for real-time applications. Neural Architecture Search (NAS) combined with Pareto optimization was used to identify lightweight CNN models, ensuring high performance with minimal computational cost. Post-Training Quantization (PTQ) and Quantization-Aware Training (QAT) were applied to further compress model sizes while maintaining accuracy. Performance was evaluated using metrics such as Nash-Sutcliffe Efficiency (NSE) and Root Mean Squared Error (RMSE).As a result, the lightweight Mobilenet (8.11 MB) attained an NSE of 0.828, and quantization further reduced the model size by 91%, yielding a compact 0.74 MB model with an enhanced NSE of 0.832. This quantized TSS estimation model showed the potential for real-time TSS estimation on mobile and edge devices. The proposed lightweighting and quantization framework provides a scalable solution for real-time TSS monitoring, connecting advanced machine learning methods with practical environmental applications. This approach enables efficient, real-time water quality assessment in a variety of environmental conditions, making it suitable for use on resource-constrained platforms such as drones, unmanned aerial vehicles and satellites.
Recent achievements in the fields of deep learning and remote sensing have led to their application in monitoring river water quality. One of the most researched methods is the estimation of total suspended solid (TSS) concentrations using multispectral imagery and convolutional neural network (CNN) models. Owing to the sorption capacity of other pollutants, TSS monitoring is essential. However, despite recent advances in deep learning, the application of contemporary technologies in water quality monitoring has not yet been fully explored. This study aims to develop a framework for on-device AI that can be applied to edge devices through quantization using a lightweight deep learning model. Lightweight CNN models were identified using neural architecture search (NAS) in conjunction with Pareto optimization, achieving high performance (0.806 of Nash-Sutcliffe efficiency (NSE)) while minimizing computational burden (8.118 MB). The model sizes were further compressed (0.736 MB) through the application of post-training quantization (PTQ) and quantization aware training (QAT), ensuring that accuracy (0.831 of NSE) was preserved. This provides a scalable approach for real-time TSS monitoring, bridging the gap between advanced deep learning techniques and practical environmental applications. These applications indicate that it is possible to estimate other water quality indices using multispectral imagery. It enables the tracing of the source of contamination and facilitates rapid responses by identifying changes in real time.
Advanced suspect and non-target screening (SNTS) approach can identify a large number of potential hazardous micropollutants in groundwater, underscoring the need for pinpointing priority pollutants among detected chemicals. This present study therefore demonstrates a novel multi-criteria decision making (MCDM) framework utilizing machine learning (ML) algorithms coupled with toxicological prioritization index tool (i.e., ml_ToxPi) to rank 252 chemicals of interest in groundwater for subsequent targeted analysis. The MCDM framework integrated chemical analysis data (i.e., peak area and detection frequency), toxicity profiles (i.e., bioactivity ratio, human exposure metadata, and carcinogenicity metadata), as well as the environmental fate and transport information (i.e., octanol-water partition coefficient (log Kow), water solubility, biodegradation half-life, and soil adsorption coefficient (Koc)) for ranking the identified pollutants, and the random forest machine learning model was useful for systematically determining the weighting factors of each variable according to their variable importance scores (R2 = 0.808 and 0.778 for training and testing datasets, respectively, while RMSE = 0.042 in both cases). A total of 47 unique high priority compounds (i.e., ml_ToxPi score >= 0.55) were identified across the investigated sampling regions, which constituted diverse groups of compounds classified according to their chemical uses, such as alkylated polycyclic aromatic hydrocarbons (alkyl-PAHs), organophosphate flame retardants (OPFRs), parent PAHs, personal care products (PCPs), pesticides, pharmaceuticals, phenols, plasticizers, transformation product (TPs), and other industrial use chemicals. By incorporating relevant variables into the proposed ML-optimized ToxPi MCDM framework, the prioritization approach described here may be adopted in future SNTS assessment of environmental and biological media.
To ensure millennial isolation of high-level waste, Deep geological repositories (DGR) safety assessments depend on process-based simulators such as Parallel flow and reactive transport model (PFLOTRAN). Yet, without feasible long-term monitoring, the high computational burden of these models becomes a bottleneck for iterative scenario testing and policy decisions. To overcome this, we developed a graph attention-based physics-guided deep learning (GAT-PGDL) surrogate that embeds decay-diffusion-sorption equations. U-238 and Th-230 transport was simulated for 5000 years in DGRs, with release commencing at year 2000 across ten monitoring nodes. On a single-node workstation, PFLOTRAN requires ∼5760 min per scenario, whereas the GAT-PGDL trains once (∼94 min) and infers in seconds, delivering ∼61 × per-scenario speedup. To assess reliability and interpretability, we employed split conformal prediction for uncertainty quantification, perturbation-based sensitivity analysis, and feature importance analysis. The GAT-PGDL produced 95 % prediction intervals and highlighted sorption and bulk density as key transport controls. For generalization, performance was compared with a data-driven surrogate under scenarios with altered material properties and earlier release times, where the GAT-PGDL outperformed the data-driven model, maintaining R² and NSE > 0.98. These results establish GAT-PGDL as a fast, accurate, and physically reliable surrogate for long-term DGR safety assessments.
Fecal coliforms are thermotolerant bacteria excreted from warm-blooded animals into soil and water, contaminating water bodies through runoff and resuspension of sediments. This contamination poses significant public health risks, especially during summer recreational activities, leading to waterborne diseases like diarrhea, typhoid, cholera, and dysentery. Monitoring and managing fecal coliform levels in recreational waters are crucial for public health and environmental safety. However, variability in fecal coliform concentrations due to human and wildlife activities complicates the management. This study aims to enhance water safety and public health by utilizing sentinel-2 band reflectance data and backscattering albedo to understand the relationship between fecal coliform reflectance in the rivers to generalize the fecal coliform management model.In this study, we constructed Sentinel-2 dataset covering the period from January 2017 to December 2022 for the Han, Nakdong, Geum, and Yeongsan Rivers in South Korea. To accurately align the water quality monitoring stations with the Sentinel-2 data, we ensured that the latitude and longitude coordinates were free from clouds and not located on bridges. Therefore, monitoring stations that did not meet the specified conditions, with an above NDWI (Normalized Difference Water Index) of 0.1, and a below HOT (Hazed-Optimized Transformation) of 0.05 were preprocessed. For the preprocessed data points, this study converted the reflectance values of 10 Sentinel-2 bands (2, 3, 4, 5, 6, 7, 8, 8A, 11, and 12) into backscattering albedo. This approach was taken to account for the characteristics of fecal coliform, which is colorless. Model training was performed using CNN (Convolutional Neural Network), ANN (Artificial Neural Network), Random Forest, and XGBoost. As a result, CNN successfully predicted the trend of fecal coliform in the all the rivers and showed superior performance compared to other models. The results of this study are expected to provide a basis for fecal coliform management using Sentinel-2 band reflectance data in the four major rivers of South Korea and other regions around the world.
Numerous gridded precipitation (P) datasets have been developed to address a variety of needs and challenges. However, selecting the most suitable and reliable dataset remains a challenge for users. We conducted the most comprehensive global evaluation to date of gridded (sub-)daily $P$ datasets using hydrological modeling. A total of 23 datasets, derived from satellite, model, gauge sources, or their combinations thereof, were assessed. To evaluate their performance, we calibrated the conceptual hydrological model HBV against observed daily streamflow for 16,295 catchments (each
Deep geological repositories (DGRs) are designed for the permanent disposal of spent nuclear fuel, necessitating precise radionuclide transport predictions. Owing to the impracticality of large-scale physical experiments, computational simulations are a key alternative. Although the Parallel Flow and Reactive Transport Model (PFLOTRAN) is widely used for radionuclide transport simulations, its high computational demands limit its practical application. This study employs Graph Convolutional Long Short-Term Memory (GCLSTM) as a surrogate model for PFLOTRAN to simulate radionuclide transport and significantly reduce computational costs while maintaining predictive accuracy. GCLSTM was trained using time-series data from PFLOTRAN simulations over a 5,000-year period. The model achieved a coefficient of determination above 0.99 and a Nash-Sutcliffe efficiency exceeding 0.97 at all observation nodes. Combined uncertainty quantification and sensitivity analyses demonstrate that over 95 % of GCLSTM predictions fall within PFLOTRAN-derived confidence intervals and that permeability and inter-node distance are the primary drivers of predictive variance. Additionally, scenario-based simulations validated the adaptability of GCLSTM to varying prediction lengths and release conditions. By reducing the computational time by approximately 576 times compared to that of PFLOTRAN while maintaining predictive accuracy, GCLSTM demonstrated its potential as an efficient and reliable alternative. This approach enhances modeling efficiency by utilizing GCLSTM as a surrogate for PFLOTRAN, offering a practical solution for long-term radionuclide transport simulations.
Harmful algal blooms (HABs) pose a serious threat to aquatic life in surface water. Several early warning systems have been developed to mitigate problems related to HABs by predicting HABs using machine learning methods. However, HABs are dynamic phenomena that depend on interactions between weather conditions, hydrodynamic parameters, and weir operations. This requires a real-time control strategy that uses hydrodynamic parameters to determine an optimal strategy for weir operation. In this study, reinforcement learning (RL) was employed to determine the optimum weir operation for mitigating the occurrence of HABs in six weirs of the Nakdong River, Republic of Korea. The impact of weir operation on HABs was simulated using the Soil and Water Assessment Tool (SWAT), and the outflow from the reservoir was predicted using RL agents. The RL agents were trained to predict the outflow strategy from the reservoir by minimizing the chlorophyll-a (Chl-a) concentration. Our results showed that the trained RL agents could reduce the Chl-a concentration by manipulating weir operations. The trained agents were especially effective in minimizing the peak Chl-a concentrations during summer. They reduced concentrations by an average of 35 % at six weir stations. Our study demonstrates the successful integration of RL with SWAT to improve surface water quality. The developed model has the potential to be employed as a tool for managing weir operations and water resources.
High concentrations of chlorophyll-a (Chl-a) in aquatic systems pose serious environmental and public health concerns. Chl-a, a primary marker of phytoplankton biomass, is often associated with the proliferation of harmful algal blooms (HABs). These blooms produce toxins that not only threaten marine organisms but also have far-reaching impacts on human health and aquatic ecosystems. These toxins can degrade water quality, disrupt food webs, and result in significant fish mortality. When these harmful substances contaminate drinking water sources, they can cause a range of health problems, from short-term illnesses to chronic diseases.Despite the importance of predicting Chl-a levels, earlier research has largely focused on water quality parameters without adequately considering the dynamic nature of river hydrology. This study bridges that gap by leveraging satellite data to enhance predictive accuracy. Sentinel-2 imagery was utilized to monitor water quality, while Sentinel-1 data captured the hydrological characteristics of rivers. To forecast Chl-a, four machine learning models were deployed, with their performance evaluated through Nash-Sutcliffe Efficiency (NSE) and Root Mean Square Error (RMSE) metrics. Additionally, the study used Shapley Additive Explanations (SHAP) to unravel the contribution of individual water quality variables and satellite-derived data to the prediction process.By integrating hydrological factors with water quality predictions, this research provides a more holistic understanding of river systems. Such insights are vital for optimizing the operation of water management structures like dams and weirs. Moreover, the incorporation of retention time analysis offers a proactive approach to monitoring and preventing HABs, enabling more effective management of aquatic ecosystems under varying environmental conditions worldwide.
Harmful Algal Blooms (HABs) threaten aquatic ecosystems, necessitating effective monitoring strategies in water resource management. Satellite-based remote sensing has emerged as a popular method to address the limitations of in-situ monitoring. However, cloud covers can obstruct optical imagery, causing data loss. Synthesis Aperture Radar (SAR) imagery, with its capabilities, can penetrate through any weather conditions. We applied SAR imagery with the Faster Regional Convolutional Neural Networks (Faster R-CNN) model to detect the algal bloom. The dataset of the Geum River Basin was obtained from 2020 to 2022. The sigma naught values (dB) were analyzed from the SAR imagery to clarify the reflectance properties of algae in VH and VV polarizations. The values ranged between -12 dB and -33 dB and -5 dB and -27 dB for VH and VV polarization, respectively. The model was developed with hyperparameter optimization to detect the algal bloom by splitting the training from 2020 to 2022, and the testing dataset 2022. Evaluation metrics including precision, recall, and F1 scores yielded values of 0.600, 0.692, and 0.643, respectively. The developed model was simulated to identify the seasonal outbreak. The result illustrated that the algal blooms were detected only in the summer of 2021 and 2022. Furthermore, the model was validated in supporting an existing algal alert report, demonstrating the potential for real-time monitoring. Finally, this study highlights the effectiveness of employing SAR imagery with the Faster R-CNN model to develop an algorithm for detecting algal blooms, offering advancements in water management practices.
This study assessed the U.S. Environmental Protection Agency's Storm Water Management Model (SWMM) for urban water management challenges. This study conducted a sensitivity analysis to identify the most influential factors in the SWMM. Moreover, the performance of SWMM was evaluated with the HYDRUS-1D module in simulating infiltration rates. The sensitivity results showed field capacity as the most significant factor, highlighting the need for advanced modeling techniques to consider factors like field capacity. The SWMM was evaluated by the HYDRUS-1D that SWMM consistently underestimated peak infiltration rates and commenced infiltration calculations only when soil moisture exceeded field capacity. It reveals its limitations in handling unsaturated soil conditions and highlights the consideration of the matric head of the soil during the infiltration calculation in soil media. Moreover, the evaluation of bioretention areas showed larger areas resulting in more substantial flow reductions but with significant variability under different rainfall conditions. Accordingly, this result emphasizes the importance of careful consideration for environmental factors in bioretention design. This study contributes by enhancing understanding of SWMM's limitations in simulating urban water management challenges. Thus, this research will offer technical assistance to stakeholders focused on challenges such as runoff, hydrologic cycle, and urban flooding in urban areas.
Conventional environmental health research is primarily focused on isolated chemical exposures, neglecting the complex interactions between multiple pollutants that may synergistically or antagonistically influence toxicity, thereby posing unexpected health risks. In this study, we address this knowledge gap by introducing an explainable machine learning (ML) approach with Feature Localized Intercept Transformed-Shapley Additive Explanations (FLIT-SHAP) designed to extract the dose-response relationships of specific pollutants in mixtures. In contrast to traditional SHAP, FLIT-SHAP can localize the global intercept to elucidate mixture effects, which is crucial for understanding the oxidative potential (OP) of ambient particulate matter (PM). Assessing multipollutant OP using FLIT-SHAP revealed both synergistic (55-63 %) and antagonistic (25-42 %) effects in laboratory-controlled OP data, but an antagonistic (33-66 %; lower OP) effect in ambient PM. Notably, the FLITSHAP approach demonstrated higher prediction accuracy (R2 = 0.99) compared to the additive model (R2 = 0.89) when evaluated against real-world PM samples. Quinones, such as phenanthrenequinone, play a more significant role in PM2.5 than previously recognized. Through this study, we highlighted the potential of FLITSHAP to enhance toxicity predictions and aid decision-making in the field of environmental health.
The impacts of climate change on hydrology underscore the urgency of understanding watershed hydrological patterns for sustainable water resource management. The conventional physics-based fully distributed hydrological models are limited due to computational demands, particularly in the case of large-scale watersheds. Deep learning (DL) offers a promising solution for handling large datasets and extracting intricate data relationships. Here, we propose a DL modeling framework, incorporating convolutional neural networks (CNNs) to efficiently replicate physics-based model outputs at high spatial resolution. The goal was to estimate groundwater head and surface water depth in the Sabgyo Stream Watershed, South Korea. The model datasets consisted of input variables, including elevation, land cover, soil type, evapotranspiration, rainfall, and initial hydrological conditions. The initial conditions and target data were obtained from the fully distributed hydrological model HydroGeoSphere (HGS), whereas the other inputs were actual measurements in the field. By optimizing the training sample size, input design, CNN structure, and hyperparameters, we found that CNNs with residual architectures (ResNets) yielded superior performance. The optimal DL model reduces computation time by 45 times compared to the HGS model for monthly hydrological estimations over five years (RMSE 2.35 and 0.29 m for groundwater and surface water, respectively). In addition, our DL framework explored the predictive capabilities of hydrological responses to future climate scenarios. Although the proposed model is cost-effective for hydrological simulations, further enhancements are needed to improve the accuracy of long-term predictions. Ultimately, the proposed DL framework has the potential to facilitate decision-making, particularly in large-scale and complex watersheds.
Eutrophication is a major cause of water quality degradation in South Korea, owing to severe algal blooms. To manage eutrophication, the South Korean government provided the Trophic State Index (TSIko), which was revised according to Carlson's TSI. The TSIko levels were simulated using mechanistic water quality modeling. However, the computational complexity of model parameter calibration and the nonlinearity of water quality kinetics complicate analyzing accurate eutrophication conditions. Deep learning models have been considered alternatives to numerical model approaches because they directly extract water quality variables without prior knowledge. In particular, the convolutional neural network (CNN) model showed robust feature extraction from the complex datasets. This study constructed and optimized a CNN model using water quality data from the Han, Guem, Yeongsan, and Nakdong Rivers in South Korea over nine years from 2014 to 2022 to classify the TSIko. The CNN model provided validation results using the statistical measurement of classification accuracy, known as the F1 score, which is the harmonic mean of precision and recall. The F1 scores were 0.922, 0.950, 0.964, and 0.896 for oligotrophic, mesotrophic, eutrophic, and hypertrophic statuses, respectively. The CNN model outperformed conventional machine learning models. Subsequently, a eutrophication map for the four major rivers was generated using the CNN model to simulate the spatial and temporal variations of the eutrophication index, mimicking high spatio-temporal eutrophic dynamics with respect to the mainstream and tributaries of the Yeongsan and Nakdong Rivers. Therefore, this study demonstrates the capability of the CNN model to analyze eutrophication conditions at various spatial and temporal scales of major rivers in South Korea.
Estimation of aquatic ecosystem health indices can assist in reducing the burden of time-consuming, labor-intensive, and cost-effective fieldwork for the sustainable evaluation of freshwater ecosystem status. In this study, we developed a deep neural network to estimate the trophic diatom index (TDI), benthic macroinvertebrate index (BMI), and fish assessment index (FAI) using water quality and hydraulic and hydrological data. A convolutional neural network (CNN) model was built to estimate health indices. In addition, an autoencoder was adopted to produce manifold features that were used as inputs for the CNN model. Conventional machine learning models, including artificial neural networks, support vector machines, random forests, and extreme gradient boosting, have been developed to estimate the TDI, BMI, and FAI. The results showed that the CNN with an autoencoder exhibited the best performance, with validation accuracies of Nash Sutcliffe Efficiency (NSE) and root mean squared error (RMSE) values of 0.592 and 17.249 for TDI, 0.669 and 12.282 for BMI, and 0.638 and 13.897 for FAI, respectively. The autoencoder enhanced the nonlinear feature learning of the time series and static input data, which contributed to improving the CNN feature extraction for accurate estimation of aquatic ecosystem health indices compared to other data-driven approaches. Therefore, deep learning techniques can be used to investigate aquatic ecosystem health by successfully reflecting the quantitative and qualitative features of health indices.