Flood damage assessment remains challenging, as conventional flood risk management mainly relies on hydraulic hazard maps that do not explicitly reproduce observed damage patterns. Recent advances in remote sensing and machine learning (ML) enable the integration of environmental and socio-economic data with historical impact information to improve flood damage modeling. This study proposes an explainable machine learning framework for flood damage susceptibility mapping, using observed institutional damage records from the 2011 and 2013 flood events combined with 17 geospatial flood risk factors (FRFs) representing hazard, exposure, and vulnerability. This approach enables the capture of non-linear relationships between flood damage and FRFs. For comparison purposes, the same framework was also applied using hydraulically modeled flood extents corresponding to return periods of 30, 200, and 500 years. The framework was tested along the Basilicata Ionian coast in southern Italy, a Mediterranean region characterized by complex geomorphology, intense rainfall events, and recurrent flood impacts. An eXtreme Gradient Boosting (XGBoost) model was trained using 17 FRFs related to hazard, exposure, and vulnerability at a spatial resolution of 20 m. The model achieved high performance with an accuracy of 0.988, an F1-score for the minority class of 0.860, and an ROC-AUC (test) of 0.996. High to very high flood damage probability was predicted in approximately 4.1% of the study area, mainly in low-lying floodplains near river corridors and infrastructure. SHAP-based explainability analysis revealed that damage susceptibility was predominantly driven by hazard and exposure factors: Drainage density (17.10%), Railway distance (16.33%), and Elevation (15.42%), extreme precipitation (Max rainfall, 10.66%) and Street distance (7.51%), with socio-economic vulnerability contributing less than 4%. The observed damage target exhibited clear threshold-like patterns (e.g., sharp risk increases below ~25/35 m elevation or within ~150/200 m of road infrastructure), contrasting with the smoother, continuous gradients produced by hydraulic scenarios. This analysis identified the most influential predictors and their response ranges. The proposed framework complements hydraulic hazard mapping by explicitly modeling observed flood damage, supporting flood risk assessment in flood-prone coastal regions.
This study presents an approach based on machine learning (ML) techniques to analyze the relationship between emergency room (ER) admissions for cardiorespiratory diseases (CRDs) and environmental factors. The aim of this study is the development and verification of an interpretable machine learning framework applied to environmental and health data to assess the relationship between environmental factors and daily emergency room admissions for cardiorespiratory diseases. The model’s predictive accuracy was evaluated by comparing simulated values with observed historical data, thereby identifying the most influential environmental variables and critical exposure thresholds. This approach supports public health surveillance and healthcare resource management optimization. The health and environmental data, collected through meteorological sensors and air quality monitoring stations, cover eleven years (2013–2023), including meteorological conditions and atmospheric pollutants. Four ML models were compared, with XGBoost showing the best predictive performance (R2 = 0.901; MAE = 0.047). A 10-fold cross-validation was applied to improve reliability. Global model interpretability was assessed using SHAP, which highlighted that high levels of carbon monoxide and relative humidity, low atmospheric pressure, and mild temperatures are associated with an increase in CRD cases. The local analysis was further refined using LIME, whose application—followed by experimental verification—allowed for the identification of the critical thresholds beyond which a significant increase in the risk of hospital admission (above the 95th percentile) was observed: CO > 0.84 mg/m3, P_atm ≤ 1006.81 hPa, Tavg ≤ 17.19 °C, and RH > 70.33%. The findings emphasize the potential of interpretable ML models as tools for both epidemiological analysis and prevention support, offering a valuable framework for integrating environmental surveillance with healthcare planning.
This study proposes an ensemble machine learning model to analyse the association between respiratory Emergency Room (ER) admissions and environmental factors, such as air pollution and weather-climatic conditions. The analysed climatic variables include air temperatures Tmin, Tmax, Taverage, atmospheric pressure P, relative humidity RH, and levels of CO, O₃, PM₁₀, and NO₂. The data were processed as daily averages to ensure consistency and comparability in the analyses. Data on ER, provided by the Policlinico of Bari, cover the period from 2013 to 2023. The analysis was conducted using ensemble learning techniques, applying three regression models: Random Forest, XGBoost, and Adaboost. The models were trained on a pre-processed database using a 7-day exponential moving average (EMA7) to obtain a more stable time series. Model hyperparameters were optimized through Bayesian optimization. Among the analysed models, XGBoost showed high predictive capacity in test sets. In particular, the R2 value was 0.772, while the MAE was 0.049 cases/day. Applying SHAP (SHapley Additive exPlanations) analysis to the XGBoost model allowed us to identify the most important variables influencing hospital admissions and their related patterns. The most relevant features, ranked by importance, were: low values of average air temperature and atmospheric pressure, and high values of CO. The SHAP method, and in particular the use of Bee Swarm plots, were used to globally interpret the results obtained by the model and allowed us to reach the above results. Furthermore, in order to determine for the most important features the values that cause an increase in admission to the emergency room for respiratory diseases, a local analysis was carried out by applying the LIME model which allowed us to say that the greater onset of respiratory diseases is associated with average temperatures lower than 12.28 ℃, atmospheric pressure values lower than or equal to 1006.81 hPa and CO concentrations greater than 0.84 mg/m3.
The COVID-19 pandemic has generated significant global impacts on health and society, imposing a comprehensive analysis of its influencing factors, including weather variables. This study investigates the interaction between meteorological conditions and the spread of COVID-19 in three Italian regions: Lombardia, Emilia-Romagna, and Puglia. Effects of weather variables, such as air temperature, relative humidity, dew point, solar radiation, wind speed, and barometric pressure, are explored in the incidence of disease. Observed meteorological and health data are taken from various sources, such as the citizen-science Meteonetwork Association and the National Department of Civil Protection, respectively, and they are analyzed with statistical methods and machine learning algorithms. The study emphasizes the necessity of carefully considering key meteorological quantities as primary drivers in illness diffusion and prevention strategies, offering valuable insights to address challenges to the pandemic and ensure the safety of global communities. The results reveal a significant correlation between specific atmospheric variables and the spread of COVID-19, with dew point temperature as the most influential parameter at low air temperature values.
Floods and landslides are two distinct natural phenomena influenced by different conditioning factors, though some environmental triggers may overlap. This study applied eXtreme Gradient Boosting (XGBoost) to develop susceptibility maps for both phenomena, using a unified approach based on the same geospatial predictors. The approach integrated topographical, geological, and remote sensing datasets. Flood event data were collected from institutional sources using multi-source and high-resolution remotely sensed data. The landslide inventory was compiled based on historical records and geomorphological analysis. Key conditioning factors such as elevation, slope, lithology, and land cover were analyzed to identify areas prone to floods and landslides. The methodology was applied to the Basento River basin in Southern Italy, a region frequently impacted by both hazards, to assess its vulnerability and inform risk management strategies. While flood susceptibility is primarily associated with low-lying areas near river networks, landslides are more influenced by steep slopes and geological instability. The XGBoost model achieved a classification accuracy close to 1 for flood-prone areas and 0.92 for landslide-prone areas. Results showed that flood susceptibility was primarily associated with low Elevation and Relative Elevation, and high Drainage Density, whereas landslide susceptibility was more influenced by a broader and balanced set of factors, including Elevation, Drainage Density, Relative Elevation, Distance and Lithology. The resulting susceptibility maps offered critical approaches for land use planning, emergency management, and risk mitigation. Overall, the results demonstrated the effectiveness of XGBoost in multi-hazard assessments, offering a scalable and transferable approach for similar at-risk regions worldwide.
Background: Several studies suggest that environmental and climatic factors are linked to the risk of mortality due to cardiovascular and respiratory diseases; however, it is still unclear which are the most influential ones. This study sheds light on the potentiality of a data-driven statistical approach by providing a case study analysis. Methods: Daily admissions to the emergency room for cardiovascular and respiratory diseases are jointly analyzed with daily environmental and climatic parameter values (temperature, atmospheric pressure, relative humidity, carbon monoxide, ozone, particulate matter, and nitrogen dioxide). The Random Forest (RF) model and feature importance measure (FMI) techniques (permutation feature importance (PFI), Shapley Additive exPlanations (SHAP) feature importance, and the derivative-based importance measure (κALE)) are applied for discriminating the role of each environmental and climatic parameter. Data are pre-processed to remove trend and seasonal behavior using the Seasonal Trend Decomposition (STL) method and preliminary analyzed to avoid redundancy of information. Results: The RF performance is encouraging, being able to predict cardiovascular and respiratory disease admissions with a mean absolute relative error of 0.04 and 0.05 cases per day, respectively. Feature importance measures discriminate parameter behaviors providing importance rankings. Indeed, only three parameters (temperature, atmospheric pressure, and carbon monoxide) were responsible for most of the total prediction accuracy. Conclusions: Data-driven and statistical tools, like the feature importance measure, are promising for discriminating the role of environmental and climatic factors in predicting the risk related to cardiovascular and respiratory diseases. Our results reveal the potential of employing these tools in public health policy applications for the development of early warning systems that address health risks associated with climate change, and improving disease prevention strategies.
The Mediterranean basin is one of those areas where the impact of climate change is showing its most alarming consequences. Many regions in this area, both woodlands and croplands, have been suffering from droughts and water deficits due to the intense summer heatwaves of the last decades. Monitoring these phenomena is key to understanding how they are evolving and what could be done to mitigate their effects. Emissivity is a useful parameter in identifying the presence (or absence) of water. Surface and dew point temperatures are extremely useful not only in measuring the intensity of the heatwave but also in accounting for how much water content the surface is losing as humidity to the atmosphere. This paper presents a climatological study of Southern Italy's water loss for the period 2015-2023 based on daily observations acquired by the Infrared Atmospheric Sounding Interferometer (IASI), mounted on top of EUMETSAT's MetOp satellites. The Water Deficit Index (WDI) and the Emissivity Contrast Index (ECI) were estimated: monthly averages of each quantity were produced for the period of interest. Moreover, a validation with in situ measurements was conducted to better understand how these heatwave-induced droughts have been impacting the surface on different types of land covers.
The objective of this study was to determine the relationship between weather conditions and hospital admissions for cardiovascular diseases (CVD). The analysed data of CVD hospital admissions were part of the database of the Policlinico Giovanni XXIII of Bari (southern Italy) within a reference period of 4 years (2013–2016). CVD hospital admissions have been aggregated with daily meteorological recordings for the reference time interval. The decomposition of the time series allowed us to filter trend components; consequently, the non-linear exposure–response relationship between hospitalizations and meteo-climatic parameters was modelled with the application of a Distributed Lag Non-linear model (DLNM) without smoothing functions. The relevance of each meteorological variable in the simulation process was determined by means of machine learning feature importance technique. The study employed a Random Forest algorithm to identify the most representative features and their respective importance in predicting the phenomenon. As a result of the process, the mean temperature, maximum temperature, apparent temperature, and relative humidity have been determined to be the most suitable meteorological variables as the best variables for the process simulation. The study examined daily admissions to emergency rooms for cardiovascular diseases. Using a predictive analysis of the time series, an increase in the relative risk associated with colder temperatures was found between 8.3 °C and 10.3 °C. This increase occurred instantly and significantly 0–1 days after the event. The increase in hospitalizations for CVD has been shown to be correlated to high temperatures above 28.6 °C for lag day 5.
Although the Mediterranean Sea is a crucial hotspot in marine biodiversity, it has been threatened by numerous anthropogenic pressures. As flagship species, Cetaceans are exposed to those anthropogenic impacts and global changes. Assessing their conservation status becomes strategic to set effective management plans. The aim of this paper is to understand the habitat requirements of cetaceans, exploiting the advantages of a machine-learning framework. To this end, 28 physical and biogeochemical variables were identified as environmental predictors related to the abundance of three odontocete species in the Northern Ionian Sea (Central-eastern Mediterranean Sea). In fact, habitat models were built using sighting data collected for striped dolphins Stenella coeruleoalba, common bottlenose dolphins Tursiops truncatus, and Risso's dolphins Grampus griseus between July 2009 and October 2021. Random Forest was a suitable machine learning algorithm for the cetacean abundance estimation. Nitrate, phytoplankton carbon biomass, temperature, and salinity were the most common influential predictors, followed by latitude, 3D-chlorophyll and density. The habitat models proposed here were validated using sighting data acquired during 2022 in the study area, confirming the good performance of the strategy. This study provides valuable information to support management decisions and conservation measures in the EU marine spatial planning context.
Background: Cardiovascular diseases (CVD) remain the predominant global cause of mortality, with both low and high temperatures increasing CVD-related mortalities. Climate change impacts human health directly through temperature fluctuations and indirectly via factors like disease vectors. Elevated and reduced temperatures have been linked to increases in CVD-related hospitalizations and mortality, with various studies worldwide confirming the significant health implications of temperature variations and air pollution on cardiovascular outcomes. Methods: A database of daily Emergency Room admissions at the Giovanni XIII Polyclinic in Bari (Southern Italy) was developed, spanning from 2013 to 2019, including weather and air quality data. A Random Forest (RF) supervised machine learning model was used to simulate the trend of hospital admissions for CVD. The Seasonal and Trend decomposition using Loess (STL) decomposition model separated the trend component, while cross-validation techniques were employed to prevent overfitting. Model performance was assessed using specific metrics and error analysis. Additionally, the SHapley Additive exPlanations (SHAP) method, a feature importance technique within the eXplainable Artificial Intelligence (XAI) framework, was used to identify the feature importance. Results: An R2 of 0.97 and a Mean Absolute Error of 0.36 admissions were achieved by the model. Atmospheric pressure, minimum temperature, and carbon monoxide were found to collectively contribute about 74% to the model’s predictive power, with atmospheric pressure being the dominant factor at 37%. Conclusions: This research underscores the significant influence of weather-climate variables on cardiovascular diseases. The identified key climate factors provide a practical framework for policymakers and healthcare professionals to mitigate the adverse effects of climate change on CVD and devise preventive strategies.
Global circulation models (GCMs) are routinely used to project future climate conditions worldwide, such as temperature and precipitation. However, inputs with a finer resolution are required to drive impact-related models at local scales. The non-homogeneous hidden Markov model (NHMM) is a widely used algorithm for the precipitation statistical downscaling for GCMs. To improve the accuracy of the traditional NHMM in reproducing spatiotemporal precipitation features of specific geographic sites, especially extreme precipitation, we developed a new precipitation downscaling framework. This hierarchical model includes two levels: (1) establishing an ensemble learning model to predict the occurrence probabilities for different levels of daily precipitation aggregated at multiple sites and (2) constructing a NHMM downscaling scheme of daily amount at the scale of a single rain gauge using the outputs of ensemble learning model as predictors. As the results obtained for the case study in the central-eastern China (CEC), show that our downscaling model is highly efficient and performs better than the NHMM in simulating precipitation variability and extreme precipitation. Finally, our projections indicate that CEC may experience increased precipitation in the future. Compared with similar to 26 years (1990-2015), the extreme precipitation frequency and amount would significantly increase by 21.9%-48.1% and 12.3%-38.3%, respectively, by the late century (2075-2100) under the Shared Socioeconomic Pathway 585 climate scenario.
Cetaceans are species indicator of the ecosystem health status. It is necessary to increase knowledge on them to support the conservation of these species and their habitat. New strategy, based on machine learning techniques, has been adopted to estimate cetacean group size. Starting from Risso's dolphin sighting data in the Ionian Sea collected between 2009 and 2019, the aim of this work is to build a correlative model which can help in estimating Risso's dolphin group size, by using a set of physical and biogeochemical features.
A recent report “The Future is Now: Science for Achieving Sustainable Development” Global Sustainable Development Report 2019 - SDG Summit’ as part of the activity of Agenda 2030 of UN, highlights the opportunity to develop Early warning system for drought, floods and other meteorological events, that by providing timely information can be used by vulnerable countries to build resilience, reduce risks and prepare more effective responses. Following the suggestion, combining outputs from Global Circulation models, remote sensing, hydraulic models and machine learning tools, a local scale flooding Early Warning System (EWS) is proposed for the St. Lucia island ( Caribbean). The objective of the EWS is to provide forecasts of potentially dangerous flooding phenomena at different time scale: a) 0-2 hours, nowcasting; b) 24-48 hours, short range; c) 3-10 days, middle to long range. Data used to build the model are: Geopotential Height (GPH) fields at 850 hPa and Integrated Vapor Transport (IVT) fields from European Centre for Medium-range Weather Forecasts (ECMWF) - Reanalysis v5 (ERA5); Tropical Cyclone tracks from NOAA-NHC; 18 weather stations homogeneously distributed in the island; rainfall map data from the weather radar in Saint Lucia. GPH and IVT fields were defined between 110°W - 10°W and 45°N - 10°S. The EWS is constituted by an ensemble of flooding risk forecast subsystems which is potentially applicable to Atlantic tropical and extra-tropical regions. Different approaches are used for each subsystem to link large scale atmospheric features to local rainfall and flooding: a) Non-homogeneous Hidden Markov and Event Synchronization models to translate IVT and GPH at 850 hPa fields (from ECMWF-Set II- Atmospheric Model Ensemble) in local daily rainfall amount and probability of exceedance of a prefixed heavy rainfall threshold; b) a physical based cyclone/rainfall model to convert Tropical cyclone attributes – position and maximum wind velocity (provided from National Hurricane Center)- in rainfall intensity spatial distribution on the island; c) a surrogate model for a fast and accurate prediction of flooding events that is obtained from a multi-layer perceptron neural network (MLPNN), which is trained on a high-fidelity dataset relying on solution of the full two-dimensional shallow water equations with direct rainfall application. Results show an excellent ability of the models to identify the climatic configurations that determine the occurrence of extreme events and the exceeding of threshold values that generate floods. In particular, during the late hurricane season September-October-November, when is highest the probability of flood events, the EWS was able to forecast the occurrence of critical climatic configurations 86% of the times they occurred. The EWS was able to predict the exceeding of the rainfall threshold that generated floods 80% of times.
Climate change increasingly affects every aspect of human life. Recent studies report a close correlation with human health and it is estimated that global death rates will increase by 73 per 100,000 by 2100 due to changes in temperature. In this context, the present work aims to study the correlation between climate change and human health, on a global scale, using artificial intelligence techniques. Starting from previous studies on a smaller scale, that represent climate change and which at the same time can be linked to human health, four factors were chosen. Four causes of mortality, strongly correlated with the environment and climatic variability, were subsequently selected. Various analyses were carried out, using neural networks and machine learning to find a correlation between mortality due to certain diseases and the leading causes of climate change. Our findings suggest that anthropogenic climate change is strongly correlated with human health; some diseases are mainly related to risk factors while others require a more significant number of variables to derive a correlation. In addition, a forecast of victims related to climate change was formulated. The predicted scenario confirms that a prevalently increasing trend in climate change factors corresponds to an increase in victims.
Meteorological conditions play a crucial role in air pollution by affecting both directly and indirectly the emissions, transport, formation, and deposition of air pollutants. Extreme weather events can strongly affect surface air quality. Understanding relations between air pollutant concentrations and extreme weather events is a fundamental step toward improving the knowledge of how excessive heat impacts on air quality. In this work, we developed a statistical procedure for investigating the variations in the correlation structure of four air pollutants (NOx, O3, PM10, PM2.5) during extreme temperature events measured in monitoring sites located of Emilia Romagna region, Northern Italy, in summer (June–August) from 2015 to 2017. For the selected stations, Hot Days (HDs) and Heat Waves (HWs) were identified with respect to historical series of maximum temperature measured for a 30-year period (1971–2000). This method, based on multivariate techniques, allowed us to highlight the variations in air quality of study area due to the occurrence of HWs. The examined data, including PM concentrations, show higher values, whereas NOx and O3 concentrations seem to be not influenced by HWs. This operative procedure can be easily exported in other geographical areas for studying effects of climate change on a local scale.
Flow slow-down in rivers and artificial canals is a basic aspect to be monitored and kept strictly under control. Flow slow-downs can become a major concern in the event of extreme phenomena. The paper illustrates an advanced image processing method that uses particle tracking velocimetry in conjunction with a monadic approach to better characterize water flow in the presence of waste or debris that block normal water flow within a river. An high-speed camera installed beneath a bridge takes periodic images of the water flow. The measured water level and the images taken by the camera are sent to a central system in real-time. Results demonstrate the capability of the proposed method to accurately detect the presence of debris from the measured water flow.
Rural pipelines dedicated to water distribution, that is, waterworks, are essential for agriculture, notably plantations and greenhouse cultivation. Water is a primary resource for agriculture, and its optimized management is a key aspect. Saving water dispersion is not only an economic problem but also an environmental one. Spectral estimation of leakage is based on processing signals captured from sensors and/or transducers generally mounted on pipelines. There are different techniques capable of processing signals and displaying the actual position of leaks. Not all algorithms are suitable for all signals. That means, for pipelines located underground, for example, external vibrations affect the spectral response quality; then, depending on external vibrations/noises and flow velocity within pipeline, one should choose a suitable algorithm that fits better with the expected results in terms of leak position on the pipeline and expected time for localizing the leak. This paper presents findings related to the application of a decimated linear prediction (DLP) algorithm for agriculture and rural environments. In a certain manner, the application also detects the hydrodynamics of the water transportation. A general statement on the issue, DLP illustration, a real application and results are also included.
Control and monitoring of pipelines, dedicated to liquids transportation and distribution, are generally performed by active and passive techniques. With active techniques, we intend the possibility of using external sources to target pipelines and/or liquid inside the infrastructure. The used radiation can be partly reflected by the liquid and pipeline. Passive techniques are related to spontaneous emissions from liquids within the pipelines. However, capturing water displacements and fluctuations within the pipeline, by means of sensors and transducers, is somehow a passive technique. This latter is here illustrated thanks to pressure sensors mounted on an experimental pipeline. The basic idea is to detect leaks but also to monitor water flow regime by means of imaging without using a camera. That is possible thanks to the algorithm here developed and based on 2D/3D Decimated Signal Diagonalization (DSD). It is a technique used in nuclear magnetic resonance, and susceptible to bring to excellent results if water flow within a pipeline is considered as blood flow inside an artery or a vein. The approach has been applied to an experimental hydraulic circuit.
Artificial intelligence, in particular a supervised and unsupervised machine learning approach, has been becoming an interest in the field of measurement and instrumentation. Many problems of classification can be faced by a machine learning approach. We know machine learning is a broad area of artificial intelligence that comprises some other lines of research and activities such as deep learning. Synthetic aperture radar (SAR) measurements by means of its sensors are of great interest in environmental monitoring, in particular in land classification. This paper presents findings related to measurements and characterization through land classification of an environmentally sensitive area in Italy over two different time periods in order to assess changing parameters. A deep learning algorithm has been designed and implemented, and a comparison has been established with a spectral density approach.
In the last decades, Mediterranean rural landscapes have undergone significant changes, with relevant considerable environmental and socio-economic impacts. These phenomena are often triggered by agricultural abandonment, especially in environmentally-sensitive areas, which are usually located in marginal and less profitable regions, and which could indeed irremediably compromise the identity and role of these Mediterranean landscapes. On the other hand, the progressive increase of available multi-source geodata allows to reconstruct the landscape original structure, providing new tools able to prevent negative impacts on environment. Hence, thanks to the development of increasingly advanced and open-source GIS tools, it is possible to implement several geodata typologies that can be mutually integrated in an increasingly efficient approach. In this paper the process of landscape reshaping pattern is analyzed in a study area of Basilicata region (Southern Italy) using remote sensing. In particular, the vegetation component of a landscape has been assessed by means of SAR images by using an artificial intelligence approach, that is machine learning to understand landscape dynamics in two different time periods. In this way, it has been possible to integrate data of different source and composition into landscape analysis methodologies, hence developing a suitable tool for planning and managing the rural landscape.