Wildfires cause extensive damage to physical assets exposed to them. So far, assessing the risk of these events remains an understudied area of global disaster risk assessment. Probabilistic risk estimates covering the range and likelihood of devastating events are crucial for various applications such as prioritising adaptation measures and determining insurance pricing. Quantifying tail risks such as a one-in-a-hundred-year impact has important implications for disaster risk management, including the pricing of insurance. However, short observational time series render modelling efforts indispensable for risk assessments on a global scale. In parallel, increasing data availability allows for the use of machine learning techniques to predict wildfire behaviour. In this context, an open-source wildfire risk model based on globally available data would facilitate the accessibility of such analysis to stakeholders from both the public and private sector. Here, we present such a machine learning model that estimates wildfire probabilities and we integrate these within a global socio-economic risk framework. We determine burning probabilities based on MODIS burnt area, a set of predictors and a country-and-biome specific machine learning model. The chosen predictors include weather variables, land use covariates and population density. We enhance the model with spatial and temporal feature-engineered covariates, such as the count of neighbouring burnt cells and time since the last fire in each cell. The model employs XGBoost, a tree boosting system, tailored for each country and biome. The model generates stochastic, counterfactual historic wildfire seasons by leveraging the inherent randomness in its predictions, further influenced by temporal and spatial covariates.Secondly, we compute socio-economic impacts as the combination of the newly developed wildfire hazard, an exposure representing physical assets; and a vulnerability that was calibrated on historic fire damage data. We compute wildfire risks by combining the resulting impacts with their respective probabilities. This renders a globally consistent modelling approach of wildfire risk to physical assets. Our model's stochastic representation of wildfire hazards enables the analysis of extreme events with return periods extending beyond available observational data, enhancing our understanding of potential high-impact scenarios.
When extreme weather events affect large areas, their regional to sub-continental spatial scale is important for their impacts. We propose a novel machine learning (ML) framework that integrates spatial extreme-value theory to model weather extremes and to quantify probabilities associated with the occurrence, intensity, and spatial extent of these events. Our approach employs new loss functions adapted to extreme values, enabling our model to prioritize the tail rather than the bulk of the data distribution. Applied to a case study of Western European summertime heat extremes, we use daily 500-hPa geopotential height fields and local soil moisture as predictors to capture the complex interplay between local and remote physical processes. Our generative model reveals that different facets of heat extremes are influenced by individual circulation features, such as the relative position of upper-level ridges and troughs that are part of a large-scale wave pattern. This enriches our process understanding from a data-driven perspective. Our approach can extrapolate beyond the range of the data to make risk-related probabilistic statements. It applies more generally to other weather extremes and offers an alternative to traditional physical and ML-based techniques that focus less on the extremal aspects of weather data.
Rainfall return levels are guiding hazard protection, insurance models, infrastructure design, construction, and planning of cities. When deriving information about the frequency and intensity of extremes by fitting extreme value models to pointwise observations, the regionalization of these models is challenging. Rain gauges are distributed unevenly, where some regions suffer from data scarcity in space and time. To address this, topographical and/or climatological covariates are often used for the spatial interpolation. On the other hand, high-resolution climate simulations are available to provide spatial information on rainfall extremes. However, these simulations are still governed by model biases, where the bias adjustment of extremes at ungauged locations is also inducing uncertainty. In this study, we propose a combination of observations and a high-resolution convection-permitting climate model simulation in the framework of smooth spatial Generalized Extreme Value (GEV) models in order to estimate spatial rainfall return levels. We choose a study area over southern Germany with complex terrain, which is densely monitored with 1132 rain gauges providing more than 30-year daily rainfall observations. There, a 30-year simulation of the Weather and Forecasting Research (WRF) model is available at 1.5 km resolution driven by ERA5 reanalysis data. We combine observations and covariates from the WRF simulation in the spatial GEV and refer to this approach as sGEV-WRF.We want to answer three research questions to assess the added value of the proposed framework:Is it worth the effort? Does the sGEV-WRF improve the generation of rainfall return levels compared to the WRF alone? Does the WRF simulation as covariate add value? Can the sGEV-WRF outperform a topography-only spatial GEV? Does the dynamical downscaling at high resolution add value? Can sGEV-WRF outperform a spatial GEV based on observations and covariates from the coarser resolution ERA5? By evaluating the percentage bias, mean absolute error, and root-mean-square error, we show that the combination of observations and WRF can improve the representation of 10-year and 100-year return levels of daily rainfall.In addition, we aim to assess the performance of this framework under data-scarce conditions. Therefore, we devise an extensive cross-validation study. We select 5%, 10%, 20%, 50%, 80%, 90%, and 95% of all 1132 rain gauges to re-build the spatial GEV models with 1000 random folds each. We show that the performance is robust under these conditions, highlighting the potential for the application in data-scarce regions. Furthermore, in a non-stationary setup with climate model future projections, it can serve as a reliable tool to assess climate change effects on heavy rainfall.
Citizen science mobilizes many observers and gathers huge datasets but often without strict sampling protocols, resulting in observation biases due to heterogeneous sampling effort, which can lead to biased predictions. We develop a spatio-temporal Bayesian hierarchical model for bias-corrected estimation of arrival dates of the first migratory bird individuals at their breeding sites. Higher sampling effort could be correlated with earlier observed dates. We implement data fusion of two citizen-science datasets with fundamentally different protocols (Breeding Bird Survey, eBird) and obtain posterior distributions of the latent process, which contains four spatial components endowed with Gaussian process priors: species niche; sampling effort; position and scale parameters of annual first arrival date. The data layer consists of four response variables: counts of observed eBird locations (Poisson); presence-absence at observed eBird locations (Binomial); BBS occurrence counts (Poisson); first arrival dates (generalized extreme-value). We devise a Markov chain Monte Carlo scheme and check by simulation that the latent process components are identifiable. We apply our model to several migratory bird species in the northeastern US for 2001-2021 and find that the sampling effort significantly modulates the observed first arrival dates. We exploit this relationship to effectively bias-correct predictions of the true first arrivals.
Severe thunderstorms cause substantial economic and human losses in the United States. Simultaneous high values of convective available potential energy (CAPE) and storm relative helicity (SRH) are favorable to severe weather, and both they and the composite variable $\mathrm{PROD}=\sqrt{\mathrm{CAPE}} \times \mathrm{SRH}$ can be used as indicators of severe thunderstorm activity. Their extremal spatial dependence exhibits temporal non-stationarity due to seasonality and large-scale atmospheric signals such as El Ni\~no-Southern Oscillation (ENSO). In order to investigate this, we introduce a space-time model based on a max-stable, Brown--Resnick, field whose range depends on ENSO and on time through a tensor product spline. We also propose a max-stability test based on empirical likelihood and the bootstrap. The marginal and dependence parameters must be estimated separately owing to the complexity of the model, and we develop a bootstrap-based model selection criterion that accounts for the marginal uncertainty when choosing the dependence model. In the case study, the out-sample performance of our model is good. We find that extremes of PROD, CAPE and SRH are generally more localized in summer and, in some regions, less localized during El Ni\~no and La Ni\~na events, and give meteorological interpretations of these phenomena.
This paper details the approach of the team Kohrrelation in the 2021 Extreme Value Analysis data challenge, dealing with the prediction of wildfire counts and sizes over the contiguous US. Our approach uses ideas from extreme-value theory in a machine learning context with theoretically justified loss functions for gradient boosting. We devise a spatial cross-validation scheme and show that in our setting it provides a better proxy for test set performance than naive cross-validation. The predictions are benchmarked against boosting approaches with different loss functions, and perform competitively in terms of the score criterion, finally placing second in the competition ranking.
Accurate spatiotemporal modeling of conditions leading to moderate and large wildfires provides better understanding of mechanisms driving fire-prone ecosystems and improves risk management. We here develop a joint model for the occurrence intensity and the wildfire size distribution by combining extreme-value theory and point processes within a novel Bayesian hierarchical model, and use it to study daily summer wildfire data for the French Mediterranean basin during 1995–2018. The occurrence component models wildfire ignitions as a spatiotemporal log-Gaussian Cox process. Burnt areas are numerical marks attached to points and are considered as extreme if they exceed a high threshold. The size component is a two-component mixture varying in space and time that jointly models moderate and extreme fires. We capture non-linear influence of covariates (Fire Weather Index, forest cover) through component-specific smooth functions, which may vary with season. We propose estimating shared random effects between model components to reveal and interpret common drivers of different aspects of wildfire activity. This leads to increased parsimony and reduced estimation uncertainty with better predictions. Specific stratified subsampling of zero counts is implemented to cope with large observation vectors. We compare and validate models through predictive scores and visual diagnostics. Our methodology provides a holistic approach to explaining and predicting the drivers of wildfire activity and associated uncertainties.
<p>Wildfires are devastating events destroying large parts of physical assets exposed to them in many regions of the world. Therefore, a high-resolution hazard model is needed to accurately assess socio-economic impacts caused by wildfires. Moreover, a probabilistic representation of the hazard covering the range and likelihood of possible wildfire events under certain conditions allows for a more comprehensive risk assessment. This is crucial for many applications, among others the prioritization of adaptation measures and the pricing of insurance.</p><p>We determine burning probabilities based on MODIS hotspots and a set of predictors (weather variables, geography, land use) by using a country-specific machine learning model based on the efficient tree boosting system XGBoost. Subsequently, stochastic wildfire events are generated on the basis of these burning probabilities.</p><p>Lastly, the open-source climate risk assessment platform CLIMADA is used to compute socio-economic impacts as the combination of the newly developed hazard, an exposure and a vulnerability. The used exposure LitPop spatially distributes macroeconomic indicators (e.g. produced capital) as a function of night light intensity and population density. The vulnerability is represented by an impact function that was calibrated on historic fire damage data. Combining the stochastic impacts with their respective probabilities results in a globally consistent country-specific model of wildfire risk to physical assets.</p>
We review how machine learning has transformed our ability to model the Earth system, and how we expect recent breakthroughs to benefit end-users in Switzerland in the near future. Drawing from our review, we identify three recommendations. Recommendation 1: Develop Hybrid AI-Physical Models: Emphasize the integration of AI and physical modeling for improved reliability, especially for longer prediction horizons, acknowledging the delicate balance between knowledge-based and data-driven components required for optimal performance. Recommendation 2: Emphasize Robustness in AI Downscaling Approaches, favoring techniques that respect physical laws, preserve inter-variable dependencies and spatial structures, and accurately represent extremes at the local scale. Recommendation 3: Promote Inclusive Model Development: Ensure Earth System Model development is open and accessible to diverse stakeholders, enabling forecasters, the public, and AI/statistics experts to use, develop, and engage with the model and its predictions/projections.
Extreme precipitation events that occur in close succession can have important societal and economic repercussions. Here we use 42 years of reanalysis data (ERA-5) to investigate the link between Euro-Atlantic large-scale pattern of weather and climate variability and the temporal clustering of extreme rainfall events over Europe. We implicitly model the seasonal rate of extreme occurrences as part of a Poisson General Additive Model (GAM) using cyclic regression cubic splines. The smoothed seasonal rate of extreme rainfall occurrences is used to (i) infer the frequency of significant temporal clustering and (ii) implicitly serves as the baseline rate when modeling the effects of atmospheric drivers on extreme rainfall clustering. We use GAMs to model the association between the temporal clustering of extreme rainfall events and seven predominant year-round weather regimes in the Euro-Atlantic sector as well as a measure of synoptic-scale transient recurrent Rossby wave packets. Sub-seasonal clustering of precipitation events is significant at all grid-points over Europe; the proportion of extreme rainfall events that cluster in time ranges between 2% to 27%. The most relevant weather regime is the Atlantic Trough (corresponding to NAO+ with a southward shift of the jet) explaining most of the significant increase in clustering probability over Europe. The Greenland Blocking regime explains most of the clustering over the Iberian Peninsula. The Scandinavian Blocking regime is associated with a significant increase in clustering probability over the western Mediterranean, with a northwards shift in the signal to central Europe in summer.
Severe thunderstorms can have devastating impacts. Concurrently high values of convective available potential energy (CAPE) and storm relative helicity (SRH) are known to be conducive to severe weather, so high values of PROD=$\sqrt{\mathrm{CAPE}} \times$SRH have been used to indicate high risk of severe thunderstorms. We consider the extreme values of these three variables for a large area of the contiguous US over the period 1979-2015, and use extreme-value theory and a multiple testing procedure to show that there is a significant time trend in the extremes for PROD maxima in April, May and August, for CAPE maxima in April, May and June, and for maxima of SRH in April and May. These observed increases in CAPE are also relevant for rainfall extremes and are expected in a warmer climate, but have not previously been reported. Moreover, we show that the El Ni\~no-Southern Oscillation explains variation in the extremes of PROD and SRH in February. Our results suggest that the risk from severe thunderstorms in April and May is increasing in parts of the US where it was already high, and that the risk from storms in February tends to be higher over the main part of the region during La Ni\~na years. Our results differ from those obtained in earlier studies using extreme-value techniques to analyze a quantity similar to PROD.