We propose methods to enhance the predictive performance of generalized additive models (GAMs) in the context of covariate extrapolation, where predictions rely on covariates beyond their observed range. When using predictive models such as GAMs, shifts in the covariate distribution between training and prediction datasets can occur. Ignoring this issue may lead to inaccurate predictions in the tail of the covariate distributions. For example, this problem is particularly critical in climate-change scenarios, where covariates simulated from future climate scenarios are likely to contain more extreme conditions. Our approach integrates GAMs for the bulk of covariate distributions with asymptotic models from multivariate extreme-value theory at high covariate values. We consider binary responses based on a latent variable assumption, and also continuous responses. For large values of the covariates, on a specific marginal scale motivated by extreme-value theory the latent variable or continuous response is assumed to depend linearly on the covariates with an additive error term, when using an appropriate link function. In an application to wildfires in Europe, we explore how the new method can improve predictions, using environmental and meteorological covariates.
Climate change is reshaping the agroclimatic conditions for wheat, taking some past decennial events to frequent climatic features, substantially increasing the probability of crop stress and yield loss. We propose a method using ecoclimatic indicators combined with probability models, continuous-time regression and copulas, to evaluate — in non-stationary context —the future climate suitability of common wheat in France expressed as the probability of exceeding past decennial thresholds. Under a high-emission scenario, our analysis identifies four emerging risks from 2050 onward: early heat stress and warm nights during the flag-leaf-to-anthesis stage, late warm nights from anthesis to grain maturity, and devernalisation. Precipitation- and humidity-related extremes during the vegetative phase are projected to slightly increase or stagnate, with strong regional and model-specific variability. Mild winters are likely to become a dominant climatic feature by century’s end, while late cold stress events decline sharply. In contrast, reproductive-phase heat stress intensifies markedly, becoming a dominant agroclimatic constraint by century’s end, whereas risks linked to excess moisture or precipitation decrease. Drought-related stress remains mostly stable, due to shorter phenological cycles. By century’s end, under high emission scenario, compound hazards —such as simultaneous drought and heat stress during key reproductive stages—are projected to become 3 to 6 times more frequent than in the historical baseline. Compound mild-winter and wet- or damp-condition events increase 2.5- to 12-fold. Regional disparities emerge: the English Channel coast and the Paris Basin appear comparatively less exposed to these climatic risks, positioning them as potential, though not risk-free, future refuges for wheat cultivation. A low-emission mitigation pathway would reduce the frequency of these risks by a factor of 2 to 6, and reduce interannual risk variability by up to 45% relative to high emissions. These findings highlight priority targets for adaptation and underscore the urgent need for mitigation to safeguard wheat production.
Over the recent decades, Europe has experienced a significant decline in common bird species, particularly farmland species, due to anthropic pressures like agricultural intensification. Protected areas, such as the & Eacute;crins National Park (ENP) in France, can help mitigate these impacts. We evaluated whether an opportunistic presence-only dataset collected by trained ENP rangers contains biological signals strong enough to support robust statistical inference. Using a generalized additive Poisson model with spatial and spatio-temporal covariates, monthly latent spatio-temporal Gaussian random fields, and a non-spatial inter-annual effect, we estimated the relative abundance of 76 passerine species on a regular grid, with occurrences aggregated per spatio-temporal cell used as a proxy for sampling effort. The model showed good calibration for most species (AUC > 0.8) and reliably captured habitat preferences and migratory status. Relative-abundance trends in ENP were compared with relative abundance from three monitoring programs: STOM (ENP), STOC (France), and MHB (Switzerland). For most species with significant trends, model predictions aligned with survey-based trends. Forest specialists benefited most from the protected-area status, and farmland species declined more slowly in ENP than in France. High-elevation specialists generally decreased in both ENP and Switzerland. Discrepancies mostly arose for common species, likely reflecting uncorrected declines in ranger reporting rates. These results demonstrate that high-resolution opportunistic presence-only data can provide valuable insights into biological patterns and trends while reducing reliance on external data to estimate sampling effort.
Extreme events arising in georeferenced stochastic processes can take various forms, such as occurring in isolated patches or stretching contiguously over large areas, and can further vary with the spatial location and the extremeness of the events. We use excursion sets above threshold exceedances in data observed over a two-dimensional grid of rectangular pixels to propose a general family of coefficients that assess spatial-extent properties relevant for risk assessment, and study five candidate coefficients from this family. These coefficients are defined locally and interpreted as a spatial distance from a reference site where the threshold is exceeded. We develop statistical inference and discuss robustness to boundary effects and resolution of the pixel grid. To statistically extrapolate coefficients towards very high threshold levels, we formulate an asymptotically motivated semiparametric model and estimate a parameter characterizing how coefficients scale with the quantile level of the threshold. The utility of the new coefficients is illustrated through simulated data, as well as in an application to gridded daily temperature in continental France.
The problem of detecting the existence of a changepoint in a data sequence and of identifying its position is challenging when the focus is on extreme events and the distribution of data is heavy-tailed. In this setting, we recall several existing techniques, and then we propose a novel robust semi-parametric approach to changepoint identification that does not require the likelihood function. The changepoint is estimated as the position of the maximum of a statistic inspired by classical ANOVA to contrast the tail behavior of data to the left and right of all changepoint candidates. It is shown that the estimator consistently identifies the true changepoint under mild assumptions when the sample size increases. In numerical experiments, the novel method shows reliable finite-sample behavior for various simulation settings and is very competitive in comparison to the alternative changepoint identification approaches from the literature, especially for small sample sizes. Finally, the utility of the method is highlighted by identifying interpretable changepoints in three real-data applications: very large motor insurance claim amounts for a French administrative region with age as covariate; daily Bitcoin cryptocurrency price data and daily returns of stocks of the Boeing company both with time as covariate.
Modeling precipitation and its accumulation over time and space is essential for flood risk assessment. We here analyze rainfall data collected over several years through a microscale precipitation sensor network in Montpellier, France, by the OMSEV observatory. A novel spatio-temporal stochastic model is proposed for high-resolution urban rainfall and combines realistic marginal behavior and flexible extremal dependence structure. Rainfall intensities are described by the Extended Generalized Pareto Distribution (EGPD), capturing both moderate and extreme events without threshold selection. Based on spatial extreme-value theory, dependence during extreme episodes is modeled by an r-Pareto process with a non-separable variogram including episode-specific advection, allowing the displacement of rainfall cells to be represented explicitly. Parameters are estimated by a composite likelihood based on joint exceedances, and empirical advection velocities are derived from radar reanalysis. The model accurately reproduces the spatio-temporal structure of extreme rainfall observed in the Montpellier OMSEV network and enables realistic stochastic scenario generation for flood risk assessment.
Abstract. Agricultural droughts, hydrological droughts and wildfires have significant environmental and socioeconomic consequences. These hazards are physically linked because they share a number of forcings, and their space-time properties are important as impacts result from their spatial extent and duration, in addition to their intensity. This paper introduces a probabilistic model adapted to the description of multiple spatial hazards, based on the combination of simple ingredients: regressions to describe dependencies between hazards, principal component analysis to describe spatial dependence, and simple covariance functions to describe time dependence and residual spatial dependence. This results in a modular framework that decomposes a complex model into several simpler models. This model is then used to analyze the observed Soil Wetness Index, Fire Weather Index and river flows in France over the last six decades, and in particular to estimate the probability of occurrence of the remarkable 2022 summer event. Locally, the 2022 summer was extreme in terms of agricultural drought over a large part of the country, but was rather moderate in terms of hydrological drought and fire weather. However, the magnitude, spatial extent and duration of the event become extreme nearly everywhere when the three hazards are considered together. The underlying trends affecting all three hazards have more than doubled the probability of the event during the historical period, and future projections suggest that it might become common by the end of the century with global warming.
Spatial modelling of extreme values allows studying the risk of joint occurrence of extreme events at different locations and is of significant interest in climatic and other environmental sciences. A popular class of dependence models for spatial extremes is that of random location-scale mixtures, in which a spatial "baseline" process is multiplied or shifted by a random variable, potentially altering its extremal dependence behaviour. Gaussian location-scale mixtures retain benefits of their Gaussian baseline processes while overcoming some of their limitations, such as symmetry, light tails and weak tail dependence. We review properties of Gaussian location-scale mixtures and develop novel constructions with interesting features, together with a general algorithm for conditional simulation from these models. We leverage their flexibility to propose extended extreme-value models, that allow for appropriately modelling not only the tails but also the bulk of the data. This is important in many applications and avoids the need to explicitly select the events considered as extreme. We propose new solutions for likelihood inference in parametric models of Gaussian location-scale mixtures, in order to avoid the numerical bottleneck given by the latent location and scale variables that can lead to high computational cost of standard likelihood evaluations. The effectiveness of the models and of the inference methods is confirmed with simulated data examples, and we present an application to wildfire-related weather variables in Portugal. Although not detailed here, the approaches would also be straightforward to use for modelling multivariate (non spatial) data.
Precipitation modeling is of great interest for flood risk analysis. We present a framework for modeling the distribution and the spatio-temporal dependence of rainfall measured at high temporal resolution and fine spatial scale by the rain gauge network of the Montpellier urban observatory since 2019. This data is complemented by hourly radar reanalysis data from Meteo-France, available at 1 km resolution on a regular grid and for a longer time period. By applying a neural network downscaling approach from reanalysis to local point scale for the marginal distributions, we aim to obtain a finer resolution dataset, a longer data period and a better spatial coverage by leveraging information from the two data sources. For univariate modeling, at the point level, we use the Extended Generalized Pareto Distribution (EGPD). It allows us to model both moderate and intense rainfall simultaneously without explicit threshold selection, a step that is often challenging in statistics of extremes, and to reduce the complexity of parameter estimation. The spatio-temporal dependence is modeled using an r-Pareto process with an underlying gaussian dependence structure. Unlike max-stable processes, which are often limited by their focus on block maxima approaches, r-Pareto processes offer more flexibility and practicality for environmental applications by using a Peaks Over Threshold (POT) framework. By incorporating a non-separable spatio-temporal variogram with advection, we account for the horizontal movement of precipitation clouds, enabling realistic simulations of spatio-temporal rainfall patterns. A novel composite likelihood approach based on bivariate joint exceedance indicators used for variogram parameter estimation. The model is validated by simulations of the proposed process and is applied to rainfall data from Montpellier. This methodology will be at the core of a stochastic precipitation generator for the Montpellier region, which will be integrated into a mechanistic water flow model for flood risk analysis.
The concept of landslide hazard entails evaluating landslide occurrence in space (i.e., where landslides may occur), in time (i.e., when or how often landslides may occur), and their intensity (i.e., how destructive landslides may be). At regional scales, data-driven methods are implemented to separately analyze the spatial component (i.e., landslide susceptibility) and the temporal conditions leading to landslide occurrence, such as rainfall thresholds. However, assessing how large a landslide may develop once triggered is seldom conducted and poses a persistent challenge to satisfying the complete definition of landslide hazard. So far, only a few publications have addressed this issue by predicting the total areal extent of landslides based on certain mapping units, such as slope units. Limitations arise since the total areal extent of landslides within a mapping unit is strongly influenced by the size of the mapping unit, leading to larger mapping units being more likely to encompass larger total landslide areas. To tackle these challenges, this study aims to predict the landslide area proportion per slope unit in South Tyrol, Italy (7,400 km²). Our approach built upon past landslide occurrences from 2000 to 2020, systematically related to damage-causing and infrastructure-threatening landslide events. The method involved delineating slope units, filtering the landslide inventory, designing the sampling strategy, removing trivial areas, and aggregating the environmental variables (e.g., topography, lithology, land cover, and precipitation) to the slope unit partition. We tested a generalized additive beta regression model to estimate statistical relationships between the various static predictors and the target landslide areal density. The resulting spatially explicit predictions are evaluated through cross-validation from multiple perspectives. Applications and shortcomings of the approach are discussed. The proposed method is anticipated to provide valuable insights and alternatives to assessing landslide intensity and moving toward landslide hazard in a data-driven context. The outcomes associated with this research are framed within the PROSLIDE project, which has received funding from the research program Research Südtirol/Alto Adige 2019 of the Autonomous Province of Bozen/Bolzano – Südtirol/Alto Adige.
Landslides are geomorphic hazards in mountainous terrains across the globe, driven by a complex interplay of static and dynamic controls. Data‐driven approaches have been employed to assess landslide occurrence at regional scales by analyzing the spatial aspects and time‐varying conditions separately. However, the joint assessment of landslides in space and time remains challenging. This study aims to predict the occurrence of precipitation‐induced shallow landslides in space and time within the Italian province of South Tyrol (7,400 km 2 ). We introduce a functional predictor framework where precipitation is represented as a continuous time series, in contrast to conventional approaches that treat precipitation as a scalar predictor. Using hourly precipitation data and past landslide occurrences from 2012 to 2021, we implemented a functional generalized additive model to derive statistical relationships between landslide occurrence, various static scalar factors, and the preceding hourly precipitation as a functional predictor. We evaluated the resulting predictions through several cross‐validation routines, yielding performance scores frequently exceeding 0.90. To demonstrate the model predictive capabilities, we performed a hindcast for a storm event in the Passeier Valley on 4–5 August 2016, capturing the observed landslide locations and illustrating the hourly evolution of the predicted probabilities. Compared to standard early warning approaches, this framework eliminates the need to predefine fixed time windows for precipitation aggregation while inherently accounting for lagged effects. By integrating static and dynamic controls, this research advances the prediction of landslides in space and time for large areas, addressing seasonal effects and underlying data limitations.
Fire severity, or how an environment is affected by fire, can be estimated over large areas using remotely sensed indices like the Relative Burnt Ratio (RBR). RBR predictions typically rely on data from a single date just before the fire. However, predicting RBR accurately in both time and space remains challenging. To improve RBR predictability, we developed new models using time series data spanning several months before the fire. These models use fuel proxies derived from optical remote sensing and meteorological data. We applied this approach to fires in the French Mediterranean area during the summers of 2016-2021. We used a Lagged Generalized Additive Model (LGAM) and a Functional Linear Model (FLM) to estimate the influence of variables up to several months before the fire on RBR. A GAM fed with immediate pre-fire predictors served as a benchmark. Training and prediction were conducted at the fire-land-cover spatial scale using a training dataset spatially independent of the test dataset. FLM achieved the best prediction accuracy on test data (R=0.68, RMSE=0.057), outperforming LGAM (R=0.60, RMSE=0.063) and the benchmark (R=0.52, RMSE=0.069). FLM accurately predicted the highest RBR values when the Normalized Difference Vegetation Index decreased faster than the average and when the Duff Moisture Code increased faster than the average over the 65 days before the fire. The 17% decrease in the RMSE of FLM predictions compared to GAM predictions shows that understanding fuel dynamics up to two months before a fire provides valuable information for ranking areas by fire severity.
Background Weather conditions play a crucial role in driving fire activity in Mediterranean France. Previous research has demonstrated the influence of these conditions on the likelihood of large fire events over the world. However, certain limitations persist regarding the representation of fire weather in probabilistic models. Aim The objective of this paper is to develop an efficient method to rate fire danger by identifying the best representation of weather data for fire activity prediction in Mediterranean France. Methods We evaluated the performance of meteorological variables and the most common fire-weather indices (FWIs) worldwide as predictors of fire occurrence and size using the Firelihood framework, a probabilistic Bayesian model of fire activity. These models were compared to a fire activity baseline model incorporating only spatial and temporal effects but no explicit fire-weather information to allow for an in-depth study of information not captured by fire-weather indices. Key results The results indicate that relying solely on fire-weather indices is insufficient for efficient rating of fire activities. The inclusion of spatial and seasonal effects in the models is crucial for improving the indices' performance. While the Canadian FWI remains the most skillful indicator among the tested indices, using new combinations of several of its subcomponents further increases accuracy. Various performance analyses, including threshold selections, were carried out to comparatively assess those improvements. Conclusions The approach shows that probabilistic models informed with appropriately constructed fire-weather indices substantially improve various aspects of fire activity predictions in the Mediterranean area.
In France, year 2022 witnessed severe drought conditions, with very low flows in rivers starting already during the spring season and widespread wildfire occurrences in summer. In recent years, similar occurrences of consecutive droughts and wildfire hazards have been observed in other climatic regions of the world, including Greece, Portugal, Canary Islands, Canada, California, Australia, etc. These hazards can induce numerous strong socioeconomic impacts in areas such as agriculture, silviculture, energy, ecology, drinking water, civil protection, tourism, etc., and form a complex system of multiple drivers and risks interacting over space and time. Both the individual and the joint probabilities of occurrence of these multiple hazards driving the risks are expected to evolve with climate change. Characterizing the severity of such multiple hazards in probabilistic terms is challenging due to the multivariate nature of the problem, and the fact that each hazard has spatial structure and heterogeneity. In this presentation, we develop a relatively parsimonious stochastic model and estimation procedure to describe the joint space-time variability of three indices: (1) the Soil Wetness Index (SWI), used to characterize agricultural droughts (i.e. soil dryness); (2) River streamflow (Q), used to characterize hydrological droughts; (3) the Fire Weather Index (FWI), used to characterize fire-prone weather conditions. All indices are used at a monthly time step over the 1958-2023 period. SWI and FWI are derived from the SAFRAN atmospheric reanalysis and are available over Metropolitan France on a regular 8*8 km spatial grid (8597 pixels). Streamflow Q is measured at 232 streamgauging stations. The statistical model is based on a causal diagram where we postulate that agricultural drought (SWI) is a precursor for both hydrological drought (Q) and fire-prone conditions (FWI). The space-time distribution of SWI is therefore modeled first using a dimensionality-reduction method to provide a parsimonious description of the space-time variability of SWI. The distribution of Q is then modeled conditionally on the average value taken by SWI on each river catchment, using a generalized additive model for location, scale and shape (GAMLSS) regression. Similarly, the distribution of FWI is modeled conditionally on the value taken by SWI on the same pixel with a GAMLSS regression.Despite its simplicity, the stochastic model is shown to appropriately reproduce several key properties of the three studied hazards, in particular their joint probability of occurrence, their long-term trends and the distribution of the spatial extent or the duration of multi-hazard events. Future work will apply the model to future projections in order to estimate how these properties evolve under climate change. We finish by discussing the relevance of the proposed approach when extrapolated to extreme levels and whether or not this simple approach is adapted to other types of multiple hazards, such as heat + humidity or storm surge + flooding.
Extreme environmental events such as severe storms, drought, heat waves, flash floods, and abrupt species collapse have become more prevalent in the earth-atmosphere dynamic system in recent years. In order to fully understand the underlying mechanisms and enhance informed decision-making, a flexible model capable of accommodating extremes is necessary. Existing dynamic spatio-temporal statistical models exhibit limitations in capturing extremes when assuming Gaussian error distributions, whereas the current models for spatial extremes mostly assume temporal independence and are focused on joint upper tails at two or more locations. Here, we introduce a new class of dynamic spatio-temporal models that capture both high and low extremes using a mixture of heavy- and light-tailed distributions with varying tail indices. Our framework flexibly identifies extremal dependence and independence in both space and time with uncertainty quantification and supports missing data prediction, as in other dynamic spatio-temporal models. We demonstrate its effectiveness using a large reanalysis dataset of hourly particulate matter in the Central United States.
Environmental data science for spatial extremes has traditionally relied heavily on max-stable processes. Even though the popularity of these models has perhaps peaked with statisticians, they are still perceived and considered as the `state-of-the-art' in many applied fields. However, while the asymptotic theory supporting the use of max-stable processes is mathematically rigorous and comprehensive, we think that it has also been overused, if not misused, in environmental applications, to the detriment of more purposeful and meticulously validated models. In this paper, we review the main limitations of max-stable process models, and strongly argue against their systematic use in environmental studies. Alternative solutions based on more flexible frameworks using the exceedances of variables above appropriately chosen high thresholds are discussed, and an outlook on future research is given, highlighting recommendations moving forward and the opportunities offered by hybridizing machine learning with extreme-value statistics.
Citizen science mobilizes many observers and gathers huge datasets but often without strict sampling protocols, resulting in observation biases due to heterogeneous sampling effort, which can lead to biased predictions. We develop a spatio-temporal Bayesian hierarchical model for bias-corrected estimation of arrival dates of the first migratory bird individuals at their breeding sites. Higher sampling effort could be correlated with earlier observed dates. We implement data fusion of two citizen-science datasets with fundamentally different protocols (Breeding Bird Survey, eBird) and obtain posterior distributions of the latent process, which contains four spatial components endowed with Gaussian process priors: species niche; sampling effort; position and scale parameters of annual first arrival date. The data layer consists of four response variables: counts of observed eBird locations (Poisson); presence-absence at observed eBird locations (Binomial); BBS occurrence counts (Poisson); first arrival dates (generalized extreme-value). We devise a Markov chain Monte Carlo scheme and check by simulation that the latent process components are identifiable. We apply our model to several migratory bird species in the northeastern US for 2001-2021 and find that the sampling effort significantly modulates the observed first arrival dates. We exploit this relationship to effectively bias-correct predictions of the true first arrivals.
We respond to the discussion comments of the Proposer and Seconder of the vote of thanks, and to the eight other contributions discussing our work.
Evaluating the efficiency of lethal control of large carnivores such as wolves to reduce attacks on livestock is important given the controversy surrounding this measure. We used retrospective data over 10 years and an intra-site comparison approach to evaluate the effects of lethal control on the distribution of attack intensities in the French Alps. We analyzed 278 legal killings of wolves between 2011 and 2020 and the 6110 attacks that occurred during a period of ± 90 days and within 10 km around these lethal removals. This large number of attacks allowed us to perform an original framework that combined both continuous spatial and temporal scales through 3D kernel estimation. We also controlled the analysis for livestock presence, and explored different analysis subsets of removals in relation to their locations, dates and proximity to other removals. This statistical method provided an efficient visualization of attack intensity spatio-temporal distribution before and after removals. A decrease of the intensity of attacks was the most common result after the lethal removals of wolves. However, this outcome was not systematic for all subsets and depended on the scale of the analysis. In addition, attacks tended to persist after removals while showing frequent interruptions in time after but also before removals. We also observed localized positive trends of attack intensities at varying distances from removals after they occurred. To summarize, our results showed that considering the scale of the analysis is crucial and that effects should be analyzed separately for each local context. As a next step, we recommend to move forward from patterns to mechanisms by linking the effects of lethal control on wolves to their effects on attacks through analysis of fine-scaled data on wolves and livestock. ### Competing Interest Statement The authors have declared no competing interest.