With the recent development of new geometric and angular-radial frameworks for multivariate extremes, reliably simulating from angular variables in moderate-to-high dimensions is of increasing importance. Empirical approaches have the benefit of simplicity, and work reasonably well in low dimensions, but as the number of variables increases, they can lack the required flexibility and scalability. Classical parametric models for angular variables, such as the von Mises–Fisher distribution (vMF), provide an alternative. Exploiting finite mixtures of vMF distributions increases their flexibility, but there are cases where, without letting the number of mixture components grow considerably, a mixture model with a fixed number of components is not sufficient to capture the intricate features that can arise in data. Owing to their flexibility, generative deep learning methods are able to capture complex data structures; they therefore have the potential to be useful in the simulation of multivariate angular variables. In this paper, we introduce a range of deep learning approaches for this task, including generative adversarial networks, normalizing flows and flow matching. We assess their performance via a range of metrics, and make comparisons to the more classical approach of using a finite mixture of vMF distributions. The methods are also applied to a metocean data set, with diagnostics indicating strong performance, demonstrating the applicability of such techniques to real-world, complex data structures.
Environmental contours are often used in engineering applications to describe risky combinations of variables according to some definition of an exceedance probability. These contours can be used to both understand multivariate extreme events in environmental processes and mitigate against their effects, e.g. in the design of structures. Such ideas are also useful in other disciplines, with the types of extreme events of interest depending on the context. Despite clear connections with extreme value modelling, much of this methodology has so far not been exploited in the estimation of environmental contours; in this work, we provide a way to unify these areas. We focus on the bivariate case, introducing two new definitions of environmental contours. We develop techniques for their inference which exploit a non-standard radial and angular decomposition of the variables, building on previous work for estimating limit sets. Specifically, we model the upper tails of the radial distribution using a generalised Pareto distribution, with adaptable smoothing of the parameters of this distribution. Our methods work equally well for asymptotically independent and asymptotically dependent variables, so do not require us to distinguish between different joint tail forms. Simulations demonstrate reasonable success of the estimation procedure, and we apply our approach to an air pollution data set, which is of interest in the context of environmental impacts on health. Supplementary materials accompanying this paper appear online.
It is no secret that statistical modelling often involves making simplifying assumptions when attempting to study complex stochastic phenomena. Spatial modelling of extreme values is no exception, with one of the most common such assumptions being stationarity in the marginal and/or dependence features. If non-stationarity has been detected in the marginal distributions, it is tempting to try to model this while assuming stationarity in the dependence, without necessarily putting this latter assumption through thorough testing. However, margins and dependence are often intricately connected and the detection of non-stationarity in one feature might affect the detection of non-stationarity in the other. This work is an in-depth case study of this interrelationship, with a particular focus on a spatio-temporal environmental application exhibiting well-documented marginal non-stationarity. Specifically, we compare and contrast four different marginal detrending approaches in terms of our post-detrending ability to detect temporal non-stationarity in the spatial extremal dependence structure of a sea surface temperature dataset from the Red Sea.
The key to successful statistical analysis of bivariate extreme events lies in flexible modelling of the tail dependence relationship between the two variables. In the extreme value theory literature, various techniques are available to model separate aspects of tail dependence, based on different asymptotic limits. Results from Balkema and Nolde (2010) and Nolde (2014) highlight the importance of studying the limiting shape of an appropriately-scaled sample cloud when characterising the whole joint tail. We now develop the first statistical inference for this limit set, which has considerable practical importance for a unified inference framework across different aspects of the joint tail. Moreover, Nolde and Wadsworth (2022) link this limit set to various existing extremal dependence frameworks. Hence, a by-product of our new limit set inference is the first set of self-consistent estimators for several extremal dependence measures, avoiding the current possibility of contradictory conclusions. In simulations, our limit set estimator is successful across a range of distributions, and the corresponding extremal dependence estimators provide a major joint improvement and small marginal improvements over existing techniques. We consider an application to sea wave heights, where our estimates successfully capture the expected weakening extremal dependence as the distance between locations increases.
The conditional extremes framework allows for event-based stochastic modeling of dependent extremes, and has recently been extended to spatial and spatio-temporal settings. After standardizing the marginal distributions and applying an appropriate linear normalization, certain non-stationary Gaussian processes can be used as asymptotically-motivated models for the process conditioned on threshold exceedances at a fixed reference location and time. In this work, we adapt existing conditional extremes models to allow for the handling of large spatial datasets. This involves specifying the model for spatial observations at d locations in terms of a latent m≪ d dimensional Gaussian model, whose structure is specified by a Gaussian Markov random field. We perform Bayesian inference for such models for datasets containing thousands of observation locations using the integrated nested Laplace approximation, or INLA. We explain how constraints on the spatial and spatio-temporal Gaussian processes, arising from the conditioning mechanism, can be implemented through the latent variable approach without losing the computationally convenient Markov property. We discuss tools for the comparison of models via their posterior distributions, and illustrate the flexibility of the approach with gridded Red Sea surface temperature data at over 6,000 observed locations. Posterior sampling is exploited to study the probability distribution of cluster functionals of spatial and spatio-temporal extreme episodes.
This paper details a methodology proposed for the EVA 2021 conference data challenge. The aim of this challenge was to predict the number and size of wildfires over the contiguous US between 1993 and 2015, with more importance placed on extreme events. In the data set provided, over 14% of both wildfire count and burnt area observations are missing; the objective of the data challenge was to estimate a range of marginal probabilities from the distribution functions of these missing observations. To enable this prediction, we make the assumption that the marginal distribution of a missing observation can be informed using non-missing data from neighbouring locations. In our method, we select spatial neighbourhoods for each missing observation and fit marginal models to non-missing observations in these regions. For the wildfire counts, we assume the compiled data sets follow a zero-inflated negative binomial distribution, while for burnt area values, we model the bulk and tail of each compiled data set using non-parametric and parametric techniques, respectively. Cross validation is used to select tuning parameters, and the resulting predictions are shown to significantly outperform the benchmark method proposed in the challenge outline. We conclude with a discussion of our modelling framework, and evaluate ways in which it could be extended.
Recent extreme value theory literature has seen significant emphasis on the modelling of spatial extremes, with comparatively little consideration of spatio-temporal extensions. This neglects an important feature of extreme events: their evolution over time. Many existing models for the spatial case are limited by the number of locations they can handle; this impedes extension to space-time settings, where models for higher dimensions are required. Moreover, the spatio-temporal models that do exist are restrictive in terms of the range of extremal dependence types they can capture. Recently, conditional approaches for studying multivariate and spatial extremes have been proposed, which enjoy benefits in terms of computational efficiency and an ability to capture both asymptotic dependence and asymptotic independence. We extend this class of models to a spatio-temporal setting, conditioning on the occurrence of an extreme value at a single space-time location. We adopt a composite likelihood approach for inference, which combines information from full likelihoods across multiple space-time conditioning locations. We apply our model to Red Sea surface temperatures, show that it fits well using a range of diagnostic plots, and demonstrate how it can be used to assess the risk of coral bleaching attributed to high water temperatures over consecutive days.
Vine copulas are a type of multivariate dependence model, composed of a collection of bivariate copulas that are combined according to a specific underlying graphical structure. Their flexibility and practicality in moderate and high dimensions have contributed to the popularity of vine copulas, but relatively little attention has been paid to their extremal properties. To address this issue, we present results on the tail dependence properties of some of the most widely studied vine copula classes. We focus our study on the coefficient of tail dependence and the asymptotic shape of the sample cloud, which we calculate using the geometric approach of Nolde (2014). We offer new insights by presenting results for trivariate vine copulas constructed from asymptotically dependent and asymptotically independent bivariate copulas, focusing on bivariate extreme value and inverted extreme value copulas, with additional detail provided for logistic and inverted logistic examples. We also present new theory for a class of higher dimensional vine copulas, constructed from bivariate inverted extreme value copulas.
This paper details the approach of team Lancaster to the 2019 EVA data challenge, dealing with spatio-temporal modelling of Red Sea surface temperature anomalies. We model the marginal distributions and dependence features separately; for the former, we use a combination of Gaussian and generalised Pareto distributions, while the dependence is captured using a localised Gaussian process approach. We also propose a space-time moving estimate of the cumulative distribution function that takes into account spatial variation and temporal trend in the anomalies, to be used in those regions with limited available data. The team’s predictions are compared to results obtained via an empirical benchmark. Our approach performs well in terms of the threshold-weighted continuous ranked probability score criterion, chosen by the challenge organiser.
In multivariate extreme value analysis, the nature of the extremal dependence between variables should be considered when selecting appropriate statistical models. Interest often lies in determining which subsets of variables can take their largest values simultaneously while the others are of smaller order. Our approach to this problem exploits hidden regular variation properties on a collection of nonstandard cones, and provides a new set of indices that reveal aspects of the extremal dependence structure not available through existing measures of dependence. We derive theoretical properties of these indices, demonstrate their utility through a series of examples, and develop methods of inference that also estimate the proportion of extremal mass associated with each cone. We apply the methods to river flows in the U.K., estimating the probabilities of different subsets of sites being large simultaneously.
Thorne et al. (2004), Torsvik et al. (2010; 2006) and Burke et al. (2008) have suggested that the locations of melting anomalies (“hot spots”) and the original locations of large igneous provinces (“LIPs”) and kimberlite pipes, lie preferentially above the margins of two “large lower-mantle shear velocity provinces”, or LLSVPs, near the bottom of the mantle, and that the geographical correlations have high confidence levels (> 99.9999%) (Burke et al., 2008, Fig. 5). They conclude that the LLSVP margins are “PlumeGeneration Zones”, and that deep-mantle plumes cause hot spots, LIPs, and kimberlites. This conclusion raises questions about what physical processes could be responsible, because, for example, the LLSVPs are apparently dense and not abnormally hot (Trampert et al., 2004). The supposed LIP-hot spot-LLSVP correlations probably are examples of the “Hindsight Heresy” (Acton, 1959), of performing a statistical test using the same data sample that led to the initial formulation of a hypothesis. In this process, an analyst will consider and reject many competing hypotheses, but will not adjust statistical assessments correspondingly. Furthermore, an analyst will test extreme deviations of the data, , but not take this fact into account. “Hindsight heresy” errors are particularly problematical in Earth science, where it often is impossible to conduct controlled experiments. For random locations on the globe, the number of points within a specified distance of a given curve follows a cumulative binomial distribution. We use this fact to test the statistical significance of the observed hot spot-LLSVP correlation using several hot-spot catalogs and mantle models. The results indicate that the actual confidence levels of the correlations are two or three orders of magnitude smaller than claimed. The tests also show that hot spots correlate well with presumably shallowly rooted features such as spreading plate boundaries. Nevertheless, the correlations are significant at confidence levels in excess of 99%. But this is confidence that the null hypothesis of random coincidence is wrong. It is not confidence about what hypothesis is correct. The correlations probably are symptoms of as-yet-unidentified processes.
Air pollution is made up of a number of different components, including sulphur oxides, nitrogen oxides, carbon monoxide and particulates. Particulates are small particles of solid or liquid material in the air. Two particulates of interest are PM2.5 and PM10; particles smaller than 2.5 and 10 micrometres respectively. Many studies claim that PM2.5 has a huge impact on health, with particular links to respiratory diseases. Smaller particles stay in the air longer, meaning that they have more time to cause significant damage.