We analyze synoptic observational data on extreme daily minimum temperatures in winter months over Ireland from 1950-2022. We model the tail of the marginal distributions of extreme winter minima using a generalized Pareto distribution capturing temporal and spatial sources of non-stationarity. We disentangle long-term climate trends from obfuscating shorter-term (month-to-winter long) large fluctuations, caused by, for example, anomalous behavior of the jet stream. We identify extreme spatial events with respect to a carefully chosen risk function and fit an -Pareto process to extreme events exceeding a high-risk threshold. We show that long-term trends over Ireland of extremely cold winter temperatures are warming at a faster rate than both mean winter temperatures and summer extreme daily maximum temperatures. Critically, we show that if we did not account for large-scale, short-term climatic oscillations, we would incorrectly estimate that the extremely cold winter temperatures were getting colder. In terms of spatial extreme events, we find that in periods where the jet stream typically produces the coldest winter temperatures, the estimated rate at which temperatures could fall below the coldest value ever recorded has decreased by a factor of 100 over the study period.
Max-autogressive moving average (Max-ARMA) processes are powerful tools for modelling time series data with heavy-tailed behaviour; these are a non-linear version of the popular autoregressive moving average models. River flow data typically have features of heavy tails and non-linearity, as large precipitation events cause sudden spikes in the data that then exponentially decay. Therefore, stationary Max-ARMA models are a suitable candidate for capturing the unique temporal dependence structure exhibited by river flows. This paper contributes to advancing our understanding of the extremal properties of stationary Max-ARMA processes. We detail the first approach for deriving the extremal index, the lagged asymptotic dependence coefficient, and an efficient simulation for a general Max-ARMA process. We use the extremal properties, coupled with the belief that Max-ARMA processes provide only an approximation to extreme river flow, to fit such a model which can broadly capture river flow behaviour over a high threshold. We make our inference under a reparametrisation which gives a simpler parameter space that excludes cases where any parameter is non-identifiable. We illustrate results for river flow data from the UK River Thames.
Fully describing the entire data set is essential in multivariate risk assessment, since moderate levels of one variable can influence another, potentially leading it to be extreme. Additionally, modelling both non-extreme and extreme events within a single framework avoids the need to select a threshold vector used to determine an extremal region, or the requirement to add flexibility to bridge between separate models for the body and tail regions. We propose a copula model, based on a mixture of Gaussian distributions, as this model avoids the need to define an extremal region, it is scalable to dimensions beyond the bivariate case, and it can handle both asymptotic dependent and asymptotic independent extremal dependence structures. We apply the proposed model through simulations and to a 5-dimensional seasonal air pollution data set, previously analysed in the multivariate extremes literature. Through pairwise, trivariate and 5-dimensional analyses, we show the flexibility of the Gaussian mixture copula in capturing different joint distributional behaviours and its ability to identify potential graphical structure features, both of which can vary across the body and tail regions.
We derive some key extremal features for $k$th order Markov chains that can be used to understand how the process moves between an extreme state and the body of the process. The chains are studied given that there is an exceedance of a threshold, as the threshold tends to the upper endpoint of the distribution. Unlike previous studies with $k>1$, we consider processes where standard limit theory describes each extreme event as a single observation without any information about the transition to and from the body of the distribution. Our work uses different asymptotic theory which results in non-degenerate limit laws for such processes. We study the extremal properties of the initial distribution and the transition probability kernel of the Markov chain under weak assumptions for broad classes of extremal dependence structures that cover both asymptotically dependent and asymptotically independent Markov chains. For chains with $k>1$, the transition of the chain away from the exceedance involves novel functions of the $k$ previous states, in comparison to just the single value, when $k=1$. This leads to an increase in the complexity of determining the form of this class of functions, their properties and the method of their derivation in applications. We find that it is possible to derive an affine normalization, dependent on the threshold excess, such that non-degenerate limiting behaviour of the process is assured for all lags. These normalization functions have an attractive structure that has parallels to the Yule-Walker equations. Furthermore, the limiting process is always linear in the innovations. We illustrate the results with the study of $k$th order stationary Markov chains with exponential margins based on widely studied families of copula dependence structures.
Threshold selection is a fundamental problem in any threshold-based extreme value analysis. While models are asymptotically motivated, selecting an appropriate threshold for finite samples is difficult and highly subjective through standard methods. Inference for high quantiles can also be highly sensitive to the choice of threshold. Too low a threshold choice leads to bias in the fit of the extreme value model, while too high a choice leads to unnecessary additional uncertainty in the estimation of model parameters. We develop a novel methodology for automated threshold selection that directly tackles this bias-variance trade-off. We also develop a method to account for the uncertainty in the threshold estimation and propagate this uncertainty through to high quantile inference. Through a simulation study, we demonstrate the effectiveness of our method for threshold selection and subsequent extreme quantile estimation, relative to the leading existing methods, and show how the method's effectiveness is not sensitive to the tuning parameters. We apply our method to the well-known, troublesome example of the River Nidd dataset.
Environmental contours are often used in engineering applications to describe risky combinations of variables according to some definition of an exceedance probability. These contours can be used to both understand multivariate extreme events in environmental processes and mitigate against their effects, e.g. in the design of structures. Such ideas are also useful in other disciplines, with the types of extreme events of interest depending on the context. Despite clear connections with extreme value modelling, much of this methodology has so far not been exploited in the estimation of environmental contours; in this work, we provide a way to unify these areas. We focus on the bivariate case, introducing two new definitions of environmental contours. We develop techniques for their inference which exploit a non-standard radial and angular decomposition of the variables, building on previous work for estimating limit sets. Specifically, we model the upper tails of the radial distribution using a generalised Pareto distribution, with adaptable smoothing of the parameters of this distribution. Our methods work equally well for asymptotically independent and asymptotically dependent variables, so do not require us to distinguish between different joint tail forms. Simulations demonstrate reasonable success of the estimation procedure, and we apply our approach to an air pollution data set, which is of interest in the context of environmental impacts on health. Supplementary materials accompanying this paper appear online.
We analyse extreme daily minimum temperatures in winter months over the island of Ireland from 1950-2022. We model the marginal distributions of extreme winter minima using a generalised Pareto distribution (GPD), capturing temporal and spatial non-stationarities in the parameters of the GPD. We investigate two independent temporal non-stationarities in extreme winter minima. We model the long-term trend in magnitude of extreme winter minima as well as short-term, large fluctuations in magnitude caused by anomalous behaviour of the jet stream. We measure magnitudes of spatial events with a carefully chosen risk function and fit an r-Pareto process to extreme events exceeding a high-risk threshold. Our analysis is based on synoptic data observations courtesy of Met Éireann and the Met Office. We show that the frequency of extreme cold winter events is decreasing over the study period. The magnitude of extreme winter events is also decreasing, indicating that winters are warming, and apparently warming at a faster rate than extreme summer temperatures. We also show that extremely cold winter temperatures are warming at a faster rate than non-extreme winter temperatures. We find that a climate model output previously shown to be informative as a covariate for modelling extremely warm summer temperatures is less effective as a covariate for extremely cold winter temperatures. However, we show that the climate model is useful for informing a non-extreme temperature model.
Non-stationary methods of flood frequency analysis are widespread in research but rarely implemented by practitioners. One reason may be that research papers on non-stationary statistical models tend to focus on model fitting rather than extracting the sort of results needed by designers and decision makers. It can be difficult to extract useful results from non-stationary models that include stochastic covariates for which the value in any future year is unknown. We explore the motivation for including such covariates, whether on their own or in addition to a covariate based on time. We set out a method for expressing the results of non-stationary models as an integrated flow estimate, which removes the dependence on the covariates. This can be defined either for a particular year or over a longer period of time. The methods are illustrated by application to a set of 375 river gauges across England and Wales. We find annual rainfall to be a useful covariate at many gauges, sometimes in conjunction with a time-based covariate. For estimating flood frequency in future conditions, we advocate exploring hybrid approaches that combine the best attributes of non-stationary statistical models and simulation models that can represent changes in climate and river catchments.
Extreme value analysis (EVA) uses data to estimate long-term extreme environmental conditions for variables such as significant wave height and period, for the design of marine structures. Together with models for the short-term evolution of the ocean environment and for wave-structure interaction, EVA provides a basis for full probabilistic design analysis. Alternatively, environmental contours provide an approximate approach to estimating structural integrity, without requiring structural knowledge. These contour methods also exploit statistical models, including EVA, but avoid the need for structural modelling by making what are believed to be conservative assumptions about the shape of the structural failure boundary in the environment space. These assumptions, however, may not always be appropriate, or may lead to unnecessary wasted resources from over design. We demonstrate a methodology for efficient fully probabilistic analysis of structural failure. From this, we estimate the joint conditional probability density of the environment (CDE), given the occurrence of an extreme structural response. We use CDE as a diagnostic to highlight the deficiencies of environmental contour methods for design; none of the IFORM environmental contours considered characterise CDE well for three example structures.
The key to successful statistical analysis of bivariate extreme events lies in flexible modelling of the tail dependence relationship between the two variables. In the extreme value theory literature, various techniques are available to model separate aspects of tail dependence, based on different asymptotic limits. Results from Balkema and Nolde (2010) and Nolde (2014) highlight the importance of studying the limiting shape of an appropriately-scaled sample cloud when characterising the whole joint tail. We now develop the first statistical inference for this limit set, which has considerable practical importance for a unified inference framework across different aspects of the joint tail. Moreover, Nolde and Wadsworth (2022) link this limit set to various existing extremal dependence frameworks. Hence, a by-product of our new limit set inference is the first set of self-consistent estimators for several extremal dependence measures, avoiding the current possibility of contradictory conclusions. In simulations, our limit set estimator is successful across a range of distributions, and the corresponding extremal dependence estimators provide a major joint improvement and small marginal improvements over existing techniques. We consider an application to sea wave heights, where our estimates successfully capture the expected weakening extremal dependence as the distance between locations increases.
We develop two models for the temporal evolution of extreme events of multivariate kth order Markov processes. The foundation of our methodology lies in the conditional extremes model of Heffernan Tawn (2004), and it naturally extends the work of Winter Tawn (2016,2017) and Tendijck et al. (2019) to include multivariate random variables. We use cross-validation-type techniques to develop a model order selection procedure, and we test our models on two-dimensional meteorological-oceanographic data with directional covariates for a location in the northern North Sea. We conclude that the newly-developed models perform better than the widely used historical matching methodology for these data.
A key aspect where extreme values methods differ from standard statistical models is through having asymptotic theory to provide a theoretical justification for the nature of the models used for extrapolation. In multivariate extremes, many different asymptotic theories have been proposed, partly as a consequence of the lack of ordering property with vector random variables. One class of multivariate models, based on conditional limit theory as one variable becomes extreme has received wide practical usage. The underpinning value of this approach has been supported by further theoretical characterisations of the limiting relationships. However, the paper “Conditional extreme value models: fallacies and pitfalls” by Holger Drees and Anja Janßen provides a number of counterexamples to these results. This paper studies these counterexamples in a conditional extremes framework which involves marginal standardisation to a common exponentially decaying tailed marginal distribution. Our calculations show that some of the issues identified can be addressed in this way.
We investigate the changing nature of the frequency, magnitude and spatial extent of extreme temperatures in Ireland from 1931 to 2022. We develop an extreme value model that captures spatial and temporal non-stationarity in extreme daily maximum temperature data. We model the tails of the marginal variables using the generalised Pareto distribution and the spatial dependence of extreme events by a semi-parametric Brown-Resnick r-generalised Pareto process, with parameters of each model allowed to change over time. We use weather station observations for modelling extreme events since data from climate models (not conditioned on observational data) can over-smooth these events and have trends determined by the specific climate model configuration. However, climate models do provide valuable information about the detailed physiography over Ireland and the associated climate response. We propose novel methods which exploit the climate model data to overcome issues linked to the sparse and biased sampling of the observations. Our analysis identifies a temporal change in the marginal behaviour of extreme temperature events over the study domain, which is much larger than the change in mean temperature levels over this time window. We illustrate how these characteristics result in increased spatial coverage of the events that exceed critical temperatures.
We develop methods, based on extreme value theory, for analysing observations in the tails of longitudinal data, i.e., a data set consisting of a large number of short time series, which are typically irregularly and non-simultaneously sampled, yet have some commonality in the structure of each series and exhibit independence between time series. Extreme value theory has not been considered previously for the unique features of longitudinal data. Across time series the data are assumed to follow a common generalised Pareto distribution, above a high threshold. To account for temporal dependence of such data we require a model to describe (i) the variation between the different time series properties, (ii) the changes in distribution over time, and (iii) the temporal dependence within each series. Our methodology has the flexibility to capture both asymptotic dependence and asymptotic independence, with this characteristic determined by the data. Bayesian inference is used given the need for inference of parameters that are unique to each time series. Our novel methodology is illustrated through the analysis of data from elite swimmers in the men's 100m breaststroke. Unlike previous analyses of personal-best data in this event, we are able to make inference about the careers of individual swimmers - such as the probability an individual will break the world record or swim the fastest time next year.
The seminal Bradley-Terry model exhibits transitivity, i.e., the property that the probabilities of player A beating B and B beating C give the probability of A beating C, with these probabilities determined by a skill parameter for each player. Such transitive models do not account for different strategies of play between each pair of players, which gives rise to {\it intransitivity}. Various intransitive parametric models have been proposed but they lack the flexibility to cover the different strategies across $n$ players, with the $O(n^2)$ values of intransitivity modelled using $O(n)$ parameters, whilst they are not parsimonious when the intransitivity is simple. We overcome their lack of adaptability by allocating each pair of players to one of a random number of $K$ intransitivity levels, each level representing a different strategy. Our novel approach for the skill parameters involves having the $n$ players allocated to a random number of $A
There currently exist a variety of statistical methods for modeling bivariate extremes. However, when the dependence between variables is driven by more than one latent process, these methods are likely to fail to give reliable inferences. We consider situations in which the observed dependence at extreme levels is a mixture of a possibly unknown number of much simpler bivariate distributions. For such structures, we demonstrate the limitations of existing methods and propose two new methods: an extension of the Heffernan–Tawn conditional extreme value model to allow for mixtures and an extremal quantile-regression approach. The two methods are examined in a simulation study and then applied to oceanographic data. Finally, we discuss extensions including a subasymptotic version of the proposed model, which has the potential to give more efficient results by incorporating data that are less extreme. Both new methods outperform existing approaches when mixtures are present.
Reliable estimates of sea-level return-levels are crucial for coastal flooding risk assessments and for coastal flood defence design. We describe a novel method for estimating extreme sea-levels that is the first to capture seasonality, interannual variations and longer term changes. We use a joint probabilities method, with skew-surge and peak-tide as two sea-level components. The tidal regime is predictable, but skew-surges are stochastic. We present a statistical model for skew-surges, where the main body of the distribution is modelled empirically while a nonstationary generalised Pareto distribution (GPD) is used for the upper tail. We capture within-year seasonality by introducing a daily covariate to the GPD model and allowing the distribution of peak-tide to change over months and years. Skew-surge-peak-tide dependence is accounted for, via a tidal covariate, in the GPD model, and we adjust for skew-surge temporal dependence through the subasymptotic extremal index. We incorporate spatial prior information in our GPD model to reduce the uncertainty associated with the highest return-level estimates. Our results are an improvement on current return-level estimates, with previous methods typically underestimating. We illustrate our method at four U.K. tide gauges.
Although most models for rainfall extremes focus on point-wise values, it is aggregated precipitation over areas up to river catchment scale that is of the most interest. To capture the joint behaviour of precipitation aggregates evaluated at different spatial scales, parsimonious and effective models must be built with knowledge of the underlying spatial process. Precipitation is driven by a mixture of processes acting at different scales and intensities, e.g., convective and frontal, with extremes of aggregates for typical catchment sizes arising from extremes of only one of these processes, rather than a combination of them. High-intensity convective events cause extreme spatial aggregates at small scales but the contribution of lower-intensity large-scale fronts is likely to increase as the area aggregated increases. Thus, to capture small to large scale spatial aggregates within a single approach requires a model that can accurately capture the extremal properties of both convective and frontal events. Previous extreme value methods have ignored this mixture structure; we propose a spatial extreme value model which is a mixture of two components with different marginal and dependence models that are able to capture the extremal behaviour of convective and frontal rainfall and more faithfully reproduces spatial aggregates for a wide range of scales. Modelling extremes of the frontal component raises new challenges due to it exhibiting strong long-range extremal spatial dependence. Our modelling approach is applied to fine-scale, high-dimensional, gridded precipitation data. We show that accounting for the mixture structure improves the joint inference on extremes of spatial aggregates over regions of different sizes.