We construct a vector space whose defining characteristics are rooted in univariate regular variation of random variables. Specifically, the base vector space 𝕍_b consists of random variables whose limiting tail probabilities, when scaled by regularly varying functions of the form b(s)=s^αL(s), are finite. Defining a subspace N_b corresponding to random variables in 𝕍_b whose limiting tail probabilities are zero when normalized by b(s) allows the base space 𝕍_b to be partitioned into equivalence classes. We define a vector space 𝕎_b consisting of these equivalence classes, and show its nonzero elements are equivalence classes of regularly varying random variables. We show that a natural norm exists for 𝕎_b if α> 1. We show that the equivalence classes and convergence in norm are different than more familiar vector spaces of random variables. Turning our attention to extreme value modeling, we consider finite-dimensional subspaces of 𝕎_b whose basis vectors are jointly regularly varying. We show that in the case α= 2, the previously defined tail pairwise dependence measure serves as an inner product. As any finite-dimensional space is complete, we can use the projection theorem to perform linear prediction.
Weather extremes produce major impacts on society and ecosystems and are likely to change in likelihood and magnitude with climate change. However, very low-probability events are hard to characterize statistically using observations or even climate model output because of short records/runs. For precipitation, consideration of such events arises in quantifying probable maximum precipitation (PMP), namely, estimating extreme precipitation magnitudes for designing and assessing critical infrastructure. A recent National Academies report on modernizing PMP estimation proposed using very large climate model-based ensembles to estimate extreme quantiles, possibly through machine learning-based ensemble boosting. Here, we assess statistical aspects of such an approach for the contiguous United States using a huge ensemble (10 560 years) produced by a state-of-the-art emulator (ACE2) trained on ERA5 reanalysis. The results indicate that one can practically estimate very extreme precipitation and temperature quantiles, provided one uses appropriate statistical extreme value techniques. More specifically, the results provide evidence for 1) the use of threshold-exceedance methods with a sufficiently high threshold (necessary for precipitation) for reliable estimation, 2) the robustness of results to variation in extremes by season and storm type, and 3) the sufficiency of the ensemble for well-constrained statistical uncertainty. Our results also show that the emulator produces extremes outside the range of the ERA5 training data. While encouraging for emulators' potential use for quantifying the climatology of extremes, more investigation is needed to assess whether emulators are fit for this purpose. Our focus is on how to use huge ensembles to estimate very extreme statistics; we expect the results to be relevant for future improved emulators.
We propose a model to flexibly estimate joint tail properties by exploiting the convergence of an appropriately scaled point cloud onto a compact limit set. Characteristics of the shape of the limit set correspond to key tail dependence properties. We directly model the shape of the limit set using Bezier splines, which allow flexible and parsimonious specification of shapes in two dimensions. We fit the Bezier splines to data in pseudo-polar coordinates using Markov chain Monte Carlo sampling, utilizing a limiting approximation to the conditional likelihood of the radii given angles. We propose a novel prior on the shape of the limit set via constraints on the parameters of the Bezier splines. A direct advantage of our Bayesian approach is that the support of this prior guarantees that each posterior sample is a valid limit set boundary, allowing direct posterior analysis of any quantity derived from the shape of the curve. Furthermore, we obtain interpretable inference on the asymptotic dependence class by using mixture priors with point masses on the corner of the unit box. Finally, we apply our model to bivariate datasets of extremes of variables related to fire risk and air pollution.
Classifying a data set as asymptotically dependent (ADep) or asymptotically independent (AInd) is a necessary early choice in the modeling of multivariate extremes. These two dependence regimes are defined asymptotically which complicates inference as practitioners have finite samples. We perform a series of experiments to determine whether a finite sample has enough information for a neural network to reliably distinguish between these regimes in the bivariate case. Along the way we develop a new classification tool for practitioners which we call nnadic as it is a Neural Network for Asymptotic Dependence/Independence Classification. This tool accurately classifies over 95
In order to study potential impacts arising from climate change, future projections of numerical model output often must be calibrated to be comparable to observations. Rather than calibrating the data values themselves, we propose a novel statistical calibration method for extremes that assumes there exists a linear relationship between parameters associated with model output and parameters associated with observations. This approach allows us to capture uncertainty in both parameter estimates and the linear calibration, which we achieve via bootstrap. To focus on extreme behavior, we assume both model output and observations have distributions composed of a mixture model combining a Weibull distribution with a generalized Pareto distribution for the tail. A simulation study shows good coverage rates. We apply the method to project future daily-averaged river runoff at the Purgatoire River in southeastern Colorado.
Abstract An early choice in the modeling of multivariate extremes is to infer whether the data are asymptotically dependent (AD) or asymptotically independent (AI). We perform a series of experiments to determine whether a convolutional neural network can reliably distinguish between these asymptotically defined regimes in the finite sample bivariate case. Along the way we develop a new classification tool for practitioners which we call \texttt{nnadic} as it is a Neural Network for Asymptotic Dependence/ Independence Classification. This tool accurately classifies 95% of test datasets and is robust to a wide range of sample sizes. The datasets which we are unable to correctly classify tend to either be nearly exactly independent or exhibit near perfect dependence, which are boundary cases for both the AD and AI models used for training.
To capture the dependence in the upper tail of a time series, we develop non‐negative regularly varying time series models that are constructed similarly to classical non‐extreme ARMA models. Rather than fully characterizing tail dependence of the time series, we define the concept of weak tail stationarity which allows us to describe a regularly varying time series via a measure of pairwise extremal dependencies, the tail pairwise dependence function (TPDF). We state consistency requirements among the finite‐dimensional collections of the elements of a regularly varying time series and show that the TPDF's value does not depend on the dimension of the random vector being considered. So that our models take non‐negative values, we use transformed‐linear operations. We show existence and stationarity of these models, and develop their properties such as the model TPDFs. We fit models to hourly windspeed and daily fire weather index data, and we find that the fitted transformed‐linear models produce better estimates of upper tail quantities than a traditional ARMA model, classical linear regularly varying models, a max‐ARMA model, and a Markov model.
The innovations algorithm is a classical recursive forecasting algorithm used in time series analysis. We develop the innovations algorithm for a class of nonnegative regularly varying time series models constructed via transformed-linear arithmetic. In addition to providing the best linear predictor, the algorithm also enables us to estimate parameters of transformed-linear regularly-varying moving average (MA) models, thus providing a tool for modeling. We first construct an inner product space of transformed-linear combinations of nonnegative regularly-varying random variables and prove its link to a Hilbert space which allows us to employ the projection theorem, from which we develop the transformed-linear innovations algorithm. Turning our attention to the class of transformed linear MA($\infty$) models, we give results on parameter estimation and also show that this class of models is dense in the class of possible tail pairwise dependence functions (TPDFs). We also develop an extremes analogue of the classical Wold decomposition. Simulation study shows that our class of models captures tail dependence for the GARCH(1,1) model and a Markov time series model, both of which are outside our class of models.
Wildfire risk is greatest during high winds after sustained periods of dry and hot conditions. This paper is a statistical extreme-event risk attribution study that aims to answer whether extreme wildfire seasons are more likely now than under past climate. This requires modeling temporal dependence at extreme levels. We propose the use of transformed-linear time series models, which are constructed similarly to traditional autoregressive-moving-average (ARMA) models while having a dependence structure that is tied to a widely used framework for extremes (regular variation). We fit the models to the extreme values of the seasonally adjusted fire weather index (FWI) time series to capture the dependence in the upper tail for past and present climate. We simulate 10 000 fire seasons from each fitted model and compare the proportion of simulated high-risk fire seasons to quantify the increase in risk. Our method suggests that the risk of experiencing an extreme wildfire season in Grand Lake, Colorado, under current climate has increased dramatically relative to the risk under the climate of the mid-twentieth century. Our method also finds some evidence of increased risk of extreme wildfire seasons in Quincy, California, but large uncertainties do not allow us to reject a null hypothesis of no change.
Hazard event sets, a collection of synthetic extreme events over a given period, are important for catastrophe modelling. This paper addresses the issue of generating event sets of extreme river flow for northern England and southern Scotland, a region which has been particularly affected by severe flooding over the past 20 years. We start by analysing historical extreme river flow across 45 gauges, located within the study region, using methods from extreme value analysis, including the concept of extremal principal components. Our analysis reveals interesting connections between the extremal dependence structure and the region's topography/climate. We then introduce a framework which is based on modelling the distribution of the extremal principal components in order to generate synthetic events of extreme river flow. The generative framework is dimension-reducing in that it distinctly handles the principal components based on their contribution to describing the nature of extreme river flow across the study region. We also detail a data-driven approach to select the optimal dimension. Synthetic flood events are subsequently generated efficiently by sampling from the fitted distribution. Our approach for generating hazard event sets can be easily implemented by practitioners and our results indicate good agreement between the observed and simulated extreme river flow dynamics. For the considered application, we also find that our approach outperforms existing statistical approaches for generating hazard event sets.
Irrigation in the Eastern US receives little attention compared to the West, but farmers in humid states of the US, traditionally reliant on rainfall, have more than tripled irrigation since 1978. We examine this trend in Illinois where there has been a nearly threefold increase in center pivot irrigation systems (CPIS) installations since 1988. Specifically, we analyze where and when CPIS installations occur and their benefits in terms of annual crop yield, irrigated acreage, crop selection, and reduction in drought-related insurance payouts. To do so, we create a novel data set derived from a deep learning model capable of automatically identifying the location of CPIS during drought years along with annual county level crop, weather, and insurance data. The results indicate CPIS installations in Illinois are significantly more common over alluvial aquifers after droughts. Additionally, counties with a higher presence of CPIS do not have higher average crop yields, a shift to more water intensive crops, or an expansion of cropland. However, in drought years CPIS presence does have a significant positive effect on corn yield and a significant negative effect on indemnity payments for both soybeans and corn. The results provide insights into an emerging trend of irrigation in humid regions, raising potential policy considerations for crop insurance and signaling a potential need to address water rights as demand increases.Institutional subscribers to the NBER working paper series, and residents of developing countries may download this paper without additional charge at www.nber.org.
Guck, Adam PhD; Cooley, Daniel DO; Godwin, Victoria MD; Mansour, Mariam MD Author Information
Methane (CH4) emissions from oil and natural gas (O&NG) systems are an important contributor to greenhouse gas emissions. In the United States, recent synthesis studies of field measurements of CH4 emissions at different spatial scales are ~1.5-2× greater compared to official greenhouse gas inventory (GHGI) estimates, with the production-segment as the dominant contributor to this divergence. Based on an updated synthesis of measurements from component-level field studies, we develop a new inventory-based model for CH4 emissions, for the production-segment only, that agrees within error with recent syntheses of site-level field studies and allows for isolation of equipment-level contributions. We find that unintentional emissions from liquid storage tanks and other equipment leaks are the largest contributors to divergence with the GHGI. If our proposed method were adopted in the United States and other jurisdictions, inventory estimates could better guide CH4 mitigation policy priorities.
We consider the problem of performing prediction when observed values are at their highest levels. We construct an inner product space of nonnegative random variables from transformed-linear combinations of independent regularly varying random variables. The matrix of inner products corresponds to the tail pairwise dependence matrix, which summarizes tail dependence. The projection theorem yields the optimal transformed-linear predictor, which has the same form as the best linear unbiased predictor in non-extreme prediction. We also construct prediction intervals based on the geometry of regular variation. We show that these intervals have good coverage in a simulation study as well as in two applications; prediction of high pollution levels, and prediction of large financial losses.
Motivated by the widespread use of large gridded data sets in the atmospheric sciences, we propose a new model for extremes of areal data that is inspired by the simultaneous autoregressive (SAR) model in classical spatial statistics. Our extreme SAR model extends recent work on transformed‐linear operations applied to regularly varying random vectors, and is unique among extremes models in being directly analogous to a classical linear model. An additional appeal is its simplicity; given a proximity matrix W, spatial dependence is described by a single parameter ρ . We develop an estimation method that minimizes the discrepancy between the tail pairwise dependence matrix (TPDM) for the fitted model and the estimated TPDM. Applying this method to simulated data demonstrates that it is able to produce good estimates of extremal spatial dependence even in the case of model misspecification, and additionally produces reasonable estimates of uncertainty. We also apply the method to gridded precipitation observations for a study region over northeast Colorado, and find that a single‐parameter extreme SAR model paired with a neighborhood structure which accounts for longer range dependence effectively models spatial dependence in these data.
Availability and quality of administrative data on irrigation technology varies greatly across jurisdictions. Technology choice, however, will influence the parameters of coupled human-hydrological systems. Equally, changing parameters in the coupled system may drive technology adoption. Here we develop and demonstrate a deep learning approach to locate a particularly important irrigation technology-center pivot irrigation systems-throughout the Ogallala Aquifer. The model does not rely on super computers and thus provides a model for an accessible baseline to train and deploy on other geographies. We further demonstrate that accounting for the technology can improve the insights in both economic and hydrological models.
We propose a method for analyzing extremal behavior through the lens of a most efficient basis of vectors. The method is analogous to principal component analysis, but is based on methods from extreme value analysis. Specifically, rather than decomposing a covariance or correlation matrix, we obtain our basis vectors by performing an eigendecomposition of a matrix that describes pairwise extremal dependence. We apply the method to precipitation observations over the contiguous United States. We find that the time series of large coefficients associated with the leading eigenvector shows very strong evidence of a positive trend, and there is evidence that large coefficients of other eigenvectors have relationships with El Niño–Southern Oscillation.
Methane (CH4) emissions from oil and natural gas (O&NG) systems are an important contributor to greenhouse gas emissions. In the United States, recent synthesis studies of field measurements of CH4 emissions at different spatial scales are ~1.5x-2x greater compared to official greenhouse gas inventory (GHGI) estimates, with the production-segment as the dominant contributor to this divergence. Based on an updated synthesis of measurements from component-level field studies, we develop a new inventory-based model for CH4 emissions, for the production-segment only, that agrees within error with recent syntheses of site-level field studies and allows for isolation of equipment-level contributions. We find that unintentional emissions from liquid storage tanks and other equipment leaks are the largest contributors to divergence with the GHGI. If our proposed method were adopted in the United States and other jurisdictions, inventory estimates could better guide CH4 mitigation policy priorities.
Uncertainty in return level estimates for rare events, like the intensity of large rainfall events, makes it difficult to develop strategies to mitigate related hazards, like flooding. Latent spatial extremes models reduce the uncertainty by exploiting spatial dependence in statistical characteristics of extreme events to borrow strength across locations. However, these estimates can have poor properties due to model misspecification: Many latent spatial extremes models do not account for extremal dependence, which is spatial dependence in the extreme events themselves. We improve estimates from latent spatial extremes models that make conditional independence assumptions by proposing a weighted likelihood that uses the extremal coefficient to incorporate information about extremal dependence during estimation. This approach differs from, and is simpler than, directly modeling the spatial extremal dependence; for example, by fitting a max-stable process, which is challenging to fit to real, large datasets. We adopt a hierarchical Bayesian framework for inference, use simulation to show the weighted model provides improved estimates of high quantiles, and apply our model to improve return level estimates for Colorado rainfall events with 1% annual exceedance probability. Supplementary materials accompanying this paper appear online.
Noting a strong imperative to understand precipitation extremes, and that considerable uncertainty affects observational data sets, this paper compares the representation of extremes in a number of widely used daily gridded products, derived from rain gauge data, satellite retrieval and reanalysis for the conterminous United States. Analysis is based upon the concept of tail dependence arising in multivariate extreme value theory, and we infer the level of temporal dependence in the joint tail of the precipitation probability distribution for pairwise comparisons of products. In this way, we consider the range of products more like an ensemble and examine the relationships between members, and do not attempt to define, or compare products to, some ground truth. Linear correlation between products is also computed. Considerable discrepancy between groups of products, both annually and seasonally, is linked to source data and complex terrain. In particular, products based on rain gauge data showed remarkable similarity, but differed considerably, showing almost total loss of extremal dependence during DJF in mountainous regions, when compared with satellite products. Additionally, simulated re-forecasts revealed reasonable temporal agreement with large scale generated extremes. The diversity and extent of discrepancies identified across all products raises important questions about their use, and we urge caution, particularly for products derived from satellite data.