Surface winds can vary substantially from one minute to the next, so there is scope for studying its variation on this fine time scale. Restricting to the month of June to minimize seasonality, this work develops a range of machine learning models for generating realistic time series of surface wind vectors at a site in Lamont, Oklahoma based on more than 30 years of high quality measurements at the minute time scale. Such a generator could be used as an input into models from a range of disciplines, notably for wind energy, but also wildfire spread and aviation, among others. The data show complex diurnal structures in both wind speed and direction that would be challenging to capture with standard time series models, so we consider a number of machine learning approaches to producing a stochastic wind generator based on time vector-quantized variational autoencoders. We consider generating a day's worth of data at a time and generating a day of wind vectors conditional on the previous day's winds. We also study methods for incorporating a discrete weather state variable in the generator. We evaluate the generators using a wide range of formal and informal methods. The best of these generators can capture many but not all of the complex features present in the observational data. In particular, the best of our approaches accurately mimic diurnal changes in wind volatility but struggle to match the observed distribution of extreme wind speeds.
We propose an interpretable low-rank state-space model for conditional simulation of daily regional temperature fields. The central statistical question is whether empirical orthogonal functions (EOFs) can be treated as stable large-scale features of the field rather than only as sample-dependent basis vectors for dimension reduction. The framework represents the dominant temperature field through a small number of retained EOF coefficients, models their seasonal mean and variance structure, propagates the resulting low-dimensional state with stable multivariate dynamics, and uses structured innovations to represent the remaining uncertainty. The resulting reduced-rank model is interpretable, computationally efficient, and suitable for iterative ensemble generation. It is designed to support both inference on the structure of the retained temperature state and prediction through multi-horizon probabilistic field simulation.
We establish a rigorous asymptotic theory for the joint estimation of roughness and scale parameters in two-dimensional Gaussian random fields with power-law generalized covariances . Our main results are bivariate central limit theorems for a class of method-of-moments estimators under increasing-domain and fixed-domain asymptotics. The fixed-domain result follows immediately from the increasing-domain result from the self-similarity of Gaussian random fields with power-law generalized covariances . These results provide a unified distributional framework across these two classical regimes that makes the unusual behavior of the estimates under fixed-domain asymptotics intuitively obvious. Our increasing-domain asymptotic results use spatial averages of quadratic forms of (iterated) bilinear product difference filters that yield explicit expressions for the estimates of roughness and scale to which existing theorems on such averages can be readily applied. We further show that the asymptotics remain valid under modestly irregular sampling due to jitter or missing observations. For the fixed-domain setting, the results extend to models that behave sufficiently like the power-law model at high frequencies such as the often used Matérn model .
The problem of estimating the slope parameter in regression between two spatial processes under confounding by an unmeasured spatial process has received widespread attention in the recent statistical literature. Yet, a fundamental question remains unsolved: when is this slope consistently estimable under spatial confounding, with existing insights being largely empirical or estimator-specific. In this manuscript, we characterize conditions for consistent estimability of the regression slope between Gaussian random fields (GRFs). Under fixed-domain (infill) asymptotics, we give sufficient conditions for consistent estimability using a novel characterization of the regression slope as the ratio of principal irregular terms of covariances, dictating the relative local behavior of the exposure and confounder processes. When estimability holds, we provide consistent estimators of the slope using local differencing (taking discrete differences or Laplacians of the processes of suitable order). Using functional analysis results on Paley-Wiener spaces, we then provide an easy-to-verify necessary condition for consistent estimability of the slope in terms of the relative spectral tail decays of the confounder and exposure. As a by-product, we establish a novel and general spectral condition on the equivalence of measures on the paths of multivariate GRFs with component fields of varying smoothnesses, a result of independent importance. We show that for the Matérn, power-exponential, generalized Cauchy, and coregionalization families, the necessary and sufficient conditions become identical, thereby providing a complete characterization of consistent estimability of the slope under spatial confounding. The results are extended to accommodate measurement error using local-averaging-and-differencing based estimators. Finite sample behavior is explored via numerical experiments.
Max-stable processes are the most popular models for high-impact spatial extreme events, as they arise as the only possible limits of spatially-indexed block maxima. However, likelihood inference for such models suffers severely from the curse of dimensionality, since the likelihood function involves a combinatorially exploding number of terms. In this paper, we propose using the Vecchia approximation, which conveniently decomposes the full joint density into a linear number of low-dimensional conditional density terms based on well-chosen conditioning sets designed to improve and accelerate inference in high dimensions. Theoretical asymptotic relative efficiencies in the Gaussian setting and simulation experiments in the max-stable setting show significant efficiency gains and computational savings using the Vecchia likelihood approximation method compared to traditional composite likelihoods. Our application to extreme sea surface temperature data at more than a thousand sites across the entire Red Sea further demonstrates the superiority of the Vecchia likelihood approximation for fitting complex models with intractable likelihoods, delivering significantly better results than traditional composite likelihoods, and accurately capturing the extremal dependence structure at lower computational cost.
Estimates of extreme sea-level return periods guide flood hazard mitigation. Return period estimates calculated from tide gauge records, which are relatively short (typically less than 100 years), can fail to capture the rarest and most potentially impactful extreme events. Here, we employ a two-dimensional Poisson point process model to fuse water-level data from tide gauges with data from multi-century geologic records of extreme overwash events. Experiments with synthetic data show that including geologic data reduces the uncertainty of 1% and 0.1% average annual chance water levels by about half, relative to using tide gauge data alone. Similar uncertainty reductions occur with two case studies of geologic data (Mattapoisett Marsh, Massachusetts and Cheesequake, New Jersey) and their neighboring tide gauges (Woods Hole, Massachusetts and the Battery, New York). The analysis also reveals non-stationarity at Cheesequake and The Battery, arising from either climatic changes or changes in the fidelity of the geological record, with substantially higher 1–10% average annual chance water levels since 1900 compared to prior centuries.
Combining strengths from deep learning and extreme value theory can help describe complex relationships between variables where extreme events have significant impacts (e.g., environmental or financial applications). Neural networks learn complicated nonlinear relationships from large datasets under limited parametric assumptions. By definition, the number of occurrences of extreme events is small, which limits the ability of the data-hungry, nonparametric neural network to describe rare events. Inspired by recent extreme cold winter weather events in North America caused by atmospheric blocking, we examine several probabilistic generative models for the entire multivariate probability distribution of daily boreal winter surface air temperature. We propose metrics to measure spatial asymmetries, such as long-range anticorrelated patterns that commonly appear in temperature fields during blocking events. Compared to vine copulas, the statistical standard for multivariate copula modeling, deep learning methods show improved ability to reproduce complicated asymmetries in the spatial distribution of ERA5 temperature reanalysis, including the spatial extent of in-sample extreme events.
A common approach to approximating Gaussian log-likelihoods at scale exploits the fact that precision matrices can be well-approximated by sparse matrices in some circumstances. This strategy is motivated by the screening effect, which refers to the phenomenon in which the linear prediction of a process Z at a point x0 depends primarily on measurements nearest to x0. But simple perturbations, such as iid measurement noise, can significantly reduce the degree to which this exploitable phenomenon occurs. While strategies to cope with this issue already exist and are certainly improvements over ignoring the problem, in this work we present a new one based on the EM algorithm that offers several advantages. While in this work we focus on the application to Vecchia’s approximation (Vecchia), a particularly popular and powerful framework in which we can demonstrate true second-order optimization of M steps, the method can also be applied using entirely matrix-vector products, making it applicable to a very wide class of precision matrix-based approximation methods. Supplementary materials for this article are available online.
We propose to use deep learning to estimate parameters in statistical models when standard likelihood estimation methods are computationally infeasible. We show how to estimate parameters from max-stable processes, where inference is exceptionally challenging even with small datasets but simulation is straightforward. We use data from model simulations as input and train deep neural networks to learn statistical parameters. Our neural-network-based method provides a competitive alternative to current approaches, as demonstrated by considerable accuracy and computational time improvements. It serves as a proof of concept for deep learning in statistical parameter estimation and can be extended to other estimation problems.
One of the most common reasons a submission to JASA Applications and Case Studies is rejected is that it is deemed inappropriate for this section of the journal. If we look for guidance in the journal’s instructions for authors, the opening sentence of the section on Applications and Case Studies states, “The Applications and Case Studies section publishes original articles that cogently demonstrate statistical usage in applications from any research area.” In contrast, the instructions for Theory and Methods papers say, “The research reported should be motivated by a scientific or practical problem and, ideally, illustrated by application of the proposed methodology to that problem. Illustration of techniques with real data is especially welcomed and strongly encouraged.” Many potential authors may find the distinctions between these two statements difficult to discern, which perhaps partly explains the high frequency of submissions rejected for being inappropriate. This editorial is an attempt to clarify what I think these distinctions are. In particular, in what ways is a paper with new methodology motivated by a “scientific or practical problem” and illustrated with “real data” not necessarily appropriate for Applications and Case Studies? First, let me be clear that the views expressed here are my own and are not part of official journal policy. Every editor is the final arbiter of what papers should be published in the journal and I think it is appropriate and even desirable that different editors use somewhat different criteria in making these decisions. Nevertheless, it is my hope that authors, referees, associate editors, and future editors will find it helpful for me to spell out in greater detail than is appropriate for a journal’s website some of the things I look for when evaluating a submission. There is not and should not be a clear and wide dividing line between Applications and Case Studies papers and Theory and Methods papers. Nevertheless, the use of the word “illustration” in the instructions for Theory and Methods papers points at a key distinction. An illustrative example possesses a feature that a proposed methodology is meant to address. The resulting data analysis may be rather brief, focusing on how the methodology can handle this feature better than previously proposed methods. The example thus serves in a supporting role, with the novel methodology and possibly accompanying theory being the main research contributions. In contrast, the specific application plays a much more prominent role in an Applications and Case Studies paper. A typical Applications and Case Studies paper begins with a description of the applied problem, generally one of current scientific or policy interest, which then leads into a discussion of the proposed methodology. Note that while most Applications and Case Studies papers include novel methodology, that is not a requirement for publication. For example, a paper that adapts existing methodology to a new field of application may be publishable if this adaptation leads to substantial scientific insights in the application beyond what could be learned from previously used methods. Indeed, a good Applications and Case Studies paper might provide thoughtful analyses of real data to demonstrate the strengths and/or weaknesses of various statistical methods presently used to address an important substantive problem. All data analyses appearing in Applications and Case Studies should correspond to good statistical and scientific practice. Statisticians and other researchers using statistical methods should be able to look to papers in Applications and Case Studies as examples of statistical practice they can emulate. Here are some of the features of an application or case study that should generally appear in any Applications and Case Studies paper:
Although underwater drones are no novel technology, their widespread use in civil and industrial applications has not been widely accepted so far. Apart from that, the decrease in size and costs along with an increase in robustness of underwater drones and the ease of handling provide a strong basis for underwater drone technology to grow in various markets. This chapter introduces the application of underwater drone technology in maritime operations, focusing on the micro ROV class. An introductory framework evaluation based on a structured literature analysis of the current state of research is conducted in order to provide a structured outlook on areas of drone operations in the maritime domain. Furthermore, the combination of micro ROV and artificial intelligence in the form of a neural network based on deep learning is introduced. This contribution provides an introductory analysis regarding both operational sides of science and the industry in order to shed light on the existing literature gap as ground for future research.
Gaussian random fields (GRF) are a fundamental stochastic model for spatiotemporal data analysis. An essential ingredient of GRF is the covariance function that characterizes the joint Gaussian distribution of the field. Commonly used covariance functions give rise to fully dense and unstructured covariance matrices, for which required calculations are notoriously expensive to carry out for large data. In this work, we propose a construction of covariance functions that result in matrices with a hierarchical structure. Empowered by matrix algorithms that scale linearly with the matrix dimension, the hierarchical structure is proved to be efficient for a variety of random field computations, including sampling, kriging, and likelihood evaluation. Specifically, with n scattered sites, sampling and likelihood evaluation has an O(n) cost and kriging has an O(log n) cost after preprocessing, particularly favorable for the kriging of an extremely large number of sites (e.g., predicting on more sites than observed). We demonstrate comprehensive numerical experiments to show the use of the constructed covariance functions and their appealing computation time. Numerical examples on a laptop include simulated data of size up to one million, as well as a climate data product with over two million observations.
Current models for spatial extremes are concerned with the joint upper (or lower) tail of the distribution at two or more locations. Such models cannot account for teleconnection patterns of 2-m surface air temperature (T2m) in North America, where very low temperatures in the contiguous United States may coincide with very high temperatures in Alaska in the wintertime. This dependence between warm and cold extremes motivates the need for a model with opposite-tail dependence in spatial extremes. This work develops a statistical modeling framework that has flexible behav-ior in all four pairings of high and low extremes at pairs of locations. In particular, we use a mixture of rotations of common Archimedean copulas to capture various combinations of four-corner tail dependence. We study teleconnected T2m ex-tremes using ERA5 of daily average 2-m temperature during the boreal winter. The estimated mixture model quantifies the strength of opposite-tail dependence between warm temperatures in Alaska and cold temperatures in the midlatitudes of North America, as well as the reverse pattern. These dependence patterns are shown to correspond to blocked and zonal patterns of midtropospheric flow. This analysis extends the classical notion of correlation-based teleconnections to consid-ering dependence in higher quantiles.
Nonstationary Gaussian process models can capture complex spatially varying dependence structures in spatial data. However, the large number of observations in modern datasets makes fitting such models computationally intractable with conventional dense linear algebra. In addition, derivative-free or even first-order optimization methods can be slow to converge when estimating many spatially varying parameters. We present here a computational framework that couples an algebraic block-diagonal plus low-rank covariance matrix approximation with stochastic trace estimation to facilitate the efficient use of second-order solvers for maximum likelihood estimation of Gaussian process models with many parameters. We demonstrate the effectiveness of these methods by simultaneously fitting 192 parameters in the popular nonstationary model of Paciorek and Schervish using 107,600 sea surface temperature anomaly measurements.
The Matérn covariance function is ubiquitous in the application of Gaussian processes to spatial statistics and beyond. Perhaps the most important reason for this is that the smoothness parameter ν gives complete control over the mean-square differentiability of the process, which has significant implications for the behavior of estimated quantities such as interpolants and forecasts. Unfortunately, derivatives of the Matérn covariance function with respect to ν require derivatives of the modified second-kind Bessel function 𝒦_ν with respect to ν . While closed form expressions of these derivatives do exist, they are prohibitively difficult and expensive to compute. For this reason, many software packages require fixing ν as opposed to estimating it, and all existing software packages that attempt to offer the functionality of estimating ν use finite difference estimates for ∂ _ν𝒦_ν . In this work, we introduce a new implementation of 𝒦_ν that has been designed to provide derivatives via automatic differentiation (AD), and whose resulting derivatives are significantly faster and more accurate than those computed using finite differences. We provide comprehensive testing for both speed and accuracy and show that our AD solution can be used to build accurate Hessian matrices for second-order maximum likelihood estimation in settings where Hessians built with finite difference approximations completely fail.
Extreme value theory motivates estimating extreme upper quantiles of a distribution by selecting some threshold, discarding those observations below the threshold and fitting a generalized Pareto distribution to exceedances above the threshold via maximum likelihood. This sharp cutoff between observations that are used in the parameter estimation and those that are not is at odds with statistical practice for analogous problems such as nonparametric density estimation, in which observations are typically smoothly downweighted as they become more distant from the value at which the density is being estimated. By exploiting the fact that the order statistics of independent and identically distributed observations form a Markov chain, this work shows how one can obtain a natural weighted composite log-likelihood function for fitting generalized Pareto distributions to exceedances over a threshold. A method for producing confidence intervals based on inverting a test statistic calibrated via parametric bootstrapping is proposed. Some theory demonstrates the asymptotic advantages of using weights in the special case when the shape parameter of the limiting generalized Pareto distribution is known to be 0. Methods for extending this approach to observations that are not identically distributed are described and applied to an analysis of daily precipitation data in New York City. Perhaps the most important practical finding is that including weights in the composite log-likelihood function can reduce the sensitivity of estimates to small changes in the threshold.
Empirical distributions in a range of fields are often substantially non-Gaussian but smooth enough to suggest that the underlying population distribution has at least several derivatives. Linear combinations of many random variables often have smooth densities even if the random variables are not independent. The main result here generalizes the elementary result that a convolution of densities has as many derivatives as the sum of the number of derivatives of each component to certain nonlinear functions of random vectors whose components may not be independent. Linearity is replaced by an assumption that the nonlinear function has positive first partial derivatives bounded away from 0 along at least some coordinates and independence is replaced by an assumption that the joint density has certain mixed partial derivatives. This approach to justifying the smoothness of empirical distributions is contrasted with other possible approaches, including solutions of stochastic partial differential equations and ergodic distributions for chaotic dynamical systems.
The estimation of range parameters for spatial covariance functions has long been a source of theoretical and practical problems in spatial statistics. In particular, in many applications, one finds that likelihood and, especially, restricted likelihood functions, do not provide any meaningful upper bound on the range parameter of a parametric covariance function model. This work seeks to provide further insight into this phenomenon by showing that, in at least some circumstances, it can make sense to extend the domain of the inverse range parameter to negative values as long as the spatial domain of interest is bounded. This possibility is explored through numerical work, some limited theory and an application to the (Davis, 1973) elevation data, which played a role in the recognition of the difficulties in estimating range parameters.
Current models for spatial extremes are concerned with the joint upper (or lower) tail of the distribution at two or more locations. Such models cannot account for teleconnection patterns of two-meter surface air temperature ($T_{2m}$) in North America, where very low temperatures in the contiguous Unites States (CONUS) may coincide with very high temperatures in Alaska in the wintertime. This dependence between warm and cold extremes motivates the need for a model with opposite-tail dependence in spatial extremes. This work develops a statistical modeling framework which has flexible behavior in all four pairings of high and low extremes at pairs of locations. In particular, we use a mixture of rotations of common Archimedean copulas to capture various combinations of four-corner tail dependence. We study teleconnected $T_{2m}$ extremes using ERA5 reanalysis of daily average two-meter temperature during the boreal winter. The estimated mixture model quantifies the strength of opposite-tail dependence between warm temperatures in Alaska and cold temperatures in the midlatitudes of North America, as well as the reverse pattern. These dependence patterns are shown to correspond to blocked and zonal patterns of mid-tropospheric flow. This analysis extends the classical notion of correlation-based teleconnections to considering dependence in higher quantiles.
In traditional extreme value analysis, the bulk of the data is ignored, and only the tails of the distribution are used for inference. Extreme observations are specified as values that exceed a threshold or as maximum values over distinct blocks of time, and subsequent estimation procedures are motivated by asymptotic theory for extremes of random processes. For environmental data, nonstationary behavior in the bulk of the distribution, such as seasonality or climate change, will also be observed in the tails. To accurately model such nonstationarity, it seems natural to use the entire dataset rather than just the most extreme values. It is also common to observe different types of nonstationarity in each tail of a distribution. Most work on extremes only focuses on one tail of a distribution, but for temperature, both tails are of interest. This paper builds on a recently proposed parametric model for the entire probability distribution that has flexible behavior in both tails. We apply an extension of this model to historical records of daily mean temperature at several locations across the United States with different climates and local conditions. We highlight the ability of the method to quantify changes in the bulk and tails across the year over the past decades and under different geographic and climatic conditions. The proposed model shows good performance when compared to several benchmark models that are typically used in extreme value analysis of temperature.
Ji Meng Loh合作论文数AT&T Labs-Research3