Many statistical problems involve optimization over a discrete parameter space having an unknown dimension. In such settings, gradient-based methods often fail due to the non-differentiability of the objective function or a non-convex or massive search space with an objective function having many local maxima/minima. This paper presents GAReg, a unified genetic algorithm package that handles discrete optimization regression problems, which works well when standard algorithms are unjustified. GAReg provides a compact chromosome representation supporting optimal knot placement for regression splines, best-subset regression variable selection, and related problems. The package allows for uniform initialization, constraint-preserving crossover and mutation, steady-state replacement, and an optional island-model parallelization. GAReg efficiently searches high-dimensional model spaces, providing near-optimal solutions in settings where exhaustive enumeration or integer or dynamic programming approaches are infeasible.
Changes in tropical cyclone frequencies as the climate warms is a topic of significant current debate [1, 2]. There is no accepted theory of how tropical cyclogenesis might respond to a warmer ocean-atmosphere system as multiple controlling factors exist [3–7]. Anthropogenic warming of surface ocean temperatures due to increased greenhouse gas concentrations [8] increases the potential for tropical cyclogenesis [9–11]; however, realized cyclogenesis also requires an initial local disturbance [12–16] to develop. Most multi-decadal tropical cyclone permitting climate models (i.e. resolutions of 15-50km) exhibit frequency decreases in warmer climates, despite the increase in tropical cyclogenesis potential [17–25]. In this paper, we analyze the International Best Track Archive for Climate Stewardship, IBTrACS, [26] and the HURSAT-B1 [27] observed tropical storm data with mean shift changepoint and regression methods. Summarizing all of our analyses, we find that the global frequency of tropical cyclones with lifetime maximum sustained winds of 63km/hour or greater having a duration of more than two days (over all storm strengths) has at least very likely decreased. From the IBTrACS dataset, we also find at the global scale that it is extremely likely that the proportion of tropical cyclones that reach Category 3+ on the Saffir-Simpson Hurricane Wind Scale has increased and that it is virtually certain that both the proportion and frequency of Category 4+ storms has increased.
Abstract Single changepoint tests have become a staple check for the homogeneity of climate time series, suggesting how climate has changed should nonhomogeneity be declared. This paper summarizes the most prominent single changepoint tests used in contemporary climate literature, relating them to one another and unifying their presentations. Asymptotic quantiles for the individual tests are presented. Derivations of the quantiles are given, enabling the reader to tackle cases not considered within. Our work here studies both mean and linear trend shifts, covering the most common settings arising in climatology. Southern Oscillation index (SOI) and global temperature series are analyzed within to illustrate the techniques. Significance Statement This manuscript quantifies and unifies single changepoint statistical tests for changes in the mean and/or slope (linear trend) of climate time series. The most common tests used in current climate literature are linked, and their asymptotic quantiles are derived. The intent is to unify the climate practitioner’s work with single changepoint methods.
This article combines methods from existing techniques to identify multiple changepoints in non-Gaussian autocorrelated time series. A transformation is used to convert a Gaussian series into a non-Gaussian series, enabling penalized likelihood methods to handle non-Gaussian scenarios. When the marginal distribution of the data is continuous, the methods essentially reduce to the change of variables formula for probability densities. When the marginal distribution is count-oriented, Hermite expansions and particle filtering techniques are used to quantify the scenario. Simulations demonstrating the efficacy of the methods are given and two data sets are analyzed: 1) the proportion of home runs hit by Major League Baseball batters from 1920 to 2023 and 2) a six-dimensional series of tropical cyclone counts from the Earth's basins of generation from 1980 to 2023. In the first series, beta marginal distributions are used to describe the proportions; in the second, Poisson marginal distributions seem appropriate.
This manuscript studies nodal clustering in graphs having multivariate attributes at each node. The framework includes node-specific priors for low-dimensional representations, coupled with a neural decoder that bridges observed attributes with latent variables. Structural and attribute information are incorporated through a graph-fused LASSO regularization on the prior means, promoting nodal clustering. The optimization problem is solved via alternating direction method of multipliers, with Langevin dynamics for posterior inference. Simulation studies on grid graphs, and applications to real data with complex settings, demonstrate the effectiveness of the proposed clustering method.
Single changepoint tests have become a staple check for homogeneity of a climate time series, suggesting how climate has changed should non-homogeneity be declared. This paper summarizes the most prominent single changepoint tests used in today's climate literature, relating them to one and other and unifying their presentations. Asymptotic quantiles for the individual tests are presented. Derivations of the quantiles are given, enabling the reader to tackle cases not considered within. Our work here studies both mean and trend shifts, covering the most common settings arising in climatology. SOI and global temperature series are analyzed within to illustrate the techniques.
A single joinpoint changepoint model partitions a time series into two segments, joined at the changepoint time by constraining the estimated piecewise linear regression responses to be continuous. This manuscript derives the exact asymptotic distribution of the changepoint existence test statistic gauging whether or not a second segment is necessary. The identified asymptotic distribution, a supremum of a Gaussian process over the unit interval, is rather unwieldy. The work presented here provides the result and its derivation; quantiles of the asymptotic distribution are presented for the user. This addresses a subtle gap in the changepoint literature.
This work considers stationary vector count time series models defined via deterministic functions of a latent stationary vector Gaussian series. The construction is very general and ensures a pre-specified marginal distribution for the counts in each dimension, depending on unknown parameters that can be marginally estimated. The vector Gaussian series injects flexibility into the model's temporal and cross-dimensional dependencies, perhaps through a parametric model akin to a vector autoregression. We show that the latent Gaussian model can be estimated by relating the covariances of the counts and the latent Gaussian series. In a possibly high-dimensional setting, concentration bounds are established for the differences between the estimated and true latent Gaussian autocovariances, in terms of those for the observed count series and the estimated marginal parameters. The results are applied to the case where the latent Gaussian series is a vector autoregression, and its parameters are estimated sparsely through a LASSO-type procedure.
This paper studies statistical inference for a Lindley random walk model when the increment process driving the walk is strictly stationary. Lindley random walks govern customer waiting times in many queueing models and several natural and business processes, including snow depths, frozen soil depths, inventory quantities, etc. The probabilistic properties of a Lindley walk with time-correlated stationary changes are first reviewed. We provide a streamlined argument that the process has a proper limiting distribution when the mean of the incremental changes is negative, and that the Lindley process is strictly stationary when starting from this stationary distribution. Next, the Markov characteristics of the process are explored when the change process has a Markov structure of first or higher order. A derivation of the model's likelihood is given when the change process is a Gaussian autoregressive time series. An efficient particle filtering method for evaluating and optimizing the likelihood with Gaussian changes is then devised and studied via simulation.
This paper reviews and compares popular methods, some old and some very recent, that produce time series having Poisson marginal distributions. The paper begins by narrating ways where time series with Poisson marginal distributions can be produced. Modeling nonstationary series with covariates motivates consideration of methods where the Poisson parameter depends on time. Here, estimation methods are developed for some of the more flexible methods. The results are used in the analysis of 1) a count sequence of tropical cyclones occurring in the North Atlantic Basin since 1970, and 2) the number of no-hitter games pitched in major league baseball since 1893. Tests for whether the Poisson marginal distribution is appropriate are included.
This study develops methods to detect anomalous transactions linked with fraud in food stamp purchases through order statistics methods. The methods detect clusters in the order statistics of the transaction amounts that merit further scrutiny. Our techniques use scan statistics to determine when an excessive number of transactions occur (cluster), which is historically linked to fraud. A scoring paradigm is constructed that ranks the degree in which detected clusters and individual transactions are anomalous among approximately 250 million total transactions.
The global mean surface temperature is widely studied to monitor climate change. A current debate centers around whether there has been a recent (post-1970s) surge/acceleration in the warming rate. This paper addresses whether an acceleration in the warming rate is detectable from a statistical perspective. We use changepoint models, which are statistical techniques specifically designed for identifying structural changes in time series. Four global mean surface temperature records over 1850-2023 are scrutinized within. Our results show limited evidence for a warming surge; in most surface temperature time series, no change in the warming rate beyond the 1970s is detected. As such, we estimate minimum changes in the warming trend for a surge to be detectable in the near future.
This paper develops the theory and methods for modeling a stationary count time series via Gaussian transformations. The techniques use a latent Gaussian process and a distributional transformation to construct stationary series with very flexible correlation features that can have any pre-specified marginal distribution, including the classical Poisson, generalized Poisson, negative binomial, and binomial structures. Gaussian pseudo-likelihood and implied Yule-Walker estimation paradigms, based on the autocovariance function of the count series, are developed via a new Hermite expansion. Particle filtering and sequential Monte Carlo methods are used to conduct likelihood estimation. Connections to state space models are made. Our estimation approaches are evaluated in a simulation study and the methods are used to analyze a count series of weekly retail sales.
Climate changepoint (homogenization) methods abound today, with a myriad of techniques existing in both the climate and statistics literature. Unfortunately, the appropriate changepoint technique to use remains unclear to many. Further complicating issues, changepoint conclusions are not robust to small perturbations in assumptions; for example, allowing for a trend or correlation in the series can drastically change conclusions. This paper is a review of the changepoint topic, with an emphasis on illuminating the models and techniques that allow the scientist to make reliable conclusions. Pitfalls to avoid are demonstrated via actual applications. The discourse begins by narrating the salient statistical features of most climate time series. Thereafter, single and multiple changepoint problems are considered. Several pitfalls are discussed en route and good practices are recommended. While the majority of our applications involve temperature series, other settings are mentioned.
Changepoint methods have multiple uses in climatology, including stationary checks and record homogenization. There are still many open problems in the area, especially in the multiple changepoint setting, and statisticians are needed to help develop the methods and analyze the data.
Count time series are widely encountered in practice. As with continuous valued data, many count series have seasonal properties. This article uses a recent advance in stationary count time series to develop a general seasonal count time series modeling paradigm. The model constructed here permits any marginal distribution for the series and the most flexible autocorrelations possible, including those with negative dependence. Likelihood methods of inference are explored. The article first develops the modeling methods, which entail a discrete transformation of a Gaussian process having seasonal dynamics. Properties of this model class are then established and particle filtering likelihood methods of parameter estimation are developed. A simulation study demonstrating the efficacy of the methods is presented and an application to the number of rainy days in successive weeks in Seattle, Washington is given.
This paper develops a mathematical model and statistical methods to quantify trends in presence/absence observations of snow cover (not depths) and applies these in an analysis of Northern Hemispheric observations extracted from satellite flyovers during 1967-2021. A two-state Markov chain model with periodic dynamics is introduced to analyze changes in the data in a grid by grid fashion. Trends, converted to the number of weeks of snow cover lost/gained per century, are estimated for each study grid. Uncertainty margins for these trends are developed from the model and used to assess the significance of the trend estimates. Grids with questionable data quality are identified. Among trustworthy grids, snow presence is seen to be declining in almost twice as many grids as it is advancing. While Arctic and southern latitude snow presence is found to be rapidly receding, other locations, such as Eastern Canada, are experiencing advancing snow cover.
This paper presents a statistical analysis of structural changes in the Central England temperature series, one of the longest surface temperature records available. A changepoint analysis is performed to detect abrupt changes, which can be regarded as a preliminary step before further analysis is conducted to identify the causes of the changes (e.g., artificial, human-induced or natural variability). Regression models with structural breaks, including mean and trend shifts, are fitted to the series and compared via two commonly used multiple changepoint penalized likelihood criteria that balance model fit quality (as measured by likelihood) against parsimony considerations. Our changepoint model fits, with independent and short-memory errors, are also compared with a different class of models termed long-memory models that have been previously used by other authors to describe persistence features in temperature series. In the end, the optimal model is judged to be one containing a changepoint in the late 1980s, with a transition to an intensified warming regime. This timing and warming conclusion is consistent across changepoint models compared in this analysis. The variability of the series is not found to be significantly changing, and shift features are judged to be more plausible than either short- or long-memory autocorrelations. The final proposed model is one including trend-shifts (both intercept and slope parameters) with independent errors. The analysis serves as a walk-through tutorial of different changepoint techniques, illustrating what can be statistically inferred.
This article studies estimation of a stationary autocovariance structure in the presence of an unknown number of mean shifts. Here, a Yule–Walker moment estimator for the autoregressive parameters in a dependent time series contaminated by mean shift changepoints is proposed and studied. The estimator is based on first order differences of the series and is proven consistent and asymptotically normal when the number of changepoints m and the series length N satisfy $$m/N \rightarrow 0$$ as $$N \rightarrow \infty$$ .
Correlated time series data arise in many applications. This paper describes and compares several prominent single and multiple changepoint techniques for correlated time series. In the single changepoint problem, various cumulative sum (CUSUM) and likelihood ratio statistics, along with boundary cropping scenarios and scaling methods (e.g., scaling to an extreme value or Brownian Bridge limit) are compared. A recently developed test based on summing squared CUSUM statistics over all time indices is shown to have controlled Type I error and superior detection power. In the multiple changepoint setting, penalized likelihoods drive the discourse, with AIC, BIC, mBIC, and MDL penalties being considered. Binary and wild binary segmentation techniques are also compared. A new distance metric is introduced that measures differences between two multiple changepoint segmentations. Algorithmic and computational concerns are discussed and simulations are given to support all conclusions. In the end, the multiple changepoint setting admits no clear methodological winner, performance depending on the particular scenario. Nonetheless, some practical guidance emerges.