We propose a flexible and robust nonparametric framework for testing spatial dependence in two- and three-dimensional random fields. Our approach involves converting spatial data into one-dimensional time series using space-filling Hilbert curves. We then apply ordinal pattern-based tests for serial dependence to this series. Because Hilbert curves preserve spatial locality, spatial dependence in the original field manifests as serial dependence in the transformed sequence. The approach is easy to implement, accommodates arbitrary grid sizes through generalized Hilbert (“gilbert”) curves, and naturally extends beyond three dimensions. This provides a practical and general alternative to existing methods based on spatial ordinal patterns, which are typically limited to two-dimensional settings.
Motivated by an application to air quality data, we provide a comprehensive investigation of coherent forecasting techniques for ordinal time series. Here, the coherency requirement expresses that the computed forecast values have to be consistent with the ordered qualitative range of the data-generating process (DGP). Three classes of coherent forecasts are proposed, namely ordinal point forecasts, prediction intervals and probability mass function (pmf) forecasts. In addition, corresponding criteria for evaluating their predictive performance are established. An extensive simulation study analyses the performance of the coherent forecasting techniques across various ordinal DGPs, taking into account the impact of estimated model parameters and model misspecification. The proposed forecasting approaches are then applied to our central motivating application, the real-world time series on ordinal air quality levels and practical implications are discussed. In this context, it is also demonstrated that the coherent point and pmf forecasts can be combined visually in a meaningful and interpretable way. While centrally motivated by environmental monitoring, the proposed framework applies more broadly to ordinal time series arising in other settings where ordered categorical outcomes are used for assessment and decision making.
After clarifying possible misunderstandings concerning covariances of thinned random variables, we propose a refined definition of the first-order integer-valued autoregressive model for count random fields. We provide a comprehensive derivation of its autocorrelation structure, which also covers some former results. Moreover, we expand the refined model to higher-order autoregressions and study its stochastic properties.
Existing integer-valued autoregressive (INAR) models for count random fields suffer from difficulties in characterizing the stationary marginal distribution and in computing conditional probabilities (as required for likelihood inference). To overcome these drawbacks, the novel class of combined INAR (CINAR) models is proposed, which both exhibits the classical autoregressive dependence structure and allows to specify the marginal distribution within the wide class of discrete self-decomposable distributions. In particular, CINAR random fields can be equipped with a Poisson or negative-binomial marginal distribution. The CINAR's key stochastic properties are derived (including a simple expression for conditional probabilities), and special cases as well as possible extensions are discussed. Approaches for parameter estimation are developed and investigated, and the practical relevance of the novel CINAR family is demonstrated by an agricultural data application.
Except a few, the majority of the literature on monitoring ordinal data consider independent and identically distributed processes, where samples of data are collected sequentially in time. However, stationary ordinal processes exhibiting serial dependence are also common in many real-world process monitoring applications. This study proposes three classes of novel control charts for monitoring serially dependent stationary ordinal processes. Instead of sample statistics, individual observations are utilized. Exponentially weighted moving-average smoothing of a sequence of estimates is used for estimating the probability mass (cumulative distribution) function of the process. Defined real-valued functions of the probability mass (cumulative distribution) estimates are then used as the statistic plotted on the control charts. The methods are designed to be sensitive to a shift in the marginal distribution. Average run length performance of the control charts are computed under a comprehensive set of data-generating process models, which are inspired by real-world examples and exhibit quite different serial dependence structures. The performances of the proposed control charts are evaluated and compared to provide recommendations for implementations. The results show that the class of demerit-type charts generally perform better than the others. To illustrate the application and interpretation of the proposed methods, a real-world data example on monitoring of heating, ventilation, and air conditioning systems in passenger rail coaches is discussed.
An integer-valued moving average (INMA) model for count random fields is proposed and investigated. Closed-form expressions are derived for both its marginal distribution and spatial dependence structure, for arbitrary model order and also covering the multilateral case. In particular, general expressions for bivariate distributions and autocovariances are provided. It is shown that the INMA random field can be equipped (among others) with a Poisson marginal distribution. It is also demonstrated that different and well-interpretable dependence structures are possible. For illustration, we discuss a real-world data example and propose an INMA approximation to a given spatial dependence structure.
This paper proposes a novel class of first-order self-exciting hysteretic integer-valued autoregressive (SEHINAR(1)) time series models, in which the stochastic process is conditionally distributed based on past data within a hysteretic autoregressive framework. The basic probabilistic and statistical properties of the proposed model are thoroughly explored. Parameter estimation is obtained via conditional least squares, weighted conditional least squares, and maximum likelihood methods, with the corresponding asymptotic properties rigorously derived. A search algorithm is developed to determine the two boundary parameters, and the strong consistency of the estimators is formally established. Extensive numerical experiments and a real-world data application are provided to demonstrate the practical effectiveness of the proposed method.
Control charts for process monitoring are widely used in practice. Most control charts require the monitored (residuals) process to be serially independent (and to satisfy specified distributional assumptions), whereas undetected dependence (or violations of distributional assumptions) may severely affect the charts' performances. Therefore, (distribution-free) control charts for monitoring serial dependence are of utmost relevance for practice. Recently, various nonparametric control charts have been proposed for this purpose, which are based on ordinal patterns, and which showed an appealing performance in detecting different types of serial dependence. In this research, we further progress in this direction and develop novel nonparametric control charts being based on transcripts and algebraic distances (as derived from ordinal patterns). The performance of the newly proposed control charts is evaluated in a simulation study, and their application in practice is illustrated with a real-world data example from chemical industry.
The use of ordinal patterns (OPs) for analyzing the dependence structure of univariate and continuously distributed processes has gained popularity in recent years. This research goes one step further and considers the transcripts being computed from successive OPs in the time series. Transcripts constitute a kind of “difference” between successive OPs and thus naturally relate to two algebraic distances between OPs, the Cayley and Kendall edit distances. The original time series is transformed into a sequence of transcripts or distances, respectively, and important stochastic properties thereof are derived. It is shown that these properties differ substantially among different types of original processes. This motivates the development of various statistics based on transcripts and edit distances in order to investigate the dependence structure of the original process. In particular, the asymptotic distribution of these statistics under the null hypothesis of serial independence is derived, which is then used to implement nonparametric tests for serial dependence. A simulation study shows that these novel dependence tests have appealing power properties, often outperforming former OP-based dependence tests. A concluding real-world data example illustrates the application and interpretation of the proposed approaches in practice.
The derivation and application of Stein identities have received considerable research interest in recent years, especially for continuous or discrete-univariate distributions. In this paper, we complement the existing literature by deriving and investigating Stein-type characterizations for the three most common types of bivariate count distributions, namely the bivariate Poisson, binomial, and negative-binomial distribution. Then, we demonstrate the practical relevance of these novel Stein identities by a couple of applications, namely the deduction of sophisticated moment expressions, of flexible goodness-of-fit tests, and of novel tests for the symmetry of bivariate count distributions. The paper concludes with an analysis of real-world data examples.
We derive the asymptotic distribution of ordinal-pattern frequencies under weak dependence conditions and investigate the long-run covariance matrix not only analytically for moving-average, Gaussian, and the novel generalized coin-tossing processes, but also approximately by a simulation-based approach. Then, we deduce the asymptotic distribution of the entropy-complexity pair, which emerged as a popular tool for summarizing the time-series dynamics. Here, we make the necessary distinction between a uniform and a non-uniform ordinal pattern distribution and, thus, obtain two different limit theorems. On this basis, we consider a test for serial dependence and check its finite-sample performance. Moreover, we use our asymptotic results to approximate the estimation uncertainty of entropy-complexity pairs.
The thinning-based integer-valued autoregressive moving-average (INARMA) models are popular for count time series. Recently, types of INARMA models have also been developed for count random fields, i.e., for spatial count data located on a regular two-dimensional grid. This article provides a comprehensive survey on existing INARMA random fields, covering approaches with different thinning operators, first- and higher-order models, as well as unilateral and multilateral model structures.
The monitoring of serially independent or autocorrelated count processes is considered, having a Poisson or (negative) binomial marginal distribution under in-control conditions. Utilizing the corresponding Stein identities, exponentially weighted moving-average (EWMA) control charts are constructed, which can be flexibly adapted to uncover zero inflation, over- or underdispersion. The proposed Stein EWMA charts' performance is investigated by simulations, and their usefulness is demonstrated by a real-world data example from health surveillance.
Different types of omnibus goodness-of-fit tests for a binomial null hypothesis are developed, which are based on the counts’ probability generating function. For each statistic, closed-form asymptotics are derived and used to implement the tests in practice without the need for a bootstrap procedure. The finite-sample performance of the tests is analyzed by means of a comprehensive simulation study, and practical recommendations for the choice of the test statistic are given. The practical application of the tests is illustrated by a couple of real-world data examples.
In process monitoring, it is common for measurements to be taken regularly or randomly from different spatial locations in two or three dimensions. While there are nonparametric methods for process monitoring with such spatial data to detect changes in the mean, there is a gap in the literature for nonparametric control charting methods developed to monitor spatial dependence. This study considers streams of regular, rectangular datasets using spatial ordinal patterns (SOPs) as a nonparametric method to detect spatial dependencies. We propose novel, distribution-free SOP control charts. To uncover higher-order dependencies, we develop a new class of statistics that combines SOPs with the Box-Pierce approach. An extensive simulation study demonstrates the superiority and effectiveness of our proposed charts over traditional parametric approaches, particularly when the spatial dependence is nonlinear or bilateral or when the spatial data contains outliers. The proposed SOP control charts are illustrated using real-world datasets to detect (i) heavy rainfall in Germany, (ii) war-related fires in (eastern) Ukraine, and (iii) manufacturing defects in textile production. This wide range of applications and findings demonstrates the broad utility of the proposed nonparametric control charts. In addition, all methods in this study are provided as a publicly available Julia package on GitHub for further implementations.
The mollified uniform distribution is rediscovered, which constitutes a ``soft'' version of the continuous uniform distribution. Important stochastic properties are derived and used to demonstrate potential fields of applications. For example, it constitutes a model covering platykurtic, mesokurtic and leptokurtic shapes. Its cumulative distribution function may also serve as the soft-clipping response function for defining generalized linear models with approximately linear dependence. Furthermore, it might be considered for teaching, as an appealing example for the convolution of random variables. Finally, a discrete type of mollified uniform distribution is briefly discussed as well.
In this paper, several regression-type models for multivariate ordinal time series are developed. The regression equations are inspired by existing GARCH-type models for univariate discrete-valued time series and include feedback terms in addition to the usual lagged observations to model the memory behavior. The corresponding terms from other individuals (components) are represented by weighted averages which are calculated based on a proximity matrix. The marginal conditional distributions are either binomial (employing the simplifying rank-count formulation) or multinomial. The approach can be generalized to obtain VARMA-type models to allow for more specific dependence between individuals. Additionally, different copulas are considered to model possible cross-dependence explicitly. The main data example concerns the daily air quality (ordinal) in three cities in North China. Here, a spatial dimension is present, which can be exploited in the definition of the proximity matrix and the copulas.
{Asymptotic implementation of discrete-sum versions of pgf-based GoF-tests for counts}A popular approach for deriving omnibus goodness-of-fit (GoF) tests for Poisson counts are integral statistics with respect to the probability generating function (pgf). As a drawback, the corresponding asymptotics are rather complicated such that the tests require a computationally demanding bootstrap implementation in practice. We propose discrete-sum versions of popular pgf-based GoF-tests and derive the corresponding asymptotics. These closed-form asymptotics are easily implemented in practice such that the novel discrete-sum GoF-tests do not pose any computational burden. We also show how to adapt our approach to different null hypotheses than Poisson, where we again derive the closed-form asymptotics. This general approach is exemplified by two types of negative-binomial null distribution. A simulation study compares the performance of our novel discrete-sum GoF-tests to that of the classical integral statistics, and illustrative real-world data examples are presented as well.