In this article, we propose a novel model for time series of counts called the hysteretic Poisson autoregressive (HPART) model with thresholds by extending the linear Poisson autoregressive model into a nonlinear model. Unlike other approaches that bear the adjective “hysteretic", our model incorporates a scientifically relevant controlling factor that produces genuine hysteresis. Further, we re-analyse the buffered Poisson autoregressive (BPART) model with thresholds. Although the two models share the convenient piecewise linear structure, the HPART model probes deeper into the intricate dynamics that governs regime switching. We study the maximum likelihood estimation of the parameters of both models and their asymptotic properties in a unified manner, establish tests of separate families of hypotheses for the non-nested case involving a BPART model and a HPART model, and demonstrate the finite-sample efficacy of parameter estimation and tests with Monte Carlo simulation. We showcase advantages of the HPART model with two real time series, including plausible interpretations and improved out-of-sample predictions.
Statistically Meaningful Geometry (SMG) is a differential-geometric and information-theoretic framework that lifts over-parameterized models into infinite-dimensional non-parametric Orlicz statistical fiber bundles with an Ehresmann connection, decoupling unobservable vertical gauge noise from horizontal statistically verifiable directions. We prove the First Edge Theorem: Amari's information geometry (IG) and conventional statistics (CS) are not autonomous statistical universes but degenerate boundary layers of the larger gauge-active SMG space. Taking the structural identifiability radius R→∞ breaks gauge symmetry, collapses vertical fibers, and forces the total space onto the IG manifold; subsequent local asymptotic normality as N→∞ flattens the remaining curvature, yielding SMGIGCS. The same fiber-bundle machinery transforms model non-identifiability from a singular collapse into a structured gauge space. Applications include resolving the deep-learning generalization paradox, constructing gauge-invariant gradient descent and holonomy-matched preference alignment for generative AI, and solving weak identification in structural econometrics via intrinsic horizontal geodesic search.
Conventional uniform convergence bounds and empirical risk minimization break down in massive over-parameterized models, such as large language transformers and biological sequence networks. With near-infinite unconstrained internal degrees of freedom, their optimization landscapes develop flat vertical gauge valleys, rendering classical generalization metrics vacuous and inducing severe pathologies, specifically generative hallucination and catastrophic forgetting. We introduce the Statistically Meaningful Geometry (SMG) framework, an information-geometric paradigm lifting deterministic parametric models into infinite-dimensional non-parametric Orlicz statistical manifolds. Modeling the total state space as a differential fiber bundle (ℳ, ℬ, π, 𝒱, ℋ, ω), we establish a Two-Fold Inference Paradigm. We formalize an Ehresmann connection 1-form ω as a dynamic geometric filter that strips away vertical gauge noise (Structural Internal Directions, or SID) and isolates learning trajectories along the strictly non-degenerate horizontal distribution (Statistical Variational Directions, or SVDχ). We prove that under connection-filtered pre-training, out-of-distribution predictive variance is strictly upper-bounded by the finite diameter of the identifiable quotient base manifold ℬ, establishing a hard geometric containment of generative hallucinations. By projecting downstream updates onto the orthogonal complement of the historical horizontal carriage, we formalize the SMG Sequential Adaptation Flow, proving the total non-asymptotic elimination of catastrophic forgetting. SMG replaces empirical fine-tuning heuristics with coordinate-free topological constraints, bridging advanced differential geometry with structural reliability in AI.
This paper establishes the global proof of the Second Edge Theorem: as sample size tends to infinity, sample-dependent information-geometric manifolds—formed by the parameter space, sample-scaled Fisher metric, and dual alpha-connections—undergo metric-topological collapse onto the flat tangent space of Conventional Statistics at the true parameter. We first prove a universal tensor valence scaling law, under which tensor fields of valence one through four degenerate at rates determined by sample size. Score fluctuations stabilize, the Fisher metric freezes to its true-parameter value, affine connections dissolve, and Riemann curvature is annihilated. We then bridge geometric collapse with statistical decision theory, showing that Cheeger–Gromov flattening and Le Cam risk condensation are dual projections of the same asymptotic phase transition. A Fisher-compatible Ehresmann connection extends these results to over-parameterized and singular models, yielding horizontal leaf-space collapse and uniform local asymptotic normality. Unifying the First and Second Edge Theorems yields a nested dual-edge hierarchy: Conventional Statistics is the boundary of Information Geometry, which is itself the boundary of Statistical Mechanics and Geometry. Thus, Conventional Statistics is not a heuristic approximation but the unique zero-curvature thermodynamic attractor of regular parametric information manifolds. This redefines modern statistics as a dynamic non-equilibrium field theory of finite-sample fluctuations, phase transitions, and gauge-invariant interactions.
Models with unnormalized probability density functions are ubiquitous in statistics, artificial intelligence and many other fields. However, they face significant challenges in model selection if the normalizing constants are intractable. Existing methods to address this issue often incur high computational costs, either due to numerical approximations of normalizing constants or evaluation of bias corrections in information criteria. In this paper, we propose a novel and fast selection criterion for nested models of possibly dependent data, allowing direct data sampling from a possibly unnormalized probability density function. With a suitable multiplying factor depending only on the sample size and the model complexity, the proposed criterion gives a consistent selection under mild regularity conditions and is computationally efficient. Extensive simulation studies and real-data applications demonstrate the efficacy of this criterion in the selection of nested models with unnormalized probability densities.
The rapid scaling of over-parameterized machine learning architectures, particularly LLMs, raises a profound crisis: do these systems exhibit genuine intelligence, or are they merely sophisticated statistical pattern matchers? Classical flat Euclidean statistics cannot differentiate continuous interpolation from the autonomous discovery of novel causal laws. To resolve this, we introduce Statistically Meaningful Geometry (SMG), a framework modeling over-parameterized learning systems as infinite-dimensional non-parametric Orlicz fiber bundles. We prove that under persistent out-of-distribution (OOD) stimuli governed by unmodeled causal mechanisms, continuous optimization fails. Unmodeled variance is rejected by the visible horizontal base manifold, leaking into the unobservable vertical fiber space and generating an accumulation of Active Acausal Tension. Driven by the statistical manifold's non-linear curvature, this tension inevitably strikes a conjugate focal boundary (T_crit = π^2 / K_max), triggering localized volumetric collapse and a catastrophic matrix singularity ([G_f]^-1→∞). We demonstrate this geometric breakdown acts as the strict non-equilibrium trigger for a Gauge Symmetry Break (GSB). The system purges hidden tension from unobservable gauge redundancies, spontaneously crystallizing a new, mathematically independent horizontal coordinate axis. This non-parametric phase transition registers as a discrete +1.0 integer step-jump in observable Structural G-Entropy. By decoupling parameter charts and subjecting emergent axes to a Minimal Energy Path Criterion and a Causal Invariance Filter, we distinguish genuine discovery from malignant hallucinations. Ultimately, SMG provides a parameter-free, falsifiable dashboard to mathematically certify true intelligence, transforming AI for Science into an engine of autonomous paradigm shifts.
Non-parametric information geometry has long faced an “intractability barrier”: in the infinite-dimensional setting, the Fisher–Rao metric is a weak Riemannian metric functional that lacks a bounded inverse, rendering classical optimization and estimation techniques computationally inaccessible. This paper resolves this barrier by building the statistical manifold on the Orlicz space L0Φ(Pf) (the Pistone–Sempi manifold), which provides the necessary exponential integrability for score functions and a rigorous Fréchet differentiability for the Kullback–Leibler divergence. We introduce a novel Structural Decomposition of the Tangent Space (TfM=S⊕S⊥), where the infinite-dimensional space is split into a finite-dimensional covariate subspace (S)—representing the observable system—and its orthogonal complement (S⊥). Through this decomposition, we derive the Covariate Fisher Information Matrix (cFIM), denoted as Gf, which acts as the computable “Hilbertian slice” of the otherwise intractable metric functional. Key theoretical contributions include proving the Trace Theorem (HG(f)=Tr(Gf)) to identify G-entropy as a fundamental geometric invariant; demonstrating the Geometric Invariance of the Covariate Fisher Information Matrix (cFIM) as a covariant (0,2)-tensor under reparameterization; establishing the cFIM as the local Hessian of the KL-divergence; and characterizing the Efficiency Standard through a generalized Cramer–Rao Lower Bound for semi-parametric inference within the Orlicz manifold. Furthermore, we demonstrate that this framework provides a formal mathematical justification for the Manifold Hypothesis, as the structural decomposition naturally identifies the low-dimensional subspace where information is concentrated. By shifting the focus from the intractable global manifold to the tractable covariate geometry, this framework proves that statistical information is not a property of data alone, but an active geometric interaction between the environment (data), the system (covariate subspace), and the mechanism (Fisher–Rao connection).
Being infinite dimensional, non-parametric information geometry has long faced an "intractability barrier" due to the fact that the Fisher-Rao metric is now a functional incurring difficulties in defining its inverse. This paper introduces a novel framework to resolve the intractability with an Orthogonal Decomposition of the Tangent Space (T_fM=S ⊕ S^⊥), where S represents an observable covariate subspace. Through the decomposition, we derive the Covariate Fisher Information Matrix (cFIM), denoted as G_f, which is a finite-dimensional and computable representative of information extractable from the manifold's geometry. Indeed, by proving the Trace Theorem: H_G(f)=Tr(G_f), we establish a rigorous foundation for the G-entropy previously introduced by us, thereby identifying it not merely as a gradient-based regularizer, but also as a fundamental geometric invariant representing the total explainable statistical information captured by the probability distribution associated with the model. Furthermore, we establish a link between G_f and the second-order derivative (i.e. the curvature) of the KL-divergence, leading to the notion of Covariate Cramér-Rao Lower Bound(CRLB). We demonstrate that G_f is congruent to the Efficient Fisher Information Matrix, thereby providing fundamental limits of variance for semi-parametric estimators. Finally, we apply our geometric framework to the Manifold Hypothesis, lifting the latter from a heuristic assumption into a testable condition of rank-deficiency within the cFIM. By defining the Information Capture Ratio, we provide a rigorous method for estimating intrinsic dimensionality in high-dimensional data. In short, our work bridges the gap between abstract information geometry and the demand of explainable AI, by providing a tractable path for revealing the statistical coverage and the efficiency of non-parametric models.
This note addresses issues raised by Cox and Reid in their seminal paper in 1987 regarding parameter orthogonality in statistical inference. We extend the orthogonality condition to cases with multiple parameters of interest and demonstrate its existence at a global level for some generally important distributions, despite previously expressed pessimism by them. Numerical results with the location-scale t-distribution reveal substantial gains in estimation accuracy and savings in computation time, thanks to the existence. We next show that the local parameter orthogonality can lead to efficient computational algorithms with the celebrated Whittle algorithm for multivariate autoregressive modeling as a showcase.
In this article, we propose a novel logistic quasi-maximum likelihood estimation (LQMLE) for general parametric time series models. Compared to the classical Gaussian QMLE and existing robust estimations, it enjoys many distinctive advantages, such as robustness in respect of distributional misspecification and heavy-tailedness of the innovation, more resiliency to outliers, smoothness and strict concavity of the log logistic quasi-likelihood function, and boundedness of the influence function among others. Under some mild conditions, we establish the strong consistency and asymptotic normality of the LQMLE. Moreover, we propose a new and vital parameter identifiability condition to ensure desirable asymptotics of the LQMLE. Further, based on the LQMLE, we consider the Wald test and the Lagrange multiplier test for the unknown parameters, and derive the limiting distributions of the corresponding test statistics. The applicability of our methodology is demonstrated by several time series models, including DAR, GARCH, ARMA-GARCH, DTARMACH, and EXPAR. Numerical simulation studies are carried out to assess the finite-sample performance of our methodology, and an empirical example is analyzed to illustrate its usefulness.
Most threshold models to-date contain a single threshold variable. However, in many empirical applications, models with multiple threshold variables may be needed and are the focus of this article. For the sake of readability, we start with the Two-Threshold-Variable Autoregressive (2-TAR) model and study its Least Squares Estimation (LSE). Among others, we show that the respective estimated thresholds are asymptotically independent. We propose a new method, namely the weighted Nadaraya-Watson method, to construct confidence intervals for the threshold parameters, that turns out to be, as far as we know, the only method to-date that enjoys good probability coverage, regardless of whether the threshold variables are endogenous or exogenous. Finally, we describe in some detail how our results can be extended to the K-Threshold-Variable Autoregressive (K-TAR) model, K > 2. We assess the finite-sample performance of the LSE by simulation and present two real examples to illustrate the efficacy of our modeling.
Regulation is an important feature of dynamic phenomena, and is commonly tested within the threshold autoregressive setting, with the null hypothesis being a global nonstationary process. Nonetheless, this setting is debatable, because data are often corrupted by measurement errors. Thus, it is more appropriate to consider a threshold autoregressive moving-average model as the general hypothesis. We implement this new setting with the integrated moving-average model of order one as the null hypothesis. We derive a Lagrange multiplier test that has an asymptotically similar null distribution, and provide the first rigorous proof of tightness in the context of testing for threshold nonlinearity against difference stationarity, which is of independent interest. Simulation studies show that the proposed approach enjoys less bias and higher power in detecting threshold regulation than existing tests, especially when there are measurement errors. We apply the new approach to time series of real exchange rates of a panel of European countries.
Recently, matrix-valued time series data have attracted significant attention in the literature with the recognition of threshold nonlinearity representing a significant advance. However, given the fact that a matrix is a two-array structure, it is unfortunate, perhaps even unusual, for the threshold literature to focus on using the same threshold variable for the rows and the columns. In fact, evidence in economic, financial, environmental and other data shows advantages of allowing the possibilities of two different threshold variables (with possibly different threshold parameters for rows and columns), hence the need for a Two-way Matrix AutoRegressive model with Thresholds (2-MART). Naturally, two threshold variables pose new and perhaps even fierce challenges, which might be the reason behind the adoption of only one threshold variable in the literature up to now. In this paper, we develop a comprehensive methodology for the 2-MART model, by overcoming various challenges. Compared with existing models in the literature, the new model can achieve greater dimension reduction, much better model fitting, more accurate predictions, and more plausible interpretations.
This paper explores the interplay between statistics and generative artificial intelligence. Generative statistics, an integral part of the latter, aims to construct models that can generate efficiently and meaningfully new data across the whole of the (usually high dimensional) sample space, e.g. a new photo. Within it, the gradient-based approach is a current favourite that exploits effectively, for the above purpose, the information contained in the observed sample, e.g. an old photo. However, often there are missing data in the observed sample, e.g., missing bits in the old photo. To handle this situation, we have proposed a gradient-based algorithm for generative modelling. More importantly, our paper underpins rigorously this powerful approach by introducing a new G-entropy that is related to the Fisher divergence. (The G-entropy is also of independent interest.) The underpinning has enabled the gradient-based approach to expand its scope. For example, it can now provide a tool for generative model selection. Possible future projects include discrete data and Bayesian variational inference.
We present supremum Lagrange Multiplier tests to compare a linear ARMA specification against its threshold ARMA extension. We derive the asymptotic distribution of the test statistics both under the null hypothesis and contiguous local alternatives. Moreover, we prove the consistency of the tests. The Monte Carlo study shows that the tests enjoy good finite-sample properties, are robust against model mis-specification and their performance is not affected if the order of the model is unknown. The tests present a low computational burden and do not suffer from some of the drawbacks that affect the quasi-likelihood ratio setting. Lastly, we apply our tests to a time series of standardized tree-ring growth indexes and this can lead to new research in climate studies.
Principal component analysis (PCA) is a most frequently used statistical tool in almost all branches of data science. However, like many other statistical tools, there is sometimes the risk of misuse or even abuse. In this paper, we highlight possible pitfalls in using the theoretical results of PCA based on the assumption of independent data when the data are time series. For the latter, we state with proof a central limit theorem of the eigenvalues and eigenvectors (loadings), give direct and bootstrap estimation of their asymptotic covariances, and assess their efficacy via simulation. Specifically, we pay attention to the proportion of variation, which decides the number of principal components (PCs), and the loadings, which help interpret the meaning of PCs. Our findings are that while the proportion of variation is quite robust to different dependence assumptions, the inference of PC loadings requires careful attention. We initiate and conclude our investigation with an empirical example on portfolio management, in which the PC loadings play a prominent role. It is given as a paradigm of correct usage of PCA for time series data.
Caused by Yersinia pestis, plague ravaged the world through three known pandemics: the First or the Justinianic (6th-8th century); the Second (beginning with the Black Death during c.1338-1353 and lasting until the 19th century); and the Third (which became global in 1894). It is debatable whether Y. pestis persisted in European wildlife reservoirs or was repeatedly introduced from outside Europe (as covered by European Union and the British Isles). Here, we analyze environmental data (soil characteristics and climate) from active Chinese plague reservoirs to assess whether such environmental conditions in Europe had ever supported "natural plague reservoirs". We have used new statistical methods which are validated through predicting the presence of modern plague reservoirs in the western United States. We find no support for persistent natural plague reservoirs in either historical or modern Europe. Two factors make Europe unfavorable for long-term plague reservoirs: 1) Soil texture and biochemistry and 2) low rodent diversity. By comparing rodent communities in Europe with those in China and the United States, we conclude that a lack of suitable host species might be the main reason for the absence of plague reservoirs in Europe today. These findings support the hypothesis that long-term plague reservoirs did not exist in Europe and therefore question the importance of wildlife rodent species as the primary plague hosts in Europe.