
We study Spatial Logistic Gaussian Process (SLGP) models for non-parametric estimation of probability density fields using scattered samples of heterogeneous sizes. SLGPs are examined from the perspective of random measures and their densities, investigating the relationships between SLGPs and underlying processes. Our inquiries are motivated by SLGP’s abilities in delivering probabilistic predictions of conditional distributions at candidate points, allowing conditional simulations of probability densities, and jointly predicting multiple functionals of target distributions. We demonstrate that SLGP models exhibit joint Gaussianity of their log-increments, enabling us to establish theoretical results regarding spatial regularity. Additionally, we extend the notion of mean-square continuity to random measure fields and establish sufficient conditions on covariance kernels underlying SLGPs to ensure these models enjoy such regularity properties. Finally, we propose an implementation using Random Fourier Features and showcase its applicability on synthetic examples and on temperature distributions at meteorological stations.
This study develops a framework to analyze the convergence of systematic-scan and random-scan Gibbs samplers for bivariate discrete conditional models with possibly different supports. We validate Liu (Test, 5, 305-310, 1996)’s conjecture: any stationary distribution of a random-scan Gibbs sampler is a mixture of the stationary distributions of systematic-scan Gibbs samplers. Moreover, we demonstrate the guaranteed convergence of the random-scan Gibbs sampler regardless of the selection probability. This contrasts with the systematic-scan Gibbs sampler, which may fail to converge under different scan orders. Previous studies of the compatibility theory require that the conditional distributions share the same support. By relaxing this requirement, we introduce the concept of generalized compatibility and provide necessary and sufficient conditions for the convergence of systematic-scan and random-scan Gibbs samplers to a unique generalized joint distribution. Furthermore, we establish a link between generalized compatibility and Gibbs sampling and explore the challenges of higher-dimensional extensions.
Finite mixture models are ubiquitous in modern statistical modeling, and a recurring practical issue is choosing the model order. In Keribin (Sankhyā Series A, 62, 49–66, 2000), the Bayesian information criterion (BIC) was proved consistent in mixtures, but under strong regularity, including high moments and high-order derivatives of the component density. We introduce the ν -BIC and ϵ -BIC, which weight the BIC penalty by negligibly small logarithmic factors immaterial in practice. This minor modification yields consistency under substantially weaker conditions, without differentiability and with mild moment assumptions, and we also give a misspecification result: when the truth lies outside the candidate family, any vanishing-penalty IC eventually selects a Kullback–Leibler optimal order among candidates. Finally, we clarify two limitations of consistent IC-based selection in mixtures: there is no universally minimal BIC-scale penalty within our sufficient conditions, and order consistency can conflict with minimax optimality in Hellinger risk. We illustrate the theory for Gaussian mixtures, non-differentiable Laplace mixtures, heavy-tailed t-mixtures, and mixtures of regression models.
Factorization of the spectral density into the product of a causal and an anti-causal components plays a critical role in the frequency domain analysis of a second order time series, including linear forecasting and signal extraction. This paper generalizes the approach to higher order polyspectral densities with a view towards nonlinear filtering, and develops a general theory of polyspectral factorization, providing new mathematical results for polyspectral densities. New bijections between a restricted space of higher-dimensional cepstral coefficients (where the restrictions are induced by the symmetries of the polyspectra) and the autocumulants are derived. Applications to modeling nonlinear time series are developed; in particular, it is shown that semi-parametric nonlinear time series modeling can be accomplished by approximation of the cepstral representation of polyspectra.
A key step in separating signal from noise in time series by means of singular spectrum analysis (SSA) is grouping. We present a multiple testing method for the grouping step in SSA. As separability criterion, we utilize the weighted correlation between the signal and the noise component of the (reconstructed) time series, and we test whether the population version of this weighted correlation is equal to zero. This test has to be performed for several possible groupings, resulting in a multiple test problem. The null distributions of the corresponding test statistics are approximated by a wild bootstrap procedure. The performance of our proposed method is assessed in a simulation study, and we illustrate its practical application with an analysis of real world data.
In multiple treatment comparison in clinical trials, patients typically arrive sequentially, and they are assigned to one of the competing treatments using some randomized allocation rules that are often response-adaptive in nature. Our objective is to determine the treatments that are superior or inferior to the control, which leads to testing multiple hypotheses under the response-adaptive design framework. We develop two novel test procedures to control the probabilities of at least one false rejection (FWER-I) and at least one false acceptance (FWER-II) at prefixed levels. The allocation rules in our procedures minimize some appropriate cost functions, namely the expected number of treatment failures or the total sample sizes. We prove that our procedures terminate sampling in finite times and control both FWER-I and FWER-II asymptotically as the difference between the null and alternative hypotheses vanishes. Superiority of the proposed procedures are validated through detailed analyses of real and simulated datasets.
In this paper, we explore the idea of comparing two multiple linear regression (MLR) models using the multistage sequential sampling approach. In particular, we intend to compare the linear parametric functions of two MLR models with standardized predictors while fixing the upper bounds on type I and type II error probabilities. We establish that such a comparison is not possible with any fixed-sample statistical technique. Therefore, we propose an optimal three-stage sequential sampling strategy to solve this problem. We study various attractive properties of the associated stopping rule, namely first-order efficiency, second-order efficiency, and asymptotic consistency. We also provide brief analyses on composite hypotheses, power comparisons, and sensitivity of the proposed methodology. The practical implications and limitations of the model assumptions are also discussed. An extensive simulation analysis is performed to validate the theoretical findings, and an application based on Boston housing data is provided for illustration purposes.
As Gaussian random fields with a Markov property, the Ornstein–Uhlenbeck (OU) fields are widely used to model spatial and temporal dependence in areas such as computer experiments, geostatistics, physics, and finance. This paper considers parameter estimation for the covariance function of a univariate anisotropic OU field on ℝ^d for d ≥ 2 . Based on observations in a fixed domain, we propose closed-form estimators that eliminate the need for numerical optimization and are computationally more feasible than maximum likelihood estimators. The proposed estimators retain the strong consistency and asymptotic normality, and their asymptotic variances are slightly (at most 14
In statistical literature, joint modeling survival data and longitudinal covariates has been studied, but almost all existing methods rely on parametric or semiparametric assumptions on longitudinal covariate process, and the resulting inferences critically depend on the validity of these assumptions that are unjustifiable or difficult to verify in practice. Without these parametric/semiparametric assumptions, this article proposes a generalized latent proportional hazards (GLPH) model for right censored survival time with longitudinal covariates, which treats longitudinal covariate process as unknown stochastic process with unknown hidden random variable and takes into account the within-subject historic change patterns of longitudinal covariates. The asymptotic properties of empirical likelihood MLE for GLPH model are established here for intensive longitudinal covariates, which lead to a statistical analysis of a very intensive longitudinal dataset. Moreover, the established asymptotic results give the construction of a novel method for joint modeling survival time and sparse longitudinal covariates via GLPH model.
We consider the class of symmetric α -stable moving average processes with 1< α < 2 . These processes are H-self-similar ( 0< H < 1 ) with stationary increments, indexed by ℝ^d , and driven by a symmetric α -stable random measure M_α . Our objective is to characterize these processes by estimating the Hurst parameter H, utilizing estimators based on p-variations and wavelet decomposition techniques. The idea is to exploit the self-similarity structure to calculate the p-variation along fixed directions. Two main results will be presented: the first establishes a law of large numbers type estimator for H, while the second provides a central limit theorem demonstrating convergence either to a Gaussian or to a stable distribution.
Hawkes processes have been used to quantify activity in cryptocurrency markets that are either exogenously or endogenously driven. The renewal Hawkes (RHawkes) process, which allows the underlying process for exogenous events to be a renewal process, has been found to provide a better fit to financial activity data in some markets. Although substantially more flexible, the RHawkes process is stationary in nature and unsuitable when systematic trends in the event occurrence rate exist. Therefore, we propose an RHawkes process with an exogenous arrival intensity modulated by a multiplicative trend function. When a suitable parametric form of the trend function is unavailable, we approximate it using basis spline functions. We propose algorithms to evaluate the likelihood, assess model fit, and simulate the process. Simulation experiments and an analysis of the endogeneity of the Bitcoin cryptocurrency are presented. Supplementary materials contain R code to implement the proposed methodologies.
We consider estimation of hitting-time variance, i.e. the explicit solution in the inverse first-hitting time problem of a continuous martingale to a constant boundary. The nonparametric estimation is based on delta-sequences. We also consider tuning parameter estimation related to the boundary. We characterize feasible statistics induced by central limit theory for the estimation procedure. A numerical simulation corroborates the asymptotic theory. An empirical application to financial data documents that the volatility is periodic at the duration scale. This can be explained by the endogeneity of transaction times.
The generalized method of moments (GMM) estimation approach has been particularly welcomed in spatial autoregressive model, but little research focused on the GMM estimation of semiparametric spatial autoregressive model with high dimensionality, especially when a number of predictors are endogenous. This paper mainly commits to providing an estimation and variable selection method for the higher-order spatial autoregressive functional coefficient model with endogenous covariates and diverging dimension. Based on the instrumental variables and basis function approximation, a series-based GMM estimation approach is firstly proposed. Then, a novel variable selection procedure is developed by utilizing the smooth-threshold estimating equations, which is convenient for implementation. Under some regularity conditions, we establish asymptotic properties of the resulting estimators. Detailed issues on computation and turning parameters selection are discussed. Extensive numerical simulations are conducted to confirm the theories and to demonstrate the finite sample performance of the proposed method.
Consider the regression model (Y_i=b(X_i)+ σ u_i, i=1,… , n) with error process u_i=√(ρ)ε _0+ √(1-ρ)ε _i where (ε _i, i=0, … , n) are independent centered random variables (r.v.) with unit variance, σ >0 and (X_1, … , X_n) are independent and identically distributed r.v. independent of (ε _i, i=0, … , n) . Thus there is a common noise ε _0 to all u_i ’s. We study the nonparametric estimation of the regression function b on a subset A of ℝ from the observations (X_i, Y_i,,i=1, … , n) using a projection method on sieves. The standard least-squares contrast fails to provide consistent estimators. Therefore, we introduce a least-squared contrast taking into account the covariance matrix of the noise. By minimizing the contrast over a finite dimensional space S_m , we obtain a projection estimator and study its risk. The estimators are implemented on simulated data and show excellent performances.
In this work we focus on a recently introduced family of continuous univariate distributions which involves two functions, g (generator) and h (parametric part) satisfying appropriate conditions. We first discuss methods to generate new members of the family by combining different parametric parts h or different generators g. We also present a technique to generate new families by transforming g and introduce a transformed family. In the sequel, we study the failure rate and tail behaviour of the families obtained by the aforementioned techniques. Finally, we demonstrate how the new theoretical framework can be exploited for constructing new distributional models with desirable properties in terms of failure rate and tail behaviour, which can be profitably exploited in actuarial science and engineering.
Nonparametric regression with random design is considered. The L_2 error with integration with respect to the design measure is used as error criterion. Over-parametrized deep neural network estimates are defined with logistic activation function where all parameters are learned by stochastic gradient descent. It is shown that the estimates achieve a nearly optimal rate of convergence in case that the regression function is (p, C)–smooth. In case that the regression function satisfies a projection pursuit model or more generally a hierarchical composition model the estimate achieves a rate of convergence which does not depend on the input dimension.
Suppose (standardized) measurements or statistics are monitored to raise an alarm when a threshold is exceeded. Often, the underlying population is heterogenous with respect to important discrete variables and thus samples may consist of imbalanced classes. We propose to use thresholds which depend on such covariates to boost the sensitivity for rare classes, which otherwise tend to be ignored. Under mild conditions, we identify optimal threshold functions and develop a feasible procedure for their computation. Further, for the proportional rule a nonparametric estimator of the threshold function is proposed and a central limit theorem is shown, including the case that conditional mean and variance used for standardization are estimated. For feasible uncertainty quantification a bootstrap scheme is proposed. The approach is illustrated and evaluated by a real data analysis.
We consider drawing statistical inferences based on data subject to non-Gaussian measurement error. The proposed strategy exploits hypercomplex numbers to reduce bias in naive estimation that ignores non-Gaussian measurement error. We apply this new method to several widely applicable parametric regression models with error-prone covariates, and kernel density estimation using error-contaminated data. The efficacy of this method in bias reduction is demonstrated in simulation studies and a real-life application in sports analytics.
This paper deals with a nonparametric Nadaraya-Watson (NW) estimator of the transition density function computed from independent continuous observations of a diffusion process. A risk bound is established on this estimator. The paper also deals with an extension of the penalized comparison to overfitting bandwidths selection method for our NW estimator. Finally, numerical experiments are provided.