
In this paper, we consider asymptotic properties of empirical correlation of two correlated stationary first-order autoregressive processes with Gaussian innovations. The moderate deviation principle with explicit rate function is obtained. As applications, the asymptotic powers for hypothesis tests are shown to approach one at exponential rates. Simulation experiments are conducted to confirm the theoretical results. The main methods consist of asymptotic analysis technique and deviation inequality estimations.
Graphical models reveal the conditional dependence structure between random variables. We propose a new nonparametric method for learning edges in graphical models under a consolidated smoothing spline ANOVA (SS ANOVA) decomposition framework. We estimate the joint density function with an L1 penalty to interactions in the SS ANOVA decomposition. We propose an iterative procedure to compute estimates and establish convergence rates for the joint density and its interactions. Simulations indicate that the proposed method performs well under Gaussian and non-Gaussian settings. We illustrate the proposed methods using a real data example.
Strong orthogonal arrays are widely recognized as effective space-filling designs for computer experiments. Among them, those of strength three are particularly useful, since strong orthogonal arrays of strength four or higher may be too expensive for some investigations. Strong orthogonal arrays of strength three that possess some of the space-filling properties of strength four are more desirable. Such arrays with small numbers of factors have been thoroughly investigated, whereas those of large sizes remain relatively unexplored. In this paper, we develop a characterization and construction method for large-sized arrays with good space-filling properties. We further present theoretical and computational results that facilitate the implementation of our construction method. Additionally, we use a simulation study to illustrate the usefulness of the arrays produced by our method in developing statistical surrogate models.
This paper is concerned with joint statistical inference on the Hurst parameter H and the drift parameter α in a mixed fractional Ornstein–Uhlenbeck process. The process is defined as the solution to the stochastic differential equation dXt=−αXtdt+dMtH,t∈[0,T],where MtH=BtH+Wt is a mixed fractional Brownian motion with Hurst parameter H∈(3/4,1), BH denotes a fractional Brownian motion and W is an independent standard Brownian motion. When H>3/4, the driving noise MH is a semimartingale, and it admits a useful innovation (or filtering) representation.Assuming continuous-time observation of (Xt)0≤t≤T and letting T→∞, we establish the Local Asymptotic Normality (LAN) property for the family of probability measures generated by the mixed fractional Ornstein–Uhlenbeck process. The proof relies on the innovation approach, the integral equation satisfied by the corresponding kernel, and a detailed analysis of its Laplace transform via a scalar Hilbert boundary value problem. As a consequence, we obtain an explicit expression for the Fisher information matrix for the parameter vector θ=(H,α) and identify the associated Hájek local minimax lower bound for the asymptotic risk of regular estimators.
This paper investigates the asymptotic behavior of several nonparametric estimators within the current status model. We derive Cramér-type moderate deviations and a law of the iterated logarithm for the maximum smoothed likelihood estimator of the distribution function. Corresponding results are also established for plug-in estimators of the density and hazard rate functions. The analysis relies on Talagrand’s inequalities for VC-classes of functions and techniques from asymptotic analysis.
Global Sensitivity Analysis (GSA) is an important tool to better understand the behavior of black box models. Among the numerous methods for GSA, variance-based approaches have received much attention. Only a few papers focus on Quantile Oriented Sensitivity Analysis (QOSA), which can help in analyzing the behavior of the response at different quantile levels. Moreover, existing QOSA estimation methods have flaws: bias when input variables are dependent, loss of accuracy and efficiency as input space dimension increases. In this paper, we propose a new estimation procedure of QOSA indices based on the notion of projected random forest, with the initial random forest built from a criterion designed for quantiles: the pinball loss also known as quantile loss.
In this paper, we investigate a new partially linear quantile regression model with incompletely observed functional covariate and responses missing at random (MAR). First, we utilize the methods of reconstruction operator and regression imputation to recover the partially observed functional covariate and impute the missing responses, respectively. Then, we construct estimators for both the unknown slope function and the unknown scalar parameters by minimizing the proposed quantile loss function of the model. Second, the asymptotic properties of the estimators are established under reasonable conditions. Thirdly, extensive simulation studies demonstrate the superiority of the approach. Finally, a real data analysis about the relationship between economic development and energy consumption as well as carbon dioxide (CO2) emissions is finished to show the applicability and effectiveness of the proposed method.
The Poisson or negative binomial item count technique was developed to avoid the ceiling effect by replacing the list of non-sensitive items with a single non-sensitive question whose responses follow a Poisson or negative binomial distribution. However, identifying an appropriate "non-sensitive" question is often non-trivial. This motivates us to consider Benford's law as the basis for constructing the non-sensitive question. Building on this idea, we propose a new survey methodology that leverages Benford's law for estimating the prevalence of a sensitive attribute. We present the survey design, parameter estimation procedures, hypothesis testing strategies, and confidence interval construction for the prevalence parameter. Simulation studies are conducted to evaluate the performance of the proposed method. Finally, we illustrate the methodology using a real dataset on premarital sexual behaviour in a nationwide population survey.
The periodogram is a tool utilized in signal processing and spectral analysis to estimate the power spectrum of a signal. This research specifically concentrates on determining the asymptotic distribution of the periodogram for periodically correlated spatial processes. This enables us to derive the limiting distribution of the periodogram for multivariate spatial processes as well. The outcomes of this analysis hold practical significance in various applications such as anomaly detection, signal detection and classification, model selection, confidence interval estimation, and statistical hypothesis testing. To validate the theoretical concepts, a simulation experiment was conducted, and the results further confirm their validity.
The pinhole camera is the ubiquitous model for well-focused imaging systems but is physically unrealistic in the sense that it does not take into account the fact that the scene being imagined lies in front of the camera. Taking this directional information into account leads to the concepts of oriented projective shape and oriented projective shape space. Furthermore, in previous work it was shown that the resulting extrinsic statistical techniques for independent samples of twodimensional image data have greater statistical power than comparable statistical techniques which ignore directional information. In this work we develop a novel matched pairs test for two-dimensional oriented projective shapes. This methodology is applied to the problem of doppelg & auml;nger or body double identification in the case of Russian president Vladimir Putin where we use two galleries of images, pre-2015 and post-2015, with 38 pictures in each. An unpaired test fails to find evidence that the galleries are of two different people whereas our novel matched pairs test, with optimally paired images, is highly statistically significant.
High-dimensional studies often contain a small set of prominent signals embedded in a broad non-sparse background, which can undermine purely sparse modeling. This paper proposes Weighted Non-sparse and Sparse Iteration (WNSI), an iterative procedure for joint estimation of sparse and non-sparse components that incorporates adaptive reweighting to stabilize updates and improve adaptivity across iterations. WNSI alternates between precision matrix-based estimation for the non-sparse component and an adaptively weighted penalized regression step for the sparse component, balancing numerical stability and data-driven refinement. We prove that WNSI converges to the oracle solution and establish non-asymptotic error bounds, achieving optimal l2 rates under standard regularity conditions. Simulation studies show that WNSI achieves improved estimation accuracy and variable selection, and remains robust under strong dependence and distributional misspecification. An application to breast cancer gene expression data illustrates that WNSI achieves competitive prediction across multiple screening and grouping configurations, highlighting its practical utility for high-dimensional genomic systems.
This paper introduces BinGFI, a novel, fully automated computational method for conducting statistical inference in binary response models that does not rely on Markov chain Monte Carlo or explicit mathematical integration. BinGFI is based on generalized fiducial inference (GFI) and extends the AutoGFI framework (Du et al., 2025) originally developed for additive Gaussian noise models to Bernoulli noise. This paper also develops a regularized extension, BinGFI-R, which incorporates convex penalties and a de-biasing step for improved inference accuracy. BinGFI-R can be used in situations where regularization is needed, such as in high-dimensional settings. The flexibility of BinGFI and BinGFI-R enables their application to a wide range of binary models, including classical logistic regression, covariate-assisted ranking estimation, and the Rasch model for item response theory. Through extensive simulations and comparisons with existing inference methods, we demonstrate that BinGFI and BinGFI-R achieve competitive or superior performance in terms of estimation accuracy, coverage rates, and interval widths.
We investigate block designs, under the A- and MV-criteria, when each treatment can have only one or two replications due to resource constraints, as can happen, for example, in early generation varietal trials. While these are commonly known as partially replicated designs, a key new feature of the present work is that no restriction about a constant block size is imposed on the subdesign consisting of the twice replicated treatments. This makes the derivation more challenging but allows us to entertain a wider class of competing designs and hence increases the flexibility of the results. Considering all treatments as equally important, design-independent, sharp lower bounds on the A- and MV-criteria are derived, so as to find highly efficient designs over this wider class. The roles of (a) linked block designs, (b) designs in an online catalog designtheory.org, and (c) partially balanced incomplete block (PBIB) designs, or duals thereof, as adapted to our setup, are explored at length. Illustrative examples are presented.
Generalized linear model or GLM constitutes a large class of models and essentially extends the ordinary linear regression by connecting the mean of the response variable with the covariate through appropriate link functions. On the other hand, Lasso is a popular and easy-to-implement penalization method in regression when not all covariates are relevant. However, the asymptotic distributional properties the Lasso estimator in GLM is still unknown. In this paper, we show that the Lasso estimator in GLM does not have a tractable form and subsequently, we develop two Bootstrap methods, namely the Perturbation Bootstrap and Pearson's Residual Bootstrap methods, for approximating the distribution of the Lasso estimator in GLM. As a result, our Bootstrap methods can be used to draw valid statistical inferences for any sub-model of GLM. We support our theoretical findings by showing good finite-sample properties of the proposed Bootstrap methods through a moderately large simulation study. We also implement one of our Bootstrap methods on a real data set.
Marginal expected shortfall is unquestionably one of the most popular systemic risk measures. Studying its extreme behaviour is particularly relevant for risk protection against severe global financial market downturns. In this context, results of statistical inference rely on the bivariate extreme values approach, disregarding the extremal dependence among a large number of financial institutions that make up the market. In order to take it into account we propose an inferential procedure based on the multivariate regular variation theory. We derive an approximating formula for the extreme marginal expected shortfall and obtain from it an estimator and its bias-corrected version. Then, we show their asymptotic normality, which allows in turn the confidence intervals derivation. Simulations show that the new estimators greatly improve upon the performance of existing ones and confidence intervals are very accurate. An application to financial returns shows the utility of the proposed inferential procedure. Statistical results are extended to a general $\beta$-mixing context that allows to work with popular time series models with heavy-tailed innovations.
We observe n independent pairs of random variables (Wi,Yi), where the conditional distribution of Yi given Wi=wi follows a one-parameter exponential family with parameter γ∗(wi)∈R. The goal is to estimate the regression function γ∗. We start with an arbitrary collection of piecewise constant candidate estimators based on the observations and, using the same data, select an estimator from this collection. The approach is agnostic to the dependencies of the candidate estimators on the data, differing from methods like data splitting, cross-validation, and hold-out. To demonstrate the theoretical performance of this approach, a non-asymptotic risk bound is established for the selected estimator. We then discuss the application of the procedure to changepoint detection in exponential families. The practical performance of the proposed approach is illustrated using comparative simulations across different scenarios and real datasets.
In observational studies, achieving covariate balance in pair matching between treatment and control groups or exposed and unexposed groups is essential. This balance enables testing treatment effects or examining associations between exposures and multivariate response variables in pair-matched data. Paired design studies involve taking multiple measurements for the same subjects under different conditions. All these call for an effective test for multivariate paired data. However, current methods for assessing covariate balance in matched observational studies often ignore the paired structure, leading to reduced performance in some cases. The multivariate paired Hotelling's T^2 test can be used for paired data, but its power decreases rapidly as dimensions increase. To address these issues, we propose a new non-parametric test for paired data, significantly improving power across various scenarios. We also derive the test's asymptotic distribution, making it user-friendly for practical applications. Our proposed test's effectiveness is demonstrated through an analysis of real data on Alzheimer's disease research.
The identification of the network effect is based on either group size variation, the structure of the network or the relative position in the network. I provide easy-to-verify necessary conditions for identification of undirected network models based on the number of distinct eigenvalues of the adjacency matrix. Identification of network effects is possible; although in many empirical situations existing identification strategies may require the use of many instruments or instruments that could be strongly correlated with each other. The use of highly correlated instruments or many instruments may lead to weak identification or many instruments bias. This paper proposes regularized versions of the two-stage least squares (2SLS) estimators as a solution to these problems. The proposed estimators are consistent and asymptotically normal. A Monte Carlo study illustrates the properties of the regularized estimators. An empirical application, assessing a local government tax competition model, shows the empirical relevance of using regularization methods.
Change-plane analysis has emerged as an effective tool to detect subgroups with distinct effects on the response of interest. Notably, the change-plane Cox model has gained significance in the subgroup analysis of survival data. Testing for the existence of a change plane can provide valuable insights into optimal decisions for the specific subgroups. However, classical supremum testing methods often suffer from limited efficiency in practical. To address this drawback, we propose a novel testing procedure designed to enhance statistical power. Our approach calculates the weighted average of the squared score test statistic (WAST) over the space of parameter that defines the subgroup, significantly improving power in practice. Moreover, we derive the asymptotic distributions of the test statistic under both the null and local alternative hypotheses. The performance of the proposed method is evaluated by extensive simulation studies, exhibiting better size control and higher power than existing testing approaches. Additionally, we illustrate the practical application of our approach using the Lending club loan dataset and the German credit dataset, showing its ability to identify subgroups with different default risks.