
This paper introduces a penalized least squares method using the group exponential (GEXP) penalty to simultaneously perform variable selection and estimation of the coefficient regression matrix in multivariate linear regression models intended to show the relationship between multiple response variables and multiple explanatory variables. The GEXP penalty is a continuous approximation of the ℓ0 penalty, which, despite its simple form, avoids the drawbacks of ℓ0 penalized estimation methods that arise from the discontinuity of the ℓ0 penalty and thereby provides more stable estimation. We show that the estimator possesses the oracle properties under a high-dimensional asymptotic framework in which the sample size goes to infinity, and both the number of response variables and the number of explanatory variables can go to infinity. We also apply the EGCV criterion as a model selection criterion to choose the two tuning parameters required for the GEXP penalized estimation and show that the EGCV criterion maintains variable selection consistency under certain assumptions. Through numerical experiments, we confirm that the performance of the proposed method is comparable to existing methods.
The issue of estimation and unit root test has been essential and critical for analysis of time series, but it remains less addressed for spatial time series data. In this paper, we first extend the estimation and unit root test from time series to spatio-temporal unstable AR(1) model. We have proposed a local location coefficient fitting to overcome the location-wide heteroscedasticity and developed the estimation procedure and unit root test with the asymptotic properties established. Further extension to the spatio-temporal AR(p) model with unit root test is developed. Our theory and method can be seen as an extension of the popular augmented Dickey–Fuller (ADF) test in time series analysis to the spatio-temporal data setting. Both asymptotic theory and simulations clearly demonstrate that our spatio-temporal estimation and unit root test are more efficient and reliable than the time series only based estimation procedure and traditional unit root test.
This paper introduces a new test for the population mean vector in high-dimensional data with missing observations. Our method overcomes key limitations of existing approaches, including the need for finite fourth-order moments and restrictive assumptions on the probabilities of missing data indicating dense missing pattern. The proposed test statistic is standardized using the conditional variance given the observed pattern of missing values, eliminating the requirement that the probabilities of missing observations be bounded away from one and accommodating sparse missing structure.In contrast to existing methods, our approach requires only the existence of finite (2+ζ)-th moments for some ζ>0, making it applicable to heavy-tailed distributions where fourth order moment is not finite. The procedure also accommodates weak dependence across variables via geometric α-mixing and does not rely on the assumption of identically distributed observations. This significantly improves the versatility of the proposed test. Theoretical guarantees are provided through the asymptotic normality of the test statistic, and we establish consistency and correct asymptotic size under appropriate conditions.
The compound symmetric structure testing of a high-dimensional covariance matrix plays an important role in multivariate statistical analysis. Based on the Frobenius norms of the difference and ratio between two estimates of the population covariance matrix, this paper constructs two different test statistics to test if the population covariance matrix has the compound symmetric structure for elliptically distributed data. To enhance the powers of the two statistics, another three test methods are developed. In addition, under the high-dimensional framework that the data dimension is allowed to grow to infinity in proportion to the sample size, we derive the asymptotic distributions of the proposed test statistics under the null and alternative hypotheses. Notably, we establish the asymptotic normality of the difference statistic by employing a refined technique. Finally, extensive simulation results and a real data analysis demonstrate the performance of our proposed methods.
To solve the problems of both low imputation precision and high time costs, the trimmed scores regression (TSR) method has been proposed as an imputation method for clustering data with missing values, where clustering pursues homogeneity within clusters and heterogeneity between clusters. However, as a centralized method, the TSR suffers from critical computational bottlenecks in large-scale scenarios. To overcome these limitations, we propose the distributed TSR (DTSR) method within a distributed framework. Two distributed methods, namely the DEMI and the DRPCA, as well as the TSR, are used as benchmarks. Through simulation and real-data analysis, we evaluate the DTSR from the perspectives of imputation accuracy and computational efficiency. The results show that the DTSR maintains high imputation accuracy equivalent to the TSR and significantly reduces running time compared with the centralized version. Meanwhile, the DTSR exhibits similar running time to both the DEMI and the DRPCA, achieving a better balance between accuracy and efficiency.
We propose a kernel-based nonparametric estimator for a smooth coefficient panel data model with fixed effects. Without requiring a zero sum of fixed effects, we propose an estimator that is easy to construct and computationally efficient. Eliminating the fixed effects through a local within transformation, we perform a local linear estimation for the coefficient functions associated with time-varying variables and associated derivatives. We further estimate the intercept coefficient function, if present, through a difference of kernel-weighted averages. We characterize the estimator's asymptotic properties under a large-n and large-T framework. We demonstrate that the estimator is not asymptotically equivalent to the standard kernel estimator that ignores fixed effects. Through extensive simulation studies, we highlight the estimator's encouraging numerical performance and computational advantages over existing kernel estimators in the literature. We showcase the empirical applicability by estimating a smooth coefficient model for the Environmental Kuznets Curve through a panel of OECD countries.
A variance-corrected U-type test statistic is presented for testing the equality of covariance matrices when the populations are high-dimensional mixtures and the dimension can exceed the sample sizes. The proposed statistic is a modification to the state-of-art two-sample Ustatistic, which is basically designed for the case where the mixing distributions degenerate to a Dirac point measure. Combined with a nontrivial use of the theory of one-sample and two-sample U-statistics in a context of high-dimensional mixtures, we prove that the asymptotic distributions of the proposed test statistic are normal under the high-dimensional null hypothesis and scenarios of the alternatives. The power of the proposed test is also investigated based on these asymptotic distributions. Without specifying the kurtosis and skewness of data, our test is widely applicable in practice. Simulation results, under a variety of mixture settings, are used to show the accuracy and robustness of the proposed test. A real data set is used to illustrate the proposed method.
In this work, we study a random orthogonal projection based least squares estimator for the stable solution of a multivariate non-parametric regression (MNPR) problem. The proposed estimator is based on the use of a restricted tensor product of uni-variate weighted Jacobi polynomials. More precisely, given an integer d >= 1 corresponding to the dimension of the MNPR problem and under the assumption that the given n random sampling vectors Xi follow the d-variate Beta distribution Beta( 1 2, 1 ). we show that a fairly large class of d-variate 2 regression functions is stably approximated by its random projection over this special set of d-variate weighted Jacobi polynomials system. In this case and unlike most least squares based estimators, no extra regularization scheme is needed by our proposed estimator. Moreover, we provide an L2-error as well as an L2-risk error of the estimator. Also, we show that if the regression function belongs to a Sobolev type space, then our estimator has the optimal convergence rate. In the general case where the sampling distribution of the Xi is unknown, we propose a regularized version of our previous d-variates orthonormal polynomials based estimator. Finally, we illustrate the performance of our proposed multivariate non-parametric estimator by some numerical simulations on synthetic as well as on real datasets.
Spatio-temporal data collected from irregularly spaced locations frequently exhibit complex characteristics, such as spatial nonstationarity and dynamic temporal dependence. Existing approaches typically rely on two-step kernel estimation procedures involving temporal regression followed by spatial smoothing, which can be computationally expensive and statistically inefficient. To address these challenges, we propose a global one-step estimation method based on tensor product B-splines for the location-dependent Spatio-temporal Autoregressive (STAR) model. This approach simultaneously approximates the nonparametric covariate effect and the spatially varying autoregressive coefficients via a single least-squares minimisation. We establish consistency and L2 convergence rates for both components, and further derive their asymptotic distributions. Monte Carlo simulations confirm the finite-sample advantages of the proposed estimator. The methodology is illustrated through analyses of provincial inflation dynamics in China and state-level housing prices in the United States.
Large-dimensional Markov chains appear in many models and many applications. In this paper, we introduce the Multidimensional Approximative Regenerative Block Bootstrap (MARBB), a bootstrap algorithm designed for high-dimensional Markov chains. We focus, in this paper, on a vector autoregressive (VAR(1)) process with a low-rank structure. We first use a reduction algorithm that transforms the original high-dimensional time series into a lower-dimensional Markov chain. Once the chain is reduced, we leverage the regenerative properties of Harris recurrent Markov chains within a general state space, using the Nummelin splitting technique to extend existing results from the one-dimensional settings to the multidimensional case. This approach enables the identification of approximate regeneration times of the lower-dimensional Markov chain, which in turn leads to the splitting of the original high-dimensional Markov chain into approximate regenerative blocks. These blocks are then used to bootstrap relevant statistics based on regenerative blocks. Finally, we give the MARBB consistency results, and we apply our algorithm to simulation data.
Motivated by the preeminence of Poisson point processes in various imaging modalities, we address the problem of recovering the intensity of a filtered Poisson point process on Rd. That is, we observe the convolution of the empirical measure underlying the point process with a known filter h. We investigate both the case where the full process is continuously observed, or the more realistic setting where it is observed on a grid. Upon choosing an orthornormal basis, we propose two adaptive nonparametric estimators of the intensity, either through a model selection approach or a thresholding procedure, with a focus on the case of a tensorized Hermite basis. We provide convergence rates for both observation schemes based on appropriate concentration inequalities for Poisson processes. Extensive and elaborate numerical experiments show the implementability of the procedures and illustrate the theoretical results.
In this paper, we address the challenges associated with testing independence between a high-dimensional random vector and a categorical random variable. These challenges include the high dimensionality of data and robustness. To overcome these challenges, we propose a novel semi-Grothendieck’s covariance via random integration technique with Grothendieck’s identity to measure the correlation between a high-dimensional random vector and a categorical random variable. The proposed semi-Grothendieck’s covariance satisfies the independence-zero equivalence property. Furthermore, we present a nonparametric independence test for a high-dimensional random vector based on our proposed semi-Grothendieck’s covariance. The asymptotic null distribution of our proposed test statistic follows a standard normal distribution without moment restrictions on the vectors, thanks to the high dimensions of random vectors. Furthermore, specific conditions can be easily verified in practical applications due to the mixingale condition. Numerical studies and a real data analysis demonstrate the superior finite-sample performance of our proposed method.
This paper simultaneously estimates the number of factors and the number of lags in a generalized high-dimensional dynamic factor model when both cross-sectional and temporal dimension of the data go to infinity proportionally. An estimation of the variance of the idiosyncratic errors is also proposed when it is unknown. Simulation shows the validity of the results. Finally, an empirical study is conducted.
The Growth Curve Model (GCM), extensively utilized in longitudinal studies, requires precise estimation of within-subject covariance for the effective computation of mean estimates. When estimating covariance matrices in high-dimensional contexts—where the matrix dimension p exceeds the sample size n—employing sample covariance estimators becomes unsuitable. This study employs Hyper-sphere Decomposition and Alternative Hyper-sphere Decomposition to the within-subject covariance in GCM, which facilitates the joint estimation of mean and covariance. Furthermore, we explore the asymptotic properties of these estimators in various dimensional scenarios, demonstrating that the proposed methods remain consistent as p approaches infinity, while n remains fixed.
This paper proposes a nonparametric additive regression technique that can be used to analyze the interaction effects as well as the individual effects of the covariates. A powerful method of estimating the component functions that represent the individual and interaction effects is introduced and studied in a high-dimensional regime that allows the number of covariates to be much larger than the sample size. Asymptotic L2 error bounds are derived for the estimators under mild technical conditions in sparse settings where the number of nonzero components is smaller than the sample size but is allowed to increase to infinity as the sample size grows. The L2 error bounds reduce to the rate that can be achieved in bivariate smoothing, up to a logarithmic factor, when the number of significant effects is bounded. Numerical evidences are also provided via some simulation studies and a real data example.
The distribution-to-distribution regression models have received considerable attention over the past years. The existing studies on distribution-to-distribution regression models mainly focus on univariate distributions, and it is quite challenging to extend the existing methods to multivariate settings due to no analytical solution for the optimal transport problem and complicated theories involved. We propose a novel method to make inference on multidimensional distribution-to-distribution regression with both covariates and response variables following multidimensional probability distributions based on the theory of optimal transportation. The considered model links the conditional Fr & eacute;chet mean of response variable to the covariates via the optimal transport map, which is taken as the gradient of the input convex neural network (ICNN) trained with the minimax optimization. We develop the Fr & eacute;chet-least-square (FLS) estimator of the regression map by learning the optimal Kantorovich potential, and investigate the identifiability, consistency and convergence rate of the FLS estimator based on the theory of empirical process. We also construct a case-deletion diagnostic measure to identify influential observations. Simulation studies and a real example are used to illustrate the proposed methodologies.
We propose the Inner-Envelope Matrix Autoregression (IEMAR) for matrix-valued time series. IEMAR captures the necessary dynamics by restricting the MAR operators to the largest noise-reducing row/column subspaces contained in their coefficient spaces, yielding a strictly more economical parameterization while preserving the bilinear structure. We develop a Gaussian likelihood and a stable block algorithm with closed-form weighted least-squares updates for the dynamic cores, flip-flop updates for Kronecker-separable innovations, and mode-wise inner-SIMPLS envelope updates on the lattice of reducing subspaces. We prove existence and uniqueness of the inner envelopes (as subspaces), consistency of the estimated envelopes under mild mixing, asymptotic normality of the core parameters, and an efficiency ordering that favors IEMAR over EMAR (outer-envelope MAR) and unconstrained MAR whenever the necessary space is a proper subset of the sufficient one. Simulations show systematic parameter-efficiency gains with forecast risk comparable to EMAR at short/medium horizons, and a New York City taxi application illustrates how IEMAR isolates operative row/column dynamics and improves out-of without
This paper considers the problem of testing multiple hypotheses about the distribution of a random vector based on sequentially sampled observations from it. It is assumed that sampling from a different subset of dimensions incurs a different cost, and the goal is to minimize the expected total sampling cost while controlling the probabilities of misidentifying the true hypothesis below user-specified levels. The proposed testing procedure is shown to achieve this goal asymptotically as the levels go to zero. Specifically, it works by greedily following the optimal sampling distribution based on the current maximum likelihood hypothesis while guaranteeing sufficient exploration of all dimensions, and selecting the maximum likelihood hypothesis when the likelihood for it exceeds the likelihoods for all other hypotheses by sufficient margins. Numerical studies in the multivariate Gaussian case are presented for illustration.