
We study the fundamental statistical inference concerning the testing of independence between two random vectors. Existing asymptotic theories for test statistics based on distance covariance can only apply to either low-dimensional or high-dimensional settings and require stringent distributional assumptions. In this work we develop a new unified distributional theory of the sample generalized distance covariance that works for random vectors of arbitrary dimensions under fairly mild moment conditions. In particular, a Gaussian approximation result is established with a nonasymptotic error bound, and the asymptotic null distribution of the sample generalized distance covariance is shown to be distributed as a linear combination of independently and identically distributed chi-squared random variables. To estimate the asymptotic null distribution practically, we propose a half-permutation procedure and provide the theoretical justification for its validity. The exact asymptotic distribution of the resampling distribution is derived under general marginal moment conditions, and the proposed procedure is shown to be asymptotically equivalent to the oracle procedure with known marginal distributions.
This erratum contains a correction of the incorrect defined distribution on page 1109, last two lines, in Schneller (Ann. Statist. 17 (1989) 1103–1123). Additionally, some typographical errors are corrected.
We consider inference about a finite-dimensional parameter integrating samples from independent sources. A recently developed theory considers scenarios where sources align with subsets of the conditional distributions of a single factorization of the joint target distribution. While this theory applies in many settings, it falls short in important data fusion problems, such as two-sample instrumental variable analysis, settings that integrate data from epidemiological studies with diverse designs, and studies with mismeasured variables supplemented by external validation studies. In this paper we derive a comprehensive theory that, in particular, covers these settings by allowing the integration of sources aligned with conditional distributions that do not correspond to a single factorization of the target distribution. We provide a universal characterization of the influence functions of regular and asymptotically linear estimators and the efficient influence function of a target parameter, irrespective of the parameter of interest or the statistical model for the target distribution, thus paving the way for a unified theory for machine-learning debiased, semiparametric efficient estimation.
Large tensor learning algorithms are typically computationally expensive and require storing a vast amount of data. In this paper, we propose a unified online Riemannian gradient descent (oRGrad) algorithm for tensor learning, which is computationally efficient, consumes much less memory, and can handle sequentially arriving data while making timely predictions. The algorithm is applicable to both linear and generalized linear models. If the time horizon T is known, oRGrad achieves statistical optimality by choosing an appropriate fixed step size. We find that noisy tensor completion particularly benefits from online algorithms by avoiding the trimming procedure and ensuring sharp entry-wise statistical error, which is often technically challenging for offline methods. The regret of oRGrad is analyzed, revealing a fascinating trilemma concerning the computational convergence rate, statistical error, and regret bound. By selecting an appropriate constant step size, oRGrad achieves an O(T1/2) regret. We then introduce the adaptive-oRGrad algorithm, which can achieve the optimal O(logT) regret by adaptively selecting step sizes, regardless of whether the time horizon is known. The adaptive-oRGrad algorithm can attain a statistically optimal error rate without knowing the horizon. Comprehensive numerical simulations corroborate our theoretical findings. We show that oRGrad significantly outperforms its offline counterpart in predicting the solar F10.7 index with tensor predictors that monitor space weather impacts.
We study the classical problem of predicting an outcome variable, Y, using a linear combination of a d-dimensional covariate vector, X. We are interested in linear predictors whose coefficients solve inf(BeRd) (Ep(n)[Y-X-T beta (R)])(1/r) + delta rho(beta). where delta > 0 is a regularization parameter, p: R-d -> R is a convex penalty function, P is the empirical distribution of the data and r >= 1 Our main contribution is a new bound on the out-of-distribution prediction error of such estimators. The new bound is obtained by combining three new sets of results. First, we provide conditions under which linear predictors based on these estimators solve a distributionally robust optimization problem: they minimize the worst-case prediction error over distributions that are close to each other in a type of max-sliced Wasserstein metric. Second, we provide a detailed finite-sample and asymptotic analysis of the statistical properties of the balls of distributions over which the worst-case prediction error is analyzed. Third, we present an oracle recommendation for the choice of regularization param-eter, delta, that guarantees good out-of-distribution prediction error.
In a triangular array framework where n observations are randomly sampled from a p-dimensional elliptical distribution with shape matrix V-n, we consider the problem of testing the null hypothesis H-0 : theta = theta(0) against the alternative hypothesis H-1 : theta not equal theta(0), where theta is the (fixed) leading unit eigenvector of V-n and theta(0) is a given unit p-vector. The dependence of the shape matrix on the sample size allows us to consider challenging asymptotic scenarios in which the parameter of interest theta is unidentified in the limit, because the ratio between both leading eigenvalues of V-n converges to one. We carefully study the corresponding limiting experiments under such weak identifiability, and we show that these may be LAN or non-LAN. While earlier work in the framework was strictly limited to Gaussian distributions, where the study of local log-likelihood ratios could simply rely on explicit expressions, our asymptotic investigation allows for essentially arbitrary elliptical distributions. This requires original results on quadratic mean differentiable families for triangular arrays of observations, which are likely to be of interest in other models, too. Even in non-LAN experiments, our results enable us to investigate, through Le Cam's first and third lemmas, the asymptotic null and nonnull properties of multivariate rank tests. These nonparametric tests are shown to exhibit an excellent behavior under weak identifiability: not only do they maintain the target nominal size irrespective of the amount of weak identifiability, but they also keep their outstanding uniform efficiency properties under such nonstandard scenarios. In particular, Gaussian-score rank tests, under arbitrarily weak identifiability, still uniformly dominate their parametric pseudo-Gaussian competitor in terms of asymptotic relative efficiencies. Our theoretical results, which are the first ones to study rank tests in the triangular array framework allowing for weak identifiability, are supported by several Monte Carlo exercises.
This paper studies Johansen's (J. Econom. Dynam. Control 12 (1988) 231-254) trace test for cointegration in high-dimensional data. We show that when both cross-sectional and temporal dimension of the data go to infinity proportionally, the shifted and scaled modified trace statistic converges to a Gaussian random variable. We give explicit formulae for the shift and scale parameters as well as for the mean and variance of the Gaussian limit. Monte Carlo analysis shows excellent size properties of the asymptotic test, which is an improvement over the Bartlett-corrected versions of the original trace test, especially for relatively large ratios of the dimensionality to the sample size. The Monte Carlo also reveals a nonmonotonicity of the power of the test. We comment on the source of such a nonmonotonicity.
Stochastic approximation (SA) is a powerful and scalable computational method for iteratively estimating the solution of optimization problems in the presence of randomness, particularly well suited for large-scale and streaming data settings. In this work we propose a theoretical framework for stochastic approximation (SA) applied to nonparametric least squares in reproducing kernel Hilbert spaces (RKHS), enabling online statistical inference in non-parametric regression models. We achieve this by constructing asymptotically valid pointwise (and simultaneous) confidence intervals (bands) for local (and global) inference of the nonlinear regression function, via employing an online multiplier bootstrap approach to a functional stochastic gradient descent (SGD) algorithm in the RKHS. Our main theoretical contributions consist of a unified framework for characterizing the nonasymptotic behavior of the functional SGD estimator and demonstrating the consistency of the multiplier bootstrap method. The proof techniques involve the development of a higher-order expansion of the functional SGD estimator under the supremum norm metric and the Gaussian approximation of suprema of weighted and non-identically distributed empirical processes. Our theory specifically reveals an interesting relationship between the tuning of step sizes in SGD for estimation and the accuracy of uncertainty quantification.
In this article, we present a novel communication-efficient estimator for distributed high-dimensional quantile regression with folded-concave penalties. An Iterative Multi-Step Algorithm (IMSA) is employed to tackle the nonconvex challenge of the objective function, taking into account both the statistical accuracy and the communication constraints. We demonstrate that the proposed IMSA estimators share similar properties with the global folded-concave penalized estimator. To establish the theoretical results, we introduce a new concept called the distributed-oracle estimator. We prove that the proposed estimator converges to the distributed-oracle estimator with high probability. Compared to the & ell;1-penalized method, the IMSA estimator possesses a faster rate of convergence and requires milder conditions to achieve support recovery. Furthermore, we extend our framework to facilitate distributed inference for the preconceived low-dimensional components within the high-dimensional model. We derive the limiting distribution of the corresponding test statistic under the null hypothesis and the local alternatives. In addition, a new feature-splitting algorithm is devised to accommodate the high-dimensional data within the distributed system. Extensive numerical studies demonstrate the effectiveness and validity of our proposed estimation and inference methods. A real example is also presented for illustration.
Single-parameter summaries of variable effects in regression settings are desirable for ease of interpretation. However, (partially) linear models, for example, which would deliver these, may fit poorly to the data. On the other hand, an interpretable summary of the contribution of a given predictor is provided by the so-called average partial effect-the average slope of the regression function with respect to the predictor of interest. Although one can construct a doubly robust procedure for estimating this quantity, it entails estimating the derivative of the conditional mean and also the conditional score of the predictor of interest given all others, tasks which can be very challenging in moderate dimensions: in particular, popular decision tree based regression methods cannot be used. In this work, we introduce an approach for estimating the average partial effect whose accuracy depends primarily on the estimation of certain regression functions, which may be performed by user-chosen machine learning methods that produce potentially nondifferentiable estimates. Our procedure involves resmoothing a given first-stage regression estimator to produce a differentiable version, and modelling the conditional distribution of the predictor of interest through a location-scale model. We show that with the latter assumption, surprisingly the overall error in estimating the conditional score is controlled by a sum of errors of estimating the conditional mean and conditional standard deviation, and the estimation error in a much more tractable univariate score estimation problem. Our theory makes use of a new result on the sub-Gaussianity of Lipschitz score functions that may be of independent interest. We demonstrate the attractive numerical performance of our approach in a variety of settings including ones with misspecification. An R package drape implementing the methodology is available on CRAN.
Le Cam's method (or the two-point method) is a commonly used tool for obtaining statistical lower bound and especially popular for functional estimation problems. This work aims to explain and give conditions for the tightness of Le Cam's lower bound in functional estimation from the perspective of convex duality. Under a variety of settings, it is shown that the maximization problem that searches for the best two-point lower bound, upon dualizing, becomes a minimization problem minimizing an upper bound on the quadratic risk over a family of estimators. Since by the minimax theorem two problems have the same value, this value also characterizes (up to a universal factor) the optimal estimation rate. For estimating linear functionals of a distribution, our work strengthens prior results of Donoho-Liu (Ann. Statist. 19 (1991) 633-667) (for quadratic loss) by dropping the H & ouml;lderian assumption on the modulus of continuity. For exponential families, our results extend those of Juditsky-Nemirovski (Ann. Statist. 37 (2009) 2278-2300) by characterizing the minimax risk for the quadratic loss under weaker assumptions on the exponential family. We also provide an extension to the high-dimensional setting for estimating separable functionals. An application of our methodology to the area of "estimating the unseens" is provided in the companion paper (Polyanskiy and Wu (2023)), resolving the optimal rates (within logarithmic factors) of the distinct elements problem and Fisher's species problem.
Stick-breaking has a long history and is one of the most popular procedures for constructing random discrete distributions in Statistics and Machine Learning. In particular, due to their intuitive construction and computational tractability they are ubiquitous in modern Bayesian nonparametric inference. Most widely used models, such as the Dirichlet and the Pitman-Yor processes, rely on iid or independent length variables. Here we pursue a completely unexplored research direction by considering Markov length variables and investigate the corresponding general class of stick-breaking processes, which we term Markov stick-breaking processes. We establish conditions under which the associated species sampling process is proper and the distribution of a Markov stick-breaking process has full topological support, two fundamental desiderata for Bayesian nonparametric models. We also analyze the stochastic ordering of the weights and provide a new characterization of the Pitman-Yor process as the only stick-breaking process invariant under size-biased permutations, under mild conditions. Moreover, we identify two notable subclasses of Markov stick-breaking processes that enjoy appealing properties and include Dirichlet, Pitman-Yor and Geometric priors as special cases. Our findings include distributional results enabling posterior inference algorithms and methodological insights.
This paper develops a novel two-step estimating procedure for heavy-tailed AR models with nonzero median GARCH-type noises, allowing for time-varying volatility. We first establish the self-weighted quantile regression estimator (SQE) across all quantile levels tau is an element of (0, 1) for the AR parameters theta 0. We show that the SQE, less a bias, converges weakly to a Gaussian process at a rate of n-1/2. The bias is zero if and only if tau equals tau 0, the probability that the noise is less than zero. Based on the SQE, we propose an approach to estimate tau 0 in the second step and feed the estimated tau 0 back into the SQE to estimate theta 0. Both the estimated tau 0 and theta 0 are shown to be consistent and asymptotically normal. A random weighting bootstrap method is developed to approximate the complex distribution. The problem we study is nonstandard because tau 0 may not be identifiable in conventional quantile regression, and the usual methods cannot verify the existence of the SQE bias. Unlike existing procedures for heavy-tailed time series, our method does not require prior information about the symmetry, tail index, or the parametric form of the noise, nor does it require classical identification conditions, such as zero-mean or zero-median.
Recent advances in quasi-Monte Carlo integration have shown that for linearly scrambled digital net estimators, the convergence rate can be dramatically improved by taking the median rather than the mean of multiple independent replicates. In this work, we demonstrate that the quantiles of such estimators can be used to construct confidence intervals with asymptotically valid coverage for high-dimensional integrals. By analyzing the error distribution for a class of infinitely differentiable integrands, we prove that as the sample size increases, the integration error decomposes into an asymptotically symmetric component and a vanishing remainder. Consequently, the asymptotic error distribution is symmetric about zero, ensuring that a quantile-based interval constructed from independent replicates captures the true integral with probability converging to a nominal level determined by the binomial distribution.
We consider supervised learning (regression/classification) problems with tensor-valued input. We derive multi-linear sufficient reductions for the regression or classification problem by modeling the conditional distribution of the predictors given the response as a member of the quadratic exponential family. We develop estimation procedures of sufficient reductions for both continuous and binary tensor-valued predictors. We prove the consistency and asymptotic normality of the estimated sufficient reduction using manifold theory. For continuous predictors, the estimation algorithm is highly computationally efficient and is also applicable to situations where the dimension of the reduction exceeds the sample size. We demonstrate the superior performance of our approach in simulations and real-world data examples for both continuous and binary tensor-valued predictors.
In this paper, we propose a novel Bayesian approach for nonparametric estimation in Wicksell's problem. This has important applications in astronomy for estimating the distribution of the positions of the stars in a galaxy given projected stellar positions and in materials science to determine the 3D microstructure of a material, using its 2D cross-sections. We deviate from the classical Bayesian nonparametric approach, which would place a Dirichlet Process (DP) prior on the distribution function of the unobservables, by directly placing a DP prior on the distribution function of the observables. Our method offers computational simplicity due to the conjugacy of the posterior and allows for asymptotically efficient estimation by projecting the posterior onto the L-2 subspace of increasing, right-continuous functions. Indeed, the resulting Isotonized Inverse Posterior (IIP) satisfies a Bernstein-von Mises (BvM) phenomenon with minimax asymptotic variance g0(x)/2 gamma, where gamma > 1/2 reflects the degree of H & ouml;lder continuity of the true cdf at x. Since the IIP gives automatic uncertainty quantification, it eliminates the need to estimate gamma. Our results provide the first semiparametric Bernstein-von Mises theorem for projection-based posteriors with a DP prior in inverse problems.
This paper studies the construction of adaptive confidence intervals under Huber's contamination model when the contamination proportion is unknown. For the robust confidence interval of a Gaussian mean, we show that the optimal length of an adaptive interval must be exponentially wider than that of a non-adaptive one. An optimal construction is achieved through simultaneous uncertainty quantification of quantiles at all levels. The results are further extended beyond the Gaussian location model by addressing a general family of robust hypothesis testing. In contrast to adaptive robust estimation, our findings reveal that the optimal length of an adaptive robust confidence interval critically depends on the distribution's shape.
Le Cam's two-point testing method yields perhaps the simplest lower bound for estimating the mean of a distribution: roughly, if it is impossible to well distinguish a distribution centered at & micro; from the same distribution centered at & micro; + A, then it is impossible to estimate the mean by better than Delta/2. It is setting-dependent, whether or not a nearly matching upper bound is attainable. We study the conditions under which the two-point testing lower bound can be attained for univariate mean estimation; both in the setting of location estimation (where the distribution is known up to translation) and adaptive location estimation (unknown distribution). Roughly, we will say an estimate nearly attains the two-point testing lower bound if it incurs error that is at most polylogarithmically larger than the Hellinger modulus of continuity for (Omega) over tilde (n) samples. Adaptive location estimation is particularly interesting, as some distributions admit much better guarantees than sub-Gaussian rates (e.g., Unif(& micro; - 1, & micro; + 1) permit error Theta(1/n), while the sub-Gaussian rate is Theta(1/root n )), yet it is not obvious whether these rates may be adaptively attained by one unified approach. Our main result designs an algorithm that nearly attains the two-point testing rate for mixtures of symmetric, log-concave distributions with a common mean. Moreover, this algorithm runs in near-linear time and is parameter-free. In contrast, we show the two-point testing rate is not nearly attainable, even for symmetric, unimodal distributions. We complement this with results for location estimation, showing the two-point testing rate is nearly attainable for unimodal distributions but unattainable for symmetric distributions.
This paper aims to provide a versatile privacy-preserving release mechanism along with a unified approach for subsequent parameter estimation and statistical inference. We propose the ZIL privacy mechanism based on zero-inflated symmetric multivariate Laplace noise, which requires no prior specification of subsequent analysis tasks, allows for general loss functions under minimal conditions, imposes no limit on the number of analyses, and is adaptable to the increasing data volume in online scenarios. We derive the trade-off function for the proposed ZIL mechanism that characterizes its privacy protection level. Within the M-estimation framework, we propose a novel doubly random corrected loss (DRCL) for the ZIL mechanism, which provides consistent and asymptotic normal M-estimates for the parameters of the target population under differential privacy constraints. The proposed approach is easy to compute without numerical integration and differentiation for noisy data. It is applicable for a general class of loss functions, including non-smooth loss functions like check loss and hinge loss. Simulation studies, including logistic regression and quantile regression, are conducted to evaluate the performance of the proposed method.