In this paper, we study the subsampling technique for hypothesis testing in generalized linear models with large-scale datasets, focusing on testing simple null hypotheses against composite linear alternatives. We propose a subsample-based test statistic and show that it converges to non-central chi-square distributions under Pitman’s local alternatives. The optimal subsampling distribution that maximizes power requires iterative calculations on the full data, which is computationally infeasible. Furthermore, it depends on the true parameter, which cannot be consistently estimated under Pitman’s local alternatives. We maximize a lower bound of the non-central parameter to define the power enhancing probability and utilize side information under the alternative to replace the true parameter. Extensive simulations and an application to a real dataset on flight delays and cancellations show that the proposed method offers a computationally viable solution for hypothesis testing in the realm of big data.
Distributed estimation and statistical inference for linear models have drawn much attention recently, but few studies focus on robust learning in the presence of heavy-tailed/asymmetric errors and high-dimensional covariates. Based on adaptive Huber regression to achieve the bias-robustness tradeoff, two classes of sparse and debiased lasso estimators are proposed using aggregated and communication-efficient approaches. To be specific, an aggregated L1-penalized and a multi-round L1-penalized communication-efficient adaptive Huber estimators are respectively proposed in the first stage to handle the distributed data with high-dimensional covariates and heavy-tailed/asymmetric errors. To correct the biases caused by the lasso penalty, a unified debiasing framework based on the decorrelated score equations is considered in the second stage. In the third stage, hard-thresholding is used to produce the sparse and debiased lasso estimators. The convergence rates and asymptotic properties of the proposed two estimators are established. The finite-sample performance is studied through simulations and a real data application to Communities and Crime Data Set is also presented to illustrate the validity and feasibility of the proposed estimators.
The L-P-quantile regression generalizes both quantile regression and expectile regression, and has become popular for its robustness and effectiveness especially when 1 < p = 2. In this paper, we consider the data that are inherently distributed and propose two distributed L-P-quantile regression estimators for a preconceived low-dimensional parameter in the presence of high-dimensional extraneous covariates. To handle the impact of high-dimensional nuisance parameters, we first investigate regularized projection score for estimating low-dimensional parameter of main interest in L-P-quantile regression. To deal with the distributed data, we further propose two communication-efficient surrogate projection score estimators and establish their theoretical properties. The finite-sample performance of the proposed estimators is studied through simulations and an application to Communities and Crime data set is also presented.
Distributed estimation for parametric models has drawn attention in modern statistical learning, but few studies focus on semiparametric models. In this paper, we propose two communication-efficient distributed estimators for partially linear additive models with high-dimensional covariates. The commonly used B-spline basis functions are first applied to approximate the nonparametric functions and then we construct a profiled communication-efficient surrogate loss function with Lasso penalty based on one local machine solving the final optimization problem. Further, to reduce the effect of local machines and improve the stability of the algorithm, a profiled gradient-enhanced loss estimator is derived. The resulting two estimators and their theoretical convergence rates for both parametric and nonparametric components are established. The finite-sample performance of the proposed estimators is studied through simulations and an application to appliances energy prediction data set is also presented.
In this paper, we consider the unified optimal subsampling estimation and inference on the low-dimensional parameter of main interest in the presence of the nuisance parameter for low/high-dimensional generalized linear models (GLMs) with massive data. We first present a general subsampling decorrelated score function to reduce the influence of the less accurate nuisance parameter estimation with the slow convergence rate. The consistency and asymptotic normality of the resultant subsample estimator from a general decorrelated score subsampling algorithm are established, and two optimal subsampling probabilities are derived under the A- and L-optimality criteria to downsize the data volume and reduce the computational burden. The proposed optimal subsampling probabilities provably improve the asymptotic efficiency upon the subsampling schemes in the low-dimensional GLMs and perform better than the uniform subsampling scheme in the high-dimensional GLMs. A two-step algorithm is further proposed to implement and the asymptotic properties of the corresponding estimators are also given. Simulations show satisfactory performance of the proposed estimators, and two applications to census income and Fashion-MNIST datasets also demonstrate its practical applicability.