In this paper, we propose debiased machine learning strategies for estimating direct effect, indirect effect and total effect in logistic partially linear mediation models with high-dimensional confounders. To obtain asymptotically efficient estimators for the effects of interest, two Neyman-orthogonal score functions are proposed to remove regularization bias caused by the estimation of the nuisance functions. To address nonlinearity and unextractability of the logit link, double data splitting is applied to estimate nuisance functions and mitigate potential overfitting. Theoretically, we establish rigorous asymptotic properties for the proposed estimators of all three effects and derive their asymptotic normal distributions. The satisfactory performance of our proposed estimators is demonstrated by simulation results and a real-world PM2.5 concentration data from Beijing.
Robust principal component analysis (RPCA) is a widely used technique for recovering low-rank structure from matrices with missing entries and sparse, possibly large-magnitude corruptions. Although numerous algorithms achieve accurate point estimation, they offer little guidance on the uncertainty of recovered entries, limiting their reliability in practice. In this paper, we propose conformal prediction-RPCA (CP-RPCA), a practical and distribution-free framework for uncertainty quantification in robust matrix recovery. Our proposed method supports both split and full conformal implementations and incorporates weighted calibration to handle heterogeneous observation probabilities. We provide theoretical guarantees for finite-sample coverage and demonstrate through extensive simulations that CP-RPCA delivers reliable uncertainty quantification under severe outliers, missing data and model misspecification. Empirical results show that CP-RPCA can produce informative intervals and remain competitive in efficiency when the RPCA model is well specified, making it a scalable and robust tool for uncertainty-aware matrix analysis.
Short-term load forecasting for AI data centers presents new challenges because it is computing-driven, with heterogeneous job arrivals, sizes, and durations exhibiting bursty, non-stationary dynamics. Compared with traditional load types, data center loads are less researched and can pose greater threats to the efficiency and stability of power grids. To close the gap, this paper proposes a regime-adaptive ensemble learning forecasting algorithm to predict computing-driven dynamic workloads in AI data centers. A weight-learned neural network within an ensemble learning framework is developed to exploit the complementary strengths of two machine learning (ML) submodels across varying operating regimes. Furthermore, a novel feature engineering strategy is developed to incrementally learn from a non-stationary data stream. Thus, the ensemble weights are dynamically optimized to facilitate adaptive calibration of inter-submodel contributions. Comparative case studies on the MIT Supercloud dataset demonstrate that the proposed method significantly enhances load forecasting accuracy and adaptivity across various regimes, and the selected combination of ML models for ensemble learning outperforms other possible combinations. To the best of our knowledge, our method is the first to reduce minute-class forecasting errors for AI data center loads to below 1
In this paper, we propose three effect-specific optimal subsampling strategies for estimating direct, total and indirect effects in partially linear mediation models with high-dimensional confounders. We construct two inverse-probability weighted subsampling Neyman-orthogonal score functions for the direct and total effects, respectively, to simultaneously eliminate selection bias and regularization bias. The unconditional asymptotic distributions of three subsample effect estimators are established and then we derive effect-specific optimal subsampling probabilities by minimizing the traces of their asymptotic variances. A two-step procedure is proposed for practical implementation by using flexible machine learning and data splitting techniques to estimate the high-dimensional nuisance functions and address potential overfitting. Extensive simulations demonstrate the superior performance of the proposed subsample estimators and an application to a real-world air pollution dataset further confirms the practical utility.
Massive datasets are not only frequently affected by measurement errors, but also subject to model misspecification. This paper investigates optimal Poisson subsampling strategy for misspecified measurement error models with massive data. For a pre-specified target parameter, we advocate specifying two linear regression models for both the response and the covariate of interest and then a subsampling double-robust score function is formulated. We show that the resulting subsample estimator is consistent and asymptotically normal when either one of the postulated models is correct. The optimal Poisson subsampling probabilities are derived based on a unified criterion that encompasses both A- and L-optimality criteria, and a practical two-step algorithm is proposed to facilitate their application. Furthermore, in scenarios involving massive datasets distributed across multiple sites, we refine our proposed subsampling score function to be heterogeneity-aware for the common parameter within each site while controlling for other site-specific nuisance functions. The effectiveness of the proposed Poisson subsample estimators is demonstrated through both simulation studies and a real-world application.
In this paper, we focus on decentralized federated learning with missing data. Decentralized learning based on fully observed data has drawn attention in the modern statistical learning. In practice, however, missing data are often encountered such that the existing methods cannot be applied. To address the missing data issue, a novel decentralized algorithm is proposed by using inverse probability weighting in conjunction with network gradient descent method to correct bias and improve estimation together. In order to reduce the computation cost and make the algorithm converge fast, updating the nuisance propensity parameter and target parameter simultaneously at each iteration is proposed. Theoretically, the convergence guarantees of our proposed algorithm are provided and the error bound of the proposed estimator is also studied. The finite-sample performance is investigated through simulations under both independent and identically distributed (i.i.d.) and non-i.i.d. cases. An application to Communities and Crime Data is also presented.
To estimate and perform statistical inference on the total, direct and indirect effects of exposure and mediator variables with measurement error while accounting for high-dimensional confounders, we formulate a high-dimensional mediation model with measurement error. To simultaneously mitigate bias from measurement error and adjust for high-dimensional nuisance parameters related to confounders, two doubly debiased score functions are constructed for the total and direct effects based on the error-corrected loss functions and decorrelated score functions. We establish the asymptotic expressions and distributions of the three effect estimators and develop consistent estimators of their asymptotic covariance matrices. The satisfactory performance of our estimators is demonstrated through simulation studies, and an application to a real-world air pollution dataset further confirms their practical utility.
With the rapid growth in modern science and technology, distributed longitudinal data have drawn attention in a wide range of aspects. Realizing that not all effects of covariates are our parameters of interest, we focus on the distributed estimation and statistical inference of a pre-conceived low-dimensional parameter in the high-dimensional longitudinal GLMs with canonical links. To mitigate the impact of high-dimensional nuisance parameters and incorporate the within-subject correlation simultaneously, a decorrelated quadratic inference function is proposed for enhancing the estimation efficiency. Two communication-efficient surrogate decorrelated score estimators based on multi-round iterative algorithms are proposed. The error bounds and limiting distribution of the proposed estimators are established and extensive numerical experiments demonstrate the effectiveness of our method. An application to the National Longitudinal Survey of Youth Dataset is also presented.
To estimate and statistically infer the direct and indirect effects of exposure and mediator variables while accounting for high-dimensional confounding variables, we propose a partially linear mediation model to incorporate a flexible mechanism of confounders. To obtain asymptotically efficient estimators for the effects of interest under the influence of the nuisance functions with high-dimensional confounders, we construct two Neyman-orthogonal score functions to remove regularization bias. Flexible machine learning methods and data splitting with cross-fitting are employed to address the overfitting issue and estimate unknown nuisance functions efficiently. We rigorously investigate the asymptotic expressions of the proposed estimators for the direct, indirect and total effects and then derive their asymptotic normality properties. In addition, two Wald statistics are constructed to test the direct and indirect effects, respectively, and their limiting distributions are obtained. The satisfactory performance of our proposed estimators is demonstrated by simulation results and a genome-wide analysis of blood DNA methylation dataset.
In this paper, we study the subsampling technique for hypothesis testing in generalized linear models with large-scale datasets, focusing on testing simple null hypotheses against composite linear alternatives. We propose a subsample-based test statistic and show that it converges to non-central chi-square distributions under Pitman’s local alternatives. The optimal subsampling distribution that maximizes power requires iterative calculations on the full data, which is computationally infeasible. Furthermore, it depends on the true parameter, which cannot be consistently estimated under Pitman’s local alternatives. We maximize a lower bound of the non-central parameter to define the power enhancing probability and utilize side information under the alternative to replace the true parameter. Extensive simulations and an application to a real dataset on flight delays and cancellations show that the proposed method offers a computationally viable solution for hypothesis testing in the realm of big data.
Semi-supervised learning has become increasingly prevalent in recent years. In the semi-supervised setting, most of the data are unlabeled since the acquisition of high-quality labels requires expensive scientific measurements and/or laborious human labeling. Hence, it is common to employ some black-box machine learning methods as predictive models to generate outcomes on unlabeled data for subsequent statistical inference. In this paper, we consider sparse regression in semi-supervised setting with prediction assistance. Although the predictions may be imperfect and/or noisy, the empirical risk based on a rectified loss function, is unbiased with respect to the population risk, thereby yielding a relatively safe imputation strategy. Under some regularity conditions, the near-optimal statistical rate is established. It is interesting to derive the prediction error bound on the unlabeled data, rather than the labeled data. More importantly, the asymptotic normality and confidence intervals are also studied via debiasing strategy. With mild conditions on predictive models, it is shown that our proposed method, integrating unlabeled data for combined analysis, is asymptotically more efficient than supervised sparse regression. Both numerical evidence and real-world data analysis demonstrate the effectiveness of our method.
This paper investigates optimal subsampling strategies for the preconceived low-dimensional parameters of main interest in the presence of the nuisance parameters for Cox regression with massive survival data. A general subsampling decorrelated score function based on the log-partial likelihood is constructed to reduce the influence of the less accurate nuisance parameter estimation with a possibly slow convergence rate. The consistency and asymptotic normality of the resultant subsample estimators are established. We derive unified optimal subsampling probabilities based on A-and L-optimality criteria. A two-step algorithm is further proposed to implement practically, and the asymptotic properties of the resultant estimators are also given. The satisfactory performance of our proposed subsample estimators is demonstrated by simulation results and an airline dataset.
We consider kernel-based supervised learning using random Fourier features, focusing on its statistical error bounds and generalization properties with general loss functions. Beyond the least squares loss, existing results only demonstrate worst-case analysis with rate n−1/2 and the number of features at least comparable to n, and refined-case analysis where it can achieve almost n−1 rate when the kernel’s eigenvalue decay is exponential and the number of features is again at least comparable to n. For the least squares loss, the results are much richer and the optimal rates can be achieved under the source and capacity assumptions, with the number of features smaller than n. In this paper, for both losses with Lipschitz derivative and Lipschitz losses, we successfully establish faster rates with number of features much smaller than n, which are the same as the rates and number of features for the least squares loss. More specifically, in the attainable case (the true function is in the RKHS), we obtain the rate n−2ξ2ξ+γ which is the same as the standard method without using approximation, using o(n) features, where ξ characterizes the smoothness of the true function and γ characterizes the decay rate of the eigenvalues of the integral operator. Thus our results answer an important open question regarding random features.
Transfer learning aims to improve the performance of a target model by leveraging data from related source populations. In this paper, we consider the data are inherently distributed and focus on distributed transfer learning in the presence of heavy-tailed and/or asymmetric errors. When the transferable sources are known, based on adaptive Huber regression (AHR) to achieve the bias-robustness tradeoff, a two-step communication-efficient transfer learning for AHR (CTrans-AHR) estimator is proposed to deal with the distributed data. To correct the biases caused by the Lasso penalty, a debiased Lasso CTrans-AHR estimator is constructed and its asymptotic distribution is also studied. When the transferable sources are unknown, a data-driven transferable source detection algorithm is proposed with the theoretical guarantee. Simulation results show that our detection algorithm can consistently select all transferable sources with probability approaching one. Compared to the transfer learning estimator using one site alone or the distributed estimator based solely on target data, our CTrans-AHR estimator achieves the lowest ℓ _∞ and ℓ _2 estimation errors under various error distributions. The accurate coverage probabilities demonstrate the effectiveness of our proposed debiased Lasso CTrans-AHR estimator. An application to the DVL1 gene expression of brain tissues in the Genotype-Tissue Expression dataset is also presented.
In this paper, we focus on the estimation and statistical inference on a preconceived low-dimensional parameter of main interest for the high-dimensional quantile regression model with unmeasured/unobservable confounders and then construct a communication-efficient estimator for the distributed data. To handle the impact of the unmeasured confounders and high-dimensional nuisance covariates simultaneously, a two-step deconfounded and decorrelated (DD) estimator is proposed. Specifically, we impute the unmeasured confounders based on factor model analysis techniques, propose regularized projection decorrelated scores and apply the kernel smoothing techniques in quantile regression. To deal with the distributed data, we further propose the communication-efficient deconfounded and decorrelated (CDD) estimator. The theoretical properties of the proposed DD and CDD estimators are established. The finite sample performance is studied through simulations and an application to the Communities and Crime Dataset is also presented.
In practice, distributed data are often collected from different locations/times/populations /environments with non-negligible heterogeneity. In this paper, we consider inherently distributed data follow heterogeneous partially linear models with high-dimensional covariates, where each site involves a common parameter vector and site-specific nuisance functions. Three distributed estimators based on debiased machine learning are proposed to account for heterogeneity. The closed forms and asymptotic normal distributions of the proposed estimators have been explored and compared. Moreover, these estimators can be computed easily by transmitting some statistics from the local sites to the central site. The finite-sample performance is demonstrated through simulation studies and an application to Beijing multi-site air quality dataset is also provided.
In this paper, we explore optimal subsampling strategies for estimating the parametric regression coefficients in partially linear models with unknown nuisance functions involving high-dimensional and potentially endogenous covariates. To address model misspecifications and the curse of dimensionality, we leverage flexible machine learning (ML) techniques to estimate the unknown nuisance functions. By constructing an unbiased subsampling Neyman-orthogonal score function, we eliminate regularization bias. A two-step algorithm is then used to obtain appropriate ML estimators of the nuisance functions, mitigating the risk of over-fitting. Using martingale techniques, we establish the unconditional consistency and asymptotic normality of the subsample estimators. Furthermore, we derive optimal subsampling probabilities, including A-optimal and L-optimal probabilities as special cases. The proposed optimal subsampling approach is extended to partially linear instrumental variable models to account for potential endogeneity through instrumental variables. Simulation studies and an empirical analysis of the Physicochemical Properties of Protein Tertiary Structure dataset demonstrate the superior performance of our subsample estimators.
Distributed estimation and statistical inference for linear models have drawn much attention recently, but few studies focus on robust learning in the presence of heavy-tailed/asymmetric errors and high-dimensional covariates. Based on adaptive Huber regression to achieve the bias-robustness tradeoff, two classes of sparse and debiased lasso estimators are proposed using aggregated and communication-efficient approaches. To be specific, an aggregated L1-penalized and a multi-round L1-penalized communication-efficient adaptive Huber estimators are respectively proposed in the first stage to handle the distributed data with high-dimensional covariates and heavy-tailed/asymmetric errors. To correct the biases caused by the lasso penalty, a unified debiasing framework based on the decorrelated score equations is considered in the second stage. In the third stage, hard-thresholding is used to produce the sparse and debiased lasso estimators. The convergence rates and asymptotic properties of the proposed two estimators are established. The finite-sample performance is studied through simulations and a real data application to Communities and Crime Data Set is also presented to illustrate the validity and feasibility of the proposed estimators.