In this paper, we propose a high-dimensional factor-adjusted sparse partially linear regression model that integrates the linear effects of high-dimensional, significantly correlated covariates with the nonparametric effects of low-dimensional covariates. The proposed framework combines the interpretability of linear models, the flexibility of nonparametric modeling and the ability to effectively account for dependence among high-dimensional predictors. We develop a penalized estimation procedure that incorporates B-spline approximation and factor analysis, and establish error bounds for the resulting estimators. To facilitate valid inference for the linear component, we further propose a factor-adjusted projection-debiased procedure and employ a Gaussian multiplier bootstrap to obtain critical values. Theoretical guarantees are derived under suitable regularity conditions. Extensive simulation studies demonstrate the favorable finite-sample performance of the proposed method. An application to a birth weight dataset, which serves as the motivating example, further illustrates its effectiveness and practical usefulness.
Models with latent factors have recently attracted considerable attention. However, most existing studies focus on linear regression models and therefore fail to capture potential nonlinear structures. To address this limitation, we consider the factor augmented single-index model. We first examine whether the inclusion of the augmented component is necessary by introducing a score-type test statistic. Unlike existing test statistics, the proposed one does not require estimating high-dimensional regression coefficients or precision matrices, making it computationally simpler and more stable. To determine the critical value, we employ a Gaussian multiplier bootstrap, whose theoretical validity is established under mild regularity conditions. We further investigate the penalized estimation of the regression model. With estimated latent factors, we establish the error bounds of the estimators. In addition, we construct confidence intervals for individual coefficients based on a debiased estimator. Importantly, our method does not impose moment conditions on the error distribution, allowing it to perform well even when the errors are heavy-tailed or contaminated by outliers. Comprehensive simulation studies and an application to a gene expression dataset demonstrate the effectiveness and robustness of the proposed procedure.
Models with latent factors recently attract a lot of attention. However, most investigations focus on linear regression models and thus cannot capture nonlinearity. To address this issue, we propose a novel Factor Augmented Single-Index Model. We first address the concern whether it is necessary to consider the augmented part by introducing a score-type test statistic. Compared with previous test statistics, our proposed test statistic does not need to estimate the high-dimensional regression coefficients, nor high-dimensional precision matrix, making it simpler in implementation. We also propose a Gaussian multiplier bootstrap to determine the critical value. The validity of our procedure is theoretically established under suitable conditions. We further investigate the penalized estimation of the regression model. With estimated latent factors, we establish the error bounds of the estimators. Lastly, we introduce debiased estimator and construct confidence interval for individual coefficient based on the asymptotic normality. No moment condition for the error term is imposed for our proposal. Thus our procedures work well when random error follows heavy-tailed distributions or when outliers are present. We demonstrate the finite sample performance of the proposed method through comprehensive numerical studies and its application to an FRED-MD macroeconomics dataset.
In this paper, we introduce a novel high-dimensional Factor-Adjusted sparse Partially Linear regression Model (FAPLM), to integrate the linear effects of high-dimensional latent factors with the nonparametric effects of low-dimensional covariates. The proposed FAPLM combines the interpretability of linear models, the flexibility of nonparametric models, with the ability to effectively capture the dependencies among highdimensional covariates. We develop a penalized estimation approach for the model by leveraging B-spline approximations and factor analysis techniques. Theoretical results establish error bounds for the estimators, aligning with the minimax rates of standard Lasso problems. To assess the significance of the linear component, we introduce a factor-adjusted projection debiased procedure and employ the Gaussian multiplier bootstrap method to derive critical values. Theoretical guarantees are provided under regularity conditions. Comprehensive numerical experiments validate the finite-sample performance of the proposed method. Its successful application to a birth weight dataset, the motivating example for this study, highlights both its effectiveness and practical relevance.
In the realm of high-throughput genomic data, modeling with ultrahigh-dimensional covariates and censored survival outcomes is of great importance. We conduct conditional inference for the ultrahigh-dimensional additive hazards model, allowing both the covariates of interest and nuisance covariates to be ultrahigh-dimensional. The presence of right censorship with survival outcomes adds an extra layer of complexity to the original data structure, posing significant challenges for the ultrahigh-dimensional additive hazards model. To address this, we introduce an innovative test statistic based on the quadratic norm of the score function. Moreover, when there is a high correlation between the covariates of interest and nuisance covariates, we propose a decorrelated score function-based test statistic to enhance statistical power. Additionally, we establish the limiting distributions of the test statistics under both the null and local alternative hypotheses, further enhancing the computational appeal of our approach. The proposed statistics are thoroughly evaluated through extensive simulation studies and applied to two real data examples.
Along with the widespread adoption of high-dimensional data, traditional statistical methods face significant challenges in handling problems with high correlation of variables, heavy-tailed distribution, and coexistence of sparse and dense effects. In this paper, we propose a factor-augmented quantile regression (FAQR) framework to address these challenges simultaneously within a unified framework. The proposed FAQR combines the robustness of quantile regression and the ability of factor analysis to effectively capture dependencies among high-dimensional covariates, and also provides a framework to capture dense effects (through common factors) and sparse effects (through idiosyncratic components) of the covariates. To overcome the lack of smoothness of the quantile loss function, convolution smoothing is introduced, which not only improves computational efficiency but also eases theoretical derivation. Theoretical analysis establishes the accuracy of factor selection and consistency in parameter estimation under mild regularity conditions. Furthermore, we develop a Bootstrap-based diagnostic procedure to assess the adequacy of the factor model. Simulation experiments verify the rationality of FAQR in different noise scenarios such as normal and t_2 distributions.
Overall survival has been used as the primary endpoint for many randomized trials that aim to examine whether a new treatment is non-inferior to the standard treatment or placebo control. When a new treatment is indeed non-inferior in terms of survival, it may be important to assess other outcomes including health utility. However, analyzing health utility scores in a secondary analysis may have limited power since the primary objectives of the original study design may not include health utility. To comprehensively consider both survival and health utility, we developed a composite endpoint, HUS (Health Utility-adjusted Survival), which combines both survival and utility. HUS has been shown to be able to increase statistical power and potentially reduce the required sample size compared to the standard overall survival endpoint. Nevertheless, the asymptotic properties of the test statistics of the HUS endpoint have yet to be fully established. Besides that, the standard version of HUS cannot be applied to or has limited performance in certain scenarios, where extensions are needed. In this manuscript, we propose various methodological extensions of HUS and derive the asymptotic distributions of the test statistics. By comprehensive simulation studies and a data application using retrospective data based on a translational patient cohort in Princess Margaret Cancer Centre, we demonstrate the better efficiency and feasibility of HUS compared to different methods.
Recent advances in multi-omics technology highlight the need for statistical methods that account for complex dependencies among biological layers. In this paper, we propose a novel Multi-Omics Factor-Adjusted Cox (MOFA-Cox) model to handle multi-omics survival data, addressing the intricate correlation structures across different omics layers. Building upon this model, we introduce a decorrelated score test for the Cox model in high-dimensional survival analysis. We establish the theoretical properties of our test statistic, which show that it admits a closed-form asymptotic distribution, eliminating the need for resampling. We further analyze its local power under local alternatives. Importantly, our test statistic does not require a sparsity assumption on the covariates of interest, broadening its applicability. Numerical studies and an applaication to the TCGA breast cancer dataset demonstrate the effectiveness of our method.
For the supervised and semi-supervised settings, a group inference method is proposed for regression parameters in high-dimensional semi-parametric single-index models with an unknown random link function. The inference procedure is based on least squares, which can be extended to other general convex loss functions. The proposed test statistics are weighted quadratic forms of the regression parameter estimates, in which the weight could be a non-random matrix or the sample covariance matrix of the covariates. The proposed method could detect dense but weak signals and deal with high correlation of covariates inside the group. A 'contaminated test statistic' is established in the semi-supervised regime to decrease the variance. The asymptotic properties of the resulting estimators are established. The finite-sample behaviour of the proposed method is evaluated through extensive simulation studies. Applications to two genomic datasets are provided.
We propose a new functional additive hazards model to investigate the potential effects of functional and scalar predictors on mortality risks, and develop a penalized least squares estimation method for model parameters based on a pseudoscore estimating equation. A reproducing kernel Hilbert space approach is used to establish the consistency, convergence rate, and joint asymptotic distribution of the resulting estimators for finite-dimensional and infinite-dimensional parameters. Our simulation studies demonstrate that the proposed estimation procedure performs well. For illustration, we apply the proposed method to the Medical Information Mart for Intensive Care III dataset. Les auteurs de ce travail presentent un nouveau modele de risques additifs fonctionnels. Ce modele vise a examiner les effets potentiels des predicteurs fonctionnels et scalaires sur les risques de mortalite. Ils proposent une methode d'estimation par moindres carres penalises. La methode en question repose sur une equation d'estimation pseudo-score pour les parametres du modele. Une approche d'espaces de Hilbert a noyau reproduisant leur a permis d'etablir la convergence, le taux de convergence et la distribution asymptotique conjointe de parametres, tant de dimension finie qu'infinie. Enfin, les resultats des etudes de simulation confirment l'efficacite de cette procedure d'estimation, et en guise d' illustration, la methode est appliquee a l'ensemble de donnees 'Medical Information Mart for Intensive Care III'.
Inspired by the complexity of certain real-world datasets, this article introduces a novel flexible linear spline index regression model. The model posits piecewise linear effects of an index on the response, with continuous changes occurring at knots. Significantly, it possesses the interpretability of linear models, captures nonlinear effects similar to nonparametric models, and achieves dimension reduction like single-index models. In addition, the locations and number of knots remain unknown, which further enhances the adaptability of the model in practical applications. We propose a new method that combines penalized approaches and convolution techniques to simultaneously estimate the unknown parameters and determine the number of knots. Noteworthy is that the proposed method allows the number of knots to diverge with the sample size. We demonstrate that the proposed estimators can identify the number of knots with a probability approaching one and estimate the coefficients as efficiently as if the number of knots is known in advance. We also introduce a procedure to test the presence of knots. Simulation studies and two real datasets are employed to assess the finite sample performance of the proposed method.
This paper studies the global estimation in semiparametric quantile regression models. For estimating unknown functional parameters, an integrated quantile regression loss function with penalization is proposed. The first step is to obtain a vector-valued functional Bahadur representation of the resulting estimators, and then derive the asymptotic distribution of the proposed infinite-dimensional estimators. Furthermore, a resampling approach that generalizes the minimand perturbing technique is adopted to construct confidence intervals and to conduct hypothesis testing. Extensive simulation studies demonstrate the effectiveness of the proposed method, and applications to the real estate dataset and world happiness report data are provided.
Although quantile regressions are widely employed for heterogeneous data, simultaneously selecting covariates that globally affect the response and estimating the coefficients is very challenging. We introduce a novel sparse composite quantile regression screening method for the analysis of ultrahigh-dimensional heterogeneous data. The proposed method enjoys the sure screening property, provides a consistent selection path, and yields consistent estimates of the coefficients simultaneously across a continuous range of quantile levels. An extended Bayesian information criterion is employed to select the "best" candidate from the path. Extensive simulation studies demonstrate the effectiveness of the proposed method, and an application to a gene expression data set is provided.
Deriving the limiting distribution of a nonparametric estimate is rather challenging but of fundamental importance to statistical inference. For the current status data, we study a penalized nonparametric likelihood-based estimator for an unknown cumulative hazard function, and establish the pointwise asymptotic normality of the resulting nonparametric estimate. We also propose the penalized likelihood ratio tests for local and global hypotheses, derive their limiting distributions, and study the optimality of the global test. Simulation studies show that the proposed method works well compared to the classical likelihood ratio test.
Large-scale matrix linear regression models with high-dimensional responses and highdimensional variables have been widely employed in various large-scale biomedical studies. In this article, we propose an optimal minimax variable selection approach for the matrix linear regression model when the dimensions of both the response matrix and predictors diverge at the exponential rate of the sample size. We develop an iterative hard-thresholding algorithm for fast computation and establish an optimal minimax theory for the parameter estimates. The finite sample performance of the method is examined via extensive simulation studies and a real data application from the Alzheimer’s Disease Neuroimaging Initiative study is provided.
Imbalanced data, in which the data exhibit an unequal or highly-skewed distribution between its classes/categories, are pervasive in many scientific fields, with application range from bioinformatics, text classification, face recognition, fraud detection, etc. Imbalanced data in modern science are often of massive size and high dimensionality, for example, gene expression data for diagnosing rare diseases. To address this issue, a fused screening procedure is proposed for dimension reduction with large-scale high dimensional imbalanced data under repeated case-control samplings. There are several advantages of the proposed method: it is model-free without any model specification for the underlying distribution; it is relatively inexpensive in computational cost by using the subsampling technique; it is robust to outliers in the predictors. The theoretical properties are established under regularity conditions. Numerical studies including extensive simulations and a real data example confirm that the proposed method performs well in practical settings.
Sparse Composite Quantile Regression with Ultra-high Data Abstract. Quantile regression is widely employed in heterogeneous data, but to select covariates that globally affect the response and estimate coefficients simultaneously are very challenging. In this article, we introduce a novel sparse composite quantile regression screening method for the analysis of ultra-high dimensional heterogeneous data. The proposed method enjoys the sure screening property, provides a consistent selection path, and yields consistent estimates of coefficients simultaneously across a continuous range of quantile levels. An extended Bayesian information criterion is employed to select the “best” candidate from the path. Extensive simulation studies demonstrate the effective-ness of the proposed method, and an application to a gene expression dataset is provided.
This article studies penalized semiparametric maximum partial likelihood estimation and hypothesis testing for the functional Cox model in analyzing right-censored data with both functional and scalar predictors. Deriving the asymptotic joint distribution of finite-dimensional and infinite-dimensional estimators is a very challenging theoretical problem due to the complexity of semiparametric models. For the problem, we construct the Sobolev space equipped with a special inner product and discover a new joint Bahadur representation of estimators of the unknown slope function and coefficients. Using this key tool, we establish the asymptotic joint normality of the proposed estimators and the weak convergence of the estimated slope function, and then construct local and global confidence intervals for an unknown slope function. Furthermore, we study a penalized partial likelihood ratio test, show that the test statistic enjoys the Wilks phenomenon, and also verify the optimality of the test. The theoretical results are examined through simulation studies, and a right-censored data example from the Improving Care of Acute Lung Injury Patients study is provided for illustration. Supplementary materials for this article are available online.
The complexity of X-chromosome inactivation arouses the X-linked genetic association being overlooked in most of the genetic studies, especially for genetic association analysis on time to event outcomes. To fill this gap, we propose novel methods to analyze the X-linked genetic association for competing risk failure time data based on a subdistribution hazard function. Specifically, we consider two mechanisms for a single genetic variant on X-chromosome: (1) all the subjects in a population undergo the same inactivation process; (2) the subjects randomly undergo different inactivation processes. According to the assumptions, one of the proposed methods can be used to infer the unknown biological process under scenario (1), while another method can be used to estimate the proportion of a certain biological process in the population under scenario (2). Both of the two methods can infer the direction of skewness for skewed X-chromosome inactivation and derive asymptotically unbiased estimates of the model parameters. The asymptotic distributions for the parameter estimates and constructed score tests with nuisance parameters only presented under the alternative hypothesis are illustrated under both assumptions. Finite sample performance of these novel methods is examined via extensive simulation studies. An application is illustrated with implementation on a cancer genetic study with competing risk outcomes.
This study introduces a penalized nonparametric maximum likelihood estimation of the log-hazard function for analyzing right-censored data. Smoothing splines are employed for a smooth estimation. Our main discovery is a functional Bahadur representation, which serves as a key tool for nonparametric inferences of an unknown function. The asymptotic properties of the resulting smoothing-spline estimator of the unknown log-hazard function are established under regularity conditions. Moreover, we provide a local confidence interval for this function, as well as local and global likelihood ratio tests. We also discuss the asymptotic efficiency of the estimator. The theoretical results are validated using extensive simulation studies. Lastly, we demonstrate the estimator by applying it to a real data set.