Dependence within a high-dimensional profile of explanatory variables affects estimation and prediction performance of regression models. However, the strong belief that dependence should not be ignored, based on our well-proven knowledge of low-dimensional regression modeling, is not necessarily true in high dimension. To investigate this point, we introduce a new class of prediction scores defined as linear combinations of a same random vector, including the naive prediction score obtained when ignoring dependence and the Ordinary Least Squares (OLS) prediction score that, on the contrary, fully accounts for dependence by a preliminary whitening of the explanatory variables. Interestingly, the former class also contains Ridge and Partial Least Squares prediction scores, that both offer intermediate ways of dealing with dependence. Through a theoretical comparative study, it is first shown how the best handling of dependence should depend on the interplay between the structure of conditional dependence across explanatory variables and the pattern of the association signal. We also derive the closed form expression of the prediction score with best prediction performance within the proposed class, leading to an adaptive handling of dependence. Finally, it is demonstrated through simulation studies and using benchmark datasets that this prediction score outperforms existing methods in various settings. Supplementary materials for this article are available online.
Genetic interaction is considered as one of the main heritable component of complex traits. With the emergence of genome-wide association studies (GWAS), a collection of statistical methods dedicated to the identification of interaction at the SNP level have been proposed. More recently, gene-based gene-gene interaction testing has emerged as an attractive alternative as they confer advantage in both statistical power and biological interpretation. Most of the gene-based interaction methods rely on a multidimensional modeling of the interaction, thus facing a lack of robustness against the huge space of interaction patterns. In this paper, we study a global testing approaches to address the issue of gene-based gene-gene interaction. Based on a logistic regression modeling framework, all SNP-SNP interaction tests are combined to produce a gene-level test for interaction. We propose an omnibus test that takes advantage of (1) the heterogeneity between existing global tests and (2) the complementarity between allele-based and genotype-based coding of SNPs. Through an extensive simulation study, it is demonstrated that the proposed omnibus test has the ability to detect with high power the most common interaction genetic models with one causal pair as well as more complex genetic models where more than one causal pair is involved. On the other hand, the flexibility of the proposed approach is shown to be robust and improves power compared to single global tests in replication studies. Furthermore, the application of our procedure to real datasets confirms the adaptability of our approach to replicate various gene-gene interactions.
In global testing, where a large number of pointwise test statistics are aggregated to simultaneously test for a collection of null hypotheses, the handling of dependence is a crucial issue. In various fields, more particularly in genetic epidemiology and functional data analysis, many testing methods for detecting an association signal between a response and explanatory variables have been proposed. Some aggregation procedures ignore dependence across pointwise test statistics whereas others introduce a model for decorrelation, with unclear conclusions on their relative performance. Indeed, the benefit that can be expected from decorrelation highly depends on the interplay between the structure of dependence across pointwise test statistics and the pattern of the association signal. Within a large class of test statistics covering a continuum of decorrelation approaches, an optimal procedure is introduced. This procedure is based on the maximization of an ad-hoc cumulant generating function-based distance between the null and nonnull distributions of a global test statistic, in order to adapt the aggregation of the pointwise statistics to the pattern of the association signal. A comparative study including simulations and applications to genetic association studies demonstrates that the ability of this test to detect a signal is more robust to the dependence structure than existing methods.