PDF file - 106K, Spearman's correlation coefficients calculated for each time-window, separately in cases and controls; Association Between CXCL13 and CXCR5 tagSNPs and CXCL13 Serum Levels; Association between CXCL13 serum levels and HIV-associated non-Hodgkin lymphoma risk at three time-windows, stratified by SNP genotypes.
The celebrated Nadaraya-Watson kernel estimator is among the most studied method for nonparametric regression. A classical result is that its rate of convergence depends on the number of covariates and deteriorates quickly as the dimension grows, which underscores the "curse of dimensionality" and has limited its use in high dimensional settings. In this article, we show that when the true regression function is single or multi-index, the effects of the curse of dimensionality may be mitigated for the Nadaraya-Watson kernel estimator. Specifically, we prove that with K-fold cross-validation, the Nadaraya-Watson kernel estimator indexed by a positive semidefinite bandwidth matrix has an oracle property that its rate of convergence depends on the number of indices of the regression function rather than the number of covariates. Intuitively, this oracle property is a consequence of allowing the bandwidths to diverge to infinity as opposed to restricting them all to converge to zero at certain rates as done in previous theoretical studies. Our result provides a theoretical perspective for the use of kernel estimation in high dimensional nonparametric regression and other applications such as metric learning when a low rank structure is anticipated. Numerical illustrations are given through simulations and real data examples.
In this paper we introduce fuzzy forests, a novel machine learning algorithm for ranking the importance of features in high-dimensional classification and regression problems. Fuzzy forests is specifically designed to provide relatively unbiased rankings of variable importance in the presence of highly correlated features, especially when the number of features, p, is much larger than the sample size, n (p >> n). We introduce our implementation of fuzzy forests in the R package, fuzzyforest. Fuzzy forests works by taking advantage of the network structure between features. First, the features are partitioned into separate modules such that the correlation within modules is high and the correlation between modules is low. The package fuzzyforest allows for easy use of the package WGCNA (weighted gene coexpression network analysis, alternatively known as weighted correlation network analysis) to form modules of features such that the modules are roughly uncorrelated. Then recursive feature elimination random forests (RFE-RFs) are used on each module, separately. From the surviving features, a final group is selected and ranked using one last round of RFE-RFs. This procedure results in a ranked variable importance list whose size is pre-specified by the user. The selected features can then be used to construct a predictive model.
With the advent of high-throughput technologies such as multicolor flow cytometry and next-generation sequencing, high-dimensional data has become increasingly common in biomedical research. In many applications such as proteomics, genomics, and immunology, the data has become increasingly wide. That is, we know a great deal about a small number of subjects. In these applications, the number of features greatly exceeds the number of observations. This “large p, small n” problem gives rise to a number of well-known statistical issues.
OBJECTIVE:Prior hypothesis-driven studies identified immunophenotypic characteristics associated with the control of HIV replication without antiretroviral therapy (HIV controllers) as well as with the degree of CD4 T-cell recovery during ART. We hypothesized that an unbiased 'discovery-based' approach might identify novel immunologic characteristics of these phenotypes. DESIGN:We performed immunophenotyping on four 'aviremic' patient groups: HIV controllers (n = 98), antiretroviral-treated immunologic nonresponders (CD4 < 350; n = 59), antiretroviral-treated immunologic responders (CD4 > 350, n = 142), and as a control group HIV-negative adults (n = 43). We measured levels of T-cell maturation, activation, dysfunction, senescence, functionality, and proliferation. METHODS:Supervised learning assessed the relative importance of immune parameters in predicting clinical phenotypes (controller, immunologic responder, or immunologic nonresponder). Unsupervised learning clustered immune parameters and examined if these clusters corresponded to clinical phenotypes. RESULTS:HIV controllers were characterized by high percentages of HIV-specific T-cell responses and decreased percentages of cells expressing human leukocytic antigen-antigen D related in naive, central memory, and effector T-cell subsets. Immunologic nonresponders were characterized by higher percentages of CD4 T cells that were TNFα+ or INFγ+, higher percentages of activated naive and central memory T cells, and higher percentages of cells expressing programmed cell death protein 1. Unsupervised learning found two distinct clusters of controllers and two distinct clusters of immunologic nonresponders, perhaps suggesting different mechanisms for the clinical outcomes. CONCLUSION:Our discovery-based approach confirmed previously reported characteristics that distinguish aviremic individuals, but also identified novel immunologic phenotypes and distinct clinical subpopulations that should lead to more focused pathogenesis studies that might identify targets for novel therapeutic interventions.
Abstract Background: CXCL13 and CXCR5 are a chemokine and receptor pair whose interaction is critical for naïve B-cell trafficking and activation within germinal centers. We sought to determine whether CXCL13 levels are elevated before HIV-associated non-Hodgkin B-cell lymphoma (AIDS-NHL), and whether polymorphisms in CXCL13 or CXCR5 are associated with AIDS-NHL risk and CXCL13 levels in a large cohort of HIV-infected men. Methods: CXCL13 levels were measured in sera from 179 AIDS-NHL cases and 179 controls at three time-points. TagSNPs in CXCL13 (n = 16) and CXCR5 (n = 11) were genotyped in 183 AIDS-NHL cases and 533 controls. OR and 95% confidence intervals (CI) for the associations between one unit increase in log CXCL13 levels and AIDS-NHL, as well as tagSNP genotypes and AIDS-NHL, were computed using logistic regression. Mixed linear regression was used to estimate mean ratios (MR) for the association between tagSNPs and CXCL13 levels. Results: CXCL13 levels were elevated for more than 3 years (OR = 3.24; 95% CI = 1.90–5.54), 1 to 3 years (OR = 3.39; 95% CI = 1.94–5.94), and 0 to 1 year (OR = 3.94; 95% CI = 1.98–7.81) before an AIDS-NHL diagnosis. The minor allele of CXCL13 rs355689 was associated with reduced AIDS-NHL risk (ORTCvsTT = 0.65; 95% CI = 0.45–0.96) and reduced CXCL13 levels (MRCCvsTT = 0.82; 95% CI = 0.68–0.99). The minor allele of CXCR5 rs630923 was associated with increased CXCL13 levels (MRAAvsTT = 2.40; 95% CI = 1.43–4.50). Conclusions: CXCL13 levels were elevated preceding an AIDS-NHL diagnosis, genetic variation in CXCL13 may contribute to AIDS-NHL risk, and CXCL13 levels may be associated with genetic variation in CXCL13 and CXCR5. Impact: CXCL13 may serve as a biomarker for early AIDS-NHL detection. Cancer Epidemiol Biomarkers Prev; 22(2); 295–307. ©2012 AACR.
Background: Breast cancer aggregates within some families, accounting for about 25% of incident cases. Known genetic variants, such as BRCA1/BRCA2 mutations, account for less than 10% of these familial breast cancer cases, thus, research into additional risk or protective genetic factors for familial breast cancer may provide valuable information for risk stratification. Serum levels of prolactin have been associated with sporadic breast cancer risk in previous studies. We sought to test the hypothesis that variation in the genes coding for prolactin (PRL) and the prolactin receptor (PRLR) would be associated with increased or decreased risk of familial breast cancer. Methods: We designed a case-control study of probands recruited between 1998 and 2009 by the UCLA Family Cancer Registry; a population enriched with women having a strong family history of breast cancer and BRCA1/BRCA2 mutation carriers. Genotyping of tagSNPs in PRL (n=11) and PRLR (n=28) was performed in 351 unrelated women who developed breast cancer and 290 unaffected controls with similar family histories to the cases. Multivariate logistic regression models controlling for age, Ashkenazi Jewish heritage, education, and number of affected relatives were used to calculate odds ratios (ORs) and 95% confidence intervals (CIs) for the association between SNPs and familial breast cancer. Genotype specific ORs as well as per-allele (additive model) ORs were calculated. Results: Familial breast cancer risk was inversely associated with carriership of the minor alleles for three PRLR SNPs in the additive models, rs9292573 (OR=0.7, 95% CI=0.5-0.9), rs10805603 (OR=0.6, 95% CI=0.4-0.9), and rs1587607 (OR=0.7, 95% CI=0.5-1.0). PRLR SNP rs249522 was associated with increased risk of cancer in the additive model (OR=1.8, 95% CI =1.2-2.7). Minor allele carriers of the PRL SNP rs12210179, located in a putative transcription factor binding site, had a decreased risk of cancer in the additive model (OR=0.7, 95% CI=0.5-0.9). Conclusion: These data suggest that genetic variation in PRL and PRLR are associated with breast cancer in a population with a strong family history of breast cancer. Citation Format: Shehnaz K. Hussain, Mary Sehl, Daniel Conn, Janet S. Sinsheimer, Uma Dandekar, Jeanette Papp, Zuo-Feng Zhang, Patricia A. Ganz. Variation in the prolactin and prolactin receptor genes and familial breast cancer. [abstract]. In: Proceedings of the Eleventh Annual AACR International Conference on Frontiers in Cancer Prevention Research; 2012 Oct 16-19; Anaheim, CA. Philadelphia (PA): AACR; Cancer Prev Res 2012;5(11 Suppl):Abstract nr A107.