Male-pattern baldness (MPB) is a common and highly heritable trait characterized by androgen-dependent, progressive hair loss from the scalp. Here, we carry out the largest GWAS meta-analysis of MPB to date, comprising 10,846 early-onset cases and 11,672 controls from eight independent cohorts. We identify 63 MPB-associated loci (P<5 × 10-8, METAL) of which 23 have not been reported previously. The 63 loci explain ∼39% of the phenotypic variance in MPB and highlight several plausible candidate genes (FGF5, IRF4, DKK2) and pathways (melatonin signalling, adipogenesis) that are likely to be implicated in the key-pathophysiological features of MPB and may represent promising targets for the development of novel therapeutic options. The data provide molecular evidence that rather than being an isolated trait, MPB shares a substantial biological basis with numerous other human phenotypes and may deserve evaluation as an early prognostic marker, for example, for prostate cancer, sudden cardiac arrest and neurodegenerative disorders.
BACKGROUND:Lymphedema (LE) is a chronic clinical manifestation of filarial nematode infections characterized by lymphatic dysfunction and subsequent accumulation of protein-rich fluid in the interstitial space-lymphatic filariasis. A number of studies have identified single nucleotide polymorphisms (SNPs) associated with primary and secondary LE. To assess SNPs associated with LE caused by lymphatic filariasis, a cross-sectional study of unrelated Ghanaian volunteers was designed to genotype SNPs in 285 LE patients as cases and 682 infected patients without pathology as controls. One hundred thirty-one SNPs in 64 genes were genotyped. The genes were selected based on their roles in inflammatory processes, angiogenesis/lymphangiogenesis, and cell differentiation during tumorigenesis.RESULTS:Genetic associations with nominal significance were identified for five SNPs in three genes: vascular endothelial growth factor receptor-3 (VEGFR-3) rs75614493, two SNPs in matrix metalloprotease-2 (MMP-2) rs1030868 and rs2241145, and two SNPs in carcinoembryonic antigen-related cell adhesion molecule-1 (CEACAM-1) rs8110904 and rs8111171. Pathway analysis revealed an interplay of genes in the angiogenic/lymphangiogenic pathways. Plasma levels of both MMP-2 and CEACAM-1 were significantly higher in LE cases compared to controls. Functional characterization of the associated SNPs identified genotype GG of CEACAM-1 as the variant influencing the expression of plasma concentration, a novel finding observed in this study.CONCLUSION:The SNP associations found in the MMP-2, CEACAM-1, and VEGFR-3 genes indicate that angiogenic/lymphangiogenic pathways are important in LE clinical development.
Variant p.R47H of triggering receptor expressed on myeloid cells 2 (TREM2) has been associated with Parkinson's disease (PD). We screened this TREM2-variant in 821 PD patients including 261 demented PD patients (PDD) and in healthy controls (n = 919). Neither the entire PD nor the small PDD sample was associated with p.R47H.
The genetic basis of Alzheimer's disease (AD) is complex and heterogeneous. Over 200 highly penetrant pathogenic variants in the genes APP , PSEN1 , and PSEN2 cause a subset of early-onset familial AD. On the other hand, susceptibility to late-onset forms of AD (LOAD) is indisputably associated to the ɛ4 allele in the gene APOE , and more recently to variants in more than two-dozen additional genes identified in the large-scale genome-wide association studies (GWAS) and meta-analyses reports. Taken together however, although the heritability in AD is estimated to be as high as 80%, a large proportion of the underlying genetic factors still remain to be elucidated. In this study, we performed a systematic family-based genome-wide association and meta-analysis on close to 15 million imputed variants from three large collections of AD families (~3500 subjects from 1070 families). Using a multivariate phenotype combining affection status and onset age, meta-analysis of the association results revealed three single nucleotide polymorphisms (SNPs) that achieved genome-wide significance for association with AD risk: rs7609954 in the gene PTPRG (P -value=3.98 × 10 −8 ), rs1347297 in the gene OSBPL6 ( P -value=4.53 × 10 −8 ), and rs1513625 near PDCL3 ( P -value=4.28 × 10 −8 ). In addition, rs72953347 in OSBPL6 ( P -value=6.36 × 10 −7 ) and two SNPs in the gene CDKAL1 showed marginally significant association with LOAD (rs10456232, P -value=4.76 × 10 −7 ; rs62400067, P -value=3.54 × 10 −7 ). In summary, family-based GWAS meta-analysis of imputed SNPs revealed novel genomic variants in (or near) PTPRG, OSBPL6 , and PDCL3 that influence risk for AD with genome-wide significance.
BACKGROUND:A usually confronted problem in association studies is the occurrence of population stratification. In this work, we propose a novel framework to consider population matchings in the contexts of genome-wide and sequencing association studies. We employ pairwise and groupwise optimal case-control matchings and present an agglomerative hierarchical clustering, both based on a genetic similarity score matrix. In order to ensure that the resulting matches obtained from the matching algorithm capture correctly the population structure, we propose and discuss two stratum validation methods. We also invent a decisive extension to the Cochran-Armitage Trend test to explicitly take into account the particular population structure.RESULTS:We assess our framework by simulations of genotype data under the null hypothesis, to affirm that it correctly controls for the type-1 error rate. By a power study we evaluate that structured association testing using our framework displays reasonable power. We compare our result with those obtained from a logistic regression model with principal component covariates. Using the principal components approaches we also find a possible false-positive association to Alzheimer's disease, which is neither supported by our new methods, nor by the results of a most recent large meta analysis or by a mixed model approach.CONCLUSIONS:Matching methods provide an alternative handling of confounding due to population stratification for statistical tests for which covariates are hard to model. As a benchmark, we show that our matching framework performs equally well to state of the art models on common variants.
BACKGROUND:In family-based association analysis, each family is typically ascertained from a single proband, which renders the effects of ascertainment bias heterogeneous among family members. This is contrary to case-control studies, and may introduce sample or ascertainment bias. Statistical efficiency is affected by ascertainment bias, and careful adjustment can lead to substantial improvements in statistical power. However, genetic association analysis has often been conducted using family-based designs, without addressing the fact that each proband in a family has had a great influence on the probability for each family member to be affected.METHOD:We propose a powerful and efficient statistic for genetic association analysis that considered the heterogeneity of ascertainment bias among family members, under the assumption that both prevalence and heritability of disease are available. With extensive simulation studies, we showed that the proposed method performed better than the existing methods, particularly for diseases with large heritability.RESULTS:We applied the proposed method to the genome-wide association analysis of Alzheimer's disease. Four significant associations with the proposed method were found.CONCLUSION:Our significant findings illustrated the practical importance of this new analysis method.
The global demand for products that effectively prevent the development of male-pattern baldness (MPB) has drastically increased. However, there is currently no established genetic model for the estimation of MPB risk. We conducted a prediction analysis using single-nucleotide polymorphisms (SNPs) identified from previous GWASs of MPB in a total of 2725 German and Dutch males. A logistic regression model considering the genotypes of 25 SNPs from 12 genomic loci demonstrates that early-onset MPB risk is predictable at an accuracy level of 0.74 when 14 SNPs were included in the model, and measured using the area under the receiver-operating characteristic curves (AUC). Considering age as an additional predictor, the model can predict normal MPB status in middle-aged and elderly individuals at a slightly lower accuracy (AUC 0.69–0.71) when 6–11 SNPs were used. A variance partitioning analysis suggests that 55.8% of early-onset MPB genetic liability can be explained by common autosomal SNPs and 23.3% by X-chromosome SNPs. For normal MPB status in elderly individuals, the proportion of explainable variance is lower (42.4% for autosomal and 9.8% for X-chromosome SNPs). The gap between GWAS findings and the variance partitioning results could be explained by a large body of common DNA variants with small effects that will likely be identified in GWAS of increased sample sizes. Although the accuracy obtained here has not reached a clinically desired level, our model was highly informative for up to 19% of Europeans, thus may assist decision making on early MPB intervention actions and in forensic investigations.
Genetic interaction is suspected to play an important role in genetically complex diseases such as AD. However, due to computational challenges, lack of a consensus statistical analysis method, and lack of power, no compelling evidence for any kind of epistasis has yet been given for AD. We conducted an exhaustive search for interacting SNP pairs in 6,415 AD patients and 19,492 controls from the IGAP consortium. The analysis was conducted in a two-stage meta-analysis fashion. In stage I, participating groups performed a Genome-wide Interaction Analysis (GWIA) of 1010 SNP pairs with a fast pre-test. The groups interchanged their top 2,000,000 results to form a joined list of candidate pairs which all groups re-analyzed in stage II. The full 8 degrees of freedom model was applied, using sex, age and leading PCAs as covariates. Meta-Analysis of the results was done using the "Sigma-method", an extension of fixed effects meta-analysis to multiple regression models. The linkage disequilibrium region surrounding APOE was excluded from the analysis region, since interaction with APOE is investigated in an own dedicated project. A pair of SNP from the known PICALM gene and DOCK1 reaches experiment-wide significance (p < 6.5 x 10 -12). In addition, a SNP from TM4SF4, again together with PICALM, comes close to significance at the current stage of analysis. The analysis is ongoing and results from further participating groups are expected to arrive.
We present a genome-wide association study of a quantitative trait, "progression of systolic blood pressure in time," in which 142 unrelated individuals of the Genetic Analysis Workshop 18 real genotype data were analyzed. Information on systolic blood pressure and other phenotypic covariates was missing at certain time points for a considerable part of the sample. We observed that the dropout process causing missingness is not independent of the initial systolic blood pressure; that is, the data is not missing completely at random. However, after the adjustment for age, the impact of systolic blood pressure on dropouts was no longer significant. Therefore, we decided to impute missing phenotype values by using information from individuals with complete phenotypic data. Progression of systolic blood pressure (∆SBP/∆t) was defined based on the imputed phenotypes and analyzed in a genome-wide fashion. We also conducted an exhaustive genome-wide search for interaction between single-nucleotide polymorphisms (7.14 × 1010 tests) under an allelic model. The suggested data imputation and the association analysis strategy proved to be valid in the sense that there was no evidence of genome-wide inflation or increased type I error in general. Furthermore, we detected 2 single-nucleotide polymorphisms (SNPs) that met the criterion for genome-wide significance (p≤5 × 10−8), which was also confirmed via Monte-Carlo simulation. In view of the rather small sample size, however, the results have to be followed-up in larger studies.
Important methodological advancements in rare variant association testing have been made recently, among them collapsing tests, kernel methods and the variable threshold (VT) technique. Typically, rare variants from a region of interest are tested for association as a group ('bin'). Rare variant studies are already routinely performed as whole-exome sequencing studies. As an alternative approach, we propose a pipeline for rare variant analysis of imputed data and develop respective quality control criteria. We provide suggestions for the choice and construction of analysis bins in whole-genome application and support the analysis with implementations of standard burden tests (COLL, CMAT) in our INTERSNP-RARE software. In addition, three rare variant regression tests (REG, FRACREG and COLLREG) are implemented. All tests are accompanied with the VT approach which optimizes the definition of 'rareness'. We integrate kernel tests as implemented in SKAT/SKAT-O into the suggested strategies. Then, we apply our analysis scheme to a genome-wide association study of Alzheimer's disease. Further, we show that our pipeline leads to valid significance testing procedures with controlled type I error rates. Strong association signals surrounding the known APOE locus demonstrate statistical power. In addition, we highlight several suggestive rare variant association findings for follow-up studies, including genomic regions overlapping MCPH1, MED18 and NOTCH3. In summary, we describe and support a straightforward and cost-efficient rare variant analysis pipeline for imputed data and demonstrate its feasibility and validity. The strategy can complement rare variant studies with next generation sequencing data.
Cerebrospinal fluid amyloid-beta 1-42 (A beta(1-42)) and phosphorylated Tau at position 181 (pTau(181)) are biomarkers of Alzheimer's disease (AD). We performed an analysis and meta-analysis of genome-wide association study data on A beta(1-42) and pTau(181) in AD dementia patients followed by independent replication. An association was found between A beta(1-42) level and a single-nucleotide polymorphism in SUCLG2 (rs62256378) (P = 2.5 x 10(-12)). An interaction between APOE genotype and rs62256378 was detected (P = 9.5 x 10(-5)), with the strongest effect being observed in APOE-epsilon 4 noncarriers. Clinically, rs62256378 was associated with rate of cognitive decline in AD dementia patients (P = 3.1 x 10(-3)). Functional microglia experiments showed that SUCLG2 was involved in clearance of A beta(1-42).
The pathogenesis of androgenetic alopecia (AGA, male-pattern baldness) is driven by androgens, and genetic predisposition is the major prerequisite. Candidate gene and genome-wide association studies have reported that single-nucleotide polymorphisms (SNPs) at eight different genomic loci are associated with AGA development. However, a significant fraction of the overall heritable risk still awaits identification. Furthermore, the understanding of the pathophysiology of AGA is incomplete, and each newly associated locus may provide novel insights into contributing biological pathways. The aim of this study was to identify unknown AGA risk loci by replicating SNPs at the 12 genomic loci that showed suggestive association (5 x 10(-8)<P<10(-5)) with AGA in a recent meta-analysis. We analyzed a replication set comprising 2,759 cases and 2,661 controls of European descent to confirm the association with AGA at these loci. Combined analysis of the replication and the meta-analysis data identified four genome-wide significant risk loci for AGA on chromosomes 2q35, 3q25.1, 5q33.3, and 12p12.1. The strongest association signal was obtained for rs7349332 (P = 3.55 x 10(-15)) on chr2q35, which is located intronically in WNT10A. Expression studies in human hair follicle tissue suggest that WNT10A has a functional role in AGA etiology. Thus, our study provides genetic evidence supporting an involvement of WNT signaling in AGA development.
Cerebrospinal fluid (CSF) markers Aβ42 and phosphorylated Tau (pTau181) have been proposed as potential biomarkers in Alzheimer's disease (AD). Interestingly, genome-wide association studies (GWAS) using these CSF biomarkers in the Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort have identified novel candidate risk loci for AD. However, these studies were based on one single sample without replication. Here, we present results of a GWAS using CSF biomarkers (Aβ42, pTau181 and pTau181 / Aβ42) in an independent sample of AD patients of German origin. ELISA test was used to determine Aβ42 and pTau181 in 274 AD patients recruited at 12 German University hospital memory clinics within the German Dementia Competence Network. All patients gave written informed consent for participation in the entire study. The Illumina 610-quad BeadChip was used to genotype 113 AD patients. The rest AD patients were genotyped using the Illumina Omni 1M-quad BeadChip. Quality control and deviation from Hardy-Weinberg equilibrium was checked for the entire sample. We followed a meta-analysis strategy because two different maker panels were used. We imputed both samples with the IMPUTE software package using 1000 Genomes reference data. Association within each of the two samples was performed using a linear regression for imputed genotype data as provided in PLINK. Finally, we combined the results of both studies SNP by SNP using fixed and random effects meta-analysis. The computation was carried out with the METAL. In the locus of APOE-TOMM40, we replicated the association for two SNPs previously reported in the ADNI cohort for Aβ42 (rs2075650, P= 1.04e-5; rs429358, P= 1.4e-7). We found additional association signals in other AD susceptibility loci and CSF biomarkers. In addition, we identified 74 SNPs with p-values smaller than 10–5 including two SNPs with genome-wide level of significance. We expected from the meta-analysis 44 positive SNPs by chance. These results provide additional confirmation that genetic variants within the APOE-TOMM40 locus modulate CSF levels of Aβ42. Identification of more SNPs with lower p-values as expected may represent some true association signals. Replication of these results is currently underway. These results will also be included in our presentation for AAIC 2012.