Amyotrophic lateral sclerosis (ALS) is a heritable disorder where rare variants with low-to-moderate penetrance are thought to dominate genetic risk. To identify such rare variants, we harmonized and analyzed exome data from 22 cohorts, totaling 17,919 individuals with ALS and 200,703 controls across discovery and replication phases. Rare variant analyses identified several new risk genes, with replication confirming association of YKT6 and supporting HTR3C, GBGT1 and KNTC1. We also provide strong, independent validation for genes with limited previous evidence: ARPP21, DNAJC7 and CFAP410. Notably, in ARPP21, we identified a new high-effect variant (p.P747L) and confirmed that p.P563L is an ALS-associated variant leading to an aggressive disease course. Beyond new discoveries, our analyses largely recapitulated the known genetic architecture of ALS, identifying risk variants in over 20% of cases and supporting a cumulative oligogenic risk model. These findings highlight new translational targets and show that rare variant analyses capture substantially more genetic risk than common variant genome-wide association studies.
Many non-coding variants influence complex traits and diseases through gene regulation, yet the mechanisms linking these variants to downstream biology remain poorly understood. Here, we present eQTLGen Phase 2, a comprehensive genome-wide analysis of gene expression quantitative trait loci (eQTLs) in 43,301 blood samples from 52 datasets. Beyond local ciseffects, this sample size enabled the first systematic mapping of trans-eQTLs at scale. We identify cis-eQTLs for nearly all expressed genes (94.7%) and trans-eQTLs for over half (56.2%). Second, by colocalizing cis-eQTLs with trans-eQTLs, we infer a directed gene regulatory network comprising 47,554 directed gene regulatory relationships. These networks reveal how genetic perturbations in upstream regulators produce dose-dependent downstream effects, supported by Perturb-seq and ChIP-seq data. Third, integrating this network with 87 genome-wide association studies allows us to systematically prioritize trait-relevant pathways and candidate genes. Variants exerting both cis- and trans-effects are markedly more likely to colocalize with trait associations than cis-only variants, delineating a subset of functionally active cis-eQTLs from a large group with limited downstream impact. This distinction provides a conceptual framework for identifying regulatory variants that truly mediate complex trait biology. Together, these results provide a publicly available resource of cis- and trans-eQTLs and an in vivo scaffold for human gene-regulatory networks, elucidating how propagation of cis-effects modulates complex disease.
Identifying causal mechanisms from genome-wide association studies (GWAS) requires an understanding of how disease-associated genetic variants influence gene expression in specific cell types. Here, we present scMetaBrain, a large-scale single-nucleus RNA-sequencing (snRNA-seq) resource derived using 1,260 samples from 785 individuals spanning 10 brain datasets. By analyzing 3.9 million transcriptomes, we identified 19,371 unique expression quantitative trait locus (eQTL) genes (eGenes) at a major cell type level, with the largest number of eQTLs observed in excitatory neurons. Notably, 31% of the eQTLs detected were highly cell-type-specific, with most restricted to excitatory neurons (69%). We compared the eQTLs with bulk RNA-seq datasets across different tissues and with a newly generated single nucleus dataset of 123 donors from peripheral blood mononuclear cells. We observed that differences in eQTL effect sizes between brain cell types are often as large as comparing eQTLs between brain tissue and non-brain tissue from bulk RNA-seq studies. Furthermore, we observe that eQTL effect size agreement was highest for cell types with similar function, even when comparing brain to blood cells. This suggests that that bulk analyses substantially overestimate eQTL agreement, likely due to tissue-level averaging of cellular regulatory effects. Through colocalization, we prioritized 662 genes for 11 brain-related traits and prioritized a single cell type in 68% of genes. Our findings demonstrate that eQTL effects are far more cell-type-specific than previously recognized, underscoring the need to expand single-cell eQTL studies across diverse tissues and cell types to fully capture the regulatory architecture of genetic variants.
Background and Objectives Multiple sclerosis (MS) age at onset (AAO) is a clinical predictor of long-term disease outcomes, independent of disease duration. Little is known about the genetic and biological mechanisms underlying age of first symptoms. We conducted a genome-wide association study (GWAS) to investigate associations between individual genetic variation and the MS AAO phenotype. Methods The study population was comprised participants with MS in 6 clinical trials: ADVANCE (N = 655; relapsing-remitting [RR] MS), ASCEND (N = 555; secondary-progressive [SP] MS), DECIDE (N = 1,017; RRMS), OPERA1 (N = 581; RRMS), OPERA2 (N = 577; RRMS), and ORATORIO (N = 529; primary-progressive [PP] MS). Altogether, 3,905 persons with MS of European ancestry were analyzed. GWAS were conducted for MS AAO in each trial using linear additive models controlling for sex and 10 principal components. Resultant summary statistics across the 6 trials were then meta-analyzed, for a total of 8.3 x 10(-6) single nucleotide polymorphisms (SNPs) across all trials after quality control and filtering for heterogeneity. Gene-based tests of associations, pathway enrichment analyses, and Mendelian randomization analyses for select exposures were also performed. Results Four lead SNPs within 2 loci were identified (p < 5 x 10(-8)), including a) 3 SNPs in the major histocompatibility complex and their effects were independent of HLA-DRB1*15:01 and b) a LOC105375167 variant on chromosome 7. At the gene level, the top association was HLA-C (p = 1.2 x 10(-7)), which plays an important role in antiviral immunity. Functional annotation revealed the enrichment of pathways related to T-cell receptor signaling, autoimmunity, and the complement cascade. Mendelian randomization analyses suggested a link between both earlier age at puberty and shorter telomere length and earlier AAO, while there was no evidence for a role for either body mass index or vitamin D levels. Discussion Two genetic loci associated with MS AAO were identified, and functional annotation demonstrated an enrichment of genes involved in adaptive and complement immunity. There was also evidence supporting a link with age at puberty and telomere length. The findings suggest that AAO in MS is multifactorial, and the factors driving onset of symptoms overlap with those influencing MS risk.
While the genetics of MS risk susceptibility are well-described, and recent progress has been made on the genetics of disease severity, the genetics of disease progression remain elusive. We therefore investigated the genetic determinants of MS progression on longitudinal brain MRI: change in brain volume (BV) and change in T2 lesion volume (T2LV), reflecting progressive tissue loss and increasing disease burden, respectively. We performed genome-wide association studies of change in BV (N = 3401) and change in T2LV (N = 3513) across six randomized clinical trials from Biogen and Roche/Genentech: ADVANCE, ASCEND, DECIDE, OPERA I & II, and ORATORIO. Analyses were adjusted for randomized treatment arm, age, sex, and ancestry. Results were pooled in a meta-analysis, and were evaluated for enrichment of MS risk variants. Variant colocalization and cell-specific expression analyses were performed using published cohorts. The strongest peaks were in PTPRD (rs77321193-C/A, p = 3.9 × 10–7) for BV change, and NEDD4L (rs11398377-GC/G, p = 9.3 × 10–8) for T2LV change. Evidence of colocalization was observed for NEDD4L, and both genes showed increased expression in neuronal and/or glial populations. No association between MS risk variants and MRI outcomes was observed. In this unique, precompetitive industry partnership, we report putative regions of interest in the neurodevelopmental gene PTPRD, and the ubiquitin ligase gene NEDD4L. These findings are distinct from known MS risk genetics, indicating an added role for genetic progression analyses and informing drug discovery.
necrosis factor alpha pathway through the linear ubiquitin chain assembly complex. We also built a new genetic risk score associated with the risk of future AD/dementia or progression from mild cognitive impairment to AD/dementia. The improvement in prediction led to a 1.6-to 1.9-fold increase in AD risk from the lowest to the highest decile, in addition to effects of age and the APOE ε 4 allele
Muscle strength is highly heritable and predictive for multiple adverse health outcomes including mortality. Here, we present a rare protein-coding variant association study in 340,319 individuals for hand grip strength, a proxy measure of muscle strength. We show that the exome-wide burden of rare protein-truncating and damaging missense variants is associated with a reduction in hand grip strength. We identify six significant hand grip strength genes, KDM5B , OBSCN , GIGYF1 , TTN , RB1CC1 , and EIF3J . In the example of the titin ( TTN) locus we demonstrate a convergence of rare with common variant association signals and uncover genetic relationships between reduced hand grip strength and disease. Finally, we identify shared mechanisms between brain and muscle function and uncover additive effects between rare and common genetic variation on muscle strength.
Identification of therapeutic targets from genome-wide association studies (GWAS) requires insights into downstream functional consequences. We harmonized 8,613 RNA-sequencing samples from 14 brain datasets to create the MetaBrain resource and performed cis - and trans -expression quantitative trait locus (eQTL) meta-analyses in multiple brain region- and ancestry-specific datasets ( n ≤ 2,759). Many of the 16,169 cortex cis -eQTLs were tissue-dependent when compared with blood cis -eQTLs. We inferred brain cell types for 3,549 cis -eQTLs by interaction analysis. We prioritized 186 cis -eQTLs for 31 brain-related traits using Mendelian randomization and co-localization including 40 cis -eQTLs with an inferred cell type, such as a neuron-specific cis -eQTL ( CYP24A1 ) for multiple sclerosis. We further describe 737 trans -eQTLs for 526 unique variants and 108 unique genes. We used brain-specific gene-co-regulation networks to link GWAS loci and prioritize additional genes for five central nervous system diseases. This study represents a valuable resource for post-GWAS research on central nervous system diseases.
OBJECTIVE:Obesity is a significant public health concern across the globe. Research investigating epigenetic mechanisms related to obesity and obesity-associated conditions has identified differences that may contribute to cellular dysregulation that accelerates the development of disease. However, few studies include Black women, who experience the highest incidence of obesity and early onset of cardiometabolic disorders. METHODS:The association of BMI with epigenome-wide DNA methylation (DNAm) was examined using the 850K Illumina EPIC BeadChip in two Black populations (Intergenerational Impact of Genetic and Psychological Factors on Blood Pressure [InterGEN], n = 239; and The Genetic Epidemiology Network of Arteriopathy [GENOA] study, n = 961) using linear mixed-effects regression models adjusted for batch effects, cell type heterogeneity, population stratification, and confounding factors. RESULTS:Cross-sectional analysis of the InterGEN discovery cohort identified 28 DNAm sites significantly associated with BMI, 24 of which had not been previously reported. Of these, 17 were replicated using the GENOA study. In addition, a meta-analysis, including both the InterGEN and GENOA cohorts, identified 658 DNAm sites associated with BMI with false discovery rate < 0.05. In a meta-analysis of Black women, we identified 628 DNAm sites significantly associated with BMI. Using a more stringent significance threshold of Bonferroni-corrected p value 0.05, 65 and 61 DNAm sites associated with BMI were identified from the combined sex and female-only meta-analyses, respectively. CONCLUSIONS:This study suggests that BMI is associated with differences in DNAm among women that can be identified with DNA extracted from salivary (discovery) and peripheral blood (replication) samples among Black populations across two cohorts.
Objective: Despite evidence that trauma exposure is linked to higher risk of hypertension, epigenetic mechanisms (such as DNA methylation) by which trauma potentially influences hypertension risk among Black adults remain understudied. Methods: Data from a longitudinal study of Black mothers were used to test the hypothesis that direct childhood trauma (ie, personal exposure) and vicarious trauma (ie, childhood trauma experienced by their children) would interact with DNA methylation to increase blood pressure (BP). Separate linear mixed effects models were fitted at each CpG site with the DNA methylation beta-value and direct and vicarious trauma as predictors and systolic and diastolic BP modeled as dependent variables adjusted for age, cigarette smoking, and body mass index. Interaction terms between DNA methylation beta-values with direct and vicarious trauma were added. Results: The sample included 244 Black mothers with a mean age of 31.2 years (SD = ±5.8). Approximately 45% of participants reported at least one form of direct childhood trauma and 49% reported at least one form of vicarious trauma. Epigenome-wide interaction analyses found that no CpG sites passed the epigenome-wide significance level indicating the interaction between direct or vicarious trauma with DNAm did not influence systolic or diastolic BP. Conclusions: This is one of the first studies to simultaneously examine whether direct or vicarious exposure to trauma interact with DNAm to influence BP. Although findings were null, this study highlights directions for future research that investigates epigenetic mechanisms that may link trauma exposure with hypertension risk in Black women.
Introduction: Experiencing psychosocial stress is associated with poor health outcomes such as hypertension and obesity, which are risk factors for developing cardiovascular disease. African American women experience disproportionate risk for cardiovascular disease including exposure to high levels of psychosocial stress. We hypothesized that psychosocial stress, such as perceived stress overload, may influence epigenetic marks, specifically DNA methylation (DNAm), that contribute to increased risk for cardiovascular disease in African American women. Methods: We conducted an epigenome-wide study evaluating the relationship of psychosocial stress and DNAm among African American mothers from the Intergenerational Impact of Genetic and Psychological Factors on Blood Pressure (InterGEN) cohort. Linear mixed effects models were used to explore the epigenome-wide associations with the Stress Overload Scale (SOS), which examines self-reported past-week stress, event load and personal vulnerability. Results: In total, n = 228 participants were included in our analysis. After adjusting for known epigenetic confounders, we did not identify any DNAm sites associated with maternal report of stress measured by SOS after controlling for multiple comparisons. Several of the top differentially methylated CpG sites related to SOS score ( P < 1 × 10−5), mapped to genes of unknown significance for hypertension or heart disease, namely, PXDNL and C22orf42. Conclusions: This study provides foundational knowledge for future studies examining epigenetic associations with stress and other psychosocial measures in African Americans, a key area for growth in epigenetics. Future studies including larger sample sizes and replication data are warranted.
Coronary artery disease (CAD) is a preeminent cause of death, and smoking is a strong risk factor for CAD. Genetic factors contribute to the development of CAD, but the interplay between genetic predisposition and smoking history in CAD remains unclear. Using data from the UK Biobank, we constructed several genetic risk scores (GRSs) based on known CAD loci and assessed their interactions with smoking for the development of incident CAD in 307,147 participants of European ancestry who were free of CAD. We fitted Cox proportional hazard models and assessed gene-smoking interaction on both multiplicative and additive scales. Overall, we found no multiplicative interactions, but observed a synergistic additive interaction of GRS with both smoking status and pack-years of smoking, finding that the absolute CAD risk due to smoking was higher for those with high genetic risk. Trait-based sub-GRSs suggested smoking status and smoking intensity measured by pack-years might confer gene-smoking interaction effects with different intermediate risk factors for CAD. Our study results suggest that genetics could modify the effects of smoking on CAD and highlight the value of addressing gene-lifestyle interactions on both additive and multiplicative scales.
Elevated body mass index (BMI) is heritable and associated with many health conditions that impact morbidity and mortality. The study of the genetic association of BMI across a broad range of common disease conditions offers the opportunity to extend current knowledge regarding the breadth and depth of adiposity-related diseases. We identify 906 (364 novel) and 41 (6 novel) genome-wide significant loci for BMI among participants of European (N~1.1 million) and African (N~100,000) ancestry, respectively. Using a BMI genetic risk score including 2446 variants, 316 diagnoses are associated in the Million Veteran Program, with 96.5% showing increased risk. A co-morbidity network analysis reveals seven disease communities containing multiple interconnected diseases associated with BMI as well as extensive connections across communities. Mendelian randomization analysis confirms numerous phenotypes across a breadth of organ systems, including conditions of the circulatory (heart failure, ischemic heart disease, atrial fibrillation), genitourinary (chronic renal failure), respiratory (respiratory failure, asthma), musculoskeletal and dermatologic systems that are deeply interconnected within and across the disease communities. This work shows that the complex genetic architecture of BMI associates with a broad range of major health conditions, supporting the need for comprehensive approaches to prevent and treat obesity.
Genetic colocalisation is an important tool to test for shared genetic aetiology and is commonly used to strengthen causal inference in genetic studies of molecular traits and drug targets. However, the single causal variant assumption of the original colocalization method is a considerable limitation in genomic regions with multiple causal effects. We integrated conditional analyses (GCTA-COJO) and colocalisation analyses (coloc), into a novel analysis tool called Pair-Wise Conditional Colocalization (PWCoCo). PWCoCo performs conditional analyses to identify independent signals for the two tested traits in a genomic region and then conducts colocalisation of each pair of conditionally independent signals for the two traits using summary-level data. This allows for the stringent single-variant assumption to hold for each pair of colocalisation analysis. We found that the computational efficiency of PWCoCo is on average better than colocalisation with Sum of Single Effects Regression using Summary Stats (SuSiE-RSS), with greater gains in efficiency for high-throughput analysis. In a case study using GWAS data for multiple sclerosis and brain cortex-derived eQTLs (MetaBrain), we recapitulated all previously identified genes, which showcased the robustness of the method. We further found colocalisation evidence for secondary signals in nine additional loci, which was not identifiable in conventional GWAS and/or colocalisation. PWCoCo offers key improvements over existing methods, including: (1) robust colocalisation when the single variant assumption is violated; (2) independent colocalisation of secondary signals, which enables identification of novel disease-causing variants; (3) an easy-to-use and computationally efficient tool to test for colocalisation of high-dimensional omics data.
Potentially traumatic experiences have been associated with chronic diseases. Epigenetic mechanisms, including DNA methylation (DNAm), have been proposed as an explanation for this association. We examined the association of experiences of trauma with epigenome-wide DNAm among African American mothers (n = 236) and their children aged 3–5 years (n = 232; N = 500), using the Life Events Checklist-5 (LEC) and Traumatic Events Screening Inventory—Parent Report Revised (TESI-PRR). We identified no DNAm sites significantly associated with potentially traumatic experience scores in mothers. One CpG site on the ENOX1 gene was methylome-wide-significant in children (FDR-corrected q-value = 0.05) from the TESI-PRR. This protein-coding gene is associated with mental illness, including unipolar depression, bipolar, and schizophrenia. Future research should further examine the associations between childhood trauma, DNAm, and health outcomes among this understudied and high-risk group. Findings from such longitudinal research may inform clinical and translational approaches to prevent adverse health outcomes associated with epigenetic changes.
Complex traits are characterized by multiple genes and variants acting simultaneously on a phenotype. However, studying the contribution of individual pairs of genes to complex traits has been challenging since human genetics necessitates very large population sizes, while findings from model systems do not always translate to humans. Here, we combine genetics with combinatorial RNAi (coRNAi) to systematically test for pairwise additive effects (AEs) and genetic interactions (GIs) between 30 lipid genome-wide association studies (GWAS) genes. Gene-based burden tests from 240,970 exomes show that in carriers with truncating mutations in both, APOB and either PCSK9 or LPL ("human double knock-outs") plasma lipid levels change additively. Genetics and coRNAi identify overlapping AEs for 12 additional gene pairs. Overlapping GIs are observed for TOMM40/APOE with SORT1 and NCAN. Our study identifies distinct gene pairs that modulate plasma and cellular lipid levels primarily via AEs and nominates putative drug target pairs for improved lipid-lowering combination therapies.
ABSTRACTBackgroundSeveral monogenic causes for isolated dystonia have been identified, but they collectively account for only a small proportion of cases. Two genome‐wide association studies have reported a few potential dystonia risk loci; but conclusions have been limited by small sample sizes, partial coverage of genetic variants, or poor reproducibility.ObjectiveTo identify robust genetic variants and loci in a large multicenter cervical dystonia cohort using a genome‐wide approach.MethodsWe performed a genome‐wide association study using cervical dystonia samples from the Dystonia Coalition. Logistic and linear regressions, including age, sex, and population structure as covariates, were employed to assess variant‐ and gene‐based genetic associations with disease status and age at onset. We also performed a replication study for an identified genome‐wide significant signal.ResultsAfter quality control, 919 cervical dystonia patients compared with 1491 controls of European ancestry were included in the analyses. We identified one genome‐wide significant variant (rs2219975, chromosome 3, upstream of COL8A1, P‐value 3.04 × 10−8). The association was not replicated in a newly genotyped sample of 473 cervical dystonia cases and 481 controls. Gene‐based analysis identified DENND1A to be significantly associated with cervical dystonia (P‐value 1.23 × 10−6). One low‐frequency variant was associated with lower age‐at‐onset (16.4 ± 2.9 years, P‐value = 3.07 × 10−8, minor allele frequency = 0.01), located within the GABBR2 gene on chromosome 9 (rs147331823).ConclusionThe genetic underpinnings of cervical dystonia are complex and likely consist of multiple distinct variants of small effect sizes. Larger sample sizes may be needed to provide sufficient statistical power to address the presumably multi‐genic etiology of cervical dystonia. © 2021 International Parkinson and Movement Disorder Society
Severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) causes coronavirus disease-19 (COVID-19), a respiratory illness that can result in hospitalization or death. We investigated associations between rare genetic variants and seven COVID-19 outcomes in 543,213 individuals, including 8,248 with COVID-19. After accounting for multiple testing, we did not identify any clear associations with rare variants either exome-wide or when specifically focusing on (i) 14 interferon pathway genes in which rare deleterious variants have been reported in severe COVID-19 patients; (ii) 167 genes located in COVID-19 GWAS risk loci; or (iii) 32 additional genes of immunologic relevance and/or therapeutic potential. Our analyses indicate there are no significant associations with rare protein-coding variants with detectable effect sizes at our current sample sizes. Analyses will be updated as additional data become available, with results publicly browsable at https://rgc-covid19.regeneron.com.
Severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) causes coronavirus disease 2019 (COVID-19), a respiratory illness that can result in hospitalization or death. We used exome sequence data to investigate associations between rare genetic variants and seven COVID-19 outcomes in 586,157 individuals, including 20,952 with COVID-19. After accounting for multiple testing, we did not identify any clear associations with rare variants either exome wide or when specifically focusing on (1) 13 interferon pathway genes in which rare deleterious variants have been reported in individuals with severe COVID-19, (2) 281 genes located in susceptibility loci identified by the COVID-19 Host Genetics Initiative, or (3) 32 additional genes of immunologic relevance and/or therapeutic potential. Our analyses indicate there are no significant associations with rare protein-coding variants with detectable effect sizes at our current sample sizes. Analyses will be updated as additional data become available, and results are publicly available through the Regeneron Genetics Center COVID-19 Results Browser.