ABSTRACT Huntington’s disease is a rare neurodegenerative disease whose primary risk factors are inherited expansions of a CAG repeat tract in the HTT gene. Somatic expansion of these tracts leads to neuronal toxicity, neuronal death and clinical disease progression. To identify genetic factors with a major impact on disease onset and progression, we genome sequenced 18,825 individuals for the ENROLL-HD study. Our results show rare inactivating mutations in three genes, all involved in DNA damage repair, are major determinants of age of onset for motor symptoms (n=10,610) and other clinical manifestations. Heterozygote carriers of predicted loss-of-function (pLoF) variants in POLD1 and PMS1 developed motor symptoms an average 20 years (n=3; P=1×10 −5 ) and 7 years (n=6; P=2×10 −3 ) later than non-carriers, respectively. Conversely, heterozygote carriers of pLoF variants in FAN1 (n=30) developed symptoms 10 years earlier (P=2×10 −10 ). Our findings highlight therapeutic strategies and help predict age of onset for at-risk individuals.
Most genetic variants associated with complex traits are hypothesized to regulate gene expression. To understand the genetics underlying gene expression variability, we characterized 14,324 RNA-sequencing samples from the Trans-Omics for Precision Medicine program and performed expression and splicing quantitative trait locus (e/sQTL) analyses in six tissues and cell types, including whole blood (n = 6454) and lung (n = 1291). We detected tens of thousands of secondary cis-e/sQTLs, showing that secondary cis-e/sQTL discovery remains unsaturated. We fine-mapped UK Biobank-derived genome-wide association study (GWAS) signals from 164 traits and identified e/sQTL colocalizations for 10,611 GWAS signals, including 7096 that colocalize with secondary e/sQTLs. Our results suggest that even larger e/sQTL analyses will uncover additional secondary e/sQTLs, further benefiting GWAS interpretation.
Rare coding variants that alter protein function and confer beneficial health effects can suggest potential drug targets. CHRNB3 encodes the β3 subunit of nicotinic acetylcholine receptors that bind nicotine and mediate its action in the brain. Here we report an exome-wide association study of number of cigarettes smoked per day (cig per day) in 37,897 current smokers from the Mexico City Prospective Study. We identify a deleterious missense variant in CHRNB3, p.Glu284Gly, that associates with a significant reduction in daily cigarette consumption. The missense variant is enriched in people of Indigenous Mexican ancestry but rare in other ancestries. We further identify a predicted loss-of-function variant in CHRNB3 that significantly associates with reduction in number of smoked cigarettes per day in participants of Japan Biobank. This variant is enriched in people of East Asian ancestry but is rare in other ancestries. Finally, we find that rare deleterious missense and predicted loss-of-function variants in aggregate associate with a reduction in the number of smoked cigarettes per day in individuals of European ancestry from the UK Biobank. Our results suggest that loss of function of CHRNB3 significantly associates with daily cigarette smoking, proposing β3 inhibition as a potential therapeutic strategy for nicotine addiction.
Through an analysis of 2,602 genome-wide association studies (GWAS) across 830 human traits, we find that most (56% of) well-studied traits have at least two published GWAS, and many (29%) have at least five. We show that the lack of an established approach for adjudicating variant association estimates across multiple published studies can lead to uncertainty and invalid inferences: using all associations ever published for a trait increases true positives (by 12%) but also false positives (by 55%) relative to using associations from the largest published GWAS for the trait. We employ a "bottom-line" procedure for meta-analyzing published GWAS while inferring and accounting for sample overlap, which identifies a more accurate and comprehensive list of associations relative to existing approaches. Five commonly used bioinformatic methods for post-GWAS analyses produce reliable results when applied to the bottom-line associations. We present these results for 1,281 human complex traits, including 1,839 single-ancestry and 576 trans-ancestry analyses, for browsing or download via the NHGRI Association to Function Knowledge Portal. This resource of "consensus" GWAS results is intended to increase replicability, reuse, and interpretation of GWAS and downstream analyses.
Pathogenic expansions of short tandem repeats (STRs) cause over 70 neurological diseases1-3. Here we performed a population-scale survey of pathogenic repeat expansions by analysing repeat length in 37 disease-associated STR loci in a diverse set of 1,020,833 samples using short-read sequencing whole-exome and whole-genome data. Consistent with previous findings, we found that the frequency of pathogenic repeats is higher than the prevalence of corresponding diseases for most loci4,5. Associations of repeat length with 7,671 binary traits captured known locus-trait associations, including HTT and Huntington's disease, DMPK and myotonic disorders and C9orf72 and motor neuron disease, among others. Finally, we found that, even before disease diagnosis, repeat expansions in several loci strongly associate with increased levels of neurofilament light chain (NfL) and a loss of brain volume in specific disease-associated regions. For example, carriers of HTT expansions exhibited a 22.1% loss of putamen volume, and carriers of CACNA1A expansions showed a 24.6% loss of cerebellar volume. These observations suggest that both decreased brain volumes and increased NfL levels occur earlier than disease diagnosis. This study demonstrates the use of characterizing repeat expansions from short-read sequencing data in diverse population-scale cohorts and its application to epidemiology and clinical biomarker development.
Mitochondrial heteroplasmic variant has been increasingly recognized as a potential contributor to common complex diseases, yet its relationship with cardiometabolic disorders (CMDs) remains poorly understood. Leveraging deep whole-genome sequencing data from 16,882 participants across six multi-ancestry TOPMed cohorts, we systematically evaluated the associations between rare heteroplasmic variants and eight CMD traits, including body mass index (BMI), obesity, blood pressure, hypertension, blood glucose, diabetes, low-density lipoprotein (LDL), and hyperlipidemia. Using a previously developed statistical framework, we identified heteroplasmic variants according to three coding definitions and performed gene-based burden, SKAT, SKAT-O and ACAT-O tests within sixteen mitochondrial DNA (mtDNA) genes. We identified twelve significant gene-trait associations after Bonferroni correction, with consistent effect directions across coding definitions. The strongest association was observed between hyperlipidemia and heteroplasmic variants in CO1 gene (OR=0.28, 95% CI=(0.17, 0.46), p=3.4E-7) among EA (European Americans). Additional associations were detected for BMI, adjusted SBP (systolic blood pressure), BG (blood glucose), diabetes, and adjusted LDL. These findings highlight the contribution of heteroplasmic variation within mtDNA to cardiometabolic phenotypes and provide new insight into mitochondrial involvement in CMD pathophysiology.
Myostatin negatively regulates skeletal muscle size in multiple species, and therefore, myostatin blockade has been therapeutically explored to promote muscle growth in humans, including to counter the muscle loss seen in obese humans using GLP1R agonists. In this study, we present results from a large multi-cohort genetic association analysis, using data from 1.1 million individuals to examine the effects of function-disrupting mutations in the myostatin gene (MSTN) on traits relevant to body composition and cardiometabolic health. Carriers of function-disrupting variants display decreased adiposity, an increase in lean mass, and increased grip strength and creatinine levels. We further characterize the effects of these variants on body composition using whole-body MRI data from UK Biobank, leveraging deep learning models to perform automated image segmentation for 77,572 individuals. Among mutation carriers increased muscle mass is observed across multiple muscle groups, with heterozygote carriers of loss-of-function-like mutations exhibiting increases in excess of 10%. Our findings demonstrate that lifelong reduction in myostatin function enhances muscle size and strength in humans while decreasing body adiposity, providing insights into the potential benefits and safety of long-term therapeutic blockade of myostatin signaling.
Altered energy metabolism is a shared driver across cardiometabolic diseases-the leading cause of death globally1. Energy metabolism varies between individuals and is partly heritable2-9. Here, to investigate the genetic basis of energy metabolism, we perform an exome-sequencing analysis of 1,032,116 people from America, Europe and Asia, and estimate associations between rare protein-coding variants and the ratio of triglyceride to high-density-lipoprotein cholesterol (TG:HDL)-an energy-state biomarker that we associate with diverse cardiometabolic risk factors and diseases. We identify 59 independent genes (P < 1.04 × 10-7) that are enriched for liver- and adipose-expressed master regulators of energy balance, storage and metabolism; 23 (39%) of these genes encode approved or clinical-stage drug targets. Ultra-rare protein-truncating variants in FNIP1 (allele frequency, 0.01%), which encodes a suppressor of energy expenditure and mitochondrial metabolism, are associated with a lower TG:HDL ratio, lower liver fat, lower glycaemia, favourable fat distribution and around 60% lower odds of cardiometabolic disease. FNIP1 knockdown in primary human hepatocytes induces lipid breakdown and lysosomal gene expression, while combined hepatic knockdown of Fnip1 with its paralogue Fnip2 or knockdown of its interactor Flcn protect against weight gain, reduce liver fat and enhance insulin sensitivity in mice fed a high-fat diet. Our study implicates the FNIP1 pathway in human energy metabolism and highlights its inhibition as a potential therapeutic strategy in cardiometabolic disease.
Most genetic variants associated with complex traits and diseases occur in non-coding genomic regions and are hypothesized to regulate gene expression. To understand the genetics underlying gene expression variability, we characterize 14,324 ancestrally diverse RNA-sequencing samples from the NHLBI Trans-Omics for Precision Medicine (TOPMed) program and integrate whole genome sequencing data to perform cis and trans expression and splicing quantitative trait locus (cis-/trans-e/sQTL) analyses in six tissues and cell types, most notably whole blood (N=6,454) and lung (N=1,291). We show this dataset enables greater detection of secondary cis-e/sQTL signals than was achieved in previous studies, and that secondary cis-eQTL and primary trans-eQTL signal discovery is not saturated even though eGene discovery is. Most TOPMed trans-eQTL signals colocalize with cis-e/sQTL signals, suggesting many trans signals are mediated by cis signals. We fine-map European UK BioBank GWAS signals from 164 traits and colocalize the resulting 34,107 fine-mapped GWAS signals with TOPMed e/sQTL signals, finding that of 10,611 GWAS signals with a colocalization, 7,096 GWAS signals colocalize with at least one secondary e/sQTL signal. These results demonstrate that larger e/sQTL analyses will continue to uncover secondary e/sQTL signals, and that these new signals will benefit GWAS interpretation.
In studies of individuals of primarily European genetic ancestry, common and low-frequency variants and rare coding variants have been found to be associated with the risk of bipolar disorder (BD) and schizophrenia (SZ). However, less is known for individuals of other genetic ancestries or the role of rare non-coding variants in BD and SZ risk. We performed whole-genome sequencing (∼27X) of African American individuals: 1,598 with BD, 3,295 with SZ, and 2,651 unaffected controls (InPSYght study). We increased power by incorporating 14,812 jointly called psychiatrically unscreened ancestry-matched controls from the Trans-Omics for Precision Medicine (TOPMed) Program for a total of 17,463 controls (∼37X). To identify variants and sets of variants associated with BD and/or SZ, we performed single-variant tests, gene-based tests for singleton protein truncating variants, and rare and low-frequency variant annotation-based tests with conservation and universal chromatin states and sliding windows. We found suggestive evidence of the association of BD with single variants on chromosome 18 and of lower BD risk associated with rare and low-frequency variants on chromosome 11 in a region with multiple BD genome-wide association study loci, using a sliding window approach. We also found that chromatin and conservation state tests can be used to detect differential calling of variants in controls sequenced at different centers and to assess the effectiveness of sequencing metric covariate adjustments. Our findings reinforce the need for continued whole-genome sequencing in additional samples of African American individuals and more comprehensive functional annotation of non-coding variants.
Rare variant association analysis, which assesses the aggregate effect of rare damaging variants within a gene, is a powerful strategy for advancing knowledge of human biology. Numerous models have been proposed to identify damaging coding variants, with the most recent ones employing deep learning and large language models (LLMs) to predict the impact of changes in coding sequences. Here, we use newly available proteomics data on 2898 proteins across 46,665 individuals to evaluate and refine LLM predictors of damaging variants. Using one of these refined models, we evaluate the association between rare damaging variants and human phenotypes at 241 positive control gene-trait pairs. Among these gene-trait pairs, our proteomics-guided model outperforms an ensemble of conventional approaches including PolyPhen2, MutationTaster, SIFT, and LRT, as well as newer machine learning approaches for identifying damaging missense variants, such as CADD, ESM-1v, ESM-1b, and AlphaMissense. When attempting to recover known associations by correctly separating damaging singleton missense variants from other singleton variants, our approach recapitulates 36.5% of gene-trait pairs with known associations, exceeding all the alternatives we considered. Furthermore, when we apply our model to 10 example traits from the UK Biobank, we identify 177 gene-trait associations-again exceeding all other approaches. Our results demonstrate that summary statistics from large-scale human proteomics data enable evaluation and refinement of coding variant classification LLMs, improving discovery potential in human genetic studies.
Meta-analysis of gene-based tests using single-variant summary statistics is a powerful strategy for genetic association studies. However, current approaches require sharing the covariance matrix between variants for each study and trait of interest. For large-scale studies with many phenotypes, these matrices can be cumbersome to calculate, store and share. Here, to address this challenge, we present REMETA-an efficient tool for meta-analysis of gene-based tests. REMETA uses a single sparse covariance reference file per study that is rescaled for each phenotype using single-variant summary statistics. We develop new methods for binary traits with case-control imbalance, and to estimate allele frequencies, genotype counts and effect sizes of burden tests. We demonstrate the performance and advantages of our approach through meta-analysis of five traits in 469,376 samples in UK Biobank. The open-source REMETA software will facilitate meta-analysis across large-scale exome sequencing studies from diverse studies that cannot easily be combined.
Personality traits describe stable differences in how individuals think, feel, and behave and how they interact with and experience their social and physical environments. We assemble data from 46 cohorts including 611K-1.14M participants with European-like and African-like genomes for genome-wide association studies (GWAS) of the Big Five personality traits (extraversion, agreeableness, conscientiousness, neuroticism, and openness to experience), and data from 51K participants for within-family GWAS. We identify 1,257 lead genetic variants associated with personality, including 823 novel variants. Common genetic variants explain 4.8%-9.3% of the variance in each trait, and 10.5%-16.2% accounting for measurement unreliability. Genetic effects on personality are highly consistent across geography, reporter (self vs. close other), age group, and measurement instrument, and we find minimal spousal assortment for personality in recent history. In stark contrast to many other social and behavioral traits, within-family GWAS and polygenic index analyses indicate little to no shared environmental confounding in genetic associations with personality. Polygenic prediction, genetic correlation, and Mendelian randomization analyses indicate that personality genetics have widespread, potentially causal associations with a wide range of consequential behaviors and life outcomes. The genetic architecture of personality is robust and fundamental to being a human.
Clonal hematopoiesis (CH) is defined by the expansion of a lineage of genetically identical cells in blood. Genetic lesions that confer a fitness advantage, such as leukemogenic point mutations or mosaic chromosomal alterations (mCAs), are frequent mediators of CH. However, recent analyses of both single cell-derived colonies of hematopoietic cells and population sequencing cohorts have revealed CH frequently occurs in the absence of known driver genetic lesions. To characterize CH without known driver genetic lesions, we use 51,399 deeply sequenced whole genomes from the NHLBI TOPMed sequencing initiative to perform simultaneous germline and somatic mutation analyses among individuals without leukemogenic point mutations (LPM), which we term CH-LPMneg. We quantify CH by estimating the total mutation burden. Because estimating somatic mutation burden without a paired-tissue sample is challenging, we develop a novel statistical method, the Genomic and Epigenomic informed Mutation (GEM) rate, that uses external genomic and epigenomic data sources to distinguish artifactual signals from true somatic mutations. We perform a genome-wide association study of GEM to discover the germline determinants of CH-LPMneg. We identify seven genes associated with CH-LPMneg (TCL1A, TERT, SMC4, NRIP1, PRDM16, MSRA, SCARB1).Functional analyses of SMC4 and NRIP1 implicated altered hematopoietic stem cell self-renewal and proliferation as the primary mediator of mutation burden in blood. We then perform comprehensive multi-tissue transcriptomic analyses, finding that the expression levels of 404 genes are associated with GEM. Finally, we perform phenotypic association meta-analyses across four cohorts, finding that GEM is associated with increased white blood cell count, but is not significantly associated with incident stroke or coronary disease events. Overall, we develop GEM for quantifying mutation burden from WGS and use GEM to discover the genetic, genomic, and phenotypic correlates of CH-LPMneg.
Persistent opioid use after surgery is a common morbidity outcome associated with subsequent opioid use disorder, overdose, and death. While phenotypic associations have been described, genetic associations remain unidentified. Here, we conducted the largest genetic study of persistent opioid use after surgery, comprising ~40,000 non-Hispanic, European-ancestry Michigan Genomics Initiative participants (3198 cases and 36,321 surgically exposed controls). Our study primarily focused on the reproducibility and reliability of 72 genetic studies of opioid use disorder phenotypes. Nominal associations (p < 0.05) occurred at 12 of 80 unique (r2 < 0.8) signals from these studies. Six occurred in OPRM1 (most significant: rs79704991-T, OR = 1.17, p = 8.7 × 10-5), with two surviving multiple testing correction. Other associations were rs640561-LRRIQ3 (p = 0.015), rs4680-COMT (p = 0.016), rs9478495 (p = 0.017, intergenic), rs10886472-GRK5 (p = 0.028), rs9291211-SLC30A9/BEND4 (p = 0.043), and rs112068658-KCNN1 (p = 0.048). Two highly referenced genes, OPRD1 and DRD2/ANKK1, had no signals in MGI. Associations at previously identified OPRM1 variants suggest common biology between persistent opioid use and opioid use disorder, further demonstrating connections between opioid dependence and addiction phenotypes. Lack of significant associations at other variants challenges previous studies' reliability.
Gene-based burden tests are a popular and powerful approach for analysis of exome-wide association studies. These approaches combine sets of variants within a gene into a single burden score that is then tested for association. Typically, a range of burden scores are calculated and tested across a range of annotation classes and frequency bins. Correlation between these tests can complicate the multiple testing correction and hamper interpretation of the results. We introduce a method called the sparse burden association test (SBAT) that tests the joint set of burden scores under the assumption that causal burden scores act in the same effect direction. The method simultaneously assesses the significance of the model fit and selects the set of burden scores that best explain the association at the same time. Using simulated data, we show that the method is well calibrated and highlight scenarios where the test outperforms existing gene-based tests. We apply the method to 73 quantitative traits from the UK Biobank, showing that SBAT is a valuable additional gene-based test when combined with other existing approaches. This test is implemented in the REGENIE software.
Whole-genome sequencing (WGS), whole-exome sequencing (WES) and array genotyping with imputation (IMP) are common strategies for assessing genetic variation and its association with medically relevant phenotypes. To date, there has been no systematic empirical assessment of the yield of these approaches when applied to hundreds of thousands of samples to enable the discovery of complex trait genetic signals. Using data for 100 complex traits from 149,195 individuals in the UK Biobank, we systematically compare the relative yield of these strategies in genetic association studies. We find that WGS and WES combined with arrays and imputation (WES + IMP) have the largest association yield. Although WGS results in an approximately fivefold increase in the total number of assayed variants over WES + IMP, the number of detected signals differed by only 1% for both single-variant and gene-based association analyses. Given that WES + IMP typically results in savings of lab and computational time and resources expended per sample, we evaluate the potential benefits of applying WES + IMP to larger samples. When we extend our WES + IMP analyses to 468,169 UK Biobank individuals, we observe an approximately fourfold increase in association signals with the threefold increase in sample size. We conclude that prioritizing WES + IMP and large sample sizes rather than contemporary short-read WGS alternatives will maximize the number of discoveries in genetic association studies. Comparison of association signals in UK Biobank using different strategies for assessing genetic variation shows that whole-exome sequencing combined with array genotyping and imputation offers similar performance to whole-genome sequencing at a reduced cost.
Coronavirus disease 2019 (COVID-19) and influenza are respiratory illnesses caused by the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) and influenza viruses, respectively. Both diseases share symptoms and clinical risk factors(1), but the extent to which these conditions have a common genetic etiology is unknown. This is partly because host genetic risk factors are well characterized for COVID-19 but not for influenza, with the largest published genome-wide association studies for these conditions including >2 million individuals(2) and about 1,000 individuals(3-6), respectively. Shared genetic risk factors could point to targets to prevent or treat both infections. Through a genetic study of 18,334 cases with a positive test for influenza and 276,295 controls, we show that published COVID-19 risk variants are not associated with influenza. Furthermore, we discovered and replicated an association between influenza infection and noncoding variants in B3GALT5 and ST6GAL1, neither of which was associated with COVID-19. In vitro small interfering RNA knockdown of ST6GAL1-an enzyme that adds sialic acid to the cell surface, which is used for viral entry-reduced influenza infectivity by 57%. These results mirror the observation that variants that downregulate ACE2, the SARS-CoV-2 receptor, protect against COVID-19 (ref. 7). Collectively, these findings highlight downregulation of key cell surface receptors used for viral entry as treatment opportunities to prevent COVID-19 and influenza.
Rare coding variants that substantially affect function provide insights into the biology of a gene1-3. However, ascertaining the frequency of such variants requires large sample sizes4-8. Here we present a catalogue of human protein-coding variation, derived from exome sequencing of 983,578 individuals across diverse populations. In total, 23% of the Regeneron Genetics Center Million Exome (RGC-ME) data come from individuals of African, East Asian, Indigenous American, Middle Eastern and South Asian ancestry. The catalogue includes more than 10.4 million missense and 1.1 million predicted loss-of-function (pLOF) variants. We identify individuals with rare biallelic pLOF variants in 4,848 genes, 1,751 of which have not been previously reported. From precise quantitative estimates of selection against heterozygous loss of function (LOF), we identify 3,988 LOF-intolerant genes, including 86 that were previously assessed as tolerant and 1,153 that lack established disease annotation. We also define regions of missense depletion at high resolution. Notably, 1,482 genes have regions that are depleted of missense variants despite being tolerant of pLOF variants. Finally, we estimate that 3% of individuals have a clinically actionable genetic variant, and that 11,773 variants reported in ClinVar with unknown significance are likely to be deleterious cryptic splice sites. To facilitate variant interpretation and genetics-informed precision medicine, we make this resource of coding variation from the RGC-ME dataset publicly accessible through a variant allele frequency browser.