ABSTRACT Huntington’s disease is a rare neurodegenerative disease whose primary risk factors are inherited expansions of a CAG repeat tract in the HTT gene. Somatic expansion of these tracts leads to neuronal toxicity, neuronal death and clinical disease progression. To identify genetic factors with a major impact on disease onset and progression, we genome sequenced 18,825 individuals for the ENROLL-HD study. Our results show rare inactivating mutations in three genes, all involved in DNA damage repair, are major determinants of age of onset for motor symptoms (n=10,610) and other clinical manifestations. Heterozygote carriers of predicted loss-of-function (pLoF) variants in POLD1 and PMS1 developed motor symptoms an average 20 years (n=3; P=1×10 −5 ) and 7 years (n=6; P=2×10 −3 ) later than non-carriers, respectively. Conversely, heterozygote carriers of pLoF variants in FAN1 (n=30) developed symptoms 10 years earlier (P=2×10 −10 ). Our findings highlight therapeutic strategies and help predict age of onset for at-risk individuals.
Rare coding variants that alter protein function and confer beneficial health effects can suggest potential drug targets. CHRNB3 encodes the β3 subunit of nicotinic acetylcholine receptors that bind nicotine and mediate its action in the brain. Here we report an exome-wide association study of number of cigarettes smoked per day (cig per day) in 37,897 current smokers from the Mexico City Prospective Study. We identify a deleterious missense variant in CHRNB3, p.Glu284Gly, that associates with a significant reduction in daily cigarette consumption. The missense variant is enriched in people of Indigenous Mexican ancestry but rare in other ancestries. We further identify a predicted loss-of-function variant in CHRNB3 that significantly associates with reduction in number of smoked cigarettes per day in participants of Japan Biobank. This variant is enriched in people of East Asian ancestry but is rare in other ancestries. Finally, we find that rare deleterious missense and predicted loss-of-function variants in aggregate associate with a reduction in the number of smoked cigarettes per day in individuals of European ancestry from the UK Biobank. Our results suggest that loss of function of CHRNB3 significantly associates with daily cigarette smoking, proposing β3 inhibition as a potential therapeutic strategy for nicotine addiction.
Pathogenic expansions of short tandem repeats (STRs) cause over 70 neurological diseases1-3. Here we performed a population-scale survey of pathogenic repeat expansions by analysing repeat length in 37 disease-associated STR loci in a diverse set of 1,020,833 samples using short-read sequencing whole-exome and whole-genome data. Consistent with previous findings, we found that the frequency of pathogenic repeats is higher than the prevalence of corresponding diseases for most loci4,5. Associations of repeat length with 7,671 binary traits captured known locus-trait associations, including HTT and Huntington's disease, DMPK and myotonic disorders and C9orf72 and motor neuron disease, among others. Finally, we found that, even before disease diagnosis, repeat expansions in several loci strongly associate with increased levels of neurofilament light chain (NfL) and a loss of brain volume in specific disease-associated regions. For example, carriers of HTT expansions exhibited a 22.1% loss of putamen volume, and carriers of CACNA1A expansions showed a 24.6% loss of cerebellar volume. These observations suggest that both decreased brain volumes and increased NfL levels occur earlier than disease diagnosis. This study demonstrates the use of characterizing repeat expansions from short-read sequencing data in diverse population-scale cohorts and its application to epidemiology and clinical biomarker development.
Highly heritable, polygenic, and easily measured, adult height has long been the model trait in human genetics 1,2 . While the landscape of height-associated common genetic variation has been studied extensively 2 , rare variation remains relatively unexplored 1 . Using rare protein-altering variants in a discovery set of 826,066 exomes, we identify 207 height-associated genes - 98% of which replicate in an additional 624,567 individuals. The rarest and most deleterious class of variation, singleton (frequency <0.0001%) putative loss-of-function (pLoF) variants implicated 17 genes with large effects on height ranging from -17 cm ( ACAN ) to +11 cm ( FBN1 ) per allele, 52× larger than the average effect of common height-associated variants and comparable to the 1% tails of a common variant polygenic score. Several genes (e.g., TET1 , DTL , IGF2BP2 ) have effect sizes at least as large as established Mendelian height genes but lack documented stature or skeletal growth syndromes. This is particularly true for genes in which rare variants associate with increased height. We performed the largest rare-variant study of height to date, directly implicate 207 genes that broadly overlap with both GWAS associations and Mendelian height syndromes, assess the impact of rare variants on heritability and prediction, provide evidence that height is an underappreciated clinical feature of Mendelian disorders, and demonstrate the utility of large population-scale sequencing studies for classifying individual variants and dissecting complex trait architecture.
Myostatin negatively regulates skeletal muscle size in multiple species, and therefore, myostatin blockade has been therapeutically explored to promote muscle growth in humans, including to counter the muscle loss seen in obese humans using GLP1R agonists. In this study, we present results from a large multi-cohort genetic association analysis, using data from 1.1 million individuals to examine the effects of function-disrupting mutations in the myostatin gene (MSTN) on traits relevant to body composition and cardiometabolic health. Carriers of function-disrupting variants display decreased adiposity, an increase in lean mass, and increased grip strength and creatinine levels. We further characterize the effects of these variants on body composition using whole-body MRI data from UK Biobank, leveraging deep learning models to perform automated image segmentation for 77,572 individuals. Among mutation carriers increased muscle mass is observed across multiple muscle groups, with heterozygote carriers of loss-of-function-like mutations exhibiting increases in excess of 10%. Our findings demonstrate that lifelong reduction in myostatin function enhances muscle size and strength in humans while decreasing body adiposity, providing insights into the potential benefits and safety of long-term therapeutic blockade of myostatin signaling.
Altered energy metabolism is a shared driver across cardiometabolic diseases-the leading cause of death globally1. Energy metabolism varies between individuals and is partly heritable2-9. Here, to investigate the genetic basis of energy metabolism, we perform an exome-sequencing analysis of 1,032,116 people from America, Europe and Asia, and estimate associations between rare protein-coding variants and the ratio of triglyceride to high-density-lipoprotein cholesterol (TG:HDL)-an energy-state biomarker that we associate with diverse cardiometabolic risk factors and diseases. We identify 59 independent genes (P < 1.04 × 10-7) that are enriched for liver- and adipose-expressed master regulators of energy balance, storage and metabolism; 23 (39%) of these genes encode approved or clinical-stage drug targets. Ultra-rare protein-truncating variants in FNIP1 (allele frequency, 0.01%), which encodes a suppressor of energy expenditure and mitochondrial metabolism, are associated with a lower TG:HDL ratio, lower liver fat, lower glycaemia, favourable fat distribution and around 60% lower odds of cardiometabolic disease. FNIP1 knockdown in primary human hepatocytes induces lipid breakdown and lysosomal gene expression, while combined hepatic knockdown of Fnip1 with its paralogue Fnip2 or knockdown of its interactor Flcn protect against weight gain, reduce liver fat and enhance insulin sensitivity in mice fed a high-fat diet. Our study implicates the FNIP1 pathway in human energy metabolism and highlights its inhibition as a potential therapeutic strategy in cardiometabolic disease.
BACKGROUND:It has long been appreciated that multiple allergic/atopic conditions (type-2-inflammation-associated diseases, T2IDs) often appear in the same individual, suggesting a shared systemic immune deviation. Emerging molecular and clinical investigations of interleukin (IL)4/IL13 blockade with the fully-human antibody dupilumab suggest that many are driven by over-activation of the IL4/IL13 pathway, manifesting differently in distinct tissues. METHODS AND RESULTS:We performed comprehensive quantitative analyses using several large, independent datasets (including real-world and clinical trial datasets) to confirm that many T2IDs-including asthma, atopic dermatitis, eosinophilic esophagitis, nasal polyps, prurigo nodularis-are highly interrelated based on co-prevalence, genetic predispositions, and transcriptomic signatures driven by over-activation of the IL4/IL13 pathway. We also summarize extensive clinical trial data showing that these T2IDs also share uniformly impressive clinical responses to dual IL4/IL13 blockade with dupilumab. CONCLUSIONS:While highly consistent with previous data and observations, this is the first systematic comparison of these types of data across multiple T2IDs, using a rigorous quantitative assessment of the largest datasets evaluated to date. Many T2IDs are not only highly interrelated and comorbid, but share transcriptomic signatures and genetic predispositions that suggest they all share over-activation of the IL4/IL13 pathway as a key causative driver, consistent with and explaining their uniform clinical responsiveness to dual IL4/IL13 blockade. These genetic and transcriptomic signatures may also have utility for identifying other diseases mediated by IL4/IL13 or T2-mediated subpopulations of heterogenous inflammatory diseases.
Abstract Background Polygenic risk scores (PRSs) strongly discriminate for prostate cancer risk at the population level. The role of these genetic scores in determining prostate cancer survival is unclear, potentially due to methodological issues. Methods We included 19,607 men from the Malmö Diet and Cancer Study (MDCS) and the Health Professionals Follow-up Study (HPFS) and analyzed 20-year incidence and mortality (from full cohort analyses) and survival (from case-only analyses) according to a 451-variant PRS. For most of the men included, early access to prostate-specific antigen (PSA) testing was limited. Results In full cohort analyses, a PRS at or above the median (vs. below the median) shows a strong association with prostate cancer incidence (hazard ratio (HR) 3.02, 95% CI 2.78-3.28) and, somewhat stronger, with prostate cancer mortality (HR 3.26, 95% CI 2.63-4.04). As expected, case-only survival analyses of prostate cancer death show a similar direction (HR 1.21, 95% CI 0.98-1.50), which becomes stronger when excluding variants linked to PSA (HR 1.25, 95% CI, 1.01-1.54), in particular in the age group 65-74 years at diagnosis (HR 1.72, 95% CI 1.21-2.45). In the other age groups, HRs are close to or below 1. This indicates that standard case-only survival estimates may not accurately reflect risk across age, consistent with age-dependent selection mechanisms, and further points to an influence from disease detection. Conclusions Overall, these findings support that inherited genetic risk captured by the 451-variant PRS may have an influence on prostate cancer survival.
BACKGROUND Genetic deficiency of otoferlin, a protein critical to synaptic transmission by the sensory hair cells of the ear, causes congenital deafness. Medicines to treat the condition are lacking; children typically receive cochlear implants. DB-OTO is a dual adenoassociated virus 1 gene therapy that delivers human OTOF complementary DNA (encoding otoferlin) regulated by a hair cell-specific promoter. METHODS We conducted an open-label, single-group, first-in-human registrational study to evaluate DB-OTO. Children with OTOF variants and profound deafness (defined by an average audiometric threshold of >90 decibel hearing level [dB HL], indicating an inability to hear a gas-powered lawn mower) received an intracochlear infusion of DB-OTO (7.2x10(12) vector genomes per ear) in one or both ears. The primary efficacy end point was an average threshold on behavioral pure-tone audiometry (PTA) at week 24 of 70 dB HL or less, a clinical standard that generally avoids cochlear implantation and enables natural acoustic hearing. A key secondary end point was the presence of an auditory brain-stem response to a click stimulus at a threshold at or below 90 dB normalized hearing level (db nHL) at week 24. Safety assessments included adverse events, laboratory results, and vestibular testing. RESULTS A total of 12 children have been enrolled in the study. After a single infusion of DB-OTO, a PTA average threshold of 70 dB HL or less at week 24 (primary end point) and an auditory brain-stem response at or below 90 dB nHL (key secondary end point) were found in 9 of the 12 participants (75%; 95% confidence interval, 43 to 95; P = 1.1x10(-13 )for both end points). Six participants could hear soft speech without assistive devices, and 3 had average normal hearing sensitivity. A total of 67 adverse events occurred or worsened during or after treatment, none of which led to discontinued participation in the study. CONCLUSIONS DB-OTO gene therapy improved hearing in patients with OTOF-related deafness, enabling natural acoustic hearing and normalizing hearing sensitivity in 3 of 12 treated patients. (Funded by Regeneron Pharmaceuticals; ClinicalTrials.gov number, NCT05788536.)
Rare variant association analysis, which assesses the aggregate effect of rare damaging variants within a gene, is a powerful strategy for advancing knowledge of human biology. Numerous models have been proposed to identify damaging coding variants, with the most recent ones employing deep learning and large language models (LLMs) to predict the impact of changes in coding sequences. Here, we use newly available proteomics data on 2898 proteins across 46,665 individuals to evaluate and refine LLM predictors of damaging variants. Using one of these refined models, we evaluate the association between rare damaging variants and human phenotypes at 241 positive control gene-trait pairs. Among these gene-trait pairs, our proteomics-guided model outperforms an ensemble of conventional approaches including PolyPhen2, MutationTaster, SIFT, and LRT, as well as newer machine learning approaches for identifying damaging missense variants, such as CADD, ESM-1v, ESM-1b, and AlphaMissense. When attempting to recover known associations by correctly separating damaging singleton missense variants from other singleton variants, our approach recapitulates 36.5% of gene-trait pairs with known associations, exceeding all the alternatives we considered. Furthermore, when we apply our model to 10 example traits from the UK Biobank, we identify 177 gene-trait associations-again exceeding all other approaches. Our results demonstrate that summary statistics from large-scale human proteomics data enable evaluation and refinement of coding variant classification LLMs, improving discovery potential in human genetic studies.
Background: Von Willebrand factor (VWF) and coagulation factor VIII (FVIII) plasma levels are associated with increased risk for venous thromboembolism (VTE). Objectives: This study aimed to determine the thrombotic risk of rare and common variants of 27 genes linked to VWF or FVIII plasma levels in genome-wide association studies. Methods: Exon sequences of 27 genes linked to plasma levels of VWF or FVIII in genome-wide association studies were analyzed for common and rare variants in 28,794 subjects without VTE (born during 1923-1950, 60% women), who participated in the Malmö Diet and Cancer study (1991-1996), with a follow-up time until 2018. Hazard ratios (HRs) were determined. P values were Bonferroni-corrected (P value = .05/27 <.0019). Common variants were analyzed individually. Rare qualifying variants (<0.1%) were collapsed. Results: None of the 27 genes were associated with VTE in the rare variant collapsing analysis. Three common exon variants were significantly associated with VTE: rs8176719 (frameshift) in ABO (HR = 1.30; 95% CI, 1.20-1.42; P = 3.9 × 10−10), rs1800291 (p.Asp1260Glu) in F8 (HR = 1.29; 95% CI, 1.08-1.55; P = .00046 for men; HR = 1.17; 95% CI, 1.06-1.29; P = .00019 for women), and rs1063856 (p.Thr789Ala) in VWF (HR = 1.10; 95% CI, 1.04-1.17; P = .00057). A risk score of these 3 variants was dose-dependently associated with VTE (5 risk alleles): HR = 2.8; 95% CI, 1.7-4.7; and P value = .00008. The area under the curve for VTE in receiver operating characteristics for the risk score was similar to FV Leiden (0.55 vs 0.54). Conclusion: The risk score of 3 common variants in VWF, F8, and AB0 genes is associated with VTE risk similar to FV Leiden.
Gene-based burden tests are a popular and powerful approach for analysis of exome-wide association studies. These approaches combine sets of variants within a gene into a single burden score that is then tested for association. Typically, a range of burden scores are calculated and tested across a range of annotation classes and frequency bins. Correlation between these tests can complicate the multiple testing correction and hamper interpretation of the results. We introduce a method called the sparse burden association test (SBAT) that tests the joint set of burden scores under the assumption that causal burden scores act in the same effect direction. The method simultaneously assesses the significance of the model fit and selects the set of burden scores that best explain the association at the same time. Using simulated data, we show that the method is well calibrated and highlight scenarios where the test outperforms existing gene-based tests. We apply the method to 73 quantitative traits from the UK Biobank, showing that SBAT is a valuable additional gene-based test when combined with other existing approaches. This test is implemented in the REGENIE software.
Whole-genome sequencing (WGS), whole-exome sequencing (WES) and array genotyping with imputation (IMP) are common strategies for assessing genetic variation and its association with medically relevant phenotypes. To date, there has been no systematic empirical assessment of the yield of these approaches when applied to hundreds of thousands of samples to enable the discovery of complex trait genetic signals. Using data for 100 complex traits from 149,195 individuals in the UK Biobank, we systematically compare the relative yield of these strategies in genetic association studies. We find that WGS and WES combined with arrays and imputation (WES + IMP) have the largest association yield. Although WGS results in an approximately fivefold increase in the total number of assayed variants over WES + IMP, the number of detected signals differed by only 1% for both single-variant and gene-based association analyses. Given that WES + IMP typically results in savings of lab and computational time and resources expended per sample, we evaluate the potential benefits of applying WES + IMP to larger samples. When we extend our WES + IMP analyses to 468,169 UK Biobank individuals, we observe an approximately fourfold increase in association signals with the threefold increase in sample size. We conclude that prioritizing WES + IMP and large sample sizes rather than contemporary short-read WGS alternatives will maximize the number of discoveries in genetic association studies. Comparison of association signals in UK Biobank using different strategies for assessing genetic variation shows that whole-exome sequencing combined with array genotyping and imputation offers similar performance to whole-genome sequencing at a reduced cost.
Coronavirus disease 2019 (COVID-19) and influenza are respiratory illnesses caused by the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) and influenza viruses, respectively. Both diseases share symptoms and clinical risk factors(1), but the extent to which these conditions have a common genetic etiology is unknown. This is partly because host genetic risk factors are well characterized for COVID-19 but not for influenza, with the largest published genome-wide association studies for these conditions including >2 million individuals(2) and about 1,000 individuals(3-6), respectively. Shared genetic risk factors could point to targets to prevent or treat both infections. Through a genetic study of 18,334 cases with a positive test for influenza and 276,295 controls, we show that published COVID-19 risk variants are not associated with influenza. Furthermore, we discovered and replicated an association between influenza infection and noncoding variants in B3GALT5 and ST6GAL1, neither of which was associated with COVID-19. In vitro small interfering RNA knockdown of ST6GAL1-an enzyme that adds sialic acid to the cell surface, which is used for viral entry-reduced influenza infectivity by 57%. These results mirror the observation that variants that downregulate ACE2, the SARS-CoV-2 receptor, protect against COVID-19 (ref. 7). Collectively, these findings highlight downregulation of key cell surface receptors used for viral entry as treatment opportunities to prevent COVID-19 and influenza.
Rare coding variants that substantially affect function provide insights into the biology of a gene1-3. However, ascertaining the frequency of such variants requires large sample sizes4-8. Here we present a catalogue of human protein-coding variation, derived from exome sequencing of 983,578 individuals across diverse populations. In total, 23% of the Regeneron Genetics Center Million Exome (RGC-ME) data come from individuals of African, East Asian, Indigenous American, Middle Eastern and South Asian ancestry. The catalogue includes more than 10.4 million missense and 1.1 million predicted loss-of-function (pLOF) variants. We identify individuals with rare biallelic pLOF variants in 4,848 genes, 1,751 of which have not been previously reported. From precise quantitative estimates of selection against heterozygous loss of function (LOF), we identify 3,988 LOF-intolerant genes, including 86 that were previously assessed as tolerant and 1,153 that lack established disease annotation. We also define regions of missense depletion at high resolution. Notably, 1,482 genes have regions that are depleted of missense variants despite being tolerant of pLOF variants. Finally, we estimate that 3% of individuals have a clinically actionable genetic variant, and that 11,773 variants reported in ClinVar with unknown significance are likely to be deleterious cryptic splice sites. To facilitate variant interpretation and genetics-informed precision medicine, we make this resource of coding variation from the RGC-ME dataset publicly accessible through a variant allele frequency browser.
The genetic factors of stroke in South Asians are largely unexplored. Exome-wide sequencing and association analysis (ExWAS) in 75K Pakistanis identified NM_000435.3(NOTCH3):c.3691C>T, encoding the missense amino acid substitution p.Arg1231Cys, enriched in South Asians (alternate allele frequency = 0.58% compared to 0.019% in Western Europeans), and associated with subcortical hemorrhagic stroke [odds ratio (OR)=3.39, 95% confidence interval (CI)=[2.26, 5.10], p=3.87x10(-9)), and all strokes (OR [CI]=2.30 [1.77, 3.01], p=7.79x10(-10)). NOTCH3 p.Arg231Cys was strongly associated with white matter hyperintensity on MRI in United Kingdom Biobank (UKB) participants (effect [95% CI] in SD units=1.1 [0.61, 1.5], p=3.0x10(-6)). The variant is attributable for approximately 2.0% of hemorrhagic strokes and 1.1% of all strokes in South Asians. These findings highlight the value of diversity in genetic studies and have major implications for genomic medicine and therapeutic development in South Asian populations.