Genetic predisposition and alcohol consumption are risk factors for increased blood pressure (BP), but their interactions influencing BP remain understudied. We conducted population-specific and cross-population meta-analyses of genome-wide gene-alcohol (GxAlc) interactions affecting BP in >1.1M individuals from multiple populations. We identified 46 GxAlc interaction loci for BP, including 21 from one-degree-of-freedom interaction tests (PGxAlc<5×10-8; or <0.05/Meff, Meff independent BP associations at P<10-5), and 25 from two-degree-of-freedom tests of main and interaction effects (PGxAlc<0.05/M2df, M2df independent 2df-associations at P2df<5×10-8), including 7 novel and 39 known BP loci. The 12q24 locus highlights the genetic effect of BRAP-rs11066001 on BP, being ~6 times larger in current drinkers than in non-drinkers. Gene prioritization with 46 GxAlc loci identified 15 genes with ≥3 lines of evidence (location, literature, druggability, functional/regulatory annotation, or pathway analyses). Several loci showed sex- and population-specific effects and revealed biological pathways of alcohol's influence on BP, suggesting mechanisms underlying alcohol-induced hypertension.
The identification of sex-differential gene regulatory elements is essential for understanding sex-differential patterns of health and disease. We leveraged bulk and single-nucleus RNA sequencing (RNA-seq) and single-nucleus ATAC-seq data from 281 skeletal muscle biopsies to characterize sex differences in gene expression and regulation at the cell-type and whole-tissue levels. We found highly concordant sex-biased expression of over 2,100 genes across the three muscle fiber types and bulk tissue. Gene pathways related to mitochondrial activity and energy metabolism were enriched for male-biased expression, whereas those related to signal transduction and cell differentiation were enriched for female-biased expression. We found widespread sex-biased chromatin accessibility enriched in proximal and distal gene regulatory states; in gene promoters, sex-biased chromatin accessibility was positively associated with sex-biased expression. Long noncoding RNAs (lncRNAs) and microRNAs (miRNAs) also showed extensive sex-biased expression in the fiber-type and bulk data, respectively. Together, these results highlight nuclear and cytoplasmic mechanisms for sex-differential gene regulation in skeletal muscle.
Polygenic scores (PGSs) for body mass index (BMI) may guide early prevention and targeted treatment of obesity. Using genetic data from up to 5.1 million people (4.6% African ancestry, 14.4% American ancestry, 8.4% East Asian ancestry, 71.1% European ancestry and 1.5% South Asian ancestry) from the GIANT consortium and 23andMe, Inc., we developed ancestry-specific and multi-ancestry PGSs. The multi-ancestry score explained 17.6% of BMI variation among UK Biobank participants of European ancestry. For other populations, this ranged from 16% in East Asian-Americans to 2.2% in rural Ugandans. In the ALSPAC study, children with higher PGSs showed accelerated BMI gain from age 2.5 years to adolescence, with earlier adiposity rebound. Adding the PGS to predictors available at birth nearly doubled explained variance for BMI from age 5 onward (for example, from 11% to 21% at age 8). Up to age 5, adding the PGS to early-life BMI improved prediction of BMI at age 18 (for example, from 22% to 35% at age 5). Higher PGSs were associated with greater adult weight gain. In intensive lifestyle intervention trials, individuals with higher PGSs lost modestly more weight in the first year (0.55 kg per s.d.) but were more likely to regain it. Overall, these data show that PGSs have the potential to improve obesity prediction, particularly when implemented early in life.
Metabolites are small molecules that are useful for estimating disease risk and elucidating disease biology. Here, we perform two-sample Mendelian randomization to systematically infer the potential causal effects of 1099 plasma metabolites measured in 6136 Finnish men from the METSIM study on risk of 2099 binary disease endpoints measured in 309,154 Finnish individuals from FinnGen. We find evidence for 282 putative causal effects of 70 metabolites on 183 disease endpoints. We also identify 25 metabolites with potential causal effects across multiple disease domains, including ascorbic acid 2-sulfate affecting 26 disease endpoints in 12 disease domains. Our study suggests that N-acetyl-2-aminooctanoate and glycocholenate sulfate affect risk of atrial fibrillation through two distinct metabolic pathways and that N-methylpipecolate may mediate the putative causal effect of N6,N6-dimethyllysine on anxious personality disorder.
Complete characterization of the genetic effects on gene expression is needed to elucidate tissue biology and the etiology of complex traits. Here, we analyzed 2,344 subcutaneous adipose tissue samples and identified 34K conditionally distinct expression quantitative trait locus (eQTL) signals in 18K genes. Over half of eQTL genes exhibited at least two eQTL signals. Compared to primary signals, non-primary signals had lower effect sizes, lower minor allele frequencies, and less promoter enrichment; they corresponded to genes with higher heritability and higher tolerance for loss of function. Colocalization of eQTL with conditionally distinct genome-wide association study signals for 28 cardiometabolic traits identified 3,605 eQTL signals for 1,861 genes. Inclusion of non-primary eQTL signals increased colocalized signals by 46%. Among 30 genes with ≥2 pairs of colocalized signals, 21 showed a mediating gene dosage effect on the trait. Thus, expanded eQTL identification reveals more mechanisms underlying complex traits and improves understanding of the complexity of gene expression regulation.
Distinct tissue-specific mechanisms mediate insulin action in fasting and postprandial states. Previous genetic studies have largely focused on insulin resistance in the fasting state, where hepatic insulin action dominates. Here we studied genetic variants influencing insulin levels measured 2 h after a glucose challenge in >55,000 participants from three ancestry groups. We identified ten new loci (P < 5 × 10−8) not previously associated with postchallenge insulin resistance, eight of which were shown to share their genetic architecture with type 2 diabetes in colocalization analyses. We investigated candidate genes at a subset of associated loci in cultured cells and identified nine candidate genes newly implicated in the expression or trafficking of GLUT4, the key glucose transporter in postprandial glucose uptake in muscle and fat. By focusing on postprandial insulin resistance, we highlighted the mechanisms of action at type 2 diabetes loci that are not adequately captured by studies of fasting glycemic traits. Genome-wide association analyses of two oral glucose tolerance test-derived measures of postprandial insulin resistance discover ten new loci. Functional characterization identifies nine candidate genes implicated in the regulation of GLUT4.
Insulin secretion is critical for glucose homeostasis, and increased levels of the precursor proinsulin relative to insulin indicate pancreatic islet beta-cell stress and insufficient insulin secretory capacity in the setting of insulin resistance. We conducted meta-analyses of genome-wide association results for fasting proinsulin from 16 European-ancestry studies in 45,861 individuals. We found 36 independent signals at 30 loci (p value < 5 × 10-8), which validated 12 previously reported loci for proinsulin and ten additional loci previously identified for another glycemic trait. Half of the alleles associated with higher proinsulin showed higher rather than lower effects on glucose levels, corresponding to different mechanisms. Proinsulin loci included genes that affect prohormone convertases, beta-cell dysfunction, vesicle trafficking, beta-cell transcriptional regulation, and lysosomes/autophagy processes. We colocalized 11 proinsulin signals with islet expression quantitative trait locus (eQTL) data, suggesting candidate genes, including ARSG, WIPI1, SLC7A14, and SIX3. The NKX6-3/ANK1 proinsulin signal colocalized with a T2D signal and an adipose ANK1 eQTL signal but not the islet NKX6-3 eQTL. Signals were enriched for islet enhancers, and we showed a plausible islet regulatory mechanism for the lead signal in the MADD locus. These results show how detailed genetic studies of an intermediate phenotype can elucidate mechanisms that may predispose one to disease.
Background Genome-wide association studies for glycemic traits have identified hundreds of loci associated with these biomarkers of glucose homeostasis. Despite this success, the challenge remains to link variant associations to genes, and underlying biological pathways. Methods To identify coding variant associations which may pinpoint effector genes at both novel and previously established genome-wide association loci, we performed meta-analyses of exome-array studies for four glycemic traits: glycated hemoglobin (HbA1c, up to 144,060 participants), fasting glucose (FG, up to 129,665 participants), fasting insulin (FI, up to 104,140) and 2hr glucose post-oral glucose challenge (2hGlu, up to 57,878). In addition, we performed network and pathway analyses. Results Single-variant and gene-based association analyses identified coding variant associations at more than 60 genes, which when combined with other datasets may be useful to nominate effector genes. Network and pathway analyses identified pathways related to insulin secretion, zinc transport and fatty acid metabolism. HbA1c associations were strongly enriched in pathways related to blood cell biology. Conclusions Our results provided novel glycemic trait associations and highlighted pathways implicated in glycemic regulation. Exome-array summary statistic results are being made available to the scientific community to enable further discoveries.
Resting heart rate is associated with cardiovascular diseases and mortality in observational and Mendelian randomization studies. The aims of this study are to extend the number of resting heart rate associated genetic variants and to obtain further insights in resting heart rate biology and its clinical consequences. A genome-wide meta-analysis of 100 studies in up to 835,465 individuals reveals 493 independent genetic variants in 352 loci, including 68 genetic variants outside previously identified resting heart rate associated loci. We prioritize 670 genes and in silico annotations point to their enrichment in cardiomyocytes and provide insights in their ECG signature. Two-sample Mendelian randomization analyses indicate that higher genetically predicted resting heart rate increases risk of dilated cardiomyopathy, but decreases risk of developing atrial fibrillation, ischemic stroke, and cardio-embolic stroke. We do not find evidence for a linear or non-linear genetic association between resting heart rate and all-cause mortality in contrast to our previous Mendelian randomization study. Systematic alteration of key differences between the current and previous Mendelian randomization study indicates that the most likely cause of the discrepancy between these studies arises from false positive findings in previous one-sample MR analyses caused by weak-instrument bias at lower P -value thresholds. The results extend our understanding of resting heart rate biology and give additional insights in its role in cardiovascular disease development.
Conventional measurements of fasting and postprandial blood glucose levels investigated in genome-wide association studies (GWAS) cannot capture the effects of DNA variability on ‘around the clock’ glucoregulatory processes. Here we show that GWAS meta-analysis of glucose measurements under nonstandardized conditions (random glucose (RG)) in 476,326 individuals of diverse ancestries and without diabetes enables locus discovery and innovative pathophysiological observations. We discovered 120 RG loci represented by 150 distinct signals, including 13 with sex-dimorphic effects, two cross-ancestry and seven rare frequency signals. Of these, 44 loci are new for glycemic traits. Regulatory, glycosylation and metagenomic annotations highlight ileum and colon tissues, indicating an underappreciated role of the gastrointestinal tract in controlling blood glucose. Functional follow-up and molecular dynamics simulations of lower frequency coding variants in glucagon-like peptide-1 receptor ( GLP1R ), a type 2 diabetes treatment target, reveal that optimal selection of GLP-1R agonist therapy will benefit from tailored genetic stratification. We also provide evidence from Mendelian randomization that lung function is modulated by blood glucose and that pulmonary dysfunction is a diabetes complication. Our investigation yields new insights into the biology of glucose regulation, diabetes complications and pathways for treatment stratification.
Increased blood lipid levels are heritable risk factors of cardiovascular disease with varied prevalence worldwide owing to different dietary patterns and medication use1. Despite advances in prevention and treatment, in particular through reducing low-density lipoprotein cholesterol levels2, heart disease remains the leading cause of death worldwide3. Genome-wideassociation studies (GWAS) of blood lipid levels have led to important biological and clinical insights, as well as new drug targets, for cardiovascular disease. However, most previous GWAS4–23 have been conducted in European ancestry populations and may have missed genetic variants that contribute to lipid-level variation in other ancestry groups. These include differences in allele frequencies, effect sizes and linkage-disequilibrium patterns24. Here we conduct a multi-ancestry, genome-wide genetic discovery meta-analysis of lipid levels in approximately 1.65 million individuals, including 350,000 of non-European ancestries. We quantify the gain in studying non-European ancestries and provide evidence to support the expansion of recruitment of additional ancestries, even with relatively small sample sizes. We find that increasing diversity rather than studying additional individuals of European ancestry results in substantial improvements in fine-mapping functional variants and portability of polygenic prediction (evaluated in approximately 295,000 individuals from 7 ancestry groupings). Modest gains in the number of discovered loci and ancestry-specific variants were also achieved. As GWAS expand emphasis beyond the identification of genes and fundamental biology towards the use of genetic variants for preventive and precision medicine25, we anticipate that increased diversity of participants will lead to more accurate and equitable26 application of polygenic scores in clinical practice. A genome-wide association meta-analysis study of blood lipid levels in roughly 1.6 million individuals demonstrates the gain of power attained when diverse ancestries are included to improve fine-mapping and polygenic score generation, with gains in locus discovery related to sample size.
Metabolites are small molecules that are useful for estimating disease risk and elucidating disease biology. Nevertheless, their causal effects on human diseases have not been evaluated comprehensively. We performed two-sample Mendelian randomization to systematically infer the causal effects of 1,099 plasma metabolites measured in 6,136 Finnish men from the METSIM study on risk of 2,099 binary disease endpoints measured in 309,154 Finnish individuals from FinnGen. We identified evidence for 282 causal effects of 70 metabolites on 183 disease endpoints (FDR<1%). We found 25 metabolites with potential causal effects across multiple disease domains, including ascorbic acid 2-sulfate affecting 26 disease endpoints in 12 disease domains. Our study suggests that N-acetyl-2-aminooctanoate and glycocholenate sulfate affect risk of atrial fibrillation through two distinct metabolic pathways and that N-methylpipecolate may mediate the causal effect of N6, N6-dimethyllysine on anxious personality disorder. This study highlights the broad causal impact of plasma metabolites and widespread metabolic connections across diseases.
Although physical activity and sedentary behavior are moderately heritable, little is known about the mechanisms that influence these traits. Combining data for up to 703,901 individuals from 51 studies in a multi-ancestry meta-analysis of genome-wide association studies yields 99 loci that associate with self-reported moderate-to-vigorous intensity physical activity during leisure time (MVPA), leisure screen time (LST) and/or sedentary behavior at work. Loci associated with LST are enriched for genes whose expression in skeletal muscle is altered by resistance training. A missense variant in ACTN3 makes the alpha-actinin-3 filaments more flexible, resulting in lower maximal force in isolated type IIA muscle fibers, and possibly protection from exercise-induced muscle damage. Finally, Mendelian randomization analyses show that beneficial effects of lower LST and higher MVPA on several risk factors and diseases are mediated or confounded by body mass index (BMI). Our results provide insights into physical activity mechanisms and its role in disease prevention.
ABSTRACTCommon SNPs are predicted to collectively explain 40-50% of phenotypic variation in human height, but identifying the specific variants and associated regions requires huge sample sizes. Here we show, using GWAS data from 5.4 million individuals of diverse ancestries, that 12,111 independent SNPs that are significantly associated with height account for nearly all of the common SNP-based heritability. These SNPs are clustered within 7,209 non-overlapping genomic segments with a median size of ~90 kb, covering ~21% of the genome. The density of independent associations varies across the genome and the regions of elevated density are enriched for biologically relevant genes. In out-of-sample estimation and prediction, the 12,111 SNPs account for 40% of phenotypic variance in European ancestry populations but only ~10%-20% in other ancestries. Effect sizes, associated regions, and gene prioritization are similar across ancestries, indicating that reduced prediction accuracy is likely explained by linkage disequilibrium and allele frequency differences within associated regions. Finally, we show that the relevant biological pathways are detectable with smaller sample sizes than needed to implicate causal genes and variants. Overall, this study, the largest GWAS to date, provides an unprecedented saturated map of specific genomic regions containing the vast majority of common height-associated variants.
COVID-19 severity has varied widely, with demographic and cardio-metabolic factors increasing risk of severe reactions to SARS-CoV-2 infection, but the underlying mechanisms for this remain uncertain. We investigated phenotypic and genetic factors associated with subcutaneous adipose tissue expression of Angiotensin I Converting Enzyme 2 (ACE2), which has been shown to act as a receptor for SARS-CoV-2 cellular entry. In a meta-analysis of three independent studies including up to 1,471 participants, lower adipose tissue ACE2 expression was associated with adverse cardio-metabolic health indices including type 2 diabetes (T2D) and obesity status, higher serum fasting insulin and BMI, and lower serum HDL levels (P<5.32x10-4). ACE2 expression levels were also associated with estimated proportions of cell types in adipose tissue; lower ACE2 expression was associated with a lower proportion of microvascular endothelial cells (P=4.25x10-4) and higher macrophage proportion (P=2.74x10-5), suggesting a link to inflammation. Despite an estimated heritability of 32%, we did not identify any proximal or distal genetic variants (eQTLs) associated with adipose tissue ACE2 expression. Our results demonstrate that at-risk individuals have lower background ACE2 levels in this highly relevant tissue. Further studies will be required to establish how this may contribute to increased COVID-19 severity.
Few studies have explored the impact of rare variants (minor allele frequency < 1%) on highly heritable plasma metabolites identified in metabolomic screens. The Finnish population provides an ideal opportunity for such explorations, given the multiple bottlenecks and expansions that have shaped its history, and the enrichment for many otherwise rare alleles that has resulted. Here, we report genetic associations for 1391 plasma metabolites in 6136 men from the late-settlement region of Finland. We identify 303 novel association signals, more than one third at variants rare or enriched in Finns. Many of these signals identify genes not previously implicated in metabolite genome-wide association studies and suggest mechanisms for diseases and disease-related traits.
Transcriptomics data have been integrated with genome-wide association studies (GWASs) to help understand disease/trait molecular mechanisms. The utility of metabolomics, integrated with transcriptomics and disease GWASs, to understand molecular mechanisms for metabolite levels or diseases has not been thoroughly evaluated. We performed probabilistic transcriptome-wide association and locus-level colocalization analyses to integrate transcriptomics results for 49 tissues in 706 individuals from the GTEx project, metabolomics results for 1,391 plasma metabolites in 6,136 Finnish men from the METSIM study, and GWAS results for 2,861 disease traits in 260,405 Finnish individuals from the FinnGen study. We found that genetic variants that regulate metabolite levels were more likely to influence gene expression and disease risk compared to the ones that do not. Integrating transcriptomics with metabolomics results prioritized 397 genes for 521 metabolites, including 496 previously identified gene-metabolite pairs with strong functional connections and suggested 33.3% of such gene-metabolite pairs shared the same causal variants with genetic associations of gene expression. Integrating transcriptomics and metabolomics individually with FinnGen GWAS results identified 1,597 genes for 790 disease traits. Integrating transcriptomics and metabolomics jointly with FinnGen GWAS results helped pinpoint metabolic pathways from genes to diseases. We identified putative causal effects of UGT1A1/UGT1A4 expression on gallbladder disorders through regulating plasma (E,E)-bilirubin levels, of SLC22A5 expression on nasal polyps and plasma carnitine levels through distinct pathways, and of LIPC expression on age-related macular degeneration through glycerophospholipid metabolic pathways. Our study highlights the power of integrating multiple sets of molecular traits and GWAS results to deepen understanding of disease pathophysiology.
The genetic determinants of fasting glucose (FG) and fasting insulin (FI) have been studied mostly through genome arrays, resulting in over 100 associated variants. We extended this work with high-coverage whole genome sequencing analyses from fifteen cohorts in NHLBI’s Trans-Omics for Precision Medicine (TOPMed) program. Over 23,000 non-diabetic individuals from five race-ethnicities/populations (African, Asian, European, Hispanic and Samoan) were included. Eight variants were significantly associated with FG or FI across previously identified regions MTNR1B, G6PC2, GCK, GCKR and FOXA2 . We additionally characterize suggestive associations with FG or FI near previously identified SLC30A8, TCF7L2 , and ADCY5 regions as well as APOB, PTPRT , and ROBO1 . Functional annotation resources including the Diabetes Epigenome Atlas were compiled for each signal (chromatin states, annotation principal components, and others) to elucidate variant-to-function hypotheses. We provide a catalog of nucleotide-resolution genomic variation spanning intergenic and intronic regions creating a foundation for future sequencing-based investigations of glycemic traits.
Few studies have explored the impact of rare variants (minor allele frequency, MAF<1%) on highly heritable plasma metabolites identified in metabolomic screens. The Finnish population provides an ideal opportunity for such explorations, given the multiple bottlenecks and expansions that have shaped its history, and the enrichment for many otherwise rare alleles that has resulted. Here, we report genetic associations for 1,391 plasma metabolites in 6,136 men from the late-settlement region of Finland. We identify 303 novel association signals, more than one third at variants rare or enriched in Finns. Many of these signals identify genes not previously implicated in metabolite genome-wide association studies and suggest mechanisms for diseases and disease-related traits.