Polygenic scores (PGSs) for body mass index (BMI) may guide early prevention and targeted treatment of obesity. Using genetic data from up to 5.1 million people (4.6% African ancestry, 14.4% American ancestry, 8.4% East Asian ancestry, 71.1% European ancestry and 1.5% South Asian ancestry) from the GIANT consortium and 23andMe, Inc., we developed ancestry-specific and multi-ancestry PGSs. The multi-ancestry score explained 17.6% of BMI variation among UK Biobank participants of European ancestry. For other populations, this ranged from 16% in East Asian-Americans to 2.2% in rural Ugandans. In the ALSPAC study, children with higher PGSs showed accelerated BMI gain from age 2.5 years to adolescence, with earlier adiposity rebound. Adding the PGS to predictors available at birth nearly doubled explained variance for BMI from age 5 onward (for example, from 11% to 21% at age 8). Up to age 5, adding the PGS to early-life BMI improved prediction of BMI at age 18 (for example, from 22% to 35% at age 5). Higher PGSs were associated with greater adult weight gain. In intensive lifestyle intervention trials, individuals with higher PGSs lost modestly more weight in the first year (0.55 kg per s.d.) but were more likely to regain it. Overall, these data show that PGSs have the potential to improve obesity prediction, particularly when implemented early in life.
Obesity is a heterogeneous condition not adequately captured by a single adiposity trait. We conducted a multi-trait genome-wide association analysis using individual-level data from 452,768 UK Biobank participants to study obesity in relation to cardiometabolic health. We defined continuous 'uncoupling phenotypes', ranging from high adiposity with healthy cardiometabolic profiles to low adiposity with unhealthy ones. We identified 266 variants across 205 genomic loci where adiposity-increasing alleles were simultaneously associated with lower cardiometabolic risk. A genetic risk score (GRSuncoupling) aggregating these variants was associated with a lower risk of cardiometabolic disorders, including dyslipidemia and ischemic heart disease, despite higher obesity risk; unlike an adiposity score based on body fat percentage-associated variants (GRSBFP). The 266 variants formed eight genetic subtypes of obesity, each with distinct risk profiles and pathway signatures. Proteomic analyses revealed signatures separating adiposity- and health-driven effects. Our findings reveal new mechanisms that uncouple obesity from cardiometabolic comorbidities and lay a foundation for genetically informed subtyping of obesity to support precision medicine.
Obesity is a major risk factor for a myriad of diseases, affecting >600 million people worldwide. Genome-wide association studies (GWASs) have identified hundreds of genetic variants that influence body mass index (BMI), a commonly used metric to assess obesity risk. Most variants are non-coding and likely act through regulating genes nearby. Here, we apply multiple computational methods to prioritize the likely causal gene(s) within each of the 536 previously reported GWAS-identified BMI-associated loci. We performed summary-data-based Mendelian randomization (SMR), FINEMAP, DEPICT, MAGMA, transcriptome-wide association studies (TWASs), mutation significance cutoff (MSC), polygenic priority score (PoPS), and the nearest gene strategy. Results of each method were weighted based on their success in identifying genes known to be implicated in obesity, ranking all prioritized genes according to a confidence score (minimum: 0; max: 28). We identified 292 high-scoring genes (≥11) in 264 loci, including genes known to play a role in body weight regulation (e.g., DGKI, ANKRD26, MC4R, LEPR, BDNF, GIPR, AKT3, KAT8, MTOR) and genes related to comorbidities (e.g., FGFR1, ISL1, TFAP2B, PARK2, TCF7L2, GSK3B). For most of the high-scoring genes, however, we found limited or no evidence for a role in obesity, including the top-scoring gene BPTF. Many of the top-scoring genes seem to act through a neuronal regulation of body weight, whereas others affect peripheral pathways, including circadian rhythm, insulin secretion, and glucose and carbohydrate homeostasis. The characterization of these likely causal genes can increase our understanding of the underlying biology and offer avenues to develop therapeutics for weight loss.
Oral azacitidine (oral-Aza) treatment results in longer median overall survival (OS) (24.7 vs. 14.8 months in placebo) in patients with acute myeloid leukemia (AML) in remission after intensive chemotherapy. The dosing schedule of oral-Aza (14 days/28-day cycle) allows for low exposure of Aza for an extended duration thereby facilitating a sustained therapeutic effect. However, the underlying mechanisms supporting the clinical impact of oral-Aza in maintenance therapy remain to be fully understood. In this preclinical work, we explore the mechanistic basis of oral-Aza/extended exposure to Aza through in vitro and in vivo modeling. In cell lines, extended exposure to Aza results in sustained DNMT1 loss, leading to durable hypomethylation, and gene expression changes. In mouse models, extended exposure to Aza, preferentially targets immature leukemic cells. In leukemic stem cell (LSC) models, the extended dose of Aza induces differentiation and depletes CD34+CD38- LSC. Mechanistically, LSC differentiation is driven in part by increased myeloperoxidase (MPO) expression. Inhibition of MPO activity either by using an MPO-specific inhibitor or blocking oxidative stress, a known mechanism of MPO, partly reverses the differentiation of LSC. Overall, our preclinical work reveals novel mechanistic insights into oral-Aza and its ability to target LSC.
Genome-wide association studies have identified over five hundred loci that contribute to variation in type 2 diabetes (T2D), an established risk factor for many diseases. However, the mechanisms and extent through which these loci contribute to subsequent outcomes remain elusive. We hypothesized that combinations of T2D-associated variants acting on tissue-specific regulatory elements might account for greater risk for tissue-specific outcomes, leading to diversity in T2D disease progression. We searched for T2D-associated variants acting on regulatory elements and expression quantitative trait loci (eQTLs) in nine tissues. We used T2D tissue-grouped variant sets as genetic instruments to conduct 2-Sample Mendelian Randomization (MR) in ten related outcomes whose risk is increased by T2D using the FinnGen cohort. We performed PheWAS analysis to investigate whether the T2D tissue-grouped variant sets had specific predicted disease signatures. We identified an average of 176 variants acting in nine tissues implicated in T2D, and an average of 30 variants acting on regulatory elements that are unique to the nine tissues of interest. In 2-Sample MR analyses, all subsets of regulatory variants acting in different tissues were associated with increased risk of the ten secondary outcomes studied on similar levels. No tissue-grouped variant set was associated with an outcome significantly more than other tissue-grouped variant sets. We did not identify different disease progression profiles based on tissue-specific regulatory and transcriptome information. Bigger sample sizes and other layers of regulatory information in critical tissues may help identify subsets of T2D variants that are implicated in certain secondary outcomes, uncovering system-specific disease progression.
Although physical activity and sedentary behavior are moderately heritable, little is known about the mechanisms that influence these traits. Combining data for up to 703,901 individuals from 51 studies in a multi-ancestry meta-analysis of genome-wide association studies yields 99 loci that associate with self-reported moderate-to-vigorous intensity physical activity during leisure time (MVPA), leisure screen time (LST) and/or sedentary behavior at work. Loci associated with LST are enriched for genes whose expression in skeletal muscle is altered by resistance training. A missense variant in ACTN3 makes the alpha-actinin-3 filaments more flexible, resulting in lower maximal force in isolated type IIA muscle fibers, and possibly protection from exercise-induced muscle damage. Finally, Mendelian randomization analyses show that beneficial effects of lower LST and higher MVPA on several risk factors and diseases are mediated or confounded by body mass index (BMI). Our results provide insights into physical activity mechanisms and its role in disease prevention.
A major challenge of genome-wide association studies (GWASs) is to translate phenotypic associations into biological insights. Here, we integrate a large GWAS on blood lipids involving 1.6 million individuals from five ancestries with a wide array of functional genomic datasets to discover regulatory mechanisms underlying lipid associations. We first prioritize lipid-associated genes with expression quantitative trait locus (eQTL) colocalizations and then add chromatin interaction data to narrow the search for functional genes. Polygenic enrichment analysis across 697 annotations from a host of tissues and cell types confirms the central role of the liver in lipid levels and highlights the selective enrichment of adipose-specific chromatin marks in high-density lipoprotein cholesterol and triglycerides. Overlapping transcription factor (TF) binding sites with lipid-associated loci identifies TFs relevant in lipid biology. In addition, we present an integrative framework to prioritize causal variants at GWAS loci, producing a comprehensive list of candidate causal genes and variants with multiple layers of functional evidence. We highlight two of the prioritized genes, CREBRF and RRBP1, which show convergent evidence across functional datasets supporting their roles in lipid biology.
Summary We present the results of the largest genome wide association study (GWAS) performed so far in dilated cardiomyopathy (DCM), a leading cause of systolic heart failure and cardiovascular death, with 2,719 cases and 4,440 controls in the discovery population. We identified and replicated two new DCM-associated loci, one on chromosome 3p25.1 (lead SNP rs62232870, p = 8.7 × 10 −11 and 7.7 × 10 −4 in the discovery and replication step, respectively) and the second on chromosome 22q11.23 (lead SNP rs7284877, p = 3.3 × 10 −8 and 1.4 × 10 −3 in the discovery and replication step, respectively) while confirming two previously identified DCM loci on chromosome 10 and 1, BAG3 and HSPB7 . The genetic risk score constructed from the number of lead risk-alleles at these four DCM loci revealed that individuals with 8 risk-alleles were at a 27% increased risk of DCM compared to individuals with 5 risk alleles (median of the referral population). We estimated the genome wide heritability at 31% ± 8%. In silico annotation and functional 4C-sequencing analysis on iPSC-derived cardiomyocytes strongly suggest SLC6A6 as the most likely DCM gene at the 3p25.1 locus. This gene encodes a taurine and beta-alanine transporter whose involvement in myocardial dysfunction and DCM is supported by recent observations in humans and mice. Although less easy to discriminate the better candidate at the 22q11.23 locus, SMARCB1 appears as the strongest one. This study provides both a better understanding of the genetic architecture of DCM and new knowledge on novel biological pathways underlying heart failure, with the potential for a therapeutic perspective.
Even though physical activity and sedentary behavior are moderately heritable, little is known about the mechanisms that influence these traits. Here, we combine data for up to 674,980 individuals from 51 studies in a trans-ancestry meta-analysis of genome-wide association studies for self-reported moderate-to-vigorous intensity physical activity during leisure time (MVPA); leisure screen time (LST); sedentary commuting; and sedentary behavior at work. We identify 99 loci that associate with at least one trait. Loci associated with LST are enriched for genes whose expression in skeletal muscle is altered by resistance training. Molecular dynamics simulations suggest that the Glu to Ala substitution encoded by rs2229456 (ACTN3) – associated with more MVPA – disrupts salt bridge interactions and makes the alpha actinin 3 filaments more flexible. In isolated type IIA muscle fibers, the Ala-encoding allele is associated with lower maximal force and power during an isometric contraction, suggesting protection from exercise-induced muscle damage. Finally, Mendelian Randomization analyses show that the causal effect of LST on BMI is 2-3 times larger than the effect of body mass index (BMI) on LST, and that beneficial effects of LST and MVPA on several risk factors and diseases are mediated or confounded by BMI. Taken together, our results provide mechanistic insights into the regulation of MVPA and into the role of LST and MVPA in disease prevention. These insights may facilitate the development of tailored physical activity interventions.
We present the results of the largest genome wide association study (GWAS) performed so far in dilated cardiomyopathy (DCM), a leading cause of systolic heart failure and cardiovascular death, with 2,719 cases and 4,440 controls in the discovery population. We identified and replicated two new DCM-associated loci, one on chromosome 3p25.1 (lead SNP rs62232870, p = 8.7 × 10−11 and 7.7 × 10−4 in the discovery and replication step, respectively) and the second on chromosome 22q11.23 (lead SNP rs7284877, p = 3.3 × 10−8 and 1.4 × 10−3 in the discovery and replication step, respectively) while confirming two previously identified DCM loci on chromosome 10 and 1, BAG3 and HSPB7. The genetic risk score constructed from the number of lead risk-alleles at these four DCM loci revealed that individuals with 8 risk-alleles were at a 27% increased risk of DCM compared to individuals with 5 risk alleles (median of the referral population). We estimated the genome wide heritability at 31% ± 8%. In silico annotation and functional 4C-sequencing analysis on iPSC-derived cardiomyocytes strongly suggest SLC6A6 as the most likely DCM gene at the 3p25.1 locus. This gene encodes a taurine and beta-alanine transporter whose involvement in myocardial dysfunction and DCM is supported by recent observations in humans and mice. Although less easy to discriminate the better candidate at the 22q11.23 locus, SMARCB1 appears as the strongest one. This study provides both a better understanding of the genetic architecture of DCM and new knowledge on novel biological pathways underlying heart failure, with the potential for a therapeutic perspective.
A typical task arising from main effect analyses in a Genome Wide Association Study (GWAS) is to identify single nucleotide polymorphisms (SNPs), in linkage disequilibrium with the observed signals, that are likely causal variants and the affected genes. The affected genes may not be those closest to associating SNPs. Functional genomics data from relevant tissues are believed to be helpful in selecting likely causal SNPs and interpreting implicated biological mechanisms, ultimately facilitating prevention and treatment in the case of a disease trait. These data are typically used post GWAS analyses to fine-map the statistically significant signals identified agnostically by testing all SNPs and applying a multiple testing correction. The number of tested SNPs is typically in the millions, so the multiple testing burden is high. Motivated by this, in this study we investigated an alternative workflow, which consists in utilizing the available functional genomics data as a first step to reduce the number of SNPs tested for association. We analyzed GWAS on electrocardiographic QRS duration using these two workflows. The alternative workflow identified more SNPs, including some residing in loci not discovered with the typical workflow. Moreover, the latter are corroborated by other reports on QRS duration. This indicates the potential value of incorporating functional genomics information at the onset in GWAS analyses.
Background: Regulatory elements may be involved in the mechanisms by which 52 loci influence myocardial mass, reflected by abnormal amplitude and duration of the QRS complex on the ECG. Functional annotation thus far did not take into account how these elements are affected in disease context. Methods: We generated maps of regulatory elements on hypertrophic cardiomyopathy patients (ChIP-seq N=14 and RNA-seq N=11) and nondiseased hearts (ChIP-seq N=4 and RNA-seq N=11). We tested enrichment of QRS-associated loci on elements differentially acetylated and directly regulating differentially expressed genes between hypertrophic cardiomyopathy patients and controls. We further performed functional annotation on QRS-associated loci using these maps of differentially active regulatory elements. Results: Regions differentially affected in disease showed a stronger enrichment ( P =8.6×10 −5 ) for QRS-associated variants than those not showing differential activity ( P =0.01). Promoters of genes differentially regulated between hypertrophic cardiomyopathy patients and controls showed more enrichment ( P =0.001) than differentially acetylated enhancers ( P =0.8) and super-enhancers ( P =0.025). We also identified 74 potential causal variants overlapping these differential regulatory elements. Eighteen of the genes mapped confirmed previous findings, now also pinpointing the potentially affected regulatory elements and candidate causal variants. Fourteen new genes were also mapped. Conclusions: Our results suggest differentially active regulatory elements between hypertrophic cardiomyopathy patients and controls can offer more insights into the mechanisms of QRS-associated loci than elements not affected by disease.
Genome-wide association studies (GWAS) of quantitative electrocardiographic (ECG) traits in large consortia have identified more than 130 loci associated with QT interval, QRS duration, PR interval, and heart rate (RR interval). In the current study, we meta-analyzed genome-wide association results from 30,000 mostly Dutch samples on four ECG traits: PR interval, QRS duration, QT interval, and RR interval. SNP genotype data was imputed using the Genome of the Netherlands reference panel encompassing 19 million SNPs, including millions of rare SNPs (minor allele frequency < 5%). In addition to many known loci, we identified seven novel locus-trait associations: KCND3, NR3C1, and PLN for PR interval, KCNE1, SGIP1, and NFKB1 for QT interval, and ATP2A2 for QRS duration, of which six were successfully replicated. At these seven loci, we performed conditional analyses and annotated significant SNPs (in exons and regulatory regions), demonstrating involvement of cardiac-related pathways and regulation of nearby genes.
Cardiovascular diseases (CVDs) are the leading cause of death in the world. Genome-wide association (GWAS) studies have identified many genetic loci robustly associated to CVDs. Because most CVD-associated loci are non-coding, one of the main challenges in the post-GWAS era is interpretation of these statistical signals. This thesis presents bioinformatics applications that integrate genome, regulome and transcriptome information to address this challenge. Integrative approaches such as the ones presented in this thesis can help expand our knowledge of the biological mechanisms involved in CVDs, which in turn can be translated into better prevention and treatment.
Worldwide over 5 million children have been conceived using assisted reproductive technology, and research has concentrated on increasing the likelihood of ongoing pregnancy. However, studies using animal models have indicated undesirable effects of in vitro embryo culture on offspring development and health. In vivo, the oviduct hosts a period in which the early embryo undergoes complete reprogramming of its (epi) genome in preparation for the reacquisition of (epi) genetic marks. We designed an oviduct-on-a-chip platform to better investigate the mechanisms related to (epi) genetic reprogramming and the degree to which they differ between in vitro and in vivo embryos. The device supports more physiological (in vivo-like) zygote genetic reprogramming than conventional IVF. This approach will be instrumental in identifying and investigating factors critical to fertilization and preimplantation development, which could improve the quality and (epi) genetic integrity of IVF zygotes with likely relevance for early embryonic and later fetal development.
Introduction: Cell damage causes the release of a significant amount of ATP and subsequent activation of purinergic system.Purinergic signaling is involved in the regulation of a variety of physiological processes, including inflammation, cell proliferation and migration.A specific set of activated pathways is mainly determined by the expression of target receptors and ectonucleotidases, which regulate levels of extracellular ATP and its metabolic products, primarily adenosine.Clarification of the role of purinergic signaling in the vascular wall response to injury is important for understanding of the fine mechanisms of atherosclerotic plaque formation and restenosis development.Purpose: The aim of the study was to determine temporal expression profiles of genes involved in purinergic signaling during the healing response to arterial injury in rat.Methods: Total RNA was isolated from the rat carotid arteries at seven timepoints ranging from 2 hours to 12 weeks after balloon injury.Transcriptome profiling was performed using microarrays.Results: Several ectonucleotidases, participating in the balance of extracellular ATP and its metabolic products, showed differential expression.In particular, CD39 catabolizing pro-inflammatory ATP to ADP and AMP, decreased from 20 hours of observation.CD73, which metabolizes AMP to anti-inflammatory adenosine, and adenosine deaminase, which destructs extracellular adenosine, increased at 20 hours and remain upregulated until day 2 and day 5, respectively.Several purinergic receptors of ATP and ADP from P2X and P2Y families demonstrate differential expression.Specifically, P2X1, known as regulator of vSMC contractile phenotype, decreased from 20 hours of observation, while P2Y2 and P2Y6, known as regulators of vSMC synthetic phenotype, upregulated immediately after injury.Low-affinity adenosine receptors A2b and A3 were upregulated from day 2 to day 5. Several downstream targets of purinergic signaling also showed differential expression.Among them protein kinase A, involved in vSMC proliferation, and genes associated with inflammatory response -NLRP3 subunit of inflammasome and its targets IL1b and IL18, as well as adhesion molecules ICAM1 and VCAM1.Conclusions: Analysis of time-course expression profiles demonstrated that purinergic signaling pathways are dynamically controlled at mRNA level in rat carotid artery balloon injury model.This indicates the potential involvement of purinergic signaling in the regulation of local inflammation and vSMC phenotype during response to vascular injury.
In the last decade, over 175 genetic loci have robustly been associated to levels of major circulating blood lipids. Most loci are specific to one or two lipids, whereas some (SUGP1, ZPR1, TRIB1, HERPUD1, and FADS1) are associated to all. While exposing the polygenic architecture of circulating lipids and the underpinnings of dyslipidaemia, these genome-wide association studies (GWAS) have provided further evidence of the critical role that lipids play in coronary heart disease (CHD) risk, as indicated by the 2.7-fold enrichment for macrophage gene expression in atherosclerotic plaques and the association of 25 loci (such as PCSK9, APOB, ABCG5-G8, KCNK5, LPL, HMGCR, NPC1L1, CETP, TRIB1, ABO, PMAIP1-MC4R, and LDLR) with CHD. These GWAS also confirmed known and commonly used therapeutic targets, including HMGCR (statins), PCSK9 (antibodies), and NPC1L1 (ezetimibe). As we head into the post-GWAS era, we offer suggestions for how to move forward beyond genetic risk loci, towards refining the biology behind the associations and identifying causal genes and therapeutic targets. Deep phenotyping through lipidomics and metabolomics will refine and increase the resolution to find causal and druggable targets, and studies aimed at demonstrating gene transcriptional and regulatory effects of lipid associated loci will further aid in identifying these targets. Thus, we argue the need for deeply phenotyped, large genetic association studies to reduce costs and failures and increase the efficiency of the drug discovery pipeline. We conjecture that in the next decade a paradigm shift will tip the balance towards a data-driven approach to therapeutic target development and the application of precision medicine where human genomics takes centre stage.
BACKGROUND:Genome-wide association studies have identified multiple loci associated with coronary artery disease and myocardial infarction, but only a few of these loci are current targets for on-market medications. To identify drugs suitable for repurposing and their targets, we created 2 unique pipelines integrating public data on 49 coronary artery disease/myocardial infarction-genome-wide association studies loci, drug-gene interactions, side effects, and chemical interactions. METHODS:We first used publicly available genome-wide association studies results on all phenotypes to predict relevant side effects, identified drug-gene interactions, and prioritized candidates for repurposing among existing drugs. Second, we prioritized gene product targets by calculating a druggability score to estimate how accessible pockets of coronary artery disease/myocardial infarction-associated gene products are, then used again the genome-wide association studies results to predict side effects, excluded loci with widespread cross-tissue expression to avoid housekeeping and genes involved in vital processes and accordingly ranked the remaining gene products. RESULTS:These pipelines ultimately led to 3 suggestions for drug repurposing: pentolinium, adenosine triphosphate, and riociguat (to target CHRNB4, ACSS2, and GUCY1A3, respectively); and 3 proteins for drug development: LMOD1 (leiomodin 1), HIP1 (huntingtin-interacting protein 1), and PPP2R3A (protein phosphatase 2, regulatory subunit b-double prime, α). Most current therapies for coronary artery disease/myocardial infarction treatment were also rediscovered. CONCLUSIONS:Integration of genomic and pharmacological data may prove beneficial for drug repurposing and development, as evidence from our pipelines suggests.
High blood pressure or hypertension is an established risk factor for a myriad of cardiovascular diseases. Genome-wide association studies have successfully found over nine hundred loci that contribute to blood pressure. However, the mechanisms through which these loci contribute to disease are still relatively undetermined as less than 10% of hypertension-associated variants are located in coding regions. Phenotypic cell-type specificity analyses and expression quantitative trait loci show predominant vascular and cardiac tissue involvement for blood pressure-associated variants. Maps of chromosomal conformation and expression quantitative trait loci (eQTL) in critical tissues identified 2,424 genes interacting with blood pressure-associated loci, of which 517 are druggable. Integrating genome, regulome and transcriptome information in relevant cell-types could help to functionally annotate blood pressure associated loci and identify drug targets.
Nature Communications 8: Article number: 15805 (2017); Published: 14 June 2017; Updated: 2 August 2017 In Supplementary Fig. 10 of this Article, images for panels a and b were inadvertently omitted. The correct version of Supplementary Fig. 10 is provided as Supplementary Information associated withthis Erratum.