Pulmonary arterial hypertension (PAH) is characterised by pulmonary vascular remodelling causing premature death from right heart failure. Established DNA variants influence PAH risk, but susceptibility from epigenetic changes is unknown. We addressed this through epigenome-wide association study (EWAS), testing 865,848 CpG sites for association with PAH in 429 individuals with PAH and 1226 controls. Three loci, at Cathepsin Z (CTSZ, cg04917472), Conserved oligomeric Golgi complex 6 (COG6, cg27396197), and Zinc Finger Protein 678 (ZNF678, cg03144189), reached epigenome-wide significance (p < 10(-7)) and are hypermethylated in PAH, including in individuals with PAH at 1-year follow-up. Of 16 established PAH genes, only cg10976975 in BMP10 shows hypermethylation in PAH. Hypermethylation at CTSZ is associated with decreased blood cathepsin Z mRNA levels. Knockdown of CTSZ expression in human pulmonary artery endothelial cells increases caspase-3/7 activity (p < 10(-4)). DNA methylation profiles are altered in PAH, exemplified by the pulmonary endothelial function modifier CTSZ, encoding protease cathepsin Z.
Polycystic ovary syndrome (PCOS) is a very common endocrine condition in women in India. Gut microbiome alterations were shown to be involved in PCOS, yet it is remarkably understudied in Indian women who have a higher incidence of PCOS as compared to other ethnic populations. During the regional PCOS screening program among young women, we recruited 19 drug naive women with PCOS and 20 control women at the Sher-i-Kashmir Institute of Medical Sciences, Kashmir, North India. We profiled the gut microbiome in faecal samples by 16S rRNA sequencing and included 40/58 operational taxonomic units (OTUs) detected in at least 1/3 of the subjects with relative abundance (RA) ≥ 0.1%. We compared the RAs at a family/genus level in PCOS/non-PCOS groups and their correlation with 33 metabolic and hormonal factors, and corrected for multiple testing, while taking the variation in day of menstrual cycle at sample collection, age and BMI into account. Five genera were significantly enriched in PCOS cases: Sarcina, Megasphaera, and previously reported for PCOS Bifidobacterium, Collinsella and Paraprevotella confirmed by different statistical models. At the family level, the relative abundance of Bifidobacteriaceae was enriched, whereas Peptococcaceae was decreased among cases. We observed increased relative abundance of Collinsella and Paraprevotella with higher fasting blood glucose levels, and Paraprevotella and Alkalibacterium with larger hip, waist circumference, weight, and Peptococcaceae with lower prolactin levels. We also detected a novel association between Eubacterium and follicle-stimulating hormone levels and between Bifidobacterium and alkaline phosphatase, independently of the BMI of the participants. Our report supports that there is a relationship between gut microbiome composition and PCOS with links to specific reproductive health metabolic and hormonal predictors in Indian women.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
Background Multi-phenotype genome-wide association studies (MP-GWAS) of correlated traits have greater power to detect genotype–phenotype associations than single-trait GWAS. However, no multi-phenotype analysis method exists for epigenome-wide association studies (EWAS).Results We extended the SCOPA approach developed by us to “methylSCOPA” software in C++ by ‘reversely’ regressing DNA hyper/hypo-methylation information on a linear combination of phenotypes. We evaluated two models of association between DNA methylation and fasting glucose (FG) and insulin (FI) levels: Model 1, including FG, FI, and three measured potential confounders (body mass index [BMI], fasting serum triglyceride levels [TG], and waist/hip ratio [WHR]), and Model 2, including FG and FI corrected for the effects of BMI, TG, and WHR. Both models were additionally corrected for participant sex and smoking status (current/ever/never). We meta-analyzed the cohort-specific MP-EWAS results with our novel software META-methylSCOPA, mapped genomic locations to CGCh37/hg19, and adopted P <1×10−7 to denote epigenome-wide significance. We used the Illumina Infinium HumanMethylation450K BeadChip array data from the Northern Finland Birth Cohorts (NFBC) 1966/1986. We quality-controlled the data, regressed out the effects of measured potential confounders, and normalized the methylation signal intensity and FI data. The MP-EWAS included data for 643/457 individuals from NFBC1966 and NFBC1986, respectively (total N=1,100).In Model 1, we detected epigenome-wide significant association in the MP-EWAS meta-analysis at cg13708645 (chr12:121,974,305; P =1.2×10−8) within KDM2B gene. Single-trait effects within KDM2B were on FI, BMI, and WHR. Model with effect on BMI and WHR showed the strongest association at this locus, while effect on FI in single-phenotype analysis was driven by the effect of adiposity. In Model 2, the strongest association was at cg05063096 (chr3:143,689,810; P =2.3×10−7) annotated to C3orf58 with strongest effect on FI in single-trait analysis and multi-phenotype effect on FI and WHI within Model 1.We characterized the effects of established EWAS loci for diabetes and its risk factors and detected suggestive (p<0.01) associations at six markers including PHGDH, TXNIP, SLC7A11, CPT1A, MYO5C and ABCG1 , through the dissection of the multi-phenotype effects in Model 1.Conclusions We implemented MP-EWAS in methylSCOPA and demonstrated its enhanced power over single-trait EWAS for correlated phenotypes in large-scale data.
Metabolomics examines the small molecules involved in cellular metabolism. Approximately 50% of total phenotypic differences in metabolite levels is due to genetic variance, but heritability estimates differ across metabolite classes and lipid species. We performed a review of all genetic association studies, and identified > 800 class-specific metabolite loci that influence metabolite levels. In a twin-family cohort ( N = 5,117), these metabolite loci were leveraged to simultaneously estimate total heritability ( h 2 total ), and the proportion of heritability captured by known metabolite loci ( h 2 Metabolite-hits ) for 309 lipids and 52 organic acids. Our study revealed significant differences in h 2 Metabolite-hits among different classes of lipids and organic acids. Furthermore, phosphatidylcholines with a high degree of unsaturation had higher h 2 Metabolite-hits estimates than phosphatidylcholines with a low degree of unsaturation. This study highlights the importance of common genetic variants for metabolite levels, and elucidates the genetic architecture of metabolite classes and lipid species.
Telomere shortening has been associated with multiple age-related diseases such as cardiovascular disease, diabetes, and dementia. However, the biological mechanisms responsible for these associations remain largely unknown. In order to gain insight into the metabolic processes driving the association of leukocyte telomere length (LTL) with age-related diseases, we investigated the association between LTL and serum metabolite levels in 7,853 individuals from seven independent cohorts. LTL was determined by quantitative polymerase chain reaction and the levels of 131 serum metabolites were measured with mass spectrometry in biological samples from the same blood draw. With partial correlation analysis, we identified six metabolites that were significantly associated with LTL after adjustment for multiple testing: lysophosphatidylcholine acyl C17:0 (lysoPC a C17:0, p-value = 7.1 × 10 −6 ), methionine ( p-value = 9.2 × 10 −5 ), tyrosine ( p-value = 2.1 × 10 −4 ), phosphatidylcholine diacyl C32:1 (PC aa C32:1, p-value = 2.4 × 10 −4 ), hydroxypropionylcarnitine (C3-OH, p-value = 2.6 × 10 −4 ), and phosphatidylcholine acyl-alkyl C38:4 (PC ae C38:4, p-value = 9.0 × 10 −4 ). Pathway analysis showed that the three phosphatidylcholines and methionine are involved in homocysteine metabolism and we found supporting evidence for an association of lipid metabolism with LTL. In conclusion, we found longer LTL associated with higher levels of lysoPC a C17:0 and PC ae C38:4, and with lower levels of methionine, tyrosine, PC aa C32:1, and C3-OH. These metabolites have been implicated in inflammation, oxidative stress, homocysteine metabolism, and in cardiovascular disease and diabetes, two major drivers of morbidity and mortality.
We aimed to detect Attention-deficit/hyperactivity (ADHD) risk-conferring genes in adults. In children, ADHD is characterized by age-inappropriate levels of inattention and/or hyperactivity-impulsivity and may persists into adulthood. Childhood and adulthood ADHD are heritable, and are thought to represent the clinical extreme of a continuous distribution of ADHD symptoms in the general population. We aimed to leverage the power of studies of quantitative ADHD symptoms in adults who were genotyped. Within the SAGA (Study of ADHD trait genetics in adults) consortium, we estimated the single nucleotide polymorphism (SNP)-based heritability of quantitative self-reported ADHD symptoms and carried out a genome-wide association meta-analysis in nine adult population-based and case-only cohorts of adults. A total of n = 14,689 individuals were included. In two of the SAGA cohorts we found a significant SNP-based heritability for self-rated ADHD symptom scores of respectively 15% (n = 3656) and 30% (n = 1841). The top hit of the genome-wide meta-analysis (SNP rs12661753; p -value = 3.02 × 10 −7 ) was present in the long non-coding RNA gene STXBP5-AS1 . This association was also observed in a meta-analysis of childhood ADHD symptom scores in eight population-based pediatric cohorts from the Early Genetics and Lifecourse Epidemiology (EAGLE) ADHD consortium (n = 14,776). Genome-wide meta-analysis of the SAGA and EAGLE data (n = 29,465) increased the strength of the association with the SNP rs12661753. In human HEK293 cells, expression of STXBP5-AS1 enhanced the expression of a reporter construct of STXBP5 , a gene known to be involved in “SNAP” (Soluble NSF attachment protein) Receptor” (SNARE) complex formation. In mouse strains featuring different levels of impulsivity, transcript levels in the prefrontal cortex of the mouse ortholog Gm28905 strongly correlated negatively with motor impulsivity as measured in the five choice serial reaction time task (r 2 = − 0.61; p = 0.004). Our results are consistent with an effect of the STXBP5-AS1 gene on ADHD symptom scores distribution and point to a possible biological mechanism, other than antisense RNA inhibition, involved in ADHD-related impulsivity levels.
Background Identification of imprinted genes, demonstrating a consistent preference towards the paternal or maternal allelic expression, is important for the understanding of gene expression regulation during embryonic development and of the molecular basis of developmental disorders with a parent-of-origin effect. Combining allelic analysis of RNA-Seq data with phased genotypes in family trios provides a powerful method to detect parent-of-origin biases in gene expression. Results We report findings in 296 family trios from two large studies: 165 lymphoblastoid cell lines from the 1000 Genomes Project and 131 blood samples from the Genome of the Netherlands (GoNL) participants. Based on parental haplotypes, we identified > 2.8 million transcribed heterozygous SNVs phased for parental origin and developed a robust statistical framework for measuring allelic expression. We identified a total of 45 imprinted genes and one imprinted unannotated transcript, including multiple imprinted transcripts showing incomplete parental expression bias that was located adjacent to strongly imprinted genes. For example, PXDC1 , a gene which lies adjacent to the paternally expressed gene FAM50B , shows a 2:1 paternal expression bias. Other imprinted genes had promoter regions that coincide with sites of parentally biased DNA methylation identified in the blood from uniparental disomy (UPD) samples, thus providing independent validation of our results. Using the stranded nature of the RNA-Seq data in lymphoblastoid cell lines, we identified multiple loci with overlapping sense/antisense transcripts, of which one is expressed paternally and the other maternally. Using a sliding window approach, we searched for imprinted expression across the entire genome, identifying a novel imprinted putative lncRNA in 13q21.2. Overall, we identified 7 transcripts showing parental bias in gene expression which were not reported in 4 other recent RNA-Seq studies of imprinting. Conclusions Our methods and data provide a robust and high-resolution map of imprinted gene expression in the human genome.
Type 2 diabetes (T2D) is a global health burden that will benefit from personalised risk prediction and targeted prevention programmes. Omics data have enabled more detailed risk prediction; however, most studies have focussed on directly on the ability of DNA variants predicting T2D onset with less attention given to epigenetic regulation and glycaemic trait variability. By applying machine learning to the longitudinal Northern Finland Birth Cohort 1966 (NFBC 1966) at 31 (T1) and 46 (T2) years old, we predicted fasting glucose (FG) and insulin (FI), glycated haemoglobin (HbA1c) and 2-hour glucose and insulin from oral glucose tolerance test (2hGlu, 2hIns) at T2 in 513 individuals from 1,001 variables at T1 and T2, including anthropometric, metabolic, metabolomic and epigenetic variables. We further tested whether the information obtained by the machine learning models in NFBC could be used to predict glycaemic traits in the independent French study with 48 matching predictors (DESIR, N=769, age range 30-65 years at recruitment, interval between data collections: 9 years). In this study, FG and FI were best predicted, with average R2 values of 0.38 and 0.53. Sex, branched-chain and aromatic amino acids, HDL-cholesterol, glycerol, ketone bodies, blood pressure at T2 and measurements of adiposity at T1, as well as multiple methylation marks at both time points were amongst the top predictors. In the validation analysis, we reached R2 values of 0.41/0.55 for FG/FI when trained and tested in NFBC1966 and 0.17/0.30 when trained in NFBC1966 and tested in DESIR. We identified clinically relevant sets of predictors from a large multi-omics dataset and highlighted the potential of methylation markers and longitudinal changes in prediction.