Genetic regulation of DNA methylation in immune cells may mediate complex disease risk. However, current epigenomic studies are constrained by microarray CpG coverage, mixed-cell tissues, and limited representation of diverse ancestries. Thus, we generated a whole-genome, multi-ancestry atlas of genetic effects on the purified monocyte methylome. We first performed whole-genome bisulfite sequencing (WGBS) of purified peripheral blood monocytes and whole-genome sequencing (WGS) from 160 African American (AA) and 298 European American (EA) participants, profiling around 25 million CpG sites. Next, we identified cis-methylation quantitative trait loci (meQTLs), estimated cis-heritability, and evaluated replication against large external meQTL resources. We further trained population-specific DNAm imputation models and applied them to methylome-wide association studies (MWAS) of 41 traits using genome-wide association study summary statistics from the Million Veteran Program. Type 2 diabetes signals were further evaluated using Mendelian randomization and Bayesian colocalization. We also conducted exploratory trans-meQTL mapping. We identified 1,480,064 and 1,527,480 CpG sites with at least one cis-meQTL in AA and EA populations, respectively, including 543,869 shared sites and extensive population-specific regulation attributable to both allele-frequency differences and effect-size heterogeneity. Cis-meQTL effects replicated robustly in external datasets: effect sizes correlated strongly with prior studies (EA Pearson’s r = 0.76; 90.8
ABSTRACT Objective A number of susceptibility genes in prostate tissue have been identified to be associated with prostate cancer (PCa) risk. However, the reported genes based on assessing prostate tissue could not fully explain PCa genetic susceptibility. It is believed that genes functioning in the immune system may fill in the gap of some missing heritability. Methods To study potential susceptibility genes acting in such pathways, we performed a transcriptome‐wide association study (TWAS) of 79,194 PCa cases and 61,112 control of European ancestry by using three sets of gene expression prediction models of blood tissue. Results A total of 470 genes were associated at false discovery rates‐corrected p ‐value < 0.05, of which 51 were implicated as likely causal genes based on fine‐mapping analysis. Compared with previous literature, 133 novel genes were reported for the first time. Of the identified genes, five ( CREB3L4 , GSTP1 , MAPK3 , NKX3‐ 1, and PIK3C2B ) were enriched in a PCa signaling pathway, and 128 genes were enriched in five PCa categories. Importantly, 13 genes ( SCP2 , LMNA , ZNF148 , H2AFV , TACC1 , FLII , SUPT4H1 , CD300LF , MYO9B , COX6B1 , CTSA , EP300 , and TSPO ) showed consistent effect directions for the measured levels in circulating immune cells between PCa cases and controls, and 14 genes ( SLC39A1 , ZBTB7B , TRIM59 , NCEH1 , N4BP2 , TAGAP , TACC1 , TRAF1 , AIP , SECTM1 , C18orf54 , ZNF793 , YIF1B , and TSPO ) showed consistency for levels in blood exosomes between PCa patients and controls. Conclusion The identified blood‐based candidate susceptibility genes provide further insights into the genetic basis of PCa risk.
Proteome-wide association studies (PWAS) decode the intricate proteomic landscape of biological mechanisms for complex diseases. Traditional PWAS model training relies heavily on individual-level reference proteomes, thereby restricting its capacity to harness the emerging summary-level protein quantitative trait loci (pQTL) data in the public domain. Here we introduced a novel framework to train PWAS models directly from pQTL summary statistics. By leveraging extensive pQTL data from the UK Biobank, deCODE, and ARIC studies, we applied our approach to train large-scale European PWAS models (total n = 88,838 subjects). Furthermore, we developed PWAS models tailored for Asian and African ancestries by integrating multi-ancestry summary and individual-level data resources (total n = 914 for Asian and 3,042 for African ancestries). We validated the performance of our PWAS models through a systematic multi-ancestry analysis of over 700 phenotypes across five major genetic data resources. Our results bridge the gap between genomics and proteomics for drug discovery, highlighting novel protein-phenotype links and their transferability across diverse ancestries. The developed PWAS models and data resources are freely available at www.gcbhub.org .
Transcriptome-wide association studies (TWAS) integrate gene expression prediction models and genome-wide association studies (GWAS) to identify gene-trait associations. The power of TWAS is determined by the sample size of GWAS and the accuracy of the expression prediction model. Here, we present a new method, the Summary-level Unified Method for Modeling Integrated Transcriptome using Functional Annotations (SUMMIT-FA), which improves gene expression prediction accuracy by leveraging functional annotation resources and a large expression quantitative trait loci (eQTL) summary-level dataset. We build gene expression prediction models in whole blood using SUMMIT-FA with the comprehensive functional database MACIE and eQTL summary-level data from the eQTLGen consortium. We apply these models to GWAS for 24 complex traits and show that SUMMIT-FA identifies significantly more gene-trait associations and improves predictive power for identifying "silver standard" genes compared to several benchmark methods. We further conduct a simulation study to demonstrate the effectiveness of SUMMIT-FA.
BACKGROUND:Specific peripheral proteins have been implicated to play an important role in the development of Alzheimer's disease (AD). However, the roles of additional novel protein biomarkers in AD etiology remains elusive. The availability of large-scale AD GWAS and plasma proteomic data provide the resources needed for the identification of causally relevant circulating proteins that may serve as risk factors for AD and potential therapeutic targets. METHODS:We established and validated genetic prediction models for protein levels in plasma as instruments to investigate the associations between genetically predicted protein levels and AD risk. We studied 71,880 (proxy) cases and 383,378 (proxy) controls of European descent. RESULTS:We identified 69 proteins with genetically predicted concentrations showing associations with AD risk. The drugs almitrine and ciclopirox targeting ATP1A1 were suggested to have a potential for being repositioned for AD treatment. CONCLUSIONS:Our study provides additional insights into the underlying mechanisms of AD and potential therapeutic strategies.
Although DNA methylation (DNAm) has been implicated in the pathogenesis of numerous complex diseases, from cancer to cardiovascular disease to autoimmune disease, the exact methylation sites that play key roles in these processes remain elusive. One strategy to identify putative causal CpG sites and enhance disease etiology understanding is to conduct methylome-wide association studies (MWASs), in which predicted DNA methylation that is associated with complex diseases can be identified. However, current MWAS models are primarily trained using the data from single studies, thereby limiting the methylation prediction accuracy and the power of subsequent association studies. Here, we introduce a new resource, MWAS Imputing Methylome Obliging Summary-level mQTLs and Associated LD matrices (MIMOSA), a set of models that substantially improve the prediction accuracy of DNA methylation and subsequent MWAS power through the use of a large summary-level mQTL dataset provided by the Genetics of DNA Methylation Consortium (GoDMC). Through the analyses of GWAS (genome-wide association study) summary statistics for 28 complex traits and diseases, we demonstrate that MIMOSA considerably increases the accuracy of DNA methylation prediction in whole blood, crafts fruitful prediction models for low heritability CpG sites, and determines markedly more CpG site-phenotype associations than preceding methods. Finally, we use MIMOSA to conduct a case study on high cholesterol, pinpointing 146 putatively causal CpG sites.
Prostate cancer (PCa) represents a huge public health burden among men. Many susceptibility genetic factors for PCa still remain unknown. In this study, we performed a large splicing transcriptome-wide association study (spTWAS) using three modeling strategies to develop alternative splicing genetic prediction models for identifying novel susceptibility loci and splicing introns for PCa risk by assessing 79,194 cases and 61,112 controls of European ancestry in the PRACTICAL, CRUK, CAPS, BPC3, and PEGASUS consortia. We identified 120 splicing introns of 97 genes showing an association with PCa risk at false discovery rate (FDR)-corrected threshold (FDR <0.05). Of them, 33 genes were enriched in PCa-related diseases and function categories. Fine-mapping analysis suggested that 21 splicing introns of 19 genes were likely causally associated with PCa risk. Thirty-five splicing introns of 34 novel genes were identified to be related to PCa susceptibility for the first time, and 11 of the genes were enriched in a cancer-related network. Our study identified novel loci and splicing introns associated with PCa risk, which can improve our understanding of the etiology of this common malignancy.
Alzheimer's disease (AD) is a common neurodegenerative disease in aging individuals. Alternative splicing is reported to be relevant to AD development while their roles in etiology of AD remain largely elusive. We performed a comprehensive splicing transcriptome-wide association study (spTWAS) using intronic excision expression genetic prediction models of 12 brain tissues developed through three modelling strategies, to identify candidate susceptibility splicing introns for AD risk. A total of 111,326 (46,828 proxy) cases and 677,663 controls of European ancestry were studied. We identified 343 associations of 233 splicing introns (143 genes) with AD risk after Bonferroni correction (0.05/136,884 = 3.65 × 10−7). Fine-mapping analyses supported 155 likely causal associations corresponding to 83 splicing introns of 55 genes. Eighteen causal splicing introns of 15 novel genes (EIF2D, WDR33, SAP130, BYSL, EPHB6, MRPL43, VEGFB, PPP1R13B, TLN2, CLUHP3, LRRC37A4P, CRHR1, LINC02210, ZNF45-AS1, and XPNPEP3) were identified for the first time to be related to AD susceptibility. Our study identified novel genes and splicing introns associated with AD risk, which can improve our understanding of the etiology of AD.
Alzheimer disease (AD) is a common neurodegenerative disease with a late onset. It is critical to identify novel blood-based DNA methylation biomarkers to better understand the extent of the molecular pathways affected in AD. Two sets of blood DNA methylation genetic prediction models developed using different reference panels and modelling strategies were leveraged to evaluate associations of genetically predicted DNA methylation levels with AD risk in 111,326 (46,828 proxy) cases and 677,663 controls. A total of 1,168 cytosine-phosphate-guanine (CpG) sites showed a significant association with AD risk at a false discovery rate (FDR) < 0.05. Methylation levels of 196 CpG sites were correlated with expression levels of 130 adjacent genes in blood. Overall, 52 CpG sites of 32 genes showed consistent association directions for the methylation-gene expression-AD risk, including nine genes ( CNIH4 , THUMPD3 , SERPINB9 , MTUS1 , CISD1 , FRAT2 , CCDC88B , FES , and SSH 2) firstly reported as AD risk genes. Nine of 32 genes were enriched in dementia and AD disease categories ( P values ranged from 1.85 × 10 -4 to 7.46 × 10 -6 ), and 19 genes in a neurological disease network (score = 54) were also observed. Our findings improve the understanding of genetics and etiology for AD.
Genes with moderate to low expression heritability may explain a large proportion of complex trait etiology, but such genes cannot be sufficiently captured in conventional transcriptome-wide association studies (TWASs), partly due to the relatively small available reference datasets for developing expression genetic prediction models to capture the moderate to low genetically regulated components of gene expression. Here, we introduce a method, the Summary-level Unified Method for Modeling Integrated Transcriptome (SUMMIT), to improve the expression prediction model accuracy and the power of TWAS by using a large expression quantitative trait loci (eQTL) summary-level dataset. We apply SUMMIT to the eQTL summary-level data provided by the eQTLGen consortium. Through simulation studies and analyses of genome-wide association study summary statistics for 24 complex traits, we show that SUMMIT improves the accuracy of expression prediction in blood, successfully builds expression prediction models for genes with low expression heritability, and achieves higher statistical power than several benchmark methods. Finally, we conduct a case study of COVID-19 severity with SUMMIT and identify 11 likely causal genes associated with COVID-19 severity.