The scPrediXcan framework enables cell-type-specific transcriptome-wide association studies (TWASs) by integrating deep learning-based prediction of gene expression from DNA sequence and epigenetic features. We present a protocol for scPrediXcan: training cell-type-specific models for expression prediction, predicting personalized expression, and testing associations with genome-wide association study (GWAS) summary statistics. This framework produces scalable TWAS models for different cellular contexts with minimal computational burden.For complete details on the use and execution of this protocol, please refer to Zhou et al.1
BACKGROUNDWe constructed multi-trait polygenic risk scores (PRSs) predicting chronic obstructive pulmonary disease (COPD) and exacerbations, validated their performance in diverse cohorts, and identified PRS-related proteins for potential therapeutic targeting.METHODSPRSmix+, a multi-trait PRS framework, is used to train a composite PRS (PRSmulti) in COPDGene non-Hispanic White participants (n = 6,647). Associations of PRSmulti with COPD status (GOLD 2-4 vs. GOLD 0 or ICD) and exacerbation frequency were tested in COPDGene African American (n = 2,466), ECLIPSE (n = 1,858), Mass General Brigham Biobank (n = 15,152), and All of Us (n = 118,566). Protein prediction models were applied to GWAS summary statistics from traits contributing to PRSmulti and were validated with proteomic data in COPDGene (n = 5,173) and UK Biobank (n = 5,012).RESULTSPRSmix+ selected 7 traits for PRSmulti. In multivariable models, PRSmulti was associated with COPD status (meta-analysis random effects [RE] OR 1.58 [95% CI: 1.28-1.94]) and exacerbation frequency (meta-analysis RE β 0.21 [95% CI: 0.11-0.31]), with higher effect sizes observed in smoking-enriched cohorts. PRSmulti outperformed traditional single-trait PRS in all tested cohorts. Using protein prediction models, we identified 73 proteins associated with the PRSs that were also validated with measured protein levels in COPDGene and UK Biobank. Of these proteins, 25 were linked to approved or investigational drugs. Notable targets include RAGE/sRAGE, IL1RL1, and SCARF2, all implicated in COPD pathogenesis and exacerbations.CONCLUSIONSMulti-trait PRS improves prediction of COPD and exacerbation risk. Integration with proteomic data identifies druggable protein targets, offering a promising avenue for precision medicine in COPD management.TRIAL REGISTRATIONCOPDGene: ClinicalTrials.gov NCT00608764; ECLIPSE: ClinicalTrials.gov NCT00292552.
Transcriptome-wide association studies (TWASs) and related methods (xWASs) have been widely adopted in genetic studies to understand molecular traits as mediators between genetic variation and disease. However, the effect of polygenicity on the validity of these mediator-trait association tests has largely been overlooked. Given the widespread polygenicity of complex traits, it is necessary to assess the accuracy of these mediator-trait association tests. We found that, for highly polygenic target traits, the standard test based on linear regression is inflated, leading to dramatically increased false-positive rates that grow linearly with sample size and heritability. To address this inflation, we propose an effective variance-control method—similar to genomic control but allowing for a different correction factor for each gene. Using simulated and real data, as well as theoretical derivations, we show that our method yields calibrated false-positive rates, outperforming existing approaches. We further demonstrate that methods analogous to TWASs, namely those that associate genetic predictors of mediating traits with target traits, suffer from similar inflation issues. We advise developers of genetic predictors for molecular traits (including polygenic risk scores, PRSs) to compute and provide the necessary inflation parameters to ensure proper false-positive control. Finally, we have updated our PrediXcan software package and resources to facilitate this correction for end users.
Introduction:Coronary artery disease (CAD) is a leading cause of death and disability worldwide. Although genome-wide association studies (GWAS) have identified over 300 loci associated with CAD risk, the molecular mechanisms linking these variants to disease and subclinical atherosclerosis are not fully understood. Methods:We performed integration of multi-ancestry CAD GWAS with transcriptomic data from the Multi-Ethnic Study of Atherosclerosis (MESA) obtained through the Trans-Omics for Precision Medicine (TOPMed) program. For integration, we applied Bayesian colocalization analysis with and without statistical fine-mapping to identify genes whose expression levels colocalize with CAD-associated loci. We further applied causal weighted gene co-expression network analysis (cWGCNA) to identify gene co-expression modules and key driver genes associated with subclinical atherosclerosis traits in MESA. Results:We identified 108 genes showing evidence of colocalization with CAD loci, including 24 shared between the two colocalization approaches and 48 novel genes not previously reported in CAD GWAS. Follow-up replication and validation analyses prioritized 5 novel ( CCDC30, ZEB1-AS1, ZPR1, PLEKHJ1 and AC018816.3 ) and 8 previously reported genes ( DHDDS, DDX59, LNPEP, DAGLA, ZKSCAN1, LIPA, OPRL1 and EIF2B2 ) with putative roles in both CAD and subclinical atherosclerosis. cWGCNA identified five gene modules significantly associated with subclinical atherosclerosis in MESA. Additionally, three key driver genes ( ATG9B, PRAM1 and ZBTB46 ) identified by cWGCNA were also identified as CAD-colocalized genes. Discussion:Our integrative analysis highlights key genetic drivers and regulatory networks underlying CAD and subclinical atherosclerosis. These findings underscore the value of incorporating statistical fine-mapping in colocalization studies and demonstrate the utility of combining colocalization with co-expression network analysis to prioritize functional genes and pathways.
Reliable reference transcriptome prediction models are key to accurate multi-ancestry transcriptome-wide association studies (TWASs). We propose three methods leveraging functionally informed variants (FIVs) for transcriptome prediction models to improve multi-ancestry TWASs. We trained models on 1,287 multi-ancestry participants from the Trans-Omics for Precision Medicine (TOPMed) program Multi-Ethnic Study of Atherosclerosis (MESA) with RNA sequencing (RNA-seq) data from peripheral blood mononuclear cells (PBMCs). We validated models’ prediction accuracy on two external independent datasets, Geuvadis and Jackson Heart Study. To test robustness of our methods for TWASs, we integrated models with three multi-ancestry GWASs from blood cell, lipid, and pulmonary function traits, respectively. Our methods presented similar prediction accuracy while using a smaller and functionally informed set of variants compared to the benchmark method, elastic net (EN). Overall, our methods achieved higher power and accuracy (with average improved accuracy of 24% over EN) for TWASs. However, no single proposed method outperformed all GWAS traits. To further improve TWAS performance, we propose an omnibus approach that aggregates TWAS summary statistics from our methods. The omnibus approach yielded the highest number of Bonferroni-significant TWAS genes for all GWAS traits, and it further improved TWAS power and accuracy for blood cell traits. Additionally, the omnibus approach detected some trait-relevant important genes that the EN missed. Our study demonstrates the value of including FIVs in multi-ancestry transcriptome prediction models for improving TWAS performance. Further, the observed TWAS improvement depends on the GWAS trait’s relevance to the PBMCs used to build our transcriptome prediction models.
Proteomic predictive models are predominantly trained on cis-acting variants in European-ancestry cohorts, limiting power and predictive accuracy in ancestrally diverse populations. We performed cis- and trans-protein quantitative trait locus (pQTL) mapping and developed protein-prediction models using whole-genome sequencing (WGS) and plasma protein levels (Olink) across four ancestry groups from the Trans-omics for Precision Medicine (TOPMed) Multi-Ethnic Study of Atherosclerosis (MESA): European (EUR, n=1270), African (AFR, n=675), Hispanic (HIS, n=642), and Chinese (CHN, n=366), and a combined population (ALL, n=2953). African-ancestry samples demonstrated improved fine-mapping resolution relative to cohort size, yielding significantly smaller cis-credible sets than European-ancestry samples, consistent with shorter linkage disequilibrium (LD) blocks and greater allele frequency diversity in African-ancestry populations. For the first time, we benchmarked fine-mapping models SuSiE, SuShiE, MultiSuSiE, and SuSiEx with multi-ancestral cohorts, revealing a precision-recall tradeoff driven by model assumptions. Comparing protein-prediction models, multivariate adaptive shrinkage (MASHR) and ultimate deconvolution in R (UDR) outperformed elastic net (EN) regression, with trans-pQTL inclusion and fine-mapping improving prediction performance and proteome-wide association study (PWAS) discovery. Applying our models in PWAS of 10 phenotypes, we discovered 68 protein-phenotype associations in All of Us (AoU) that also replicated in Pan-UK Biobank. MASHR and UDR models identified 60% more protein-phenotype associations than EN. Notably, 32 of these associations were not previously reported in the GWAS Catalog. Overall, our study demonstrates the importance of including multiple ancestries in genomic studies to capture the full spectrum of regulatory variation and improve cross-ancestry generalizability.
Heart failure (HF) is a leading global cause of morbidity and mortality, yet the regulatory molecular mechanisms that link genetic variation to cardiac dysfunction remain elusive. To bridge this gap, we created the Trans-Omics for Precision Medicine in Congestive Heart Failure (TOPCHeF) resource, a multi-omics dataset comprising >700 human left-ventricular tissue samples, including dilated cardiomyopathy (DCM), ischemic cardiomyopathy (ICM), and non-failing controls, with paired whole-genome and RNA sequencing. By mapping expression- (eQTL) and splicing- (sQTL) quantitative trait loci directly in diseased human hearts, we identified over 10,000 transcripts with significant eQTL and 8,600 isoforms with significant sQTL, across both coding and non-coding genes, many of which overlap loci previously associated with HF and emerging novel gene associations. Single-locus colocalization with a largescale DCM genome-wide association study revealed 21 expression and 17 splicing-QTL that share causal variants with disease risk. These include known Mendelian cardiomyopathy risk genes such as FLNC and ACTN2, and novel regulatory candidates like CAMK2D, LMF1, MYOZ1, SKI, SYNPO2L, and TKT. Several loci also showed coordinated effects on both gene expression and RNA splicing, implicating calcium signaling, cytoskeletal organization, and metabolic pathways in HF pathogenesis. Together, these results help define the regulatory landscape of the failing human heart and establish TOPCHeF as a foundational resource for connecting genetic variation to transcriptional and splicing molecular mechanisms in HF research.
Despite a growing body of evidence implicating genetic variants and proteins encoded by them with risk and pathogenesis of Alzheimer's disease (AD), this knowledge has not been successfully translated into effective AD treatments. We integrated current genomic, transcriptomic and proteomic profiles of AD into a network pharmacology framework that leverages comprehensive gene-gene and drug-target interactions. This approach allowed us to screen 2,413 drugs for repurposing opportunities in AD. Computational validation and drug prioritization was followed by experimental validation in 33 cell culture-based phenotypic assays combined with Bayesian hypothesis testing. Our network-based screen rediscovered drugs in clinical trials for AD, providing computational validation. Besides many cancer drugs, the screen identified three drugs previously implicated in AD-related endophenotypes: the primary bile acid chenodiol, arundine (3,3'-diindolylmethane), and cysteamine. In analysis of results from culture-based phenotypic assays, large Bayes factors supported the hypothesized benefits of arundine and the chenodiol derivative, tauroursodeoxycholic acid (TUDCA), in amyloid- β clearance and release and neuroinflammation. Follow-up network analyses mechanistically implicated Regulator of G protein signaling 4 (RGS4) in the plausible therapeutic actions of arundine and TUDCA. A network pharmacology approach identified TUDCA and arundine as promising repurposing candidates in AD that rescue disease-relevant molecular phenotypes by acting on AD-associated genes through regulation of G protein signaling.
Genetic prediction of multi-omic data has emerged as a cost-effective alternative to direct omics profiling, particularly useful for identifying molecular features associated with disease susceptibility. However, despite its popularity, multi-omic imputation models are fragmented across studies, hindering findability, accessibility, interoperability and re-use. To address this, we developed OmicsPred (https://www.omicspred.org), a centralised platform for the deposition and dissemination of genetic prediction models of multi-omic traits. OmicsPred unifies the most commonly used molecular imputation models (e.g. from PredictDB) and other published studies totalling 3,339,469 prediction models spanning transcriptomic, proteomic, and metabolomic traits (as of May 2026). Each model is accompanied by metadata describing score development and predictive performance, and distributed in formats compatible with popular analytic tools, such as PGS Catalog Calculator and MetaXcan. To demonstrate the utility of the resource for systematic target discovery, we perform a multi-omic phenome-wide association analysis in Million Veterans Program data.
Table S1. Summary statistics of 7 cancer types tested with ARIC plasma proteome models, FDR<0.05, with single- and multi-variant COLOC results.
Genetic mutation and drift, coupled with natural and human-mediated selection and migration, have produced a wide variety of genotypes and phenotypes in farmed animals. We here introduce the Farm Animal Genotype-Tissue Expression (FarmGTEx) Project, which aims to elucidate the genetic determinants of gene expression across 16 terrestrial and aquatic domestic species under diverse biological and environmental contexts. For each species, we aim to collect multiomics data, particularly genomics and transcriptomics, from 50 tissues of 1,000 healthy adults and 200 additional animals representing a specific context. This Perspective provides an overview of the priorities of FarmGTEx and advocates for coordinated strategies of data analysis and resource-sharing initiatives. FarmGTEx aims to serve as a platform for investigating context-specific regulatory effects, which will deepen our understanding of molecular mechanisms underlying complex phenotypes. The knowledge and insights provided by FarmGTEx will contribute to improving sustainable agriculture-based food systems, comparative biology and eventual human biomedicine.
We present multi-integration of transcriptome-wide association studies and colocalization (Multi-INTACT), an algorithm that models multiple gene products (e.g. encoded RNA transcript and protein levels) to implicate causal genes and relevant gene products. In simulations, Multi-INTACT achieves higher power than existing methods, maintains calibrated false discovery rates, and detects the true causal gene product(s). We apply Multi-INTACT to GWAS on 1,408 metabolites, integrating the GTEx expression and UK Biobank protein QTL datasets. Multi-INTACT infers 52% to 109% more metabolite causal genes than protein-alone or expression-alone analyses and indicates both gene products are relevant for most gene nominations.
RATIONALE: Chronic Obstructive Pulmonary Disease (COPD) is a leading cause of mortality worldwide. We previously published the polygenic transcriptome risk score (PTRS) which uses the cumulative effect of predicted gene expression to construct genetic predictors for complex diseases, and we demonstrated the value of PTRS for COPD built on Genotype-Tissue Expression (GTEx) Lung tissue with significantly improved cross-ancestry portability. However, the postmortem collection and lack of lung disease status on GTEx samples likely affect results. We hypothesized that the performance of PTRS will be improved by using disease-relevant expression quantitative trait loci (eQTLs) to construct the score. Here, we aimed to improve performance of PTRS by leveraging disease-relevant eQTLs. METHODS: We constructed two transcriptome prediction models, Elastic Net (EN) and Prediction Using Models Informed by Chromatin conformation and Epigenomics (PUMICE), using TOPMed lung RNA-seq data from the Lung Tissue Research Consortium (LTRC), which included COPD patients undergoing lung surgery. We compared their performance with two published models built on RNA-seq from GTEx Lung tissue, GTEx-EN and GTEx-PUMICE. We first integrated each of four transcriptome models with multi-ancestry GWAS of FEV1/FVC ratio (Shrine et al. 2023) to generate transcriptome-wide association study (TWAS) results, which produced trait-associated genes. Pathway analysis was then conducted on significant TWAS genes. Finally, the PTRS was computed on TWAS genes that were included in the top 20 significantly enriched pathways. We tested performance of PTRS in race-stratified analysis for moderate-to-severe and severe COPD in COPDGene. RESULTS: Our models produced more significant TWAS genes (FDR<0.05) and had more overlaps with genes identified for FEV1/FVC ratio from Shrine et al. 2023 compared to two GTEx models (Fig. A). Additionally, our models had significantly higher AUC for predicting both moderate-to-severe and severe COPD in race-stratified analysis. The mean AUC from LTRC-EN for predicting severe COPD was 0.57 and 0.55 respectively for EUR and AFA, which was significantly higher than mean AUC of 0.54 and 0.51 from GTEx-EN (Delong p-value=8.15x10-10 for EUR and 1.42x10-3 for AFA) (Fig. B). Similarly, the PTRS derived from our models produced higher Odds Ratios (OR) per standard deviation. The mean OR from LTRC-PUMICE for moderate-to-severe COPD was 1.25 and 1.18 respectively for EUR and AFA, while it was 1.17 and 1.12 from GTEx-PUMICE (Fig. C). CONCLUSIONS: Our study demonstrates the value of leveraging disease-relevant RNA-seq to construct transcriptome prediction models and their improvement on identifying lung function genes and predicting COPD across ancestries.
Essential tremor (ET) is the most common movement disorder, yet its genetic basis remains poorly understood. We performed a genome-wide association meta-analysis including 20,268 ET cases and 723,761 neurologically healthy controls of European ancestry from the Million Veteran Program, 23andMe Research Institute and All of Us. We identified 50 independent genome-wide significant loci, including 47 novel loci. We estimated the SNP-based heritability to be 24%. Genetic correlation analyses revealed considerable overlap between ET and Parkinson’s disease, myoclonus, and systemic traits such as cardiovascular and metabolic conditions. Inverse correlations were observed with cerebellar and diencephalic volumes. Integrative analyses prioritised candidate genes through transcriptome-wide association studies across 13 brain tissues, and spatial transcriptomics highlighted enrichment of ET heritability in hippocampal and cortical excitatory neurons as well as astrocytes and microglia. Polygenic risk scores significantly predicted ET risk and age at diagnosis across European and non-European cohorts, with the strongest transferability to admixed American ancestry. These findings substantially expand the catalogue of ET risk loci, implicate excitatory neurons and cerebellar–hippocampal circuit mechanisms, and provide a foundation for biomarker discovery and therapeutic development.
Gene implication methods (GIMs) are crucial tools for analyzing genome-wide association studies, but are often ambiguous. We present the LocusCompare2 platform to incorporate six popular GIMs and hundreds of quantitative trait loci datasets, enabling validation across several GIMs and window settings to improve accuracy and reproducibility.