The Meta-Analysis of Glucose and Insulin-related traits Consortium (MAGIC) identified 242 loci associated with glycaemic traits fasting insulin (FI), fasting glucose (FG), 2 h-Glucose (2hGlu), and glycated haemoglobin (HbA1c). However, for the majority, the causal variant(s) remain(s) unknown. Modelling multiple traits and integrating functional annotations have each been shown to improve fine-mapping resolution. Here, we aimed to determine whether combining these techniques would further improve fine-mapping resolution. Using single-trait fine-mapping results from FINEMAP as input, we performed multi-trait fine-mapping with flashfm at 50 loci significantly associated with more than one glycaemic trait. We used fGWAS to build models of enriched annotations by considering 32 cell-type specific and 28 static annotations. We used these models to define prior probabilities to perform annotation informed fine-mapping with both FINEMAP (single-trait) and flashfm (multi-trait). Multi-trait fine-mapping of 106 locus-trait associations significantly (P = 1.23 × 10−17) reduced the median size of the credible sets accounting for 99% of the posterior probability of being causal (99CS) to 21.5 variants compared to the 60.5 variants in single-trait fine-mapping. Annotation informed single-trait fine-mapping of 211 locus-trait associations reduced (P = 4.24 × 10−12) the median 99CS size from 72 in agnostic single-trait fine-mapping to 52 variants. Annotation informed multi-trait fine-mapping of 110 locus-trait associations led to a further significant (P = 2.69 × 10−18) decrease in median 99CS size to 14.5 variants compared to 51.0 in annotation informed single-trait fine-mapping. In conclusion, by applying combined multi-trait and annotation informed fine-mapping to 50 loci, we refined the number of potential causal variants by 71.1% compared to single-trait agnostic fine-mapping.
Abstract Background Immune-mediated inflammatory diseases (IMIDs) are associated with increased risk of cardiometabolic diseases. Investigating genetic overlap among these conditions can provide insights into their clinical management. Methods Genetic correlation was assessed using linkage disequilibrium score regression (LDSC). Then, a meta-analysis was conducted using Association Analysis Based on SubSETs (ASSET) to pinpoint independent single nucleotide polymorphisms (SNPs) shared across the diseases. Each independent SNP was then used to define a genomic window (+/-500KB) for colocalisation analysis and Local Analysis of [co]Variant Association (LAVA) to offer multiple layers of regional pleiotropic evidence. Over-representation analysis was then run to identify enriched biological pathways, which then were used for drug target analysis. Results The LDSC analysis showed a significant global genetic correlation for rheumatoid arthritis (RA) and cardiometabolic diseases including hypertension, coronary artery disease (CAD), heart failure (HF), stroke, atrial fibrillation (AF), and type two diabetes mellitus (T2DM) ranging from rg = 0.09 to 0.24. ASSET meta-analysis identified 164 independent SNPs shared across RA and the cardiometabolic diseases with P < 5×10-8 in the overall one-sided meta-analysis P -value, FDR<0.05 in both individual GWASs, and TRUE phenotype matrix. Colocalisation analysis revealed multiple loci with strong evidence (Posterior probabilities ≥ 80) of single causal SNPs between the trait pairs. LAVA analysis was then used as an additional layer of confirmation for the findings generated by ASSET and colocalisation and thus several loci were highlighted. Over-representation analysis showed significant enriched immune-related pathways across RA-hypertension, RA-CAD, RA-AF, and RA-T2DM trait pairs. Drug target analysis highlighted several drugs which could be further tested for their effectiveness in RA and its common comorbidities. Conclusion The findings revealed a shared genetic architecture and key immune-related biological pathways underlying RA and its associated cardiometabolic comorbidities. The identified genes and drugs provide opportunities for further therapeutic assessment which could improve clinical management strategies.
Type 2 diabetes (T2D) and dementia frequently co-occur, yet the biological mechanisms underlying this comorbidity remain incompletely understood. Here, we systematically investigate shared genetic signals between T2D and three forms of neurodegenerative dementia (Alzheimer’s disease, Lewy body dementia, and sporadic frontotemporal dementia) using large-scale genome-wide association studies of clinically diagnosed cases. We identify five genomic regions harbouring shared association signals between T2D and at least one dementia subtype. Among these, the APOE locus was common to all dementia subtypes, whereas the remaining four loci (GBA, CRY2/PEX16/MAPK8IP1, INO80E, and NSF) were each shared exclusively between T2D and one dementia subtype. Integrating multi-omics data across several disease-relevant tissues and orthogonal lines of functional evidence, we prioritize 26 candidate genes, through which these shared genetic loci potentially mediate their effect. Pathway enrichment highlights lipid and lipoprotein regulatory biology as a central shared axis. Mendelian randomization analyses using genetically regulated gene expression in relevant tissues indicate pleiotropic mechanisms with divergent phenotypic consequences. Our findings identify shared genetic loci between T2D and neurodegenerative dementia, revealing systemic metabolic-neurodegenerative trade-offs and highlighting key genes that underpin the comorbidity, providing a framework for improved understanding of age-related multimorbidity.
BACKGROUND:Serum creatine kinase (CK) is a routinely measured biomarker of muscle damage, yet the genetic factors underlying inter-individual variation in CK levels remain poorly defined. METHODS:Here we present a large multi-ancestry genome-wide association meta-analysis of serum CK, comprising 237,255 participants spanning Admixed American, African American, East Asian, European and Middle Eastern populations. FINDINGS:We identify 107 independent loci at genome-wide significance (P< 5 × 10-8), 98 of which are previously unreported, with pronounced enrichment for genes expressed in skeletal and cardiac muscle and overlap with pathways related to muscle structure and function. Notably, eight loci map to genes implicated in Mendelian myopathies, underscoring a continuum from common regulatory variation to rare pathogenic mutations. Integrative quantitative trait locus (QTL)-based Mendelian randomisation and colocalisation implicate several genes in CK regulation, most prominently SMAD3, KLF5 and STAT3 within the transforming growth factor beta signalling pathway. CK levels show positive genetic correlations with traits reflecting tissue damage as well as muscle mass and strength, and negative correlations with C-reactive protein, indicating pleiotropic effects from muscle biology and enzyme clearance. INTERPRETATION:These findings delineate the genetic architecture of serum CK across diverse populations and highlight muscle-related pathways contributing to CK variation. FUNDING:No funding was received for this study.
Skeletal muscle capillary density is correlated with physical performance and whole-body metabolic properties. Thus, we performed a genome-wide association study of skeletal muscle capillary-to-fiber ratio (C:F) (n = 603 males) and found that the rs115660502 G allele was associated (p < 5 × 10-8) with increased C:F and reduced skeletal muscle expression of RAB3 GTPase-activating non-catalytic protein subunit 2 (RAB3GAP2). The capillary-increasing G allele was more prevalent in elite endurance athletes than in power athletes and non-athlete controls in two independent cohorts. Low-muscle-expressing RAB3GAP2 expression quantitative trait locus (eQTL) alleles were associated with muscle damage in athletes. In healthy individuals, RAB3GAP2 expression was reduced by high-intensity intermittent training. RAB3GAP2 protein was not uniformly expressed in muscle but predominantly expressed in the endothelium and capillaries. RAB3GAP2 expression was lower in endurance compared with power athletes and was negatively associated with type I (oxidative) muscle fiber density. Experimental reduction of RAB3GAP2 in human endothelial cells led to (1) increased proliferation and tube formation in vitro, (2) regulation of secreted factors (e.g., CD70 and TNC) promoting angiogenesis and T cell activation, and (3) increased in vivo endothelial cell density in mice. RAB3GAP2 expression in skeletal muscle was negatively correlated with exercise-induced release of TNC in vivo in humans. In conclusion, RAB3GAP2 is expressed in the microvascular endothelium and is suggested to be a negative regulator of angiogenesis through a decrease in endothelial cell proliferation, possibly mediated by RAB18, with its low-expressing variant associated with higher muscle C:F and elite endurance performance.
Abstract Background Rheumatoid arthritis (RA) is an autoimmune inflammatory disease with complex and incompletely understood molecular mechanisms. Understanding circulating proteins associated with RA may improve understanding of disease biology and clarify its pathological links with cardiometabolic comorbidities. Methods A proteome-wide two-sample Mendelian randomisation (MR) drug target analysis was conducted using plasma proteins measured in 54,219 participants from the UK Biobank Pharma Proteomics Project as exposures and RA and cardiometabolic diseases as the outcomes. Summary statistics for RA included 53,663 cases and 1,070,200 controls. Colocalisation analysis was performed to confirm shared single causal variants and prioritise RA proteins supported by both MR and colocalisation. The prioritised proteins were then evaluated in the Accelerating Medicines Partnership RA Phase II synovial single-cell dataset for cell-type expression patterns. Druggability was then assessed followed by analysis of genetic overlap between RA-associated proteins and cardiometabolic diseases. Results 37 plasma proteins had a causal effect on RA risk, supported by combined evidence from MR and conditional colocalisation. In synovial tissue, TPPP3, RARRES2, AKAP12, and GGT5 were predominantly expressed in stromal and endothelial cell clusters. Druggability assessment identified IFNGR2, IL6R, CD40, and FCGR2B as Tier 1 targets. However, several biologically relevant proteins, including RARRES2, AKAP12, TPPP3, and SNX2, had limited available druggability data. Genetic overlap analysis demonstrated shared protein signals between RA and cardiovascular diseases, including overlap of RARRES2 and TPPP3 with coronary artery disease (CAD) and FCGR2B with atrial fibrillation (AF). To approximate the therapeutic effect of target inhibition, the direction of effect estimates for proteins showing overlap between RA-CAD and RA-AF was reversed. Conclusion This study identified circulating proteins involved in RA pathogenesis and reveals shared mechanisms between RA and cardiovascular diseases. While some proteins showed clear translational potential targets, several prioritised proteins had limited available druggability information and could not be confidently classified. Addressing these gaps may help identify new targets relevant to RA management. Future work should also use phenome-wide MR studies to evaluate potential on-target adverse effects of protein inhibition across RA-CAD and RA-AF.
Abstract Standard polygenic risk scores (PRSs) are constructed based on additive genome-wide association study summary statistics. Nonlinear machine learning methods have been increasingly applied to construct PRSs from individual-level data, aiming to improve predictive performance over standard PRSs through modelling nonadditive genetic effects. However, their superiority across studies has been inconsistent. The conditions under which they provide meaningful improvements remain unclear. We combined theory, simulations and a real-world application to investigate when two widely used nonlinear machine learning methods, random forest and XGBoost, outperform standard PRSs. Theoretically, we showed that standard PRSs can implicitly capture some genetic variance attributable to nonadditive genetic effects through their contributions to marginal SNP effects. Although nonlinear models have a higher theoretical potential, their bias-variance trade-off can limit predictive gains at finite sample sizes. Simulations showed that XGBoost outperformed standard PRSs only when the genetic architecture involves a large proportion of interaction genetic variance concentrated across relatively few interactions and training sample sizes are large. Random forest consistently underperformed standard PRSs. In risk prediction of ischemic heart disease using UK Biobank data, XGBoost showed little improvement in predictive performance over standard PRSs, whereas random forest again performed worse. Together, these findings suggest that nonlinear machine learning do not uniformly outperform standard PRSs; rather, their relative performance depends jointly on genetic architecture and training sample size. Our study helps to reconcile the inconsistent results reported across previous studies and provides a framework for identifying settings in which more complex PRS models are likely to be beneficial.
Kidney disease disproportionately affects populations of African ancestry, yet most genetic studies have focused on Europeans. Here, we present a three-stage genome-wide association study meta-analysis of estimated glomerular filtration rate in ~26,000 individuals across Eastern, Western, and Southern Africa and ~81,000 African-ancestry individuals in the diaspora. Continental African meta-analysis identifies four independent genome-wide significant loci, including two previously unreported loci. Pan-African meta-analysis identifies 19 independent loci, including three previously unreported loci. Fine-mapping reveals four loci with high causality probability, and phenome-wide analyses demonstrate pleiotropic effects on cardiometabolic and immunological traits. Notably, APOL1 high-risk variants strongly associated with kidney disease in African Americans show markedly lower frequency and attenuated effects in continental Africa, indicating potential distinct genetic architectures. Polygenic scores from genetically similar populations significantly outperformed those from distant cohorts. These findings demonstrate the necessity of conducting genomic research across diverse African populations to enable equitable health outcomes.
BACKGROUND:Polygenic risk scores (PRSs) improve prediction of the development of type 2 diabetes over the use of clinical risk factors alone; however, they perform poorly in populations of non-European ancestry, limiting their global clinical utility. We aimed to deliver comprehensive and rigorously tested multi-ancestry PRSs for prediction in type 2 diabetes. METHODS:We conducted meta-analyses using data from type 2 diabetes genome-wide association studies (GWAS) across cohorts from five major global ancestries: European, African or African American, Admixed American, South Asian, and East Asian. We used summary statistics from the GWAS to construct single-ancestry PRSs (using the continuous-shrinkage PRS-CS method) and multi-ancestry PRSs (using the PRS-CSx method), and constructed ancestry-specific linkage disequilibrium panels to model pairwise correlations between single-nucleotide polymorphisms in GWAS during PRS construction. Models were validated for association with type 2 diabetes in at least four independent cohorts per ancestry. The effect sizes of PRSs were estimated as the odds ratio (OR) per SD of the PRS, and ORs for individuals at the 90th, 95th, and 97·5th PRS percentiles were compared with the IQR as a reference. We also tested our PRS models for prediction of diabetes incidence with or without additional clinical factors, as well as microvascular complications and comorbidities. FINDINGS:Our analysis used data from 409 959 individuals with type 2 diabetes and 1 983 345 controls: respectively, 359 819 and 1 825 729 indivduals were included in the GWAS dataset, with 10 992 and 31 792 individuals in the training dataset and 39 148 and 125 824 individuals in the validation dataset. The best predictive performance for the single-ancestry PRSs was in European (incremental AUC 0·07-0·14) and East Asian (0·02-0·16) ancestries, whereas prediction was poorer for African or African American (0·02-0·03), Admixed American (0·02-0·04), and South Asian (0·02-0·04) ancestries, correlating with sample sizes in the GWAS. Compared with single-ancestry PRSs, our multi-ancestry PRSs showed higher effect sizes and smaller 95% CIs across all ancestries: OR per SD 1·73 (95% CI 1·67-1·80) in African or African American, 2·82 (2·67-2·97) in Admixed American, 2·45 (2·36-2·54) in East Asian, 2·36 (2·32-2·41) in European, and 2·23 (2·05-2·42) in South Asian ancestries. Individuals in the 97·5th PRS percentile had a 3-7 times increased risk of type 2 diabetes compared with those in the IQR (OR 3·43 [95% CI 2·80-4·21] in African or African American, 7·47 [5·64-9·89] in Admixed American, 6·62 [5·58-7·85] in East Asian, 6·25 [5·72-6·82] in European, and 4·50 [2·70-7·53] in South Asian ancestries). These PRSs were also associated with earlier onset of type 2 diabetes, higher risk of developing microvascular complications, and provide additional predictive value beyond clinical factors. In individuals with type 2 diabetes, the association between multi-ancestry PRSs and risk of microvascular complications and comorbidity was studied in populations of African, Admixed American, and European ancestries and was significant in all three ancestry groups for diabetic retinopathy (ORs per SD 1·28-1·57), diabetic nephropathy (1·25-1·58), proliferative diabetic retinopathy (1·39-2·08), and end-stage diabetic nephropathy (1·44-1·87); PRS was associated with coronary artery disease in the Admixed American ancestry group only (1·16 [95% CI 1·08-1·25]). INTERPRETATION:These validated, publicly available PRSs can improve risk stratification for type 2 diabetes onset and complications across diverse ancestries, supporting their further evaluation in clinical settings. FUNDING:The National Human Genome Research Institute of the US National Institutes of Health.
Abstract Type 2 diabetes (T2D) is a complex metabolic disorder characterized by hyperglycemia and insulin resistance. Although genome-wide association studies (GWAS) have identified >600 T2D risk loci, the causal genes and the relevant tissues mediating these associations remain largely unresolved. To address this challenge, we performed tissue-specific, ancestry-aware transcriptome-wide association studies (TWAS) across six T2D-relevant tissues: subcutaneous adipose, visceral adipose, brain hypothalamus, liver, skeletal muscle, and pancreas. We conducted ancestry-specific multi-tissue TWAS in European ancestry (EUR) data using summary statistics from the largest EUR GWAS (242,283 cases and 1,569,734 controls) and pre-trained gene expression prediction models derived from 689 EUR individuals from the Genotype-Tissue Expression (GTEx) Project. Conditional analyses were performed to identify independent TWAS signals. We identified 684-750 significant gene-T2D associations per tissue (P < 1.919 × 10 −6 ), implicating both established and novel candidate genes. Among these, JAZF1 and IDE showed consistent association signals across all six tissues, whereas TCF7L2 and WSF1 exhibited heterogeneous effects restricted to a subset of T2D-relevant tissues. Conditional analyses further refined these signals to 289–322 independent TWAS signals per tissue. Together, these finding highlight substantial regulatory heterogeneity in the genetic architecture of T2D and underscore the importance of tissue context in interpreting disease-associated loci. Cross-ancestry replication of EUR-derived TWAS signals was evaluated in African American (AFA) individuals. We conducted an AFA-TWAS using summary statistics from the largest AFA GWAS (50,251 cases and 103,909 controls) in combination with gene expression prediction models trained in 111 AFA individuals from GTEx. We observed significant enrichment of EUR-derived T2D TWAS signals in the AFA TWAS across subcutaneous adipose, visceral adipose, skeletal muscle, and pancreas, whilst enrichment was weaker in liver, likely reflecting limited sample size. Overall, our findings demonstrate that integrating tissue-specific and ancestry-aware TWAS refines the identification of causal genes for T2D, with cross-ancestry replication supporting the robustness of these signals and cross-tissue analyses revealing context-specific effects. However, they also highlight the limited availability of non-EUR datasets and the need for larger, more diverse ancestry-specific transcriptomic resources.
Gestational diabetes mellitus (GDM) affects ~14% of pregnancies and increases maternal type 2 diabetes mellitus (T2DM) risk. The GenDiP Consortium presents trans-generational, multi-ancestry genome-wide association study meta-analyses of GDM and pregnancy glycemic traits in up to 38,305 GDM cases and 776,145 controls. We identify 37 GDM-associated loci (7 novel) and five novel loci for pregnancy glycemic traits, all operating through the maternal genome. We classify 12 GDM variants with stronger effects in GDM than T2DM into five biologically informed categories, revealing pleiotropy patterns, pregnancy-dependent effect modification, and diagnostic heterogeneity. While all these loci overlap with T2DM and/or non-pregnant glycaemic traits, four (G6PC2, CAST-PCSK1, HKDC1, FOXA2) lack genome-wide-significant T2DM associations; GCK shows distinct causal variants for GDM, and MTNR1B exhibits pregnancy-amplified effects. Our findings provide new genetic insights into GDM and highlight the need for larger, ancestrally diverse studies of GDM and glycaemic traits during pregnancy to understand potential pregnancy-specific effects.
Type 2 diabetes is associated with a range of non-cardiovascular non-oncologic comorbidities. To move beyond associations and evaluate causal effects between type 2 diabetes genetic predisposition and 21 comorbidities, we apply Mendelian randomization analysis using genome-wide association studies across multiple genetic ancestries. Additionally, leveraging eight mechanistic clusters of type 2 diabetes genetic profiles, each representing distinct biological pathways, we investigate causal links between cluster-stratified type 2 diabetes genetic predisposition and comorbidity risk. We identify causal effects of type 2 diabetes genetic predisposition driven by distinct genetic clusters. For example, the risk-increasing effects of type 2 diabetes genetic predisposition on cataracts and erectile dysfunction are primarily attributed to adiposity and glucose regulation mechanisms, respectively. We observe opposing effect directions across different genetic ancestries for depression, asthma and chronic obstructive pulmonary disease. Our findings leverage the heterogeneity underpinning type 2 diabetes genetic predisposition to prioritize biological mechanisms underlying causal relationships with comorbidities.
Diabetes has a large medical and public health impact in American Indians. Studies have used genetic data to distinguish type 1 (T1D) and type 2 diabetes (T2D) and uncover biologic mechanisms underlying T2D clinical heterogeneity. We applied a T1D polygenic score (PS) to 3,084 American Indians (mean age 56 years, 58% female, 39% diabetes). We also calculated partitioned PS for eight clusters of T2D-associated variants and evaluated their association with twenty cardiometabolic traits and five clinical outcomes. The profile of T1D PS for individuals with diabetes was consistent with T2D. A total T2D PS was significantly associated with early age of onset of T2D (p-value = 3.5x10-11). Partitioned PS for T2D clusters were significantly associated with cardiometabolic traits for the obesity cluster (increased measures of body fat and total triglycerides but lower HDL cholesterol), while the lipodystrophy cluster was associated with increased fasting insulin, waist/hip ratio, triglycerides, and blood pressure and lower body fat % and HDL cholesterol. T2D clusters were not associated with cardiovascular and kidney outcomes. Our findings support a relationship of cluster-specific T2D partitioned PS with cardiometabolic traits described in other populations but there are opportunities for developing improved clustering methods using genetic variation from American Indians.
AIMS:This study retrospectively investigates the association between polygenic risk scores (PRS) derived from SNP clusters and glycaemic response to metformin in patients with newly diagnosed T2D. MATERIALS AND METHODS:Utilizing a dataset from the Taiwan Precision Medicine Initiative, we evaluated alterations in fasting glucose (FBG) and glycated haemoglobin (HbA1c) in individuals newly diagnosed with T2D who underwent metformin monotherapy for a duration of 6 months. Glycaemic responses between those in the bottom 20% of PRS (Q1) and the top 20% of PRS (Q5) for each of the SNP clusters and for the combination of two clusters were analysed. RESULTS:In responses to metformin monotherapy, significant differences of FBG levels were detected in Q1 as compared to Q5 in individuals of PRS derived from the cluster of beta-cell dysfunction with a positive association with proinsulin (Beta cell +PI) (p = 0.005) and the cluster of beta-cell dysfunction with a negative association with proinsulin (Beta cell -PI) (p = 0.003). Moreover, lower FBG levels on treatment were observed in those with both Q1 than those with both Q5 in the PRS derived from the two clusters of beta cell dysfunction (p = 0.002). Significantly reduced HbA1c values were documented in the Q1 in comparison to the Q5 within the cluster of Beta cell -PI (p = 0.002). CONCLUSION:These findings suggest that PRS derived from beta-cell dysfunction clusters may help predict glycaemic response to metformin and support the potential for genetically guided treatment in T2D.
There have been recent efforts by the human genetics research community to increase the genetic diversity of participants contributing to genome-wide association studies (GWAS) of complex human traits and diseases. The traditional multi-ancestry GWAS approach is to first assign participants to continental ancestry labels based on their genetic similarity to individuals in reference datasets. Ancestry-specific GWAS are then conducted separately for each continental label, the results of which are aggregated through multi-ancestry meta-analysis. However, with this approach, a participant may be assigned to an ancestry group that does not reflect their personal view of ethnicity/race or may be excluded because their genetic ancestry is not sufficiently similar to individuals in reference datasets to be assigned to a single group. Here, we present a novel pipeline (PANACEA) for fully inclusive multi-ancestry meta-analysis that employs a continuous and multi-dimensional representation of ancestry that maximises the genetic diversity of GWAS. Through application to multi-ancestry GWAS of type 2 diabetes susceptibility and simulations, we demonstrate that the inclusive pooled analysis provides equivalent protection against population structure to a traditional ancestry-stratified analysis but, importantly, offers increased power to detect association through increased sample size by not excluding participants with outlying ancestry. The pooled inclusive analysis also enables assessment of ancestry-correlated heterogeneity in allelic effects without the need to assign participants to continental labels that may not sufficiently reflect genetic diversity within ancestry groups.
Background Transcriptome-wide association studies (TWAS) integrate genetically regulated gene expression with genome-wide association study (GWAS) data to identify gene–trait associations. These analyses rely on accurate gene expression prediction models trained using genotype and transcriptomic data from the same set of individuals. However, most publicly available TWAS models are based on resources that predominantly consist of participants of European ancestry. This European ancestry bias limits the transferability and generalizability of gene expression models across diverse populations. Description Using genotype and gene expression data from the Genotype-Tissue Expression (GTEx) Project, we curated datasets comprising 689 European ancestry individuals and 111 African American individuals after quality control and confounder adjustment. We then developed ancestry-specific cross-tissue gene expression imputation models using the UTMOST framework. These models are accessible through AGEMdb, a web-based database that allows users to browse, search, and download gene expression imputation models with flexible filtered options (e.g. by tissue and gene). By incorporating African American-specific models, AGEMdb addresses the European ancestry bias present in existing TWAS resources and supports more equitable gene–trait association analyses. Conclusion AGEMdb provides a valuable resource for ancestry-specific TWAS, which promotes more inclusive studies that support the advancement of equitable genomic research.
Type 2 diabetes (T2D) is epidemiologically associated with a wide range of non-cardiovascular comorbidities, yet their shared etiology has not been fully elucidated. Leveraging eight non-overlapping mechanistic clusters of T2D genetic profiles, each representing distinct biological pathways, we investigate putative causal links between cluster-stratified T2D genetic predisposition and 21 non-cardiovascular comorbidities. Most of the identified putative causal effects are driven by distinct T2D genetic clusters. For example, the risk-increasing effects of T2D genetic predisposition on cataracts and erectile dysfunction are primarily attributed to obesity and glucose regulation mechanisms, respectively. When surveyed in populations across the globe, we observe opposing effect directions for depression, asthma and chronic obstructive pulmonary disease between populations. We identify a putative causal link between T2D genetic predisposition and osteoarthritis. To underscore the translational potential of our findings, we intersect high-confidence effector genes for osteoarthritis with targets of T2D-approved drugs and identify metformin as a potential candidate for drug repurposing in osteoarthritis.
Transcriptome-wide association studies (TWAS) investigate the links between genetically regulated gene expression and complex traits. TWAS involves imputing gene expression using expression quantitative trait loci (eQTL) as predictors and testing the association between the imputed expression and the trait. The effectiveness of TWAS depends on the accuracy of these imputation models, which require genotype and gene expression data from the same samples. However, publicly accessible resources, such as the Genotype Tissue Expression (GTEx) Project, are biased toward individuals of European ancestry, potentially reducing prediction accuracy into other ancestry groups. This study explored eQTL transferability across ancestry groups by comparing two imputation models: PrediXcan (tissue-specific) and UTMOST (cross-tissue). Both models were trained on tissues from the GTEx Project using European ancestry individuals and then tested on data sets of European ancestry and African American individuals. Results showed that both models performed best when the training and testing data sets were from the same ancestry group, with the cross-tissue approach generally outperforming the tissue-specific approach. This study underscores that eQTL detection is influenced by ancestry and tissue context. Developing ancestry-specific reference panels across tissues can improve prediction accuracy, enhancing TWAS analysis and our understanding of the biological processes contributing to complex traits.