The human metabolome reflects complex metabolic states affected by genetic and environmental factors. However, metabolites associated with type 2 diabetes (T2D) risk and their determinants remain insufficiently characterized. Here we integrated blood metabolomic, genomic and lifestyle data from up to 23,634 initially T2D-free participants from ten cohorts. Of 469 metabolites examined, 235 were associated with incident T2D during up to 26 years of follow-up, including 67 associations not previously reported across bile acid, lipid, carnitine, urea cycle and arginine/proline, glycine and histidine pathways. Further genetic analyses linked these metabolites to signaling pathways and clinical traits central to T2D pathophysiology, including insulin resistance, glucose/insulin response, ectopic fat deposition, energy/lipid regulation and liver function. Lifestyle factors-particularly physical activity, obesity and diet-explained greater variations in T2D-associated versus non-associated metabolites, with specific metabolites revealed as potential mediators. Finally, a 44-metabolite signature improved T2D risk prediction beyond conventional factors. These findings provide a foundation for understanding T2D mechanisms and may inform precision prevention targeting specific metabolic pathways.
Most genetic variants associated with complex traits are hypothesized to regulate gene expression. To understand the genetics underlying gene expression variability, we characterized 14,324 RNA-sequencing samples from the Trans-Omics for Precision Medicine program and performed expression and splicing quantitative trait locus (e/sQTL) analyses in six tissues and cell types, including whole blood (n = 6454) and lung (n = 1291). We detected tens of thousands of secondary cis-e/sQTLs, showing that secondary cis-e/sQTL discovery remains unsaturated. We fine-mapped UK Biobank-derived genome-wide association study (GWAS) signals from 164 traits and identified e/sQTL colocalizations for 10,611 GWAS signals, including 7096 that colocalize with secondary e/sQTLs. Our results suggest that even larger e/sQTL analyses will uncover additional secondary e/sQTLs, further benefiting GWAS interpretation.
Both short and long sleep duration have been associated with poor glycemic control and an increased risk of developing type 2 diabetes mellitus. Although sleep duration may differentially modify the effects of genetic risk factors for type 2 diabetes, this has not been systematically investigated. In the present study, we conducted genome-wide gene by sleep duration meta-analyses, separately assessing interactions of short and long sleep, for fasting glucose, fasting insulin, and hemoglobin A1c in up to 489,309 individuals without diabetes from seven different population groups. In total, 16 loci were identified to interact with sleep duration - six with short sleep and ten with long sleep. Of these, four loci were identified through cross-population meta-analysis. Mapped genes exhibit pathway connections to pericyte apoptosis, NMDA receptor activity, the GLUT1 receptor, neurological health, and sleep architecture. Eleven loci (VRK2, PCDH7, TFAP2A, CAP2, PAPPA, ZCCHC2, MYH9, SGIP1, JAKMIP3, RRAS2, MAPT) have not been reported in previous glycemic trait genome-wide association studies. Interaction loci identify divergent biological mechanisms for short and long sleep duration influencing glycemic control, suggesting specific pathways of intervention for precision medicine approaches to diabetes prevention and management.
Gene-based rare variant analyses often lack statistical power and may overlook transcript-specific effects. Here, we present a transcript-aware aggregation framework. In simulation studies, the framework maintains appropriate false-positive rates and shows competitive power relative to standard single-transcript analyses, approaching the performance of the ideal case of knowing the most informative transcript in advance. We then apply the approach to 129 cardiopulmonary traits in over 240,000 whole-genome-sequenced All of Us participants. By leveraging transcript-specific annotations, we identify 11 novel associations and recover 47 reported associations, including potentially pleiotropic genes linked to plasma lipid traits (PPARG) and body habitus (TCF12). Notably, for TTN, a gene known for its transcript-specific effects in cardiomyopathy, our framework strengthens the association signal and pinpoints the N2B isoform, which shows a stronger association with cardiomyopathy than other transcripts. These findings highlight the value of a transcript-aware framework for improving rare variant association studies.
Genetic predisposition and alcohol consumption are risk factors for increased blood pressure (BP), but their interactions influencing BP remain understudied. We conducted population-specific and cross-population meta-analyses of genome-wide gene-alcohol (GxAlc) interactions affecting BP in >1.1M individuals from multiple populations. We identified 46 GxAlc interaction loci for BP, including 21 from one-degree-of-freedom interaction tests (PGxAlc<5×10-8; or <0.05/Meff, Meff independent BP associations at P<10-5), and 25 from two-degree-of-freedom tests of main and interaction effects (PGxAlc<0.05/M2df, M2df independent 2df-associations at P2df<5×10-8), including 7 novel and 39 known BP loci. The 12q24 locus highlights the genetic effect of BRAP-rs11066001 on BP, being ~6 times larger in current drinkers than in non-drinkers. Gene prioritization with 46 GxAlc loci identified 15 genes with ≥3 lines of evidence (location, literature, druggability, functional/regulatory annotation, or pathway analyses). Several loci showed sex- and population-specific effects and revealed biological pathways of alcohol's influence on BP, suggesting mechanisms underlying alcohol-induced hypertension.
Omics data, such as transcriptomics, proteomics, and metabolomics, provide critical insights into disease mechanisms and clinical outcomes. However, their high dimensionality, small sample sizes, and intricate biological networks pose major challenges for reliable prediction and meaningful interpretation. Graph Neural Networks (GNNs) offer a promising way to integrate prior knowledge by encoding feature relationships as graphs. Yet, existing methods typically rely solely on either an externally curated feature graph or a data-driven generated one, which limits their ability to capture complementary information. To address this, we propose the external and generated Graph Neural Network (engGNN), a dual-graph framework that jointly leverages both external known biological networks and data-driven generated graphs. Specifically, engGNN constructs a biologically informed undirected feature graph from established network databases and complements it with a directed feature graph derived from tree-ensemble models. This dual-graph design produces more comprehensive embeddings, thereby improving predictive performance and interpretability. Through extensive simulations and real-world applications to gene expression data, engGNN consistently outperforms state-of-the-art baselines. Beyond classification, engGNN provides interpretable feature importance scores that facilitate biologically meaningful discoveries, such as pathway enrichment analysis. Taken together, these results highlight engGNN as a robust, flexible, and interpretable framework for disease classification and biomarker discovery in high-dimensional omics contexts.
Large-scale multiancestry genome-wide association studies have identified hundreds of loci associated with type 2 diabetes (T2D) and glycemic traits, yet imputed genotyping arrays limit the detection of low-frequency and rare variants. Whole-genome sequencing (WGS) offers a more complete view of genetic variation, especially across diverse populations. We analyzed high-coverage (38×) WGS data from 21,913 T2D case subjects, 61,036 control subjects, and up to 50,011 individuals with no diabetes with fasting glucose, fasting insulin, and HbA1c from the National Heart, Lung, and Blood Institute Trans-Omics for Precision Medicine Program. We performed single-variant association testing, conditional analysis, fine-mapping, and Bayesian colocalization to identify genetic signals and assess regulatory relevance in diabetes-related tissues. We identified 76 distinct association signals across 34 loci, including novel variants at DUSP9 for T2D, and ROBO1, NDN, and MYT1 for HbA1c. Fine-mapping narrowed credible sets and improved causal variant resolution. Colocalization highlighted 80 expression signals in diabetes-related tissues, linking genetic associations to functional regulatory mechanisms. Our findings demonstrate the utility of WGS to uncover novel variants in diverse populations, enhance locus resolution, and link regulatory variation to disease-relevant tissues. This work refines the genetic architecture of T2D and glycemic traits and supports precision medicine efforts targeting diverse populations. ARTICLE HIGHLIGHTS:We aimed to improve understanding of the genetic architecture of type 2 diabetes and glycemic traits by leveraging whole-genome sequencing in diverse populations. Our goal was to identify novel variants, refine known loci, and link genetic signals to regulatory mechanisms through colocalization with expression quantitative trait loci. We discovered novel variants, significantly improved fine-mapping resolution, and identified 80 regulatory colocalization signals in diabetes-relevant tissues. These findings support precision medicine approaches by connecting genetic variation to functional biology in type 2 diabetes.
Background Ten‐year atherosclerotic cardiovascular disease (ASCVD) risk prediction models include the pooled cohort equations (PCE) and the Predicting Risk of Cardiovascular Disease Events (PREVENT) models. We evaluated the relative contributions of predictors in these models, along with social determinants and emerging biomarkers. Methods We pooled data from 13 108 participants (58.4% female, 22.6% Black participants, median age 61 years) across 3 prospective cohorts: ARIC (Atherosclerosis Risk in Communities), FOS (Framingham Offspring Study), and MESA (Multi‐Ethnic Study of Atherosclerosis). Of these participants, 873 (6.7%) developed ASCVD within 10 years. Candidate predictors included variables from PCE and PREVENT‐ASCVD, alongside small dense low‐density lipoprotein cholesterol, hs‐CRP (high‐sensitivity C‐reactive protein), lipoprotein(a), and education level. We fit Cox proportional hazards models and applied stepwise selection, elastic net, and random forest for variable selection. Selected predictors were integrated into an exploratory model, Expanded ASCVD Non‐traditional Determinants (EXPAND), which was compared with PCE and PREVENT‐ASCVD both as originally published and after refitting in our sample. Results Most predictors shared by PCE and PREVENT‐ASCVD were selected across methods; small dense low‐density lipoprotein cholesterol, hs‐CRP, and education were also selected, whereas self‐reported race was not. Using published coefficients, PCE overestimated 10‐year ASCVD risk, whereas PREVENT‐ASCVD underestimated risk. Compared with PCE and PREVENT‐ASCVD models refitted in our sample, EXPAND showed modest calibration advantages, particularly among Black men, and consistently achieved higher C‐statistics. Conclusions In our sample, self‐reported race did not improve ASCVD risk prediction, whereas small dense low‐density lipoprotein cholesterol, hs‐CRP, and education level added predictive value, suggesting potential utility in including additional biomarkers and social factors.
OBJECTIVE:To gain insight into higher fracture risk in individuals with type 2 diabetes, we determined the association of type 2 diabetes glycemic status and severity with longitudinal changes in peripheral bone density and microarchitecture. RESEARCH DESIGN AND METHODS:We conducted a longitudinal study of 769 participants from the Framingham Study who underwent high-resolution, peripheral, quantitative computed tomography (HR-pQCT) at the tibia and radius, in 2012-2016 and 2021-2023 (mean 8-year follow-up). Linear regression models estimated mean 8-year percent changes in bone measures, across indicators of diabetes severity, adjusting for age, sex, weight, and height. RESULTS:The mean age was 67 ± 7 years, and 59% of participants were women. More than half (57%) were normoglycemic (fasting plasma glucose [FPG] <100 mg/dL, not on any treatment), 31% had prediabetes (100 ≤ FPG ≤125 mg/dL), and 12% had type 2 diabetes (FPG >125 mg/dL or on treatment). Adjusted mean percent changes in HR-pQCT bone measures were similar across diabetes severity, including glycemic status, use of diabetes medications, duration of diabetes, and HbA1c. For example, cortical volumetric bone mineral density at the radius changed by -1.50% (95% CI -2.43, -0.56) in type 2 diabetes and -1.96% (-2.53, -1.39), in prediabetes, compared with -2.42% (-2.86, -1.97) in normoglycemia (reference group; all P > 0.05). CONCLUSIONS:The magnitude of peripheral bone loss over 8 years did not differ between individuals with type 2 diabetes and those with normoglycemia, suggesting that bone deterioration alone does not explain the higher fracture risk in older adults with type 2 diabetes. Future studies should address other contributors to skeletal fragility.
Gene-environment interactions may enhance our understanding of blood pressure (BP) biology. We conducted a meta-analysis of multi-population genome-wide association studies (GWASs) of BP traits accounting for gene-depressive symptomatology (DEPR) interactions. Our study included 564,680 adults from 67 cohorts and four population backgrounds: African (5%), Asian (7%), European (85%), and Hispanic (3%). We discovered seven previously unreported BP loci showing gene-DEPR interaction. These loci mapped to genes implicated in neurogenesis (TGFA and CASP3), lipid metabolism (ACSL1), neuronal apoptosis (CASP3), and synaptic activity (CNTN6 and DBI). We also showed evidence for gene-DEPR interaction at nine known BP loci, further suggesting links between mood disturbance and BP regulation. Of the 16 identified loci, 11 were derived from non-European populations. Post-GWAS analyses prioritized 36 genes, including genes involved in synaptic functions (DOCK4 and MAGI2) and neuronal signaling (CCK, UGDH, and SLC01A2). Integrative druggability analyses identified 11 druggable candidate gene targets linked to pathways involved in mood disorders as well as known anti-hypertensive drugs. Our findings emphasize the importance of considering gene-DEPR interactions on BP, particularly in non-European populations. Our prioritized genes and druggable targets highlight biological pathways connecting mood disorders and hypertension and suggest opportunities for BP drug repurposing and risk factor prevention, especially in individuals with DEPR.
BACKGROUND AND AIMS:Metabolic dysfunction-associated steatotic liver disease (MASLD) can progress from hepatic steatosis to liver fibrosis and cirrhosis. Advanced fibrosis is associated with increased mortality. Physical activity (PA) is important in treatment and prevention, but its association with hepatic fibrosis is not well characterised. METHODS:We examined the cross-sectional association between accelerometer-measured PA and fibrosis measured by vibration-controlled transient elastography. Primary covariates included age, sex, cohort, smoking status, alcohol use and accelerometer wear time. Secondary covariates included body mass index (BMI) and hepatic steatosis. The primary dependent variable was continuous liver stiffness measurement (LSM), and the secondary dependent variable was dichotomous liver fibrosis (LSM > 8.2 kPa). We performed sex-specific, age-adjusted Pearson correlation coefficients, as well as multivariable-adjusted linear and logistic regression models. RESULTS:In our study sample (n = 2201, 54.1% women, average age 54 years old, average BMI 27.1 kg/m2) the prevalence of fibrosis was 7.7%. Each additional 30 min spent in moderate-vigorous PA (MVPA)/day was associated with a lower odds of fibrosis (odds ratio 0.57; 95% confidence interval 0.41, 0.79) even after adjusting for BMI (0.72; 0.53, 0.99) or steatosis (0.70; 0.51, 0.96). Those who achieved at least 150 min/week of MVPA had the lowest odds of hepatic fibrosis (0.51; 0.35, 0.75) even when adjusted for BMI (0.64; 0.44, 0.95) or steatosis (0.61; 0.42, 0.90). CONCLUSIONS:In our community-based cross-sectional cohort study, there was an inverse association between time spent in MVPA and liver fibrosis, even when adjusting for BMI or steatosis. Additional interventional studies are needed to determine if MVPA can reverse liver fibrosis.
Obesity is a major public health crisis associated with high mortality rates. Previous genome-wide association studies (GWAS) investigating body mass index (BMI) have largely relied on imputed data from European individuals. This study leveraged whole-genome sequencing (WGS) data from 88,873 participants from the Trans-Omics for Precision Medicine (TOPMed) Program, of which 51% were of non-European population groups. We discovered 18 BMI-associated signals (P < 5 × 10-9). Notably, we identified and replicated a novel low frequency single nucleotide polymorphism (SNP) in MTMR3 that was common in individuals of African descent. Using a diverse study population, we further identified two novel secondary signals in known BMI loci and pinpointed two likely causal variants in the POC5 and DMD loci. Our work demonstrates the benefits of combining WGS and diverse cohorts in expanding current catalog of variants and genes confer risk for obesity, bringing us one step closer to personalized medicine.
Gaussian Graphical Models (GGMs) are a type of network modeling that uses partial correlation rather than correlation for representing complex relationships among multiple variables. The advantage of using partial correlation is to show the relation between two variables after "adjusting" for the effects of other variables and leads to more parsimonious and interpretable models. There are well established procedures to build GGMs from a sample of independent and identical distributed observations. However, many studies include clustered and longitudinal data that result in correlated observations and ignoring this correlation among observations can lead to inflated Type I error. In this paper, we propose a cluster-based bootstrap algorithm to infer GGMs from correlated data. We use extensive simulations of correlated data from family-based studies to show that the proposed bootstrap method does not inflate the Type I error while retaining statistical power compared to alternative solutions when there are sufficient number of clusters. We apply our method to learn the GGM that represents complex relations between 47 Polygenic Risk Scores generated using genome-wide genotype data from the Long Life Family Study. By comparing it to the conventional methods that ignore within-cluster correlation, we show that our method controls the Type I error well without power loss.
There has been increasing discussion regarding the legalization of hallucinogens in recent years. However, literature remains limited on the associations of hallucinogen use with prescription drug misuse and illicit drug use. This study aimed to address this knowledge gap by analyzing data from the 2021-2022 National Survey on Drug Use and Health, focusing on adults aged 18 years and older (unweighted n = 90,503). The primary independent variable was hallucinogen use status (never, lifetime, or past 12 months), and the outcomes included prescription drug misuse (pain relievers, tranquilizers, or stimulants) and illicit drug use (cocaine, crack, inhalants, heroin, or methamphetamine). A series of logistic regression models were conducted. In the final sample, 14.77% of respondents reported lifetime hallucinogen use, and 4.11% reported past 12-month use. Additionally, 10.99% reported prescription drug misuse, and 3.66% reported illicit drug use in the past 12 months. Regression analysis showed that, compared to those who had never used hallucinogens, participants reporting use in the past 12 months had significantly higher odds of prescription drug misuse (AOR = 1.43, 99% CI: 1.36, 1.50) and illicit drug use (AOR = 1.64, 99% CI: 1.57, 1.72) in the past 12 months. Given that recent hallucinogen use was associated with significantly higher odds of both prescription drug misuse and illicit drug use, policymakers and public health practitioners should consider hallucinogen use, particularly past-year use, as a potential indicator of broader substance misuse risks.
Genome-wide association studies (GWAS) have identified numerous body mass index (BMI) loci. However, most underlying mechanisms from risk locus to BMI remain unknown. Leveraging omics data through integrative analyses could provide more comprehensive views of biological pathways on BMI. We analyzed genotype and blood gene expression data from up to 5619 samples in the Framingham Heart Study (FHS). Using 3992 single-nucleotide polymorphisms (SNPs) at 97 BMI loci and 1408 transcripts within 1 Mb, we performed separate association analyses of transcript with BMI and SNP with transcript (PBMI and PSNP, respectively) and then a correlated meta-analysis between the full summary data sets (PMETA). Transcripts were prioritized if we identified transcripts that met Bonferroni-corrected significance within each omic, showed stronger associations in the correlated meta-analysis than each omic, and had corresponding SNPs in the SNP-transcript-BMI association that were at least nominally associated with BMI in FHS data. We tested for generalization of identified association in a Hispanic ancestry sample of blood gene expression data and other samples in hypothalamus, nucleus accumbens, liver, and visceral adipose tissue (VAT) with significant threshold: PMETA < 0.05 & PMETA < PSNP & PMETA < PBMI. Among 308 significant SNP-transcript-BMI associations, we identified seven genes (NT5C2, GSTM3, SNAPC3, SPNS1, TMEM245, YPEL3, and ZNF646) in five association regions. We generalized results for SNAPC3 and YPEL3 in Hispanic ancestry sample, for YPEL3 in the nucleus accumbens, ZNF646 and GSTM3 in VAT, and NT5C2, SNAPC3, TMEM245, YPEL3, and ZNF646 in liver. The identified genes help link the genetic variation at obesity-risk loci to biological mechanisms and health outcomes, thus translating GWAS findings to function.
Most genetic variants associated with complex traits and diseases occur in non-coding genomic regions and are hypothesized to regulate gene expression. To understand the genetics underlying gene expression variability, we characterize 14,324 ancestrally diverse RNA-sequencing samples from the NHLBI Trans-Omics for Precision Medicine (TOPMed) program and integrate whole genome sequencing data to perform cis and trans expression and splicing quantitative trait locus (cis-/trans-e/sQTL) analyses in six tissues and cell types, most notably whole blood (N=6,454) and lung (N=1,291). We show this dataset enables greater detection of secondary cis-e/sQTL signals than was achieved in previous studies, and that secondary cis-eQTL and primary trans-eQTL signal discovery is not saturated even though eGene discovery is. Most TOPMed trans-eQTL signals colocalize with cis-e/sQTL signals, suggesting many trans signals are mediated by cis signals. We fine-map European UK BioBank GWAS signals from 164 traits and colocalize the resulting 34,107 fine-mapped GWAS signals with TOPMed e/sQTL signals, finding that of 10,611 GWAS signals with a colocalization, 7,096 GWAS signals colocalize with at least one secondary e/sQTL signal. These results demonstrate that larger e/sQTL analyses will continue to uncover secondary e/sQTL signals, and that these new signals will benefit GWAS interpretation.
Background: Diabetes is a multifactorial disease with significant genetic predisposition. Polygenic risk scores (PRS) have been developed to estimate an individual's genetic risk of a disease. Traditionally, PRS utilize sex-combined genome-wide association studies (GWAS) due to the limited availability of sex-stratified summary statistics. This study explores sex-dimorphic genetic effects and evaluates the potential benefits of incorporating sex-stratified effects in PRS for type 2 diabetes mellitus (T2DM) and glycemic traits by comparing PRS performance derived from sex-combined versus sex-stratified GWAS. Methods: We performed a sex-heterogeneity test across sex-specific GWAS and identified nine signals with sex-dimorphic effects for T2DM. PRS[sex-combined] and PRS[sex-stratified] were developed using sex-combined and sex-stratified GWAS results for T2DM (41,444 cases and 354,539 controls), fasting glucose (n= 120,595) and fasting insulin (n= 98,210). We evaluated these PRS models in 8,379 participants (1,303 cases and 7,076 controls) from the Framingham Heart Study not included in the PRS derivation. Results: Our findings suggest that sex-combined PRS currently offer better predictive performance for T2DM and glycemic traits. Conclusion: These results highlight the need for larger sex-stratified studies and the optimization of sex-stratified risk models for clinical practice.
Elevated fasting insulin levels (FI), indicative of altered insulin secretion and sensitivity, may precede type 2 diabetes (T2D) and cardiovascular disease onset. In this study, we group FI-associated genetic variants based on their genetic and phenotypic similarities and identify seven clusters with distinct mechanisms contributing to elevated FI levels. Clusters fall into two types: "non-diabetogenic hyperinsulinemia," where clusters are not associated with increased T2D risk, and "diabetogenic hyperinsulinemia," where T2D associations are driven by body fat distribution, liver function, circulating lipids, or inflammation. In over 1.1 million multi-ancestry individuals, we demonstrated that diabetogenic hyperinsulinemia cluster-specific polygenic scores exhibit varying risks for cardiovascular conditions, including coronary artery disease, myocardial infarction (MI), and stroke. Notably, the visceral adiposity cluster shows sex-specific effects for MI risk in males without T2D. This study underscores processes that decouple elevated FI levels from T2D and cardiovascular risk, offering new avenues for investigating process-specific pathways of disease.
AIMS:This study explored dietary trends and urban-rural disparities in food consumption among older adults in China. METHODS:A repeated cross-sectional study of adults aged 65+ years using data from four waves (2008-2018) of the Chinese Longitudinal Healthy Longevity Survey. Multiple logistic regression models assessed daily vs. non-daily consumption of fresh fruit, vegetables, meat, eggs, and milk products. RESULTS:Among 20 945 older adults, over half were female (51.44%) and 54.22% resided in rural communities. Most participants did not consume fresh fruit (83.23%), meat (67.21%), eggs (65.13%), or dairy products (79.89%) daily, although 64.84% consumed vegetables daily. Urban adults had significantly higher odds of daily consumption of fruit (OR = 2.06), meat (OR = 1.56), eggs (OR = 1.20), and dairy (OR = 2.01). CONCLUSION:The study highlights urban-rural disparities in dietary behaviours, emphasising the need for public health initiatives to improve healthier diets and expand dietary options, especially in rural populations.