Genetic predisposition and alcohol consumption are risk factors for increased blood pressure (BP), but their interactions influencing BP remain understudied. We conducted population-specific and cross-population meta-analyses of genome-wide gene-alcohol (GxAlc) interactions affecting BP in >1.1M individuals from multiple populations. We identified 46 GxAlc interaction loci for BP, including 21 from one-degree-of-freedom interaction tests (PGxAlc<5×10-8; or <0.05/Meff, Meff independent BP associations at P<10-5), and 25 from two-degree-of-freedom tests of main and interaction effects (PGxAlc<0.05/M2df, M2df independent 2df-associations at P2df<5×10-8), including 7 novel and 39 known BP loci. The 12q24 locus highlights the genetic effect of BRAP-rs11066001 on BP, being ~6 times larger in current drinkers than in non-drinkers. Gene prioritization with 46 GxAlc loci identified 15 genes with ≥3 lines of evidence (location, literature, druggability, functional/regulatory annotation, or pathway analyses). Several loci showed sex- and population-specific effects and revealed biological pathways of alcohol's influence on BP, suggesting mechanisms underlying alcohol-induced hypertension.
Background:Diabetic kidney disease (DKD) is a leading cause of kidney failure in individuals with type 2 diabetes (T2D), yet risk identification in routine clinical practice remains incomplete. A critical and often overlooked barrier is risk observability: how much of a patient's underlying risk is actually captured in their clinical record at the time of screening. Existing prediction models evaluate performance using model-specific thresholds, making it difficult to understand how additional data sources alter real-world screening behavior or which individuals benefit when models are expanded. Methods:We developed a series of five nested machine learning models evaluated at a one-year landmark following T2D diagnosis using data from the All of Us Research Program (N = 39,431; cases = 16,193). Each successive model added a distinct information layer -- intrinsic risk, laboratory snapshots, medication exposure, longitudinal care trajectories, and social determinants of health (SDOH) -- while retaining all prior features. All models were evaluated under a fixed screening policy targeting 90% specificity, so that the false positive rate remained constant as the information available to the model grew. External validation was conducted in the BioMe Biobank (N = 9,818) without retraining. Results:Discrimination improved consistently across layers, from AUROC 0.673 (M1) to 0.797 (M5). Under the fixed screening policy, sensitivity nearly doubled from 0.27 to 0.49, with a cumulative recovery of 30.4% of cases missed by the base model. Gains were driven by distinct subgroups at each transition: laboratory features identified biologically high-risk individuals; medication features captured those with high treatment intensity reflecting advanced cardiometabolic burden; longitudinal care trajectory features rescued cases with biological instability observable only through repeated measurements; and SDOH features recovered individuals with limited clinical observability, with rescue probability highest among those with the fewest recorded monitoring domains. Sparse data in the clinical record indicated low observability, not low risk. Social and genetic features each contributed most when downstream physiologic signal was limited, supporting a contextual rather than universal role for each. In BioMe, discrimination was attenuated (M4 AUROC 0.659), but the relative ordering of information layers was fully preserved, and a systematic upward shift in predicted probability distributions underscored the need for recalibration before deployment in a new setting. Conclusions:DKD risk detection in T2D is substantially improved by integrating complementary information layers under a fixed clinical screening policy, with gains arising from distinct domains that identify at-risk individuals in different clinical contexts. The layered landmark framework introduced here reveals how risk observability -- shaped by monitoring intensity, healthcare engagement, and access -- determines what a screening model can detect, and provides a foundation for context-aware EHR-based screening that accounts for data availability at the time of risk assessment.
BACKGROUND:Irritable bowel syndrome (IBS) is a complex disorder of gut-brain interaction, with heterogeneous symptoms, no available biomarkers and limited pathogenetic insight. OBJECTIVE:To identify genetic risk factors and actionable mechanisms for future clinical translation in IBS. DESIGN:We conducted a genome-wide association study (GWAS) meta-analysis of IBS in 2 775 539 individuals from 22 biobanks. IBS genetics was studied across multiple ancestries, different case definitions and symptom-related subtypes. Heritability and genetic correlations with other traits were estimated, and Mendelian randomisation was used to test causal relationships. GWAS data were functionally annotated and fine-mapped to prioritise tissues, cell types, pathways, candidate genes, specific mechanisms and druggable targets. RESULTS:Significant heritability was only detected in individuals of European ancestry, with near-identical genetic architecture across case definitions. Genetic correlations with GI, psychiatric and cardiometabolic traits were observed, including causal relationships with triglyceride (TG) levels. Functional annotation of IBS risk loci highlighted cell types and pathways relevant to brain, enteric neuro-glial and cardiometabolic domains, as well as actionable targets like GCKR, a regulator of TG metabolism. Druggability analyses converged on cardiometabolic mechanisms, including TG modulation. IBS polygenic risk scores were derived and showed a significant association with case status in an independent case-control dataset, supporting further evaluation in external population-based and clinically ascertained cohorts. CONCLUSIONS:This study provides the most comprehensive assessment of IBS genetics to date, demonstrating reproducible polygenic inheritance. We link IBS risk to convergent neurogastrointestinal and novel cardiometabolic mechanisms, highlight specific biological pathways and actionable mechanisms and outline translational opportunities emerging from integrated computational analyses.
STUDY OBJECTIVES:Sleep duration influences metabolic health, but the impact of daily sleep variation on next-day glycaemia in healthy individuals under real-world conditions is poorly understood. MATERIALS AND METHODS:We studied 206 adolescents (18 years) from the Copenhagen Prospective Studies on Asthma in Childhood 2000 (COPSAC2000) cohort with 2245 person-days of overlapping accelerometer-derived sleep and continuous glucose monitoring (CGM) recordings (median 13 days, IQR 9-13). Sleep duration was assessed using wrist-worn accelerometry, and glycaemic concentration, variability, and risk indices were derived from CGM during the accelerometer-defined waking period. Associations were examined using linear mixed-effects models adjusted for sociodemographic, behavioural, circadian, and cardiometabolic factors, with random effects of individuals across repeated days. RESULTS:Each additional hour of sleep was associated with higher next-day glycaemic concentration (median β 0.39 mg/dL [0.15, 0.63], p = .002), lower variability (standard deviation (SD) mg/dL β -0.12 [-0.23, -0.01], p = .036), and reduced deviation risk (Average Daily Risk Range (ADRR), indicating lower risk of extreme glucose excursions; β -0.27 [-0.43, -0.10], p = .002). Within-person deviations in sleep predicted next-day glycaemic concentration and deviation risk, whereas habitual between-person differences were more strongly associated with variability. Higher daytime glycaemic variability predicted shorter subsequent sleep (β -0.11 h [-0.18, -0.05], p < 0.001). The early-morning pre-wake glucose rise partly mediated the link between longer sleep and higher next-day median glucose (indirect effect 5.0%, p = 0.036). CONCLUSIONS:These findings indicate dynamic, bidirectional coupling between sleep and glucose regulation in free-living adolescents, with longer sleep associated with lower next-day glycaemic variability and reduced risk of extreme glucose excursions.
Large-scale multiancestry genome-wide association studies have identified hundreds of loci associated with type 2 diabetes (T2D) and glycemic traits, yet imputed genotyping arrays limit the detection of low-frequency and rare variants. Whole-genome sequencing (WGS) offers a more complete view of genetic variation, especially across diverse populations. We analyzed high-coverage (38×) WGS data from 21,913 T2D case subjects, 61,036 control subjects, and up to 50,011 individuals with no diabetes with fasting glucose, fasting insulin, and HbA1c from the National Heart, Lung, and Blood Institute Trans-Omics for Precision Medicine Program. We performed single-variant association testing, conditional analysis, fine-mapping, and Bayesian colocalization to identify genetic signals and assess regulatory relevance in diabetes-related tissues. We identified 76 distinct association signals across 34 loci, including novel variants at DUSP9 for T2D, and ROBO1, NDN, and MYT1 for HbA1c. Fine-mapping narrowed credible sets and improved causal variant resolution. Colocalization highlighted 80 expression signals in diabetes-related tissues, linking genetic associations to functional regulatory mechanisms. Our findings demonstrate the utility of WGS to uncover novel variants in diverse populations, enhance locus resolution, and link regulatory variation to disease-relevant tissues. This work refines the genetic architecture of T2D and glycemic traits and supports precision medicine efforts targeting diverse populations. ARTICLE HIGHLIGHTS:We aimed to improve understanding of the genetic architecture of type 2 diabetes and glycemic traits by leveraging whole-genome sequencing in diverse populations. Our goal was to identify novel variants, refine known loci, and link genetic signals to regulatory mechanisms through colocalization with expression quantitative trait loci. We discovered novel variants, significantly improved fine-mapping resolution, and identified 80 regulatory colocalization signals in diabetes-relevant tissues. These findings support precision medicine approaches by connecting genetic variation to functional biology in type 2 diabetes.
OBJECTIVE:To delineate organ-specific and systemic drivers of metabolic dysfunction-associated steatotic liver disease (MASLD), we applied integrative causal inference across clinical, imaging, and proteomic domains in individuals with and without type 2 diabetes (T2D). METHODS:Bayesian network analyses and complementary two-sample Mendelian randomization were used to quantify causal pathways linking adipose distribution, glycemia, and insulin dynamics with liver fat in the IMI-DIRECT prospective cohort study. Data included frequently sampled metabolic challenge tests, MRI-derived abdominal and hepatic fat content, serological biomarkers, and Olink plasma proteomics from 331 adults with new-onset T2D and 964 adults without diabetes, with harmonized protocols enabling replication. RESULTS:High basal insulin secretion rate (BasalISR), estimated via C-peptide deconvolution, emerged as the primary potential causal driver of liver fat accumulation in both cohorts. BasalISR, a clearance-independent measure of β-cell insulin output distinct from peripheral insulin levels, was independently linked to hepatic steatosis. Visceral adipose tissue exhibited bidirectional associations with liver fat, suggesting a self-reinforcing metabolic loop. Of 446 analyzed proteins, 34 mapped to these metabolic networks (27 in the non-diabetes network, 18 in the T2D network, and 11 shared). Key proteins directly associated with liver fat included GUSB, ALDH1A1, LPL, IGFBP1/2, CTSD, HMOX1, FGF21, AGRP, and ACE2. Sex-stratified analyses identified GUSB in females and LEP in males as the strongest protein predictors of liver fat. CONCLUSIONS:BasalISR may better capture early β-cell-driven disturbances contributing to MASLD. These findings outline a multifactorial, sex- and disease stage-specific proteo-metabolic architecture of hepatic steatosis and identify potential biomarkers or therapeutic targets.
BACKGROUND:Polygenic risk scores (PRSs) improve prediction of the development of type 2 diabetes over the use of clinical risk factors alone; however, they perform poorly in populations of non-European ancestry, limiting their global clinical utility. We aimed to deliver comprehensive and rigorously tested multi-ancestry PRSs for prediction in type 2 diabetes. METHODS:We conducted meta-analyses using data from type 2 diabetes genome-wide association studies (GWAS) across cohorts from five major global ancestries: European, African or African American, Admixed American, South Asian, and East Asian. We used summary statistics from the GWAS to construct single-ancestry PRSs (using the continuous-shrinkage PRS-CS method) and multi-ancestry PRSs (using the PRS-CSx method), and constructed ancestry-specific linkage disequilibrium panels to model pairwise correlations between single-nucleotide polymorphisms in GWAS during PRS construction. Models were validated for association with type 2 diabetes in at least four independent cohorts per ancestry. The effect sizes of PRSs were estimated as the odds ratio (OR) per SD of the PRS, and ORs for individuals at the 90th, 95th, and 97·5th PRS percentiles were compared with the IQR as a reference. We also tested our PRS models for prediction of diabetes incidence with or without additional clinical factors, as well as microvascular complications and comorbidities. FINDINGS:Our analysis used data from 409 959 individuals with type 2 diabetes and 1 983 345 controls: respectively, 359 819 and 1 825 729 indivduals were included in the GWAS dataset, with 10 992 and 31 792 individuals in the training dataset and 39 148 and 125 824 individuals in the validation dataset. The best predictive performance for the single-ancestry PRSs was in European (incremental AUC 0·07-0·14) and East Asian (0·02-0·16) ancestries, whereas prediction was poorer for African or African American (0·02-0·03), Admixed American (0·02-0·04), and South Asian (0·02-0·04) ancestries, correlating with sample sizes in the GWAS. Compared with single-ancestry PRSs, our multi-ancestry PRSs showed higher effect sizes and smaller 95% CIs across all ancestries: OR per SD 1·73 (95% CI 1·67-1·80) in African or African American, 2·82 (2·67-2·97) in Admixed American, 2·45 (2·36-2·54) in East Asian, 2·36 (2·32-2·41) in European, and 2·23 (2·05-2·42) in South Asian ancestries. Individuals in the 97·5th PRS percentile had a 3-7 times increased risk of type 2 diabetes compared with those in the IQR (OR 3·43 [95% CI 2·80-4·21] in African or African American, 7·47 [5·64-9·89] in Admixed American, 6·62 [5·58-7·85] in East Asian, 6·25 [5·72-6·82] in European, and 4·50 [2·70-7·53] in South Asian ancestries). These PRSs were also associated with earlier onset of type 2 diabetes, higher risk of developing microvascular complications, and provide additional predictive value beyond clinical factors. In individuals with type 2 diabetes, the association between multi-ancestry PRSs and risk of microvascular complications and comorbidity was studied in populations of African, Admixed American, and European ancestries and was significant in all three ancestry groups for diabetic retinopathy (ORs per SD 1·28-1·57), diabetic nephropathy (1·25-1·58), proliferative diabetic retinopathy (1·39-2·08), and end-stage diabetic nephropathy (1·44-1·87); PRS was associated with coronary artery disease in the Admixed American ancestry group only (1·16 [95% CI 1·08-1·25]). INTERPRETATION:These validated, publicly available PRSs can improve risk stratification for type 2 diabetes onset and complications across diverse ancestries, supporting their further evaluation in clinical settings. FUNDING:The National Human Genome Research Institute of the US National Institutes of Health.
Objective The aim of this study was to assess whether daily step counts and genetic risk interact to influence the risk of developing type 2 diabetes.Research Design and Methods We analyzed data from 9501 participants in the All of Us Research Program with both genetic and wearable device-derived physical activity data and without diabetes at baseline and a median age of 56 years (42-66). Physical activity was quantified using daily step counts. Genetic risk was assessed using a global polygenic score. Incident type 2 diabetes was identified using electronic health record-linked diagnostic codes. Multivariable Cox proportional hazards models estimated hazard ratios (HRs) for type 2 diabetes across genetic risk and physical activity levels. We tested for additive interaction using the relative excess risk due to interaction (RERI). In secondary analyses, we used physical-activity intensity measures using wearable-derived and self-reported intensity levels.Results Type 2 diabetes incidence rates ranged from 4.1 per 1000 person-years (95% CI, 2.5-5.7) in individuals with high physical activity and low genetic risk to 18.4 (95% CI, 15.2-21.6) in those with low physical activity and high genetic risk (HR, 6.2 (95% CI: 3.97, 9.6)). A significant additive interaction was observed (RERI, 0.20; 95% CI, 0.04-0.36; P = .007), with 15% (95% CI, 2-27) of excess risk attributed to the interaction. Similar interaction patterns were found using device-based intensity metrics and self-reported physical activity measures.Conclusion These findings provide evidence of additive interactions between genetic risk and physical activity, underscoring the potential value of integrating genomic and device-derived data to identify individuals who would more likely benefit from increasing physical activity.
Body mass index (BMI), type 2 diabetes (T2D) and associated cardiometabolic features modify Alzheimer's disease (AD) risk, yet shared mechanisms remain poorly understood. Using sex- and age-stratified genotyping data for BMI and T2D, we investigate how these traits converge on shared genetic pathways to AD risk. Employing multi-trait, machine learning and single-cell transcriptomics, we identify sex-specific cardiometabolic liability linked to higher BMI-associated risk in women and T2D-driven risk in men. Variant-level analyses reveal AD risk associates with genetically-driven hypotension and hypoglycaemia. We identify 35 putative effector genes in seven independent loci colocalizing between BMI/T2D and AD, mapping to peripheral immune and metabolic tissues and cell-types. Pathway enrichment identifies druggable targets in calcium and potassium channel signaling. Across 81 approved drugs modulating shared risk genes, levosimendan - a calcium sensitizer for heart failure - inhibits tau oligomerization and emerges as a repurposing candidate. These findings elucidate sex-specific cardiometabolic drivers of AD, identify actionable biological pathways, and reveal drug candidates for AD prevention and treatment.
Rare coding variants can have large effects on disease risk and provide direct routes from human genetics to disease mechanisms and therapeutic targets, but their discovery is constrained by sample size, particularly for low-prevalence diseases. Here we establish the Biobank Rare Variant Analysis (BRaVa) consortium, a global rare variant association resource that integrates sequencing and linked health-record data from ten biobanks and cohorts comprising over 1.2 million individuals across diverse ancestries. We performed gene-based meta-analyses of rare coding variation across 33 clinical endpoints and 11 quantitative traits. Aggregating evidence across biobanks and ancestries identified 514 gene-trait associations, including 31 not previously reported in prior studies or curated association resources following systematic literature review. Notably, 36.1% of gene-level associations were undetectable in any individual biobank, and 91 emerged only through cross-ancestry meta-analysis, demonstrating that federated integration enables discovery beyond the reach of single cohorts. Similar gains were observed at the variant level, where 25.0% of phenotype-locus associations were detectable only through meta-analysis. Effect size estimates were correlated across ancestries with concordant directions of effect, supporting the generalizability of rare variant associations. The identified signals implicate pathways involved in transcriptional and epigenetic regulation, metabolism, vascular and epithelial biology, and immune function, highlighting rare coding variation as an engine for biological discovery across medical record phenotypes. For example, damaging variation in ANKRD12 implicates inflammatory transcriptional dysregulation in asthma and chronic obstructive pulmonary disease, and ultra-rare predicted loss-of-function variants in NAA15 link protein acetylation processes to type 2 diabetes risk. BRaVa establishes a scalable framework and freely available community resource for rare variant meta-analysis across global biobanks. Public release of gene- and variant-level association summary statistics provides a reference map of rare coding variant associations to support disease gene discovery, biological interpretation, and therapeutic target prioritization as sequencing-linked health-record resources continue to expand.
Type 2 diabetes (T2D) prevention efforts have largely focused on intervening when dysglycemia is already established. We propose that T2D prevention be reframed around prediabetes remission, with preservation and restoration of normoglycemia as the optimal clinical goal. The transition from normoglycemia through increasing dysglycemia to T2D is progressive and cumulatively shaped by biological, behavioral and environmental exposures across the life course. Prediabetes (intermediate hyperglycemia) remission is an achievable, pragmatic and measurable prevention target. Here we provide a life-course risk architecture for T2D integrating developmental, transitional and contextual determinants, defining critical windows of amplified metabolic vulnerability and potential restoration of normoglycemia. Precision prevention should target mechanistic heterogeneity, with aligned interventions that remain scalable, affordable and adaptable across socioeconomic settings. Our framework identifies ten priorities in T2D prevention, moving beyond traditional approaches toward context-specific, actionable interventions capable of altering the natural history of disease early in the life course and restoring metabolic health.
Polygenic scores (PGSs) have promising clinical applications for risk stratification, disease screening, and personalized medicine. However, most PGSs are trained on predominantly European ancestry cohorts and have limited portability to external populations. While cross-population PGSs have demonstrated greater generalizability than single-ancestry PGSs, they fail to properly account for individuals with recent admixture between continental ancestry groups. GAUDI, a recently proposed PGS method, overcomes this gap by leveraging local ancestry to estimate ancestry-specific effects, penalizing but allowing ancestry-differential effects. However, the modified fused LASSO approach used by GAUDI is computationally expensive and does not readily accommodate more than two-way admixture. To address these limitations, we introduce HAUDI, an efficient LASSO framework for admixed PGS construction. HAUDI reparameterizes the GAUDI model as a standard LASSO problem, allowing for extension to multiway admixture settings and far superior computational speed than GAUDI. In extensive simulations, HAUDI compares favorably to GAUDI while dramatically reducing computation time. In real data applications, HAUDI uniformly outperforms GAUDI across 18 clinical phenotypes, including total triglycerides, C-reactive protein, and mean corpuscular hemoglobin concentration, and shows substantial benefits over ancestry-agnostic PGSs for white blood cell count and chronic kidney disease. It is also substantially faster and more accurate than the recently proposed SDPR_admix method.
Rare coding genetic variants may exert large effects on risk of common disease, yet their contribution to disease architecture and their utility in gene prioritization remain limited by inadequate sample sizes. Here, we performed a massive-scale rare variant association study (RVAS), analyzing over 1.1 million sequenced participants among which 130,000 had atrial fibrillation (AF). Through a multi-mask burden testing approach, we identified 15 genes significantly associated with AF through rare large-effect variation. Integrative analyses revealed strong convergence between genes implicated by rare and common variation, and highlighted instances where RVAS data may aid in GWAS prioritization. Nevertheless, several RVAS genes were not among GWAS loci ( FAM189A2 , ACTC1 , FNIP1 , FBN1 ), or were not nominated through contemporary GWAS prioritization ( KDM5B , ZFP36L2 ). Finally, we observed that ultra-rare protein-disrupting variants - concentrated in a small number of large-effect size genes - explained at least 2% of AF susceptibility across European and African ancestry groups. These findings refine the genetic architecture of AF, while highlighting the value and cost of RVAS for genomic discovery in common disease.
ABSTRACT Alzheimer’s disease (AD) is defined by progressive neurodegeneration, yet a substantial fraction of its genetic risk maps to non-neuronal processes. While brain microglia are recognized contributors to AD pathogenesis, the extent to which AD genetic risk operates through peripheral tissues remains unclear. Here, we systematically partitioned AD polygenic risk across 166 tissue-level annotations and 4.4 million cells spanning 28 peripheral tissues and 100 brain regions. AD genetic risk was consistently enriched in peripheral immune, barrier, and metabolic tissues, not explained by brain or microglial transcriptional programs. The strongest signals localized to circulating and tissue-resident myeloid cells, hepatocytes, cardiac muscle cells, and intestinal epithelial populations. Immune-cell enrichment was independent of the APOE locus, whereas a substantial proportion of metabolic enrichment was attributable to APOE -region genes. Colocalization and Mendelian randomization analyses implicated peripheral regulatory mechanisms and highlighted predominantly protective immune-gene effects. These findings extend current models of AD genetics beyond brain-resident mechanisms.
To broaden our understanding of bradyarrhythmias and conduction disease, we performed common variant genome-wide association analyses in up to 1.3 million individuals and rare variant burden testing in 460,000 individuals for sinus node dysfunction (SND), distal conduction disease (DCD) and pacemaker (PM) implantation. We identified 13, 31 and 21 common variant loci for SND, DCD and PM, respectively. Four well-known loci (SCN5A/SCN10A, CCDC141, TBX20 and CAMK2D) were shared for SND and DCD, while others were more specific for SND or DCD. SND and DCD showed a moderate genetic correlation (rg = 0.63). Cardiomyocyte-expressed genes were enriched for contributions to DCD heritability. Rare-variant analyses implicated LMNA for all bradyarrhythmia phenotypes, SMAD6 and SCN5A for DCD and TTN, MYBPC3 and SCN5A for PM. These results show that variation in multiple genetic pathways (for example, ion channel function, cardiac developmental programs, sarcomeric structure and cellular homeostasis) appear critical to the development of bradyarrhythmias. Genome-wide analyses identify variants associated with sinus node dysfunction, distal conduction disease and pacemaker implantation, implicating ion channel function, cardiac developmental programs and sarcomeric structure in bradyarrhythmia susceptibility.
BACKGROUND: Genome-wide association studies have identified several hundred susceptibility single nucleotide variants for coronary artery disease (CAD). Despite single nucleotide variant-based genome-wide association studies improving our understanding of the genetics of CAD, the contribution of structural variants (SVs) to the risk of CAD remains largely unclear. METHOD AND RESULTS: We leveraged SVs detected from high-coverage whole genome sequencing data in a diverse group of participants from the National Heart Lung and Blood Institute's Trans-Omics for Precision Medicine program. Single variant tests were performed on 58 706 SVs in a study sample of 11 556 CAD cases and 42 907 controls. Additionally, aggregate tests using sliding windows were performed to examine rare SVs. One genome-wide significant association was identified for a common biallelic intergenic duplication on chromosome 6q21 (P=1.54E-09, odds ratio=1.34). The sliding window-based aggregate tests found 1 region on chromosome 17q25.3, overlapping USP36, to be significantly associated with coronary artery disease (P=1.03E-10). USP36 is highly expressed in arterial and adipose tissues while broadly affecting several cardiometabolic traits. CONCLUSIONS: Our results suggest that SVs, both common and rare, may influence the risk of coronary artery disease.
Obesity is a major public health crisis associated with high mortality rates. Previous genome-wide association studies (GWAS) investigating body mass index (BMI) have largely relied on imputed data from European individuals. This study leveraged whole-genome sequencing (WGS) data from 88,873 participants from the Trans-Omics for Precision Medicine (TOPMed) Program, of which 51% were of non-European population groups. We discovered 18 BMI-associated signals (P < 5 × 10-9). Notably, we identified and replicated a novel low frequency single nucleotide polymorphism (SNP) in MTMR3 that was common in individuals of African descent. Using a diverse study population, we further identified two novel secondary signals in known BMI loci and pinpointed two likely causal variants in the POC5 and DMD loci. Our work demonstrates the benefits of combining WGS and diverse cohorts in expanding current catalog of variants and genes confer risk for obesity, bringing us one step closer to personalized medicine.
Polygenic risk scores (PRS) hold prognostic value for identifying individuals at higher risk of type 2 diabetes (T2D). However, further characterization is needed to understand the generalizability of T2D PRS in diverse populations across various contexts. We characterized a multi-ancestry T2D PRS among 244,637 cases and 637,891 controls across eight populations from the Population Architecture Genomics and Epidemiology (PAGE) Study and 13 additional biobanks and cohorts. PRS performance was context dependent, with better performance in those who were younger, male, with a family history of T2D, without hypertension, and not obese or overweight. Additionally, the PRS was associated with various diabetes-related cardiometabolic traits and T2D complications, suggesting its utility for stratifying risk of complications and identifying shared genetic architecture between T2D and other diseases. These findings highlight the need to account for context when evaluating PRS as a tool for T2D risk prognostication and potentially generalizable associations of T2D PRS with diabetes-related traits despite differential performance in T2D prediction across diverse populations.