Proteomics holds great promise for identifying potentially druggable effectors of common diseases, yet its application at population-scale across diverse ancestries, remains challenging. Here, we developed genetic imputation models for 2,594 plasma proteins using proteomic and genetic data from 54,219 UK Biobank participants, validating their performance across multiple ancestry groups and in an independent cohort. Plasma proteomes were then imputed for over 640,000 participants in the UK Biobank and the All of Us Research Program. To assess its aetiological value at population-scale, a further proteome-wide association study of cardiovascular diseases was performed across six genetic ancestries. We identified ∼9000 protein-disease associations across 89 cardiovascular conditions (PheCodes), the majority of which show consistent effects across ancestries and biobanks, with many comprising known targets of drugs either approved or under development. The associations reveal both shared and distinct proteomic signatures across cardiovascular conditions and defined clusters of distinct pathophysiology with shared underlying molecular pathways. Integration of data on tissue specificity and single-cell transcriptomics prioritised liver-derived proteins in circulation as candidate effectors of coronary artery disease, highlighting inter-alpha-trypsin inhibitor heavy chain H4 (ITIH4) as a putative effector. Using a liver-targeted CRISPR gene-editing platform, we show that in vivo disruption of ITIH4 reduces plasma cholesterol and pro-atherogenic lipid species in a preclinical model, consistent with a causal role in cardiovascular disease. Our study enables study of large-scale proteomics in diverse populations, provides a systematic map of protein associations of cardiovascular diseases, and demonstrates the utility of genetically imputed proteomes for target discovery and experimental validation. To facilitate proteomic analyses for the research community, the resultant models and association results have been made freely available through the OmicsPred platform.
Abstract Background and Aims Despite treatment, patients with established atherosclerotic cardiovascular disease (ASCVD) are at high risk of recurrent events. Existing clinical risk scores for recurrence provide only moderate predictive performance and rely largely on the same conventional risk factors used to predict disease onset. Proteomics is a promising source of new biomarkers but the technologies need focused use cases in order to achieve utility and implementation. We aimed to determine whether plasma proteomics improves prediction of recurrent cardiovascular events beyond established clinical risk models in secondary prevention in a population-scale cohort. Methods Plasma proteomic profiles from ∼9,300 participants in the UK Biobank with established ASCVD at baseline were analysed using machine learning methods to derive and evaluate proteomic predictors of recurrent cardiovascular events. The top performing model comprised proteins with non-zero weights (full protein score). Predictive performance of the proteomic predictors, an established clinical risk score (SMART2), and their combination was evaluated across six pre-defined testing datasets representing multiple ethnic and geographic groups. A parsimonious set of proteins with existing clinical-grade enzyme-linked immunosorbent assays (ELISAs) available was then derived. Results The full protein score achieved higher performance for recurrent ASCVD than the SMART2 risk score across all ethnic and geographic subgroups (mean C-index 0.743 vs 0.653). Adding the full protein score to SMART2 improved discrimination, with the largest increase in White Irish participants (ΔC-index, 0.140; 95% CI, 0.074–0.205; P<0.001). However, adding SMART2 to the protein score provided minimal additional value. The parsimonious score preserved most of the discrimination of the full protein model with C-indices of the recurrent ASCVD risk model comprising age, sex and the parsimonious protein score being nearly identical to the full protein model in the largest testing set (0.723 vs 0.728 for White British in England and Wales). The parsimonious protein score showed a marked gradient of risk with the top, middle and bottom quintiles showing 10-year recurrent ASCVD rates of ∼27.4%, ∼9.6% and ∼2.4%, respectively. Conclusions In patients with established ASCVD, plasma protein measurements substantially improved prediction of recurrent events beyond conventional clinical risk factors, supporting their potential as a complementary tool to guide secondary prevention of cardiovascular disease.
Abstract Cardiovascular diseases (CVDs) are highly heritable, but pathogenesis at the organ and physiological level is still poorly defined. Polygenic risk scores (PRSs), which estimate individual genetic susceptibility to a disease, may allow for the identification of associated abnormal organ structures. Ultimately, identifying where cardiovascular polygenic risk manifests can guide early interventions, shape mechanistic hypotheses, and motivate prevention trials for cardiac remodelling. This study investigated the association between PRSs for five common CVDs [heart failure (HF), coronary artery disease (CAD), atrial fibrillation (AF), abdominal aortic aneurysm (AAA) and ischaemic stroke (IS)] and 28 imaging-derived phenotypes (IDPs) from cardiac magnetic resonance imaging of ∼62,000 participants in UK Biobank. To investigate the cardiac features associated with elevated polygenic risk of CVDs, we tested CVD PRSs against cardiac IDPs and identified 97 significant associations (FDR ≤ 0.05). We further identified 32 significant putative mediators between CVD PRSs and incident disease events, revealing that across CVDs, polygenic risk manifested as distinct patterns in cardiac structures. HF implicated all cardiac chambers, including left ventricular and left atrial dysfunction alongside enlarged aorta. AF was characterised by biatrial enlargement and reduced ejection fractions, most prominently in the left atrium but also involving left ventricular wall thickness. IS exhibited left ventricular hypertrophy and left atrial dysfunction, while CAD predominantly involved left ventricular hypertrophy. AAA was primarily characterised by enlarged descending aorta. Overall, cardiac IDPs mediated a substantial proportion of polygenic risk for CVDs, in particular for HF. Taken together, our results show that cardiac structure and function lie on the pathway between polygenic risk and cardiovascular events.
Genetic prediction of multi-omic data has emerged as a cost-effective alternative to direct omics profiling, particularly useful for identifying molecular features associated with disease susceptibility. However, despite its popularity, multi-omic imputation models are fragmented across studies, hindering findability, accessibility, interoperability and re-use. To address this, we developed OmicsPred (https://www.omicspred.org), a centralised platform for the deposition and dissemination of genetic prediction models of multi-omic traits. OmicsPred unifies the most commonly used molecular imputation models (e.g. from PredictDB) and other published studies totalling 3,339,469 prediction models spanning transcriptomic, proteomic, and metabolomic traits (as of May 2026). Each model is accompanied by metadata describing score development and predictive performance, and distributed in formats compatible with popular analytic tools, such as PGS Catalog Calculator and MetaXcan. To demonstrate the utility of the resource for systematic target discovery, we perform a multi-omic phenome-wide association analysis in Million Veterans Program data.
CONTEXT:Type 2 diabetes (T2D) is a major global concern, with Asia at its epicenter in recent years. Proteins, products of gene transcription, serve as dynamic biomarkers for pinpointing perturbed pathways in disease development. Previous T2D proteomic association studies primarily focused on European populations. OBJECTIVE:The aim of this study was to investigate the relationship between plasma proteins and the incidence of T2D in Asian individuals. METHODS:We examined the association of 4775 plasma proteins with incident T2D in a Singapore multi-ethnic cohort of 1659 Asian individuals (539 cases and 1120 controls) using logistic regression. We used 2-sample mendelian randomization and colocalization analysis to evaluate the causal relationship between proteins and T2D. RESULTS:Our analysis revealed 522 proteins that were associated with incident T2D after adjusting for age, sex, and ethnicity, and 17 proteins that remained statistically significantly associated after adjusting for other T2D risk factors such as fasting glucose, waist circumference, and triglycerides. Among the 522 proteins associated with incident T2D, the change in 205 plasma proteins, observed in parallel with the development of T2D at baseline and 6-year follow-up, were further associated with incident T2D. The associated proteins showed enrichment in neuron generation, glycosaminoglycan binding, and insulin-like growth factor binding. Two-sample mendelian randomization analysis suggested 3 plasma proteins, GSTA1, INHBC, and FGL1, play causal roles in the development of T2D, with colocalization evidence supporting GSTA1 and INHBC. CONCLUSION:Our findings reveal plasma protein profiles linked to the onset of T2D in Asian populations, offering insights into the biological mechanisms of T2D development.
1 Current cardiovascular disease (CVD) risk prediction models place many individuals in an intermediate risk category where clinical decision-making remains uncertain, highlighting a critical gap in precision prevention. Polygenic risk scores (PRS) represent a promising solution to enhance risk stratification in intermediate-risk groups by identifying individuals with high genetic risk; however, observed differences in performance across ancestry groups may cause health disparities. The emerging field of algorithmic fairness offers a principled frame-work to assess equity in model performance among relevant subgroups, but have rarely been applied to clinical risk tools and PRS. To evaluate the fairness of incorporating PRS in CVD risk prediction, both as a standalone risk factor and as a risk-enhancing factor for individuals at intermediate risk (recommended in current clinical consensus statements). Using data from the UK Biobank (N = 327,923), we calculated 10-year CVD risk using QRISK3 (a guideline-endorsed prediction model) and quantified genetic risk using a validated PRS. We assessed fairness among population characteristics relevant for health equity (age, sex, ethnicity, and area-level deprivation) using four algorithmic fairness metrics relevant for prevention (accuracy equality, equal opportunity, conditional use accuracy equality, and treatment equality). PRS, when used as a stand-alone risk factor, demonstrated fairness levels similar to or better than traditional clinical predictors (age, sex, blood pressure, cholesterol). Some variation in fairness was observed across ethnic groups, especially at extreme risk thresholds. When integrated as a risk-enhancing factor for reclassifying intermediate-risk individuals into high-risk categories, PRS improved sensitivity of CVD risk prediction with minimal impact on fairness metrics across demographic groups. This study demonstrates that PRS, when incorporated into existing risk prediction frameworks, are unlikely to meaningfully exacerbate disparities in CVD risk stratification. Applying algorithmic fairness metrics provides insight into the equitable implementation of PRS and supports current recommendations for their use in risk-stratification for intermediate-risk individuals.
Drug targets that are supported by human genetics are more likely to lead to approved therapies. Research now identifies several promising drug targets and therapeutic repurposing opportunities for heart failure and its clinical sub-types.
BACKGROUND AND AIMS:Clinical biomarkers, nuclear magnetic resonance (NMR) metabolomics biomarker scores, and polygenic risk scores (PRS) have shown promise for improving cardiovascular disease (CVD) prediction but have not yet been evaluated in the context of current prediction models (SCORE2) and ESC recommendations for 10-year prediction of fatal and non-fatal CVD. METHODS:NMR metabolomic biomarker scores were constructed and compared to clinical biomarkers, PRS and SCORE2 in 297 463 UK Biobank participants (8919 incident CVD cases) aged 40-69 without previous CVD, diabetes, or lipid-lowering treatment. Improvement in risk discrimination when added to SCORE2 was assessed using Harrel's C-index. Improvement in risk stratification following ESC guideline risk thresholds was assessed using categorical net reclassification. Population modelling was subsequently applied to estimate the impact on CVD prevention if applied at scale. RESULTS:Risk discrimination provided by SCORE2 (C-index: 0.719) improved when 11 clinical biomarkers (ΔC-index: 0.014 [0.012-0.015]), NMR metabolomic biomarker scores (ΔC-index: 0.010 [0.009-0.012]) and PRSs (ΔC-index 0.009; [0.008-0.011]) were added individually. The combination of 11 clinical biomarkers, NMR metabolomic biomarker scores, and PRSs yielded the largest improvement risk discrimination, with ΔC-index 0.024 (0.022-0.027). Concomitant improvements in risk stratification were observed in categorical net reclassification index, with net case reclassification of 16.66% (15.50%-17.81%). Modelling suggested that addition of these biomarkers to SCORE2 for targeted risk reclassification would increase the number of CVD events prevented per 100 000 screened from 229 to 413 (ΔCVDprevented: 184 [174-194]) while essentially maintaining the number of statins prescribed per CVD event prevented. CONCLUSIONS:Combining NMR metabolomic, polygenic, and clinical biomarkers with SCORE2 enhanced prediction of first-onset CVD and could have substantial population health benefit if applied at scale.
Blood cell phenotypes are routinely tested in healthcare to inform clinical decisions. Genetic variants influencing mean blood cell phenotypes have been used to understand disease aetiology and improve prediction; however, additional information may be captured by genetic effects on observed variance. Here, we mapped variance quantitative trait loci (vQTL), i.e. genetic loci associated with trait variance, for 29 blood cell phenotypes from the UK Biobank (N ~ 408,111). We discovered 176 independent blood cell vQTLs, of which 147 were not found by additive QTL mapping. vQTLs displayed on average 1.8-fold stronger negative selection than additive QTL, highlighting that selection acts to reduce extreme blood cell phenotypes. Variance polygenic scores (vPGSs) were constructed to stratify individuals in the INTERVAL cohort (N ~ 40,466), where the genetically most variable individuals had increased conventional PGS accuracy (by ~19%) relative to the genetically least variable individuals. Genetic prediction of blood cell traits improved by ~10% on average combining PGS with vPGS. Using Mendelian randomisation and vPGS association analyses, we found that alcohol consumption significantly increased blood cell trait variances highlighting the utility of blood cell vQTLs and vPGSs to provide novel insight into phenotype aetiology as well as improve prediction.
Variance quantitative trait loci (vQTLs), which capture genetic contributions to phenotypic variability, remain underexplored in proteomic studies, particularly across diverse ancestries. We systematically mapped cis-vQTLs for 2,923 plasma proteins in 52,706 UK Biobank participants of European (EUR, N = 45,486), African (AFR, N = 1,336), and Central/South Asian (CSA, N = 934) ancestries, identifying 2,162 vQTLs (PVE < 5 x 10-8) for 781 proteins. We identified ancestry-specific and shared cis-vQTLs, including those for 30 proteins which were shared across all ancestries, with a few proteins, exhibiting stronger associations in non-EUR ancestry groups despite smaller sample sizes. Across ancestries, 7% (EUR), 25% (AFR), and 14% (CSA) of associations had variance effects only (vQTLonly), lacking corresponding mean effects (PME > 0.05), with chromosome X enriched for vQTLonly associations. Finally, multivariable Mendelian randomization revealed that, independent of genetically predicted mean protein levels, genetically predicted variance of three proteins influenced disease risk of coronary artery disease (Lp(a) and VAMP5) or type 2 diabetes (ANGPTL4). The MR effects for protein levels and variance were independent yet directionally consistent and significant (FDR < 0.05). Taken together, this study identifies novel protein vQTLs, highlights their transferability and demonstrates the potential therapeutic relevance of protein variance. ### Competing Interest Statement M.I. is a trustee of the Public Health Genomics (PHG) Foundation, a member of the Scientific Advisory Board of Open Targets and has research collaborations with AstraZeneca and Nightingale Health which are unrelated to this study. ### Funding Statement The authors are grateful to the participants UK Biobank for access to the data used in this study (Project #7439). This work was performed using resources provided by the Cambridge Service for Data Driven Discovery (CSD3) operated by the University of Cambridge Research Computing Service (www.csd3.cam.ac.uk), provided by Dell EMC and Intel using Tier-2 funding from the Engineering and Physical Sciences Research Council (capital grant EP/P020259/1), and DiRAC funding from the Science and Technology Facilities Council (www.dirac.ac.uk). This work was supported by core funding from the British Heart Foundation (RG/F/23/110103), NIHR Cambridge Biomedical Research Centre (NIHR203312) [*], BHF Chair Award (CH/12/2/29428), Cambridge BHF Centre of Research Excellence (RE/24/130011), and by Health Data Research UK, which is funded by the UK Medical Research Council, Engineering and Physical Sciences Research Council, Economic and Social Research Council, Department of Health and Social Care (England), Chief Scientist Office of the Scottish Government Health and Social Care Directorates, Health and Social Care Research and Development Division (Welsh Government), Public Health Agency (Northern Ireland), British Heart Foundation and the Wellcome trust. S.C.R is funded by the BHF Cambridge Centre for Research Excellence RE/24/130011. X.J. was funded by a Wellcome Trust Fellowship [227566/Z/23/Z). S.A.L. was supported by a Canadian Institutes of Health Research postdoctoral fellowship (MFE-171279). Y.X. and M.I. were supported by the UK Economic and Social Research Council (ES/T013192/1). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. *The views expressed are those of the authors and not necessarily those of the NIHR or the Department of Health and Social Care. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: All data described are available through the UK Biobank subject to approval from the UK Biobank access committee. See https://www.ukbiobank.ac.uk/enable-your-research/apply-for-access for further details I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors
Genome-wide association studies have identified thousands of variants associated with disease risk but the mechanism by which such variants contribute to disease remains largely unknown. Indeed, a major challenge is that variants do not act in isolation but rather in the framework of highly complex biological networks, such as the human metabolic network, which can amplify or buffer the effect of specific risk alleles on disease susceptibility. Here we use genetically predicted reaction fluxes to perform a systematic search for metabolic fluxes acting as buffers or amplifiers of coronary artery disease (CAD) risk alleles. Our analysis identifies 30 risk locus-reaction flux pairs with significant interaction on CAD susceptibility involving 18 individual reaction fluxes and 8 independent risk loci. Notably, many of these reactions are linked to processes with putative roles in the disease such as the metabolism of inflammatory mediators. In summary, this work establishes proof of concept that biochemical reaction fluxes can have non-additive effects with risk alleles and provides novel insights into the interplay between metabolism and genetic variation on disease susceptibility.
The biological mechanisms through which most nonprotein-coding genetic variants affect disease risk are unknown. To investigate gene-regulatory mechanisms, we mapped blood gene expression and splicing quantitative trait loci (QTLs) through bulk RNA sequencing in 4,732 participants and integrated protein, metabolite and lipid data from the same individuals. We identified cis-QTLs for the expression of 17,233 genes and 29,514 splicing events (in 6,853 genes). Colocalization analyses revealed 3,430 proteomic and metabolomic traits with a shared association signal with either gene expression or splicing. We quantified the relative contribution of the genetic effects at loci with shared etiology, observing 222 molecular phenotypes significantly mediated by gene expression or splicing. We uncovered gene-regulatory mechanisms at disease loci with therapeutic implications, such as WARS1 in hypertension, IL7R in dermatitis and IFNAR2 in COVID-19. Our study provides an open-access resource on the shared genetic etiology across transcriptional phenotypes, molecular traits and health outcomes in humans ( https://IntervalRNA.org.uk ).
Genomics can provide insight into the etiology of type 2 diabetes and its comorbidities, but assigning functionality to non-coding variants remains challenging. Polygenic scores, which aggregate variant effects, can uncover mechanisms when paired with molecular data. Here, we test polygenic scores for type 2 diabetes and cardiometabolic comorbidities for associations with 2,922 circulating proteins in the UK Biobank. The genome-wide type 2 diabetes polygenic score associates with 617 proteins, of which 75% also associate with another cardiometabolic score. Partitioned type 2 diabetes scores, which capture distinct disease biology, associate with 342 proteins (20% unique). In this work, we identify key pathways (e.g., complement cascade), potential therapeutic targets (e.g., FAM3D in type 2 diabetes), and biomarkers of diabetic comorbidities (e.g., EFEMP1 and IGFBP2) through causal inference, pathway enrichment, and Cox regression of clinical trial outcomes. Our results are available via an interactive portal ( https://public.cgr.astrazeneca.com/t2d-pgs/v1/ ).
Multiomics has shown promise in noninvasive risk profiling and early detection of various common diseases. In the present study, in a prospective population-based cohort with ~18 years of e-health record follow-up, we investigated the incremental and combined value of genomic and gut metagenomic risk assessment compared with conventional risk factors for predicting incident coronary artery disease (CAD), type 2 diabetes (T2D), Alzheimer disease and prostate cancer. We found that polygenic risk scores (PRSs) improved prediction over conventional risk factors for all diseases. Gut microbiome scores improved predictive capacity over baseline age for CAD, T2D and prostate cancer. Integrated risk models of PRSs, gut microbiome scores and conventional risk factors achieved the highest predictive performance for all diseases studied compared with models based on conventional risk factors alone. The present study demonstrates that integrated PRSs and gut metagenomic risk models improve the predictive value over conventional risk factors for common chronic diseases.
Combining information from multiple GWASs for a disease and its risk factors has proven a powerful approach for development of polygenic risk scores (PRSs). This may be particularly useful for type 2 diabetes (T2D), a highly polygenic and heterogeneous disease where the additional predictive value of a PRS is unclear. Here, we use a meta-scoring approach to develop a metaPRS for T2D that incorporated genome-wide associations from both European and non-European genetic ancestries and T2D risk factors. We evaluated the performance of this metaPRS and benchmarked it against existing genome-wide PRS in 620,059 participants and 50,572 T2D cases amongst six diverse genetic ancestries from UK Biobank, INTERVAL, the All of Us Research Program, and the Singapore Multi-Ethnic Cohort. We show that our metaPRS was the most powerful PRS for predicting T2D in European population-based cohorts and had comparable performance to the top ancestry-specific PRS, highlighting its transferability. In UK Biobank, we show the metaPRS had stronger predictive power for 10-year risk than all individual risk factors apart from BMI and biomarkers of dysglycemia. The metaPRS modestly improved T2D risk stratification of QDiabetes risk scores for 10-year risk prediction, particularly when prioritising individuals for blood tests of dysglycemia. Overall, we present a highly predictive and transferrable PRS for T2D and demonstrate that the potential for PRS to incrementally improve T2D risk prediction when incorporated into UK guideline-recommended screening and risk prediction with a clinical risk score.
Introduction Type 2 diabetes (T2D) is a heterogeneous disorder for which disease-causing pathways are incompletely understood. Here, we mapped genetic risk for T2D and its comorbidities to proteins, mechanistic pathways and clinical outcomes using proteogenomic data from a population-scale biobank and two randomized controlled trials. Methods We tested polygenic scores (PGS) for T2D and its cardiometabolic comorbidities, plus five partitioned T2D PGS (beta cell, lipodystrophy, liver lipid, obesity, and liver lipid), for association with 2,922 circulating proteins in 54,306 multi-ancestry participants (of which 42,452 were unrelated and without prevalent cardiometabolic disease) from the UK Biobank (UKB). Then, we tested the PGS-associated proteins for association with incident cardiometabolic complications in two cardiovascular outcome trials among T2D patients with proteogenomic data: EXSCEL (N=2,823) and DECLARE-TIMI 58 (N=915). We assessed causality using two-sample Mendelian randomization and mediation. Results We identified 839 unique proteins significantly associated with any T2D PGS and 1,005 proteins that were associated with at least one cardiometabolic PGS. Some PGS-associated proteins such as TFF3, EFEMP1, and MMP12 were in turn associated with renal and cardiovascular trial outcomes. PGS association patterns revealed shared pathways, e.g., complement cascade, cholesterol metabolism, IGF signaling. The proteins underlying these pathways, such as LPA, C1S, and IGFBP2, were consistently associated with clinical trial outcomes or identified via causal inference. Conclusions This proteogenomic study revealed proteins and mechanistic pathways underlying T2D and related comorbidities, advancing our understanding of T2D pathobiology and identifying putative biomarkers. All our results are available in an online data portal (<https://public.cgr.astrazeneca.com/t2d-pgs/v1/>). ### Competing Interest Statement D.P.L., M.G., D.M., D.V., X.J., I.A.G., S.P., J.O., A.N., and D.S.P. are employees of AstraZeneca and may hold AstraZeneca stock options. B.B.S. and H.R. are employees of Biogen and may hold stock options. C.D.W. is an employee of Janssen Pharmaceuticals, a Johnson & Johnson company, and may hold stock options. R.R.H. reports personal fees from Anji Pharmaceuticals, AstraZeneca and Novartis. R.J.M. received research support and honoraria from Abbott, American Regent, Amgen, AstraZeneca, Bayer, Boehringer Ingelheim, Boston Scientific, Cytokinetics, Fast BioMedical, Gilead, Innolife, Eli Lilly, Medtronic, Medable, Merck, Novartis, Novo Nordisk, Pfizer, Pharmacosmos, Relypsa, Respicardia, Roche, Rocket Pharmaceuticals, Sanofi, Verily, Vifor, Windtree Therapeutics, and Zoll. M.I. is a trustee of the Public Health Genomics (PHG) Foundation, a member of the Scientific Advisory Board of Open Targets and has research collaborations with Nightingale Health and Pfizer which are unrelated to this study. ### Funding Statement UK Biobank proteomics data was funded by a consortia of 13 participating pharmaceutical companies (Alnylam Pharmaceuticals, Amgen, AstraZeneca, Biogen, Bristol Myers Squibb, Calico, Genentech, GlaxoSmithKline, The Janssen Pharmaceutical Companies of Johnson & Johnson, Novo Nordisk, Pfizer, Regeneron and Takeda). The DECLARE-TIMI 58 and EXSCEL clinical trial were funded by AstraZeneca. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Information on EXSCEL and DECLARE-TMI 58 trials can be found on [clinicaltrials.gov][1] ([NCT01144338][2] for EXSCEL and [NCT01144338][2] for DECLARE). The trial protocols were approved by institutional review board at each participating site. UK Biobank has approval from the North West Multi-centre Research Ethics Committee (MREC) as a Research Tissue Bank (RTB). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as [ClinicalTrials.gov][3]. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors. Polygenic scores will be uploaded to the PGS Catalog. [1]: http://clinicaltrials.gov [2]: /lookup/external-ref?link_type=CLINTRIALGOV&access_num=NCT01144338&atom=%2Fmedrxiv%2Fearly%2F2024%2F03%2F19%2F2024.03.15.24304200.atom [3]: http://ClinicalTrials.gov
Polygenic scores (PGSs) have transformed human genetic research and have numerous potential clinical applications. Here we present a series of recent enhancements to the PGS Catalog and highlight the PGS Catalog Calculator, an open-source, scalable and portable pipeline for reproducibly calculating PGSs that democratizes equitable PGS applications.