Chronological age is a major risk factor for numerous diseases. However, chronological age does not capture the complex biological aging process. Biological aging can occur at a different pace in individuals of the same chronological age. Therefore, the difference between the chronological age and biologically driven aging could be more informative in reflecting health status. Metabolite levels are thought to reflect the integrated effects of both genetic and environmental factors on the rate of aging, and may thus provide a stronger signature for biological age than those previously developed using methylation and proteomics. Here, we set out to develop a metabolomic age prediction model by applying ridge regression and bootstrapping with 826 metabolites (of which 678 endogenous and 148 xenobiotics) measured by an untargeted high-performance liquid chromatography mass spectrometry platform (Metabolon) in 11,977 individuals (50.2% men) from the INTERVAL study (Cambridge, UK). Participants of the INTERVAL study are relatively healthy blood donors aged 18-75 years. After internal validation using bootstrapping, the models demonstrated high performance with an adjusted R 2 of 0.82 using the endogenous metabolites only and an adjusted R 2 of 0.83 when using the full set of 826 metabolites with age as outcome. The latter model performance could be indicative of xenobiotics predicting frailty. In summary, we developed robust models for predicting metabolomic age in a large relatively healthy population with a wide age range.
BACKGROUND AND AIMS:Clinical biomarkers, nuclear magnetic resonance (NMR) metabolomics biomarker scores, and polygenic risk scores (PRS) have shown promise for improving cardiovascular disease (CVD) prediction but have not yet been evaluated in the context of current prediction models (SCORE2) and ESC recommendations for 10-year prediction of fatal and non-fatal CVD. METHODS:NMR metabolomic biomarker scores were constructed and compared to clinical biomarkers, PRS and SCORE2 in 297 463 UK Biobank participants (8919 incident CVD cases) aged 40-69 without previous CVD, diabetes, or lipid-lowering treatment. Improvement in risk discrimination when added to SCORE2 was assessed using Harrel's C-index. Improvement in risk stratification following ESC guideline risk thresholds was assessed using categorical net reclassification. Population modelling was subsequently applied to estimate the impact on CVD prevention if applied at scale. RESULTS:Risk discrimination provided by SCORE2 (C-index: 0.719) improved when 11 clinical biomarkers (ΔC-index: 0.014 [0.012-0.015]), NMR metabolomic biomarker scores (ΔC-index: 0.010 [0.009-0.012]) and PRSs (ΔC-index 0.009; [0.008-0.011]) were added individually. The combination of 11 clinical biomarkers, NMR metabolomic biomarker scores, and PRSs yielded the largest improvement risk discrimination, with ΔC-index 0.024 (0.022-0.027). Concomitant improvements in risk stratification were observed in categorical net reclassification index, with net case reclassification of 16.66% (15.50%-17.81%). Modelling suggested that addition of these biomarkers to SCORE2 for targeted risk reclassification would increase the number of CVD events prevented per 100 000 screened from 229 to 413 (ΔCVDprevented: 184 [174-194]) while essentially maintaining the number of statins prescribed per CVD event prevented. CONCLUSIONS:Combining NMR metabolomic, polygenic, and clinical biomarkers with SCORE2 enhanced prediction of first-onset CVD and could have substantial population health benefit if applied at scale.
Case-only designs in longitudinal cohorts are a valuable resource for identifying disease-relevant genes, pathways, and novel targets influencing disease progression. This is particularly relevant in Alzheimer's disease (AD), where longitudinal cohorts measure disease "progression," defined by rate of cognitive decline. Few of the identified drug targets for AD have been clinically tractable, and phenotypic heterogeneity is an obstacle to both clinical research and basic science. In four cohorts (n = 7241), we performed genome-wide association studies (GWAS) and Mendelian randomization (MR) to discover novel targets associated with progression and assess causal relationships. We tested opportunities for patient stratification by deriving polygenic risk scores (PRS) for AD risk and severity and tested the value of these scores in predicting progression. Genome-wide association studies identified no loci associated with progression at genome-wide significance (α = 5×10-8); MR analyses provided no significant evidence of an association between cognitive decline in AD patients and protein levels in brain, cerebrospinal fluid (CSF), and plasma. Polygenic risk scores for AD risk did not reliably stratify fast from slow progressors; however, a deeper investigation found that APOE ε4 status predicts amyloid-β and tau positive versus negative patients (odds ratio for an additional APOE ε4 allele = 5.78 [95% confidence interval: 3.76-8.89], P<0.001) when restricting to a subset of patients with available CSF biomarker data. These results provided no evidence for large-effect, common-variant loci involved in the rate of memory decline, suggesting that patient stratification based on common genetic risk factors for progression may have limited utility. Where clinically relevant biomarkers suggest diagnostic heterogeneity, there is evidence that a priori identified genetic risk factors may have value in patient stratification. Mendelian randomization was less tractable due to the lack of large-effect loci, and future analyses with increased samples sizes are needed to replicate and validate our results.
For many diseases there are delays in diagnosis due to a lack of objective biomarkers for disease onset. Here, in 41,931 individuals from the United Kingdom Biobank Pharma Proteomics Project, we integrated measurements of ~3,000 plasma proteins with clinical information to derive sparse prediction models for the 10-year incidence of 218 common and rare diseases (81–6,038 cases). We then compared prediction models developed using proteomic data with models developed using either basic clinical information alone or clinical information combined with data from 37 clinical assays. The predictive performance of sparse models including as few as 5 to 20 proteins was superior to the performance of models developed using basic clinical information for 67 pathologically diverse diseases (median delta C-index = 0.07; range = 0.02–0.31). Sparse protein models further outperformed models developed using basic information combined with clinical assay data for 52 diseases, including multiple myeloma, non-Hodgkin lymphoma, motor neuron disease, pulmonary fibrosis and dilated cardiomyopathy. For multiple myeloma, single-cell RNA sequencing from bone marrow in newly diagnosed patients showed that four of the five predictor proteins were expressed specifically in plasma cells, consistent with the strong predictive power of these proteins. External replication of sparse protein models in the EPIC-Norfolk study showed good generalizability for prediction of the six diseases tested. These findings show that sparse plasma protein signatures, including both disease-specific proteins and protein predictors shared across several diseases, offer clinically useful prediction of common and rare diseases.
Pancreatic ductal adenocarcinoma (PDAC) remains a lethal malignancy, largely due to the paucity of reliable biomarkers for early detection and therapeutic targeting. Existing blood protein biomarkers for PDAC often suffer from replicability issues, arising from inherent limitations such as unmeasured confounding factors in conventional epidemiologic study designs. To circumvent these limitations, we use genetic instruments to identify proteins with genetically predicted levels to be associated with PDAC risk. Leveraging genome and plasma proteome data from the INTERVAL study, we established and validated models to predict protein levels using genetic variants. By examining 8,275 PDAC cases and 6,723 controls, we identified 40 associated proteins, of which 16 are novel. Functionally validating these candidates by focusing on 2 selected novel protein-encoding genes, GOLM1 and B4GALT1, we demonstrated their pivotal roles in driving PDAC cell proliferation, migration, and invasion. Furthermore, we also identified potential drug repurposing opportunities for treating PDAC. SIGNIFICANCE:PDAC is a notoriously difficult-to-treat malignancy, and our limited understanding of causal protein markers hampers progress in developing effective early detection strategies and treatments. Our study identifies novel causal proteins using genetic instruments and subsequently functionally validates selected novel proteins. This dual approach enhances our understanding of PDAC etiology and potentially opens new avenues for therapeutic interventions.
BACKGROUND:Specific peripheral proteins have been implicated to play an important role in the development of Alzheimer's disease (AD). However, the roles of additional novel protein biomarkers in AD etiology remains elusive. The availability of large-scale AD GWAS and plasma proteomic data provide the resources needed for the identification of causally relevant circulating proteins that may serve as risk factors for AD and potential therapeutic targets. METHODS:We established and validated genetic prediction models for protein levels in plasma as instruments to investigate the associations between genetically predicted protein levels and AD risk. We studied 71,880 (proxy) cases and 383,378 (proxy) controls of European descent. RESULTS:We identified 69 proteins with genetically predicted concentrations showing associations with AD risk. The drugs almitrine and ciclopirox targeting ATP1A1 were suggested to have a potential for being repositioned for AD treatment. CONCLUSIONS:Our study provides additional insights into the underlying mechanisms of AD and potential therapeutic strategies.
Metabolomic platforms using nuclear magnetic resonance (NMR) spectroscopy can now rapidly quantify many circulating metabolites which are potential biomarkers of cardiovascular disease (CVD). Here, we analyse ∼170,000 UK Biobank participants (5,096 incident CVD cases) without a history of CVD and not on lipid-lowering treatments to evaluate the potential for improving 10-year CVD risk prediction using NMR biomarkers in addition to conventional risk factors and polygenic risk scores (PRSs). Using machine learning, we developed sex-specific NMR scores for coronary heart disease (CHD) and ischaemic stroke, then estimated their incremental improvement of 10-year CVD risk prediction when added to guideline-recommended risk prediction models (i.e., SCORE2) with and without PRSs. The risk discrimination provided by SCORE2 (Harrell’s C-index = 0.718) was similarly improved by addition of NMR scores (ΔC-index 0.011; 0.009, 0.014) and PRSs (ΔC-index 0.009; 95% CI: 0.007, 0.012), which offered largely orthogonal information. Addition of both NMR scores and PRSs yielded the largest improvement in C-index over SCORE2, from 0.718 to 0.737 (ΔC-index 0.019; 95% CI: 0.016, 0.022). Concomitant improvements in risk stratification were observed in categorical net reclassification index when using guidelines-recommended risk categorisation, with net case reclassification of 13.04% (95% CI: 11.67%, 14.41%) when adding both NMR scores and PRSs to SCORE2. Using population modelling, we estimated that targeted risk-reclassification with NMR scores and PRSs together could increase the number of CVD events prevented per 100,000 screened from 201 to 370 (ΔCVDprevented: 170; 95% CI: 158, 182) while essentially maintaining the number of statins prescribed per CVD event prevented. Overall, we show combining NMR scores and PRSs with SCORE2 moderately enhances prediction of first-onset CVD, and could have substantial population health benefit if applied at scale. ### Competing Interest Statement During the course of this project P.S. became a full-time employee of GSK Plc. All significant contributions to this study were made prior to this role and GSK Plc had no input to the study. J.D. serves on scientific advisory boards for AstraZeneca, Novartis, and UK Biobank, and has received multiple grants from academic, charitable and industry sources outside of the submitted work. A.S.B. reports institutional grants from AstraZeneca, Bayer, Biogen, BioMarin, Bioverativ, Novartis, Regeneron and Sanofi. The remaining authors declare no competing interests. ### Funding Statement This work was performed using resources provided by the Cambridge Service for Data Driven Discovery (CSD3) operated by the University of Cambridge Research Computing Service ([www.csd3.cam.ac.uk][1]), provided by Dell EMC and Intel using Tier-2 funding from the Engineering and Physical Sciences Research Council (capital grant EP/P020259/1), and DiRAC funding from the Science and Technology Facilities Council ([www.dirac.ac.uk][2]). This work was supported by core funding from the: Cambridge BHF Centre of Research Excellence (RE/18/1/34212) and BHF Chair Award (CH/12/2/29428). S.C.R. and S.K. were funded by a British Heart Foundation (BHF) Programme Grant (RG/18/13/33946). S.C.R. was also funded by the National Institute for Health and Care Research (NIHR) Cambridge BRC (BRC-1215-20014; NIHR203312) [*]. X.J. was funded by British Heart Foundation (CH/12/2/29428) and Wellcome Trust (227566/Z/23/Z). L.P. and P.S. were supported by a Rutherford Fund Fellowship from the Medical Research Council grant MR/S003746/1. Y.X. and M.I. were supported by the UK Economic and Social Research Council (ES/T013192/1). S.A.L. was supported by a Canadian Institutes of Health Research postdoctoral fellowship (MFE-171279). E.D.A. holds a NIHR Senior Investigator Award. J.D. holds a BHF Professorship and a NIHR Senior Investigator Award. M.I. is supported by the Munz Chair of Cardiovascular Prediction and Prevention and the NIHR Cambridge Biomedical Research Centre (NIHR203312). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. *The views expressed are those of the authors and not necessarily those of the NIHR or the Department of Health and Social Care. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This study was approved under UK Biobank Projects 30418 and ethics approval was obtained from the North West Multi-Center Research Ethics Committee. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data described are available through UK Biobank subject to approval from the UK Biobank access committee. See for further details. [1]: http://www.csd3.cam.ac.uk [2]: http://www.dirac.ac.uk
Genome-wide association analyses using high-throughput metabolomics platforms have led to novel insights into the biology of human metabolism 1 – 7 . This detailed knowledge of the genetic determinants of systemic metabolism has been pivotal for uncovering how genetic pathways influence biological mechanisms and complex diseases 8 – 11 . Here we present a genome-wide association study for 233 circulating metabolic traits quantified by nuclear magnetic resonance spectroscopy in up to 136,016 participants from 33 cohorts. We identify more than 400 independent loci and assign probable causal genes at two-thirds of these using manual curation of plausible biological candidates. We highlight the importance of sample and participant characteristics that can have significant effects on genetic associations. We use detailed metabolic profiling of lipoprotein- and lipid-associated variants to better characterize how known lipid loci and novel loci affect lipoprotein metabolism at a granular level. We demonstrate the translational utility of comprehensively phenotyped molecular data, characterizing the metabolic associations of intrahepatic cholestasis of pregnancy. Finally, we observe substantial genetic pleiotropy for multiple metabolic pathways and illustrate the importance of careful instrument selection in Mendelian randomization analysis, revealing a putative causal relationship between acetone and hypertension. Our publicly available results provide a foundational resource for the community to examine the role of metabolism across diverse diseases.
Retinol is a fat-soluble vitamin that plays an essential role in many biological processes throughout the human lifespan. Here, we perform the largest genome-wide association study (GWAS) of retinol to date in up to 22,274 participants. We identify eight common variant loci associated with retinol, as well as a rare-variant signal. An integrative gene prioritisation pipeline supports novel retinol-associated genes outside of the main retinol transport complex (RBP4:TTR) related to lipid biology, energy homoeostasis, and endocrine signalling. Genetic proxies of circulating retinol were then used to estimate causal relationships with almost 20,000 clinical phenotypes via a phenome-wide Mendelian randomisation study (MR-pheWAS). The MR-pheWAS suggests that retinol may exert causal effects on inflammation, adiposity, ocular measures, the microbiome, and MRI-derived brain phenotypes, amongst several others. Conversely, circulating retinol may be causally influenced by factors including lipids and serum creatinine. Finally, we demonstrate how a retinol polygenic score could identify individuals more likely to fall outside of the normative range of circulating retinol for a given age. In summary, this study provides a comprehensive evaluation of the genetics of circulating retinol, as well as revealing traits which should be prioritised for further investigation with respect to retinol related therapies or nutritional intervention.
Background For many diseases there are delays in diagnosis due to a lack of objective biomarkers for disease onset. Whether measuring thousands of proteins offers predictive information across a wide range of diseases is unknown. Methods In 41,931 individuals from the UK Biobank Pharma Proteomics Project (UKB-PPP), we integrated ∼3000 plasma proteins with clinical information to derive sparse prediction models for the 10-year incidence of 218 common and rare diseases (81 – 6038 cases). We compared prediction models based on proteins with a) basic clinical information alone, b) basic clinical information + 37 clinical biomarkers, and c) genome-wide polygenic risk scores. Results For 67 pathologically diverse diseases, a model including as few as 5 to 20 proteins was superior to clinical models (median delta C-index = 0.07; range = 0.02 – 0.31) and to clinical models with biomarkers for 52 diseases. In multiple myeloma, for example, a set of 5 proteins significantly improved prediction over basic clinical information (delta C-index = 0.25 (95% confidence interval 0.20 – 0.29)). At a 5% false positive rate (FPR), proteomic prediction (5 proteins) identified individuals at high risk of multiple myeloma (detection rate (DR) = 50%), non-Hodgkin lymphoma (DR = 55%) and motor neuron disease (DR = 29%). At a 20% FPR, proteomic prediction identified individuals at high-risk for pulmonary fibrosis (DR= 80%) and dilated cardiomyopathy (DR = 75%). Conclusions Sparse plasma protein signatures offer novel, clinically useful prediction of common and rare diseases, through disease-specific proteins and protein predictors shared across multiple diseases. (Funded by Medical Research Council, NIHR, Wellcome Trust.) ### Competing Interest Statement J. Davitte, P. Surendran, D. Croteau-Chonka, C. Robins, T. Kanno, S. Gade, D. Freitag, F. Ziebell, J. Betts, and R. Scott are all employees of and/or shareholders for GlaxoSmithKline. None of the other authors has a competing interest. ### Funding Statement This study has been funded by the Medical Research Council, the NIHR, and the Wellcome Trust. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: All UK Biobank data was accessed in accordance with GlaxoSmithKline's UK Biobank Application 20361 and the UKB-PPP Consortium Application 65851. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All individual level data is publicly available to bona fide researchers from the UK Biobank ().
Circulating proteins have important functions in inflammation and a broad range of diseases. To identify genetic influences on inflammation-related proteins, we conducted a genome-wide protein quantitative trait locus (pQTL) study of 91 plasma proteins measured using the Olink Target platform in 14,824 participants. We identified 180 pQTLs (59 cis , 121 trans ). Integration of pQTL data with eQTL and disease genome-wide association studies provided insight into pathogenesis, implicating lymphotoxin-α in multiple sclerosis. Using Mendelian randomization (MR) to assess causality in disease etiology, we identified both shared and distinct effects of specific proteins across immune-mediated diseases, including directionally discordant effects of CD40 on risk of rheumatoid arthritis versus multiple sclerosis and inflammatory bowel disease. MR implicated CXCL5 in the etiology of ulcerative colitis (UC) and we show elevated gut CXCL5 transcript expression in patients with UC. These results identify targets of existing drugs and provide a powerful resource to facilitate future drug target prioritization.
Proteome-wide Mendelian randomization (MR) has emerged as a promising approach in uncovering novel therapeutic targets. However, genetic colocalization analysis has revealed that a third of MR associations lacked a shared causal signal between the protein and disease outcome, raising questions about the effectiveness of this approach. The impact of proteome-wide MR, stratified by cis-trans status, in the presence or absence of genetic colocalization, on therapeutic target identification remains largely unknown. In this study, we conducted genome-wide MR and cis/trans-genetic colocalization analyses using proteomic and complex trait genome-wide association studies. Using two different gold-standard datasets, we found that the enrichment of target-disease pairs supported by MR increased with more p-value stringent thresholds MR p-value, with the evidence of enrichment limited to colocalizing cis-MR associations. Using a phenome-wide proteogenetic colocalization approach, we identified 235 unique targets associated with 168 binary traits at high confidence (at colocalization posterior probability of shared signal > 0.8 and 5% FDR-corrected MR p-value). The majority of the target-trait pairs did not overlap with existing drug targets, highlighting opportunities to investigate novel therapeutic hypotheses. 42% of these non-overlapping target-trait pairs were supported by GWAS, interacting protein partners, animal models, and Mendelian disease evidence. These high confidence target-trait pairs assisted with causal gene identification and helped uncover translationally informative novel biology, especially from trans-colocalizing signals, such as the association of lower intestinal alkaline phosphatase with a higher risk of inflammatory bowel disease in FUT2 non-secretors. Beyond target identification, we used MR of colocalizing signals to infer therapeutic directions and flag potential safety concerns. For example, we found that most genetically predicted therapeutic targets for inflammatory bowel disease could potentially worsen allergic disease phenotypes, except for TNFRSF6B where we observed directionally consistent associations for both phenotypes. Our results are publicly available to download or browse in a web application enabling others to use proteogenomic evidence to appraise therapeutic targets. ### Competing Interest Statement MAK is now an employee of Variant Bio. JS is now an employee of Illumina. JM is an employee of Bristol-Myers Squibb. ESEM is an employee of Genmab. EM is an employee at Genomics PLC. MVH is an employee of 23andMe, CR, PS, SH and RAS are employees of GlaxoSmithKline. ### Funding Statement MAK, BA, JS, JH, AB, DO, MC, EMM, MG, ID were funded by Open Targets. This research was funded in part by a Wellcome Trust [Grant number 206194]. For the purpose of Open Access, the authors have applied a CC-BY public copyright license to any Author Accepted Manuscript version arising from this submission. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Summary data for both proteins and outcomes used for genetic analyses are publicly available from the GWAS catalog . I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced are available online at: [https://ftp.ebi.ac.uk/pub/databases/opentargets/publishing/mendelian\_randomisation\_results/][1] [https://ftp.ebi.ac.uk/pub/databases/opentargets/publishing/mendelian\_randomisation\_results/][1] [1]: https://ftp.ebi.ac.uk/pub/databases/opentargets/publishing/mendelian_randomisation_results/
Prostate cancer (PCa) brings huge public health burden in men. A growing number of conventional observational studies report associations of multiple circulating proteins with PCa risk. However, the existing findings may be subject to incoherent biases of conventional epidemiologic studies. To better characterize their associations, herein, we evaluated associations of genetically predicted concentrations of plasma proteins with PCa risk. We developed comprehensive genetic prediction models for protein levels in plasma. After testing 1308 proteins in 79 194 cases and 61 112 controls of European ancestry included in the consortia of BPC3, CAPS, CRUK, PEGASUS, and PRACTICAL, 24 proteins showed significant associations with PCa risk, including 16 previously reported proteins and eight novel proteins. Of them, 14 proteins showed negative associations and 10 showed positive associations with PCa risk. For 18 of the identified proteins, potential functional somatic changes of encoding genes were detected in PCa patients in The Cancer Genome Atlas. Genes encoding these proteins were significantly involved in cancer-related pathways. We further identified drugs targeting the identified proteins, which may serve as candidates for drug repurposing for treating PCa. In conclusion, this study identifies novel protein biomarker candidates for PCa risk, which may provide new perspectives on the etiology of PCa and improve its therapeutic strategies.
The Pharma Proteomics Project is a precompetitive biopharmaceutical consortium characterizing the plasma proteomic profiles of 54,219 UK Biobank participants. Here we provide a detailed summary of this initiative, including technical and biological validations, insights into proteomic disease signatures, and prediction modelling for various demographic and health indicators. We present comprehensive protein quantitative trait locus (pQTL) mapping of 2,923 proteins that identifies 14,287 primary genetic associations, of which 81% are previously undescribed, alongside ancestry-specific pQTL mapping in non-European individuals. The study provides an updated characterization of the genetic architecture of the plasma proteome, contextualized with projected pQTL discovery rates as sample sizes and proteomic assay coverages increase over time. We offer extensive insights into trans pQTLs across multiple biological domains, highlight genetic influences on ligand–receptor interactions and pathway perturbations across a diverse collection of cytokines and complement networks, and illustrate long-range epistatic effects of ABO blood group and FUT2 secretor status on proteins with gastrointestinal tissue-enriched expression. We demonstrate the utility of these data for drug discovery by extending the genetic proxied effects of protein targets, such as PCSK9, on additional endpoints, and disentangle specific genes and proteins perturbed at loci associated with COVID-19 susceptibility. This public–private partnership provides the scientific community with an open-access proteomics resource of considerable breadth and depth to help to elucidate the biological mechanisms underlying proteo-genomic discoveries and accelerate the development of biomarkers, predictive models and therapeutics 1 .
Metabolic biomarker data quantified by nuclear magnetic resonance (NMR) spectroscopy in approximately 121,000 UK Biobank participants has recently been released as a community resource, comprising absolute concentrations and ratios of 249 circulating metabolites, lipids, and lipoprotein sub-fractions. Here we identify and characterise additional sources of unwanted technical variation influencing individual biomarkers in the data available to download from UK Biobank. These included sample preparation time, shipping plate well, spectrometer batch effects, drift over time within spectrometer, and outlier shipping plates. We developed a procedure for removing this unwanted technical variation, and demonstrate that it increases signal for genetic and epidemiological studies of the NMR metabolic biomarker data in UK Biobank. We subsequently developed an R package, ukbnmr, which we make available to the wider research community to enhance the utility of the UK Biobank NMR metabolic biomarker data and to facilitate rapid analysis.
The use of omic modalities to dissect the molecular underpinnings of common diseases and traits is becoming increasingly common. But multi-omic traits can be genetically predicted, which enables highly cost-effective and powerful analyses for studies that do not have multi-omics(1). Here we examine a large cohort (the INTERVAL study(2); n = 50,000 participants) with extensive multi-omic data for plasma proteomics (SomaScan, n = 3,175; Olink, n = 4,822), plasma metabolomics (Metabolon HD4, n = 8,153), serum metabolomics (Nightingale, n = 37,359) and whole-blood Illumina RNA sequencing (n = 4,136), and use machine learning to train genetic scores for 17,227 molecular traits, including 10,521 that reach Bonferroni-adjusted significance. We evaluate the performance of genetic scores through external validation across cohorts of individuals of European, Asian and African American ancestries. In addition, we show the utility of these multi-omic genetic scores by quantifying the genetic control of biological pathways and by generating a synthetic multi-omic dataset of the UK Biobank(3) to identify disease associations using a phenome-wide scan. We highlight a series of biological insights with regard to genetic mechanisms in metabolism and canonical pathway associations with disease; for example, JAK-STAT signalling and coronary atherosclerosis. Finally, we develop a portal (https://www.omicspred.org/) to facilitate public access to all genetic scores and validation results, as well as to serve as a platform for future extensions and enhancements of multi-omic genetic scores.
Metabolite levels measured in the human population are endophenotypes for biological processes. We combined sequencing data for 3,924 (whole-exome sequencing, WES, discovery) and 2,805 (whole-genome sequencing, WGS, replication) donors from a prospective cohort of blood donors in England. We used multiple approaches to select and aggregate rare genetic variants (minor allele frequency [MAF] < 0.1%) in protein-coding regions and tested their associations with 995 metabolites measured in plasma by using ultra-high-performance liquid chromatography-tandem mass spectrometry. We identified 40 novel associations implicating rare coding variants (27 genes and 38 metabolites), of which 28 (15 genes and 28 metabolites) were replicated. We developed algorithms to prioritize putative driver variants at each locus and used mediation and Mendelian randomization analyses to test directionality at associations of metabolite and protein levels at the ACY1 locus. Overall, 66% of reported associations implicate gene targets of approved drugs or bioactive drug-like compounds, contributing to drug targets' validating efforts.
Stroke is the second leading cause of death with substantial unmet therapeutic needs. To identify potential stroke therapeutic targets, we estimate the causal effects of 308 plasma proteins on stroke outcomes in a two-sample Mendelian randomization framework and assess mediation effects by stroke risk factors. We find associations between genetically predicted plasma levels of six proteins and stroke ( P ≤ 1.62 × 10 −4 ). The genetic associations with stroke colocalize (Posterior Probability >0.7) with the genetic associations of four proteins (TFPI, TMPRSS5, CD6, CD40). Mendelian randomization supports atrial fibrillation, body mass index, smoking, blood pressure, white matter hyperintensities and type 2 diabetes as stroke risk factors ( P ≤ 0.0071). Body mass index, white matter hyperintensity and atrial fibrillation appear to mediate the TFPI, IL6RA, TMPRSS5 associations with stroke. Furthermore, thirty-six proteins are associated with one or more of these risk factors using Mendelian randomization. Our results highlight causal pathways and potential therapeutic targets for stroke.
Platelets play a key role in thrombosis and hemostasis. Platelet count (PLT) and mean platelet volume (MPV) are highly heritable quantitative traits, with hundreds of genetic signals previously identified, mostly in European ancestry populations. We here utilize whole genome sequencing (WGS) from NHLBI's Trans-Omics for Precision Medicine initiative (TOPMed) in a large multi-ethnic sample to further explore common and rare variation contributing to PLT (n = 61 200) and MPV (n = 23 485). We identified and replicated secondary signals at MPL (rs532784633) and PECAM1 (rs73345162), both more common in African ancestry populations. We also observed rare variation in Mendelian platelet-related disorder genes influencing variation in platelet traits in TOPMed cohorts (not enriched for blood disorders). For example, association of GP9 with lower PLT and higher MPV was partly driven by a pathogenic Bernard-Soulier syndrome variant (rs5030764, p.Asn61Ser), and the signals at TUBB1 and CD36 were partly driven by loss of function variants not annotated as pathogenic in ClinVar (rs199948010 and rs571975065). However, residual signal remained for these gene-based signals after adjusting for lead variants, suggesting that additional variants in Mendelian genes with impacts in general population cohorts remain to be identified. Gene-based signals were also identified at several genome-wide association study identified loci for genes not annotated for Mendelian platelet disorders (PTPRH, TET2, CHEK2), with somatic variation driving the result at TET2. These results highlight the value of WGS in populations of diverse genetic ancestry to identify novel regulatory and coding signals, even for well-studied traits like platelet traits.