Polygenic scores (PGS) can be used to predict an individual's genetic predisposition to a heritable trait or disease. In this tutorial you will learn about PGS and the PGS Catalog, an open database of existing scores that can be reused and applied in research and clinical settings.
Blood cell phenotypes are routinely tested in healthcare to inform clinical decisions. Genetic variants influencing mean blood cell phenotypes have been used to understand disease aetiology and improve prediction; however, additional information may be captured by genetic effects on observed variance. Here, we mapped variance quantitative trait loci (vQTL), i.e. genetic loci associated with trait variance, for 29 blood cell phenotypes from the UK Biobank (N ~ 408,111). We discovered 176 independent blood cell vQTLs, of which 147 were not found by additive QTL mapping. vQTLs displayed on average 1.8-fold stronger negative selection than additive QTL, highlighting that selection acts to reduce extreme blood cell phenotypes. Variance polygenic scores (vPGSs) were constructed to stratify individuals in the INTERVAL cohort (N ~ 40,466), where the genetically most variable individuals had increased conventional PGS accuracy (by ~19%) relative to the genetically least variable individuals. Genetic prediction of blood cell traits improved by ~10% on average combining PGS with vPGS. Using Mendelian randomisation and vPGS association analyses, we found that alcohol consumption significantly increased blood cell trait variances highlighting the utility of blood cell vQTLs and vPGSs to provide novel insight into phenotype aetiology as well as improve prediction.
Genome-wide association studies have identified thousands of variants associated with disease risk but the mechanism by which such variants contribute to disease remains largely unknown. Indeed, a major challenge is that variants do not act in isolation but rather in the framework of highly complex biological networks, such as the human metabolic network, which can amplify or buffer the effect of specific risk alleles on disease susceptibility. Here we use genetically predicted reaction fluxes to perform a systematic search for metabolic fluxes acting as buffers or amplifiers of coronary artery disease (CAD) risk alleles. Our analysis identifies 30 risk locus-reaction flux pairs with significant interaction on CAD susceptibility involving 18 individual reaction fluxes and 8 independent risk loci. Notably, many of these reactions are linked to processes with putative roles in the disease such as the metabolism of inflammatory mediators. In summary, this work establishes proof of concept that biochemical reaction fluxes can have non-additive effects with risk alleles and provides novel insights into the interplay between metabolism and genetic variation on disease susceptibility.
The biological mechanisms through which most nonprotein-coding genetic variants affect disease risk are unknown. To investigate gene-regulatory mechanisms, we mapped blood gene expression and splicing quantitative trait loci (QTLs) through bulk RNA sequencing in 4,732 participants and integrated protein, metabolite and lipid data from the same individuals. We identified cis-QTLs for the expression of 17,233 genes and 29,514 splicing events (in 6,853 genes). Colocalization analyses revealed 3,430 proteomic and metabolomic traits with a shared association signal with either gene expression or splicing. We quantified the relative contribution of the genetic effects at loci with shared etiology, observing 222 molecular phenotypes significantly mediated by gene expression or splicing. We uncovered gene-regulatory mechanisms at disease loci with therapeutic implications, such as WARS1 in hypertension, IL7R in dermatitis and IFNAR2 in COVID-19. Our study provides an open-access resource on the shared genetic etiology across transcriptional phenotypes, molecular traits and health outcomes in humans ( https://IntervalRNA.org.uk ).
Combining information from multiple GWASs for a disease and its risk factors has proven a powerful approach for development of polygenic risk scores (PRSs). This may be particularly useful for type 2 diabetes (T2D), a highly polygenic and heterogeneous disease where the additional predictive value of a PRS is unclear. Here, we use a meta-scoring approach to develop a metaPRS for T2D that incorporated genome-wide associations from both European and non-European genetic ancestries and T2D risk factors. We evaluated the performance of this metaPRS and benchmarked it against existing genome-wide PRS in 620,059 participants and 50,572 T2D cases amongst six diverse genetic ancestries from UK Biobank, INTERVAL, the All of Us Research Program, and the Singapore Multi-Ethnic Cohort. We show that our metaPRS was the most powerful PRS for predicting T2D in European population-based cohorts and had comparable performance to the top ancestry-specific PRS, highlighting its transferability. In UK Biobank, we show the metaPRS had stronger predictive power for 10-year risk than all individual risk factors apart from BMI and biomarkers of dysglycemia. The metaPRS modestly improved T2D risk stratification of QDiabetes risk scores for 10-year risk prediction, particularly when prioritising individuals for blood tests of dysglycemia. Overall, we present a highly predictive and transferrable PRS for T2D and demonstrate that the potential for PRS to incrementally improve T2D risk prediction when incorporated into UK guideline-recommended screening and risk prediction with a clinical risk score.
BACKGROUND:Myocardial infarction (MI) is a complex disease caused by both lifestyle and genetic factors. This study aims to investigate the predictive value of genetic risk, in addition to traditional cardiovascular risk factors, for recurrent events following early-onset MI. METHODS:The Italian Genetic Study of Early-Onset Myocardial Infarction is a cohort study enrolling patients with MI before 45 years. Monogenic variants causing familial hypercholesterolemia were identified, and a coronary artery disease polygenic score (PGS) was calculated. Ten-fold cross-validated Cox proportional hazards models were fitted sequentially including all clinical variables, the PGS, and monogenic variants on the composite outcome of cardiovascular death, recurrent MI, stroke, or revascularization. RESULTS:During a 19.9-year follow-up, 847 (50.7%) patients experienced recurrent events. Each 1-SD higher PGS was associated with a 21% higher hazard of recurrent events (hazard ratio, 1.21 [95% CI, 1.13-1.31]; P=4.04×10-6). Except for secondary prevention, PGS was the strongest determinant of recurrent event risk (C index, 0.56 [95% CI, 0.54-0.58]) compared with clinical risk factors. Overall, predictive performance of clinical risk factors (C index, 0.69 [95% CI, 0.67-0.71]) improved after adding the PGS (C index, 0.69 [95% CI, 0.68-0.71]; P=0.006). When dividing the population by PGS quintiles, the highest fifth had a 57% higher hazard of recurrent events than the lowest fifth (hazard ratio, 1.57 [95% CI, 1.26-1.96]; P=5.57×10-5). CONCLUSIONS:When compared with other clinical risk factors, PGS was the strongest predictor of event recurrence among patients with an early-onset MI. Though the discriminative power of recurrent event prediction in this cohort was modest, the addition of PGS significantly improved discrimination.
Polygenic scores (PGS) can be used for risk stratification by quantifying individuals’ genetic predisposition to disease, and many potentially clinically useful applications have been proposed. Here, we review the latest potential benefits of PGS in the clinic and challenges to implementation. PGS could augment risk stratification through combined use with traditional risk factors (demographics, disease-specific risk factors, family history, etc.), to support diagnostic pathways, to predict groups with therapeutic benefits, and to increase the efficiency of clinical trials. However, there exist challenges to maximizing the clinical utility of PGS, including FAIR (Findable, Accessible, Interoperable, and Reusable) use and standardized sharing of the genomic data needed to develop and recalculate PGS, the equitable performance of PGS across populations and ancestries, the generation of robust and reproducible PGS calculations, and the responsible communication and interpretation of results. We outline how these challenges may be overcome analytically and with more diverse data as well as highlight sustained community efforts to achieve equitable, impactful, and responsible use of PGS in healthcare.
Methods of estimating polygenic scores (PGSs) from genome-wide association studies are increasingly utilized. However, independent method evaluation is lacking, and method comparisons are often limited. Here, we evaluate polygenic scores derived via seven methods in five biobank studies (totaling about 1.2 million participants) across 16 diseases and quantitative traits, building on a reference-standardized framework. We conducted meta-analyses to quantify the effects of method choice, hyperparameter tuning, method ensembling, and the target biobank on PGS performance. We found that no single method consistently outperformed all others. PGS effect sizes were more variable between biobanks than between methods within biobanks when methods were well tuned. Differences between methods were largest for the two investigated autoimmune diseases, seropositive rheumatoid arthritis and type 1 diabetes. For most methods, cross-validation was more reliable for tuning hyperparameters than automatic tuning (without the use of target data). For a given target phenotype, elastic net models combining PGS across methods (ensemble PGS) tuned in the UK Biobank provided consistent, high, and cross-biobank transferable performance, increasing PGS effect sizes (β coefficients) by a median of 5.0% relative to LDpred2 and MegaPRS (the two best-performing single methods when tuned with cross-validation). Our interactively browsable online-results and open-source workflow prspipe provide a rich resource and reference for the analysis of polygenic scoring methods across biobanks.
Polygenic scores (PGSs) have transformed human genetic research and have numerous potential clinical applications. Here we present a series of recent enhancements to the PGS Catalog and highlight the PGS Catalog Calculator, an open-source, scalable and portable pipeline for reproducibly calculating PGSs that democratizes equitable PGS applications.
The NHGRI-EBI GWAS Catalog serves as a vital resource for the genetic research community, providing access to the most comprehensive database of human GWAS results. Currently, it contains close to 7 0 0 0 publications for > 15 0 0 0 traits, from which more than 625 0 0 0 lead associations have been curated. Additionally, 85 0 0 0 full genome-wide summary statistics datasets-containing association data for all variants in the analysis-are available for downstream analyses such as meta-analysis, fine-mapping , Mendelian randomisation or development of poly- genic risk scores. As a centralised repository for GWAS results, the GWAS Catalog sets and implements standards for data submission and harmonisation, and encourages the use of consistent descriptors for traits, samples and methodologies. We share processes and vocabulary with the PGS Catalog, improving interoperability for a growing user group. Here, we describe the latest changes in data content, improvements in our user interface, and the implementation of the GWAS-SSF standard format for summary statistics. We address the challenges of handling the rapid increase in large-scale molecular quantitative trait GWAS and the need for sensitivity in the use of population and cohort descriptors while maintaining data interoperability and reusability. [GRAPHICS] .
Metabolomic platforms using nuclear magnetic resonance (NMR) spectroscopy can now rapidly quantify many circulating metabolites which are potential biomarkers of cardiovascular disease (CVD). Here, we analyse ∼170,000 UK Biobank participants (5,096 incident CVD cases) without a history of CVD and not on lipid-lowering treatments to evaluate the potential for improving 10-year CVD risk prediction using NMR biomarkers in addition to conventional risk factors and polygenic risk scores (PRSs). Using machine learning, we developed sex-specific NMR scores for coronary heart disease (CHD) and ischaemic stroke, then estimated their incremental improvement of 10-year CVD risk prediction when added to guideline-recommended risk prediction models (i.e., SCORE2) with and without PRSs. The risk discrimination provided by SCORE2 (Harrell’s C-index = 0.718) was similarly improved by addition of NMR scores (ΔC-index 0.011; 0.009, 0.014) and PRSs (ΔC-index 0.009; 95% CI: 0.007, 0.012), which offered largely orthogonal information. Addition of both NMR scores and PRSs yielded the largest improvement in C-index over SCORE2, from 0.718 to 0.737 (ΔC-index 0.019; 95% CI: 0.016, 0.022). Concomitant improvements in risk stratification were observed in categorical net reclassification index when using guidelines-recommended risk categorisation, with net case reclassification of 13.04% (95% CI: 11.67%, 14.41%) when adding both NMR scores and PRSs to SCORE2. Using population modelling, we estimated that targeted risk-reclassification with NMR scores and PRSs together could increase the number of CVD events prevented per 100,000 screened from 201 to 370 (ΔCVDprevented: 170; 95% CI: 158, 182) while essentially maintaining the number of statins prescribed per CVD event prevented. Overall, we show combining NMR scores and PRSs with SCORE2 moderately enhances prediction of first-onset CVD, and could have substantial population health benefit if applied at scale. ### Competing Interest Statement During the course of this project P.S. became a full-time employee of GSK Plc. All significant contributions to this study were made prior to this role and GSK Plc had no input to the study. J.D. serves on scientific advisory boards for AstraZeneca, Novartis, and UK Biobank, and has received multiple grants from academic, charitable and industry sources outside of the submitted work. A.S.B. reports institutional grants from AstraZeneca, Bayer, Biogen, BioMarin, Bioverativ, Novartis, Regeneron and Sanofi. The remaining authors declare no competing interests. ### Funding Statement This work was performed using resources provided by the Cambridge Service for Data Driven Discovery (CSD3) operated by the University of Cambridge Research Computing Service ([www.csd3.cam.ac.uk][1]), provided by Dell EMC and Intel using Tier-2 funding from the Engineering and Physical Sciences Research Council (capital grant EP/P020259/1), and DiRAC funding from the Science and Technology Facilities Council ([www.dirac.ac.uk][2]). This work was supported by core funding from the: Cambridge BHF Centre of Research Excellence (RE/18/1/34212) and BHF Chair Award (CH/12/2/29428). S.C.R. and S.K. were funded by a British Heart Foundation (BHF) Programme Grant (RG/18/13/33946). S.C.R. was also funded by the National Institute for Health and Care Research (NIHR) Cambridge BRC (BRC-1215-20014; NIHR203312) [*]. X.J. was funded by British Heart Foundation (CH/12/2/29428) and Wellcome Trust (227566/Z/23/Z). L.P. and P.S. were supported by a Rutherford Fund Fellowship from the Medical Research Council grant MR/S003746/1. Y.X. and M.I. were supported by the UK Economic and Social Research Council (ES/T013192/1). S.A.L. was supported by a Canadian Institutes of Health Research postdoctoral fellowship (MFE-171279). E.D.A. holds a NIHR Senior Investigator Award. J.D. holds a BHF Professorship and a NIHR Senior Investigator Award. M.I. is supported by the Munz Chair of Cardiovascular Prediction and Prevention and the NIHR Cambridge Biomedical Research Centre (NIHR203312). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. *The views expressed are those of the authors and not necessarily those of the NIHR or the Department of Health and Social Care. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This study was approved under UK Biobank Projects 30418 and ethics approval was obtained from the North West Multi-Center Research Ethics Committee. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data described are available through UK Biobank subject to approval from the UK Biobank access committee. See for further details. [1]: http://www.csd3.cam.ac.uk [2]: http://www.dirac.ac.uk
Polygenic scores (PGS) are a measurement which represents genetic predisposition for a heritable trait or phenotype by aggregating the effects of hundreds-to-millions of genetic variants into a single number. The PGS Catalog is the world’s largest FAIR (finable, accessible, interoperable, and reusable) repository of PGS along with the relevant metadata required to evaluate and reuse them. PGS have emerged as a commonly used genomic tool that have multiple research applications and potential clinical applications. This webinar will demonstrate how to use scoring files from the PGS Catalog (or custom files) to calculate PGS in your own samples using the PGS Catalog Calculator (pgsc_calc). The pgsc_calc software reproducibly automates PGS calculation in commonly used genotyping formats (plink, VCF) alongside capabilities for adjusting PGS in the context of genetic ancestry (necessary for proper interpretation of PGS across diverse populations). This webinar introduces the software and describes the ancestry adjustment process included in the Calculator, along with information about how to run the tool and interpret results.
Polygenic scores (PGS) have transformed human genetic research and have multiple potential clinical applications, including risk stratification for disease prevention and prediction of treatment response. Here, we present a series of recent enhancements to the PGS Catalog (www.PGSCatalog.org), the largest findable, accessible, interoperable, and reusable (FAIR) repository of PGS. These include expansions in data content and ancestral diversity as well as the addition of new features. We further present the PGS Catalog Calculator (pgsc_calc, https://github.com/PGScatalog/pgsc_calc), an open-source, scalable and portable pipeline to reproducibly calculate PGS that securely democratizes equitable PGS applications by implementing genetic ancestry estimation and score normalization using reference data. With the PGS Catalog & calculator users can now quantify an individual's genetic predisposition for hundreds of common diseases and clinically relevant traits. Taken together, these updates and tools facilitate the next generation of PGS, thus lowering barriers to the clinical studies necessary to identify where PGS may be integrated into clinical practice.
Summary statistics from genome-wide association studies (GWAS) represent a huge potential for research. A challenge for researchers in this field is the access and sharing of summary statistics data due to a lack of standards for the data content and file format. For this reason, the GWAS Catalog hosted a series of meetings in 2021 with summary statistics stakeholders to guide the development of a standard format. The key requirements from the stakeholders were for a standard that contained key data elements to be able to support a wide range of data analyses, required low bioinformatics skills for file access and generation, to have easily accessible metadata, and unambiguous and interoperable data. Here, we define the specifications for the first version of the GWAS-SSF format, which was developed to meet the requirements discussed with the community. GWAS-SSF consists of a tab-separated data file with well-defined fields and an accompanying metadata file.
The NHGRI-EBIGWAS Catalog (www.ebi.ac.uk/gwas) is a FAIR knowledgebase providing detailed, structured, standardised and interoperable genome-wide association study (GWAS) data to >200 000 users per year from academic research, healthcare and industry. The Catalog contains variant-trait associations and supporting metadata for >45 000 published GWAS across >5000 human traits, and >40 000 full P- value summary statistics datasets. Content is curated from publications or acquired via author submission of prepublication summary statistics through a new submission portal and validation tool. GWAS data volume has vastly increased in recent years. We have updated our software to meet this scaling challenge and to enable rapid release of submitted summary statistics. The scope of the repository has expanded to include additional data types of high interest to the community, including sequencing-based GWAS, gene-based analyses and copy number variation analyses. Commu- nity outreach has increased the number of shared datasets from under-represented traits, e.g. cancer, and we continue to contribute to awareness of the lack of population diversity in GWAS. Interoperability of the Catalog has been enhanced through links to other resources including the Polygenic Score Catalog and the International Mouse Phenotyping Consortium, refinements to GWAS trait annotation, and the development of a standard format for GWAS data.
The aim of this patient and public involvement and engagement (PPIE) work was to explore improvised theatre as a tool for facilitating bi-directional dialogue between researchers and patients/members of the public on the topic of polygenic risk scores (PRS) use within primary or secondary care. PRS are a tool to quantify genetic risk for a heritable disease or trait and may be used to predict future health outcomes. In the United Kingdom (UK), they are often cited as a next-in-line public health tool to be implemented, and their use in consumer genetic testing as well as patient-facing settings is increasing. Despite their potential clinical utility, broader themes about how they might influence an individual’s perception of disease risk and decision-making are an active area of research; however, this has mostly been in the setting of return of results to patients. We worked with a youth theatre group and patients involved in a PPIE group to develop two short plays about public perceptions of genetic risk information that could be captured by PRS. These plays were shared in a workshop with patients/members of the public to facilitate discussions about PRS and their perceived benefits, concerns and emotional reactions. Discussions with both performers and patients/public raised three key questions: (1) can the data be trusted?; (2) does knowing genetic risk actually help the patient?; and (3) what makes a life worthwhile? Creating and watching fictional narratives helped all participants explore the potential use of PRS in a clinical setting, informing future research considerations and improving communication between the researchers and lay members of the PPIE group.
Metabolic biomarker data quantified by nuclear magnetic resonance (NMR) spectroscopy in approximately 121,000 UK Biobank participants has recently been released as a community resource, comprising absolute concentrations and ratios of 249 circulating metabolites, lipids, and lipoprotein sub-fractions. Here we identify and characterise additional sources of unwanted technical variation influencing individual biomarkers in the data available to download from UK Biobank. These included sample preparation time, shipping plate well, spectrometer batch effects, drift over time within spectrometer, and outlier shipping plates. We developed a procedure for removing this unwanted technical variation, and demonstrate that it increases signal for genetic and epidemiological studies of the NMR metabolic biomarker data in UK Biobank. We subsequently developed an R package, ukbnmr, which we make available to the wider research community to enhance the utility of the UK Biobank NMR metabolic biomarker data and to facilitate rapid analysis.