The prediction accuracy of a polygenic score (PGS) is highly determined by the size of the training sample. Although this sample is still limited for psychiatric disorders, these disorders are genetically correlated with multiple behavioral and physical phenotypes. These mostly quantitative phenotypes are much more accessible and thus currently have genome-wide association studies (GWAS) with millions of samples. Generating stand-alone PGS for publicly accessible GWAS summary statistics is nowadays possible with PGS methods that do not require a validation sample, like LDpred2-auto. There are some available methods that benefit from using genetically correlated phenotypes to increase prediction accuracy, including MTAG and wMT-SBLUP and that have been applied to psychiatric disorders. These methods require a pre-selection of the included phenotypes based on prior information about the genetic correlation estimates with the desired outcome. Here we show the results of a new method, multi-PGS, that does not require to pre-specify genetically correlated phenotypes but relies on an agnostic PGS library based on “all” publicly available GWAS summary statistics. We explore diverse applications of this multi-PGS for psychiatric disorders using the iPSYCH data. In practice, a large library of PGS including 937 scores was generated from publicly available GWAS summary statistics resources (GWAS Catalog, GWAS ATLAS, PGC) using LDpred2-auto. Then the PGS library together with covariates sex, birth year and 20 PCs were used as predictors in multivariate models. We used both penalized regression models (lasso) and gradient boosted trees (XGBoost). The out-of-sample prediction accuracy of the risk prediction models was assessed. First, we applied our multi-PGS strategy to predict ADHD, affective disorder, anorexia nervosa, autism, bipolar disorder and schizophrenia in iPSYCH. All multi-PGS models increased both R2 and logOR, with R2 increases of 4-fold on average and up to 9-fold for ADHD and autism. Increased prediction was also observed when compared to wMT-SBLUP. Interestingly, multiple PGS for the same phenotype were selected in the final model. For example, three different depression-related PGS (self-reported, medically diagnosed and broad depression) were included in the affective disorder multi-PGS. This indicates that non-overlapping signals from multiple GWAS of similar phenotypes can be combined to increase prediction accuracy. Next, we explored further the capacity of our multi-PGS to predict outcomes for which there are no available external GWAS summary statistics, as is the case for some sub-diagnoses and understudied psychiatric disorders. This question is inspired by a scenario where the studied outcome could benefit from PGS analyses, but there is still no GWAS for that outcome. Surprisingly, our results showed no decrease in prediction accuracy when the library did not include a PGS for the target disorder. Moreover, we applied generated multi-PGS for case-case predictions of highly comorbid disorders. For instance, a multi-PGS of ADHD vs. ASD explained 12% of the variance of the disjoint cases. Psychiatric disorders are very heterogeneous phenotypes, both genetically and etiologically. We exploit this feature to increase genetic prediction accuracy using multi-PGS constructed in an agnostic manner. Finally, we discuss the conflict prediction vs. explanation in the context of multi-PGS models.
Complement components have been linked to schizophrenia and autoimmune disorders. We examined the association between neonatal circulating C3 and C4 protein concentrations in 68,768 neonates and the risk of six mental disorders. We completed genome-wide association studies (GWASs) for C3 and C4 and applied the summary statistics in Mendelian randomization and phenome-wide association studies related to mental and autoimmune disorders. The GWASs for C3 and C4 protein concentrations identified 15 and 36 independent loci, respectively. We found no associations between neonatal C3 and C4 concentrations and mental disorders in the total sample (both sexes combined); however, post-hoc analyses found that a higher C3 concentration was associated with a reduced risk of schizophrenia in females. Mendelian randomization based on C4 summary statistics found an altered risk of five types of autoimmune disorders. Our study adds to our understanding of the associations between C3 and C4 concentrations and subsequent mental and autoimmune disorders.
LDpred2 is a widely used Bayesian method for building polygenic scores (PGS). LDpred2-auto can infer the two parameters from the LDpred model, the SNP heritability h 2 and polygenicity p , so that it does not require an additional validation dataset to choose best-performing parameters. The main aim of this paper is to properly validate the use of LDpred2-auto for inferring multiple genetic parameters. Here, we present a new version of LDpred2-auto that adds an optional third parameter α to its model, for modeling negative selection. We then validate the inference of these three parameters (or two, when using the previous model). We also show that LDpred2-auto provides per-variant probabilities of being causal that are well calibrated, and can therefore be used for fine-mapping purposes. We also derive a new formula to infer the out-of-sample predictive performance r 2 of the resulting PGS directly from the Gibbs sampler of LDpred2-auto. Finally, we extend the set of HapMap3 variants recommended to use with LDpred2 with 37% more variants to improve the coverage of this set, and show that this new set of variants captures 12% more heritability and provides 6% more predictive performance, on average, in UK Biobank analyses.
BACKGROUND: Single nucleotide polymorphism-based heritability is a fundamental quantity in the genetic analysis of complex traits. For case-control phenotypes, for which the continuous distribution of risk in the population is unobserved, observed-scale heritability estimates must be transformed to the more interpretable liability scale. This article describes how the field standard approach incorrectly performs the liability correction in that it does not appropriately account for variation in the proportion of cases across the cohorts comprising the meta-analysis. We propose a simple solution that incorporates cohort-specific ascertainment using the summation of effective sample sizes across cohorts. This solution is applied at the stage of single nucleotide polymorphism-based heritability estimation and does not require generating updated meta-analytic genome-wide association study summary statistics. METHODS: We began by performing a series of simulations to examine the ability of the standard approach and our proposed approach to recapture liability-scale heritability in the population. We went on to examine the differences in estimates obtained from these 2 approaches for real data for 12 major case-control genome-wide association studies of psychiatric and neurologic traits. RESULTS: We found that the field standard approach for performing the liability conversion can downwardly bias estimates by as much as approximately 50% in simulation and approximately 30% in real data. CONCLUSIONS: Prior estimates of liability-scale heritability for genome-wide association study meta-analysis may be drastically underestimated. To this end, we strongly recommend using our proposed approach of using the sum of effective sample sizes across contributing cohorts to obtain unbiased estimates.
Polygenic risk scores (PRS) trained from genome-wide association study (GWAS) results are set to play a pivotal role in biomedical research addressing multifactorial human diseases. The prospect of using these risk scores in clinical care and public health is generating both enthusiasm and controversy, with varying opinions about strengths and limitations across experts 1 . The performances of existing polygenic scores are still limited, and although it is expected to improve with increasing sample size of GWAS and the development of new powerful methods, it remains unclear how much prediction can be ultimately achieved. Here, we conducted a retrospective analysis to assess the progress in PRS prediction accuracy since the publication of the first large-scale GWASs using six common human diseases with sufficient GWAS data. We show that while PRS accuracy has grown rapidly for years, the improvement pace from recent GWAS has decreased substantially, suggesting that further increasing GWAS sample size may translate into very modest risk discrimination improvement. We next investigated the factors influencing the maximum achievable prediction using recently released whole genome-sequencing data from 125K UK Biobank participants, and state-of-the-art modeling of polygenic outcomes. Our analyses point toward increasing the variant coverage of PRS, using either more imputed variants or sequencing data, as a key component for future improvement in prediction accuracy.
The predictive performance of polygenic scores (PGS) is largely dependent on the number of samples available to train the PGS. Increasing the sample size for a specific phenotype is expensive and takes time, but this sample size can be effectively increased by using genetically correlated phenotypes. We propose a framework to generate multi-PGS from thousands of publicly available genome-wide association studies (GWAS) with no need to individually select the most relevant ones. In this study, the multi-PGS framework increases prediction accuracy over single PGS for all included psychiatric disorders and other available outcomes, with prediction R2 increases of up to 9-fold for attention-deficit/hyperactivity disorder compared to a single PGS. We also generate multi-PGS for phenotypes without an existing GWAS and for case-case predictions. We benchmark the multi-PGS framework against other methods and highlight its potential application to new emerging biobanks.
Polygenic scores (PGSs) have limited portability across different groupings of individuals (for example, by genetic ancestries and/or social determinants of health), preventing their equitable use 1 – 3 . PGS portability has typically been assessed using a single aggregate population-level statistic (for example, R 2 ) 4 , ignoring inter-individual variation within the population. Here, using a large and diverse Los Angeles biobank 5 (ATLAS, n = 36,778) along with the UK Biobank 6 (UKBB, n = 487,409), we show that PGS accuracy decreases individual-to-individual along the continuum of genetic ancestries 7 in all considered populations, even within traditionally labelled ‘homogeneous’ genetic ancestries. The decreasing trend is well captured by a continuous measure of genetic distance (GD) from the PGS training data: Pearson correlation of −0.95 between GD and PGS accuracy averaged across 84 traits. When applying PGS models trained on individuals labelled as white British in the UKBB to individuals with European ancestries in ATLAS, individuals in the furthest GD decile have 14% lower accuracy relative to the closest decile; notably, the closest GD decile of individuals with Hispanic Latino American ancestries show similar PGS performance to the furthest GD decile of individuals with European ancestries. GD is significantly correlated with PGS estimates themselves for 82 of 84 traits, further emphasizing the importance of incorporating the continuum of genetic ancestries in PGS interpretation. Our results highlight the need to move away from discrete genetic ancestry clusters towards the continuum of genetic ancestries when considering PGSs.
LDpred2 has been used to derive polygenic scores in several studies. LDpred2-auto can also be used to infer key genetic parameters, as well as provide per-variant posterior inclusions probabilities (PIP, used in fine-mapping). Thanks to its sampled effects, it can also provide estimates of the predictive performance, as well as confidence intervals for the individual polygenic scores. We will discuss challenges in deriving all these results from large GWAS summary statistics and recommendations on how to properly run LDpred2 for both prediction and inference. Finally, we will discuss remaining challenges and future developments we plan to further improve LDpred2.
Proportional hazards models have been proposed to analyse time-to-event phenotypes in genome-wide association studies (GWAS). However, little is known about the ability of proportional hazards models to identify genetic associations under different generative models and when ascertainment is present. Here we propose the age-dependent liability threshold (ADuLT) model as an alternative to a Cox regression based GWAS, here represented by SPACox. We compare ADuLT, SPACox, and standard case-control GWAS in simulations under two generative models and with varying degrees of ascertainment as well as in the iPSYCH cohort. We find Cox regression GWAS to be underpowered when cases are strongly ascertained (cases are oversampled by a factor 5), regardless of the generative model used. ADuLT is robust to ascertainment in all simulated scenarios. Then, we analyse four psychiatric disorders in iPSYCH, ADHD, Autism, Depression, and Schizophrenia, with a strong case-ascertainment. Across these psychiatric disorders, ADuLT identifies 20 independent genome-wide significant associations, case-control GWAS finds 17, and SPACox finds 8, which is consistent with simulation results. As more genetic data are being linked to electronic health records, robust GWAS methods that can make use of age-of-onset information will help increase power in analyses for common health outcomes.
In this study the authors measure the concentration of 25-hydroxyvitamin D and vitamin D binding protein (DBP) in 65,589 neonatal dried blood samples. Findings from further analyses include that the genetic correlates of DBP concentration predict the risk of vitamin D deficiency. The vitamin D binding protein (DBP), encoded by the group-specific component (GC) gene, is a component of the vitamin D system. In a genome-wide association study of DBP concentration in 65,589 neonates we identify 26 independent loci, 17 of which are in or close to the GC gene, with fine-mapping identifying 2 missense variants on chromosomes 12 and 17 (within SH2B3 and GSDMA, respectively). When adjusted for GC haplotypes, we find 15 independent loci distributed over 10 chromosomes. Mendelian randomization analyses identify a unidirectional effect of higher DBP concentration and (a) higher 25-hydroxyvitamin D concentration, and (b) a reduced risk of multiple sclerosis and rheumatoid arthritis. A phenome-wide association study confirms that higher DBP concentration is associated with a reduced risk of vitamin D deficiency. Our findings provide valuable insights into the influence of DBP on vitamin D status and a range of health outcomes.
BACKGROUND:Resource trade-off theory suggests that increased performance on a given trait comes at the cost of decreased performance on other traits. METHODS:Growth data from 1889 subjects (996 girls) were used from the GrowUp1974 Gothenburg study. Energy Trade-Off (ETO) between height and weight for individuals with extreme body types was characterized using a novel ETO-Score (ETOS). Four extreme body types were defined based on height and ETOI at early adulthood: tall-slender, short-stout, short-slender, and tall-stout; their growth trajectories assessed from ages 0.5-17.5 years.A GWAS using UK BioBank data was conducted to identify gene variants associated with height, BMI, and for the first time with ETOS. RESULTS:Height and ETOS trajectories show a two-hit pattern with profound changes during early infancy and at puberty for tall-slender and short-stout body types. Several loci (including FTO, ADCY3, GDF5, ) and pathways were identified by GWAS as being highly associated with ETOS. The most strongly associated pathways were related to "extracellular matrix," "signal transduction," "chromatin organization," and "energy metabolism." CONCLUSIONS:ETOS represents a novel anthropometric trait with utility in describing body types. We discovered the multiple genomic loci and pathways probably involved in energy trade-off.
Publicly available genome-wide association studies (GWAS) summary statistics exhibit uneven quality, which can impact the validity of follow-up analyses. First, we present an overview of possible misspecifications that come with GWAS summary statistics. Then, in both simulations and real-data analyses, we show that additional information such as imputation INFO scores, allele frequencies, and per-variant sample sizes in GWAS summary statistics can be used to detect possible issues and correct for misspecifications in the GWAS summary statistics. One important motivation for us is to improve the predictive performance of polygenic scores built from these summary statistics. Unfortunately, owing to the lack of reporting standards for GWAS summary statistics, this additional information is not systematically reported. We also show that using well-matched linkage disequilibrium (LD) references can improve model fit and translate into more accurate prediction. Finally, we discuss how to make polygenic score methods such as lassosum and LDpred2 more robust to these misspecifications to improve their predictive power.
The predictive performance of polygenic scores (PGS) is largely dependent on the number of samples available to train the PGS. Increasing the sample size for a specific phenotype is expensive and takes time, but this sample size can be effectively increased by using genetically correlated phenotypes. We propose a framework to generate multi-PGS from thousands of publicly available genome-wide association studies (GWAS) with no need to individually select the most relevant ones. In this study, the multi-PGS framework increased prediction accuracy over single PGS for all included psychiatric disorders and other available outcomes, with prediction R2 increases of up to 9-fold for attention-deficit/hyperactivity disorder (ADHD) compared to a single PGS. We also generate multi-PGS for phenotypes without an existing GWAS and for case-case predictions, with up to 15-fold increases in prediction accuracy. We benchmark the multi-PGS framework against other methods and highlight its potential application to new emerging biobanks.