Despite great progress on methods for case-control polygenic prediction (e.g. schizophrenia vs. control), there remains an unmet need for a method that genetically distinguishes clinically related disorders (e.g. schizophrenia (SCZ) vs. bipolar disorder (BIP) vs. depression (MDD) vs. control); such a method could have important clinical value, especially at disorder onset when differential diagnosis can be challenging. Here, we introduce a method, Differential Diagnosis-Polygenic Risk Score (DDx-PRS), that jointly estimates posterior probabilities of each possible diagnostic category (e.g. SCZ=50%, BIP=25%, MDD=15%, control=10%) by modeling variance/covariance structure across disorders, leveraging case-control polygenic risk scores (PRS) for each disorder (computed using existing methods) and prior clinical probabilities for each diagnostic category. DDx-PRS uses only summary-level training data and does not use tuning data, facilitating implementation in clinical settings. In simulations, DDx-PRS was well-calibrated (whereas a simpler approach that analyzes each disorder marginally was poorly calibrated), and effective in distinguishing each diagnostic category vs. the rest. We then applied DDx-PRS to Psychiatric Genomics Consortium SCZ/BIP/MDD/control data, including summary-level training data from 3 case-control GWAS ( N =41,917-173,140 cases; total N =1,048,683) and held-out test data from different cohorts with equal numbers of each diagnostic category (total N =11,460). DDx-PRS was well-calibrated and well-powered relative to these training sample sizes, attaining AUCs of 0.66 for SCZ vs. rest, 0.64 for BIP vs. rest, 0.59 for MDD vs. rest, and 0.68 for control vs. rest. DDx-PRS produced comparable results to methods that leverage tuning data, confirming that DDx-PRS is an effective method. True diagnosis probabilities in top deciles of predicted diagnosis probabilities were considerably larger than prior baseline probabilities, particularly in projections to larger training sample sizes, implying considerable potential for clinical utility under certain circumstances. In conclusion, DDx-PRS is an effective method for distinguishing clinically related disorders.
Many traits show small global sex differences in genetic correlations and heritability. However, how these differences are distributed across the genome remains unknown. Here, we use LAVA to test for local genetic sex differences in genetic correlations, heritabilities, and the magnitude of genetic effects across 157 quantitative traits in the UK Biobank. Nearly every trait shows evidence for sex-dimorphic effects in at least one locus. We find that such loci can flag biological differences between the sexes. Moreover, we test for differences in the magnitude of genetic effects on the raw and the standardized scale. We show these have complementary interpretations, where only the latter scale is informative for heritability. Our results show how average metrics of genetic correlation and heritability across the whole genome can mask important variability between loci and that the scale of genetic effects needs to be considered carefully when comparing their magnitudes.
Despite great progress on methods for case-control polygenic prediction (e.g. schizophrenia vs. control), there remains an unmet need for a method that genetically distinguishes clinically related disorders (e.g. schizophrenia (SCZ) vs. bipolar disorder (BIP) vs. depression (MDD) vs. control); such a method could have important clinical value, especially at disorder onset when differential diagnosis can be challenging. Here, we introduce a method, Differential Diagnosis-Polygenic Risk Score (DDx-PRS), that jointly estimates posterior probabilities of each possible diagnostic category (e.g. SCZ=50%, BIP=25%, MDD=15%, control=10%) by modeling variance/covariance structure across disorders, leveraging case-control polygenic risk scores (PRS) for each disorder (computed using existing methods) and prior clinical probabilities for each diagnostic category. DDx-PRS uses only summary-level training data and does not use tuning data, facilitating implementation in clinical settings. In simulations, DDx-PRS was well-calibrated (whereas a simpler approach that analyzes each disorder marginally was poorly calibrated), and effective in distinguishing each diagnostic category vs. the rest. We then applied DDx-PRS to Psychiatric Genomics Consortium (PGC) SCZ/BIP/MDD/control data, including summary-level training data from 3 case-control GWAS (N=41,917-173,140 cases; total N=1,048,683) and held-out test data from different cohorts with equal numbers of each diagnostic category (total N=11,460). DDx-PRS was well-calibrated and well-powered relative to these training sample sizes, attaining AUCs of 0.66 for SCZ vs. rest, 0.64 for BIP vs. rest, 0.59 for MDD vs. rest, and 0.68 for control vs. rest for test sample size ratios of 25% for SCZ/BIP/MDD/control, and 0.65 for SCZ vs. rest, 0.61 for BIP vs. rest and 0.64 for MDD vs. rest for case-only analyses with test sample size ratios of 33.3% for SCZ/BIP/MDD (and 0% for controls). DDx-PRS produced comparable results to methods that leverage tuning data, confirming that DDx-PRS is an effective method. True diagnosis probabilities in top deciles of predicted diagnosis probabilities were considerably larger than prior baseline probabilities, particularly in projections to larger training sample sizes, implying appreciable potential for clinical utility under certain circumstances. In conclusion, DDx-PRS is an effective method for distinguishing clinically related disorders.
Background The clinical utility of polygenic scores (PGS) in non-European populations remains limited due to the Eurocentric bias of genome-wide association studies. While methods have been developed to improve the variance explained by PGS in diverse ancestral backgrounds, method development to attain well-calibrated prediction of PGS lags. Calibration denotes the agreement between predicted and real disorder probability and is crucial for meaningful absolute probability estimates and clinical implementation. Existing research on calibration in diverse ancestral backgrounds has mainly focused on continuous traits rather than disorder traits, and methods typically rely on large and ancestry-matched case-control data that are often unavailable, especially for rare disorders. This gap is particularly relevant for psychiatric traits, where comorbidity and diagnostic ambiguity increase the importance of well-calibrated predictions. There is an urgent need for methods that enable accurate absolute risk calibration without requiring tuning sets of ancestry-matched cases and controls. Methods We introduce a new method, Bayesian polygenic score Probability Conversion for diverse ancestral backgrounds (BPCx), extending the BPC method for well-calibrated prediction in individuals of European ancestral background only. BPCx enables ancestry-aware transformation of PGS to the absolute disorder probability without case-control tuning data. BPCx first transforms the PGS to the liability scale, based on an expected value of the ancestry-specific variance explained (R2lx) and population prevalence (Kx), and an ancestry-matched population reference sample without phenotypic information (e.g., 1000 Genomes). Second, applying Bayes’ theorem, the prior disorder probability is updated to the genetically informed predicted posterior disorder probability. Calibration is evaluated using the Integrated Calibration Index (ICI). We test BPCx in simulations and real-data applications in the UK Biobank across traits of varying prevalence, including major depressive disorder, coronary artery disease, type 2 diabetes and elevated LDL. Results First, we used extensive simulations to quantify the impact of misspecifying R2lx and Kx on calibration. We found, that a sizeable degree of misspecification is tolerable, e.g., in a trait with Kx =1% and R2lx=10%, calibration remained acceptable (ICI ≤ 0.05) when the specified R2lx ranged from 7.5% to 15% and Kx from 0.5% and 4%, supporting the practical feasibility of BPCx under realistic parameter uncertainty. Second, we applied BPCx in the UK Biobank using different approaches to obtain R2lx and Kx, and confirmed that BPCx yielded well-calibrated results (mean ICI 0.017 (0.0042-0.055)) across disorders with a range of R2lx (0.013% – 10%), Kx (4.5% - 31%), and ancestries (Caribbean, Indian, Nigerian). BPCx consistently outperformed the BPC method across multiple traits and ancestries, reflected in much lower ICI values. Discussion We introduce a new method, BPCx, which achieves good calibration across ancestries, without requiring ancestry-specific tuning sets, removing a key barrier to accurate disorder probability prediction. BPCx is broadly applicable and can be integrated into existing PGS workflows for more reliable probability estimation across diverse ancestry groups. This work advances the potential of genomic medicine to benefit globally diverse populations, rather than disproportionately serving individuals of European descent.
Polygenic Scores (PGSs) summarize an individual's genetic propensity for a given trait. Bayesian methods, which improve the prediction accuracy of PGSs, are not well-calibrated for binary disorder traits in ascertained samples. This is a problem because well-calibrated PGSs are needed for future clinical implementation. We introduce the Bayesian polygenic score Probability Conversion (BPC) approach, which computes an individual's predicted disorder probability using genome-wide association study summary statistics, an existing Bayesian PGS method (e.g. PRScs, SBayesR), the individual's genotype data, and a prior disorder probability (which can be specified flexibly, based for example on literature, small reference samples, or prior elicitation). The BPC approach is practical in its application as it does not require a tuning sample with both genotype and phenotype data. Here, we show in simulated and empirical data of nine disorder traits that BPC yields well-calibrated results that are consistently better than the results of another recently published approach.
Alzheimer's disease (AD) is the most common cause of dementia, with global case numbers projected to reach 153 million in 2050 1 . AD is highly heritable, with twin-based heritability estimates of 60-80% 2 . While 1,200 causal loci are predicted to exist for AD 3 , approximately 80 have been associated with AD in two recent studies 4,5 , suggesting that many loci remain to be discovered 6 . Here, we analyzed data from 183,620 AD cases and 2.6 million controls from diverse ancestries, identifying 118 loci in a multi-ancestry analysis and 9 additional loci in ancestry-specific analyses, 48 of which are new. We identified new AD risk genes, prioritized potential drug targets, and identified microglia and, for the first time, several neuronal cell types enriched for AD-associated genetic risk. Moreover, we improved polygenic prediction and estimated a single-nucleotide polymorphism (SNP) heritability of 19%. Together, our findings offer insights into the genetic architecture and potential pathobiology of AD, as well as specific targets for future drug development research.
Background Alzheimer's disease (ALZ) is the most frequent neurodegenerative disease with a prevalence of approximately 5% in the Western population above 60 years. Worldwide more than 20 million people are affected and this is expected to double every 20 years, which places a substantial financial burden on society. According to twin studies, ALZ is highly heritable, with estimates ranging between 60 to 80%. For the late-onset type of ALZ, over 70 genetic loci have been discovered in addition to apolipoprotein E (APOE), yet these still explain only a small proportion of the phenotypic variance. The majority of gene discovery studies for ALZ have been carried out in European populations. While these have yielded several interesting loci, studying participants of non-European ancestry is expected to identify novel genetic factors. In addition, different linkage disequilibrium structures in non-European populations can aid in the search for causal variants and the prioritization of genes. ALZ is known to be more prevalent in women (2:1 female:male ratio in some Western countries) and its disease symptomology and pathology have been observed to differ between the sexes. So far, no GWAS has stratified their samples by sex to identify sex-specific genetic effects. Methods We collected samples of diverse ancestral backgrounds to conduct a large multi-ancestry GWAS of ALZ. At the time of submission, we have collected ∼150,000 cases (∼50% proxy-cases) and ∼2,300,000 controls. Of these, approximately 15% are of non-European ancestry. We will use the unique linkage disequilibrium structure of our sample to hone in on causal variants and genes within each associated locus. Polygenic Risk Score analyses will be conducted to test whether ancestrally diverse GWAS summary statistics can improve risk predictions over European-only GWAS summary statistics. Furthermore, we will run sex-stratified meta-analyses and use LDSC and LAVA to study potential genetic sex differences. Results Preliminary analyses identified 90 genomic risk loci, at least 13 of which are novel. We are currently in the process of closely investigating the biological implications of all newly identified loci. In aggregate, our results again point toward the importance of the immune response in the etiology of ALZ. Discussion We have conducted the largest multi-ancestry GWAS of ALZ to date and have identified several novel loci. At the time of abstract submission, the results are very preliminary, yet promising. We are currently running extensive post-GWAS analyses which will be presented, together with their biological implications. With the approval of novel disease-modifying drugs for ALZ, research has come closer to the treatment of the disease. However, these drugs remain controversial and their efficacy is minimal. It is clear that novel avenues, beyond targeting amyloid plaques, are required. The identification of novel genetic risk factors may be one early step in this direction.We will make several sets of summary statistics available to the wider community, including ancestry-specific, sex-stratified, and proxy and true case-stratified summary statistics, which we hope will maximize the utility of this resource.
Background: Neuroinflammation is involved in the pathophysiology of Alzheimer's disease (AD), including immune-linked genetic variants and molecular pathways, microglia and astrocytes. Multiple Sclerosis (MS) is a chronic, immune-mediated disease with genetic and environmental risk factors and neuropathological features. There are clinical and pathobiological similarities between AD and MS. Here, we investigated shared genetic susceptibility between AD and MS to identify putative pathological mechanisms shared between neurodegeneration and the immune system. Methods: We analysed GWAS data for late-onset AD (N cases = 64,549, N controls = 634,442) and MS (N cases = 14,802, N controls = 26,703). Gaussian causal mixture modelling (MiXeR) was applied to characterise the genetic architecture and overlap between AD and MS. Local genetic correlation was investigated with Local Analysis of [co]Variant Association (LAVA). The conjunctional false discovery rate (conjFDR) framework was used to identify the specific shared genetic loci, for which functional annotation was conducted with FUMA and Open Targets. Results: MiXeR analysis showed comparable polygenicities for AD and MS (approximately 1800 trait-influencing variants) and genetic overlap with 20% of shared trait-influencing variants despite negligible genetic correlation (rg = 0.03), suggesting mixed directions of genetic effects across shared variants. conjFDR analysis identified 16 shared genetic loci, with 8 having concordant direction of effects in AD and MS. Annotated genes in shared loci were enriched in molecular signalling pathways involved in inflammation and the structural organisation of neurons. Conclusions: Despite low global genetic correlation, the current results provide evidence for polygenic overlap between AD and MS. The shared loci between AD and MS were enriched in pathways involved in inflammation and neurodegeneration, highlighting new opportunities for future investigation.
Drug repurposing may provide a solution to the substantial challenges facing de novo drug development. Given that 66% of FDA-approved drugs in 2021 were supported by human genetic evidence, drug repurposing methods based on genome wide association studies (GWAS), such as drug gene-set analysis, may prove an efficient way to identify new treatments. However, to our knowledge, drug gene-set analysis has not been tested in non-psychiatric phenotypes, and previous implementations may have contained statistical biases when testing groups of drugs. Here, 1201 drugs were tested for association with hypercholesterolemia, type 2 diabetes, coronary artery disease, asthma, schizophrenia, bipolar disorder, Alzheimer’s disease, and Parkinson’s disease. We show that drug gene-set analysis can identify clinically relevant drugs (e.g., simvastatin for hypercholesterolemia [ p = 2.82E-06]; mitiglinide for type 2 diabetes [ p = 2.66E-07]) and drug groups (e.g., C10A for coronary artery disease [ p = 2.31E-05]; insulin secretagogues for type 2 diabetes [ p = 1.09E-11]) for non-psychiatric phenotypes. Additionally, we demonstrate that when the overlap of genes between drug-gene sets is considered we find no groups containing approved drugs for the psychiatric phenotypes tested. However, several drug groups were identified for psychiatric phenotypes that may contain possible repurposing candidates, such as ATC codes J02A ( p = 2.99E-09) and N07B ( p = 0.0001) for schizophrenia. Our results demonstrate that clinically relevant drugs and groups of drugs can be identified using drug gene-set analysis for a number of phenotypes. These findings have implications for quickly identifying novel treatments based on the genetic mechanisms underlying diseases.
Despite the substantial heritability of antisocial behavior (ASB), specific genetic variants robustly associated with the trait have not been identified. The present study by the Broad Antisocial Behavior Consortium (BroadABC) meta-analyzed data from 28 discovery samples ( N = 85,359) and five independent replication samples ( N = 8058) with genotypic data and broad measures of ASB. We identified the first significant genetic associations with broad ASB, involving common intronic variants in the forkhead box protein P2 (FOXP2) gene (lead SNP rs12536335, p = 6.32 × 10 −10 ). Furthermore, we observed intronic variation in Foxp2 and one of its targets (Cntnap2) distinguishing a mouse model of pathological aggression (BALB/cJ strain) from controls (BALB/cByJ strain). Polygenic risk score (PRS) analyses in independent samples revealed that the genetic risk for ASB was associated with several antisocial outcomes across the lifespan, including diagnosis of conduct disorder, official criminal convictions, and trajectories of antisocial development. We found substantial genetic correlations of ASB with mental health (depression r g = 0.63, insomnia r g = 0.47), physical health (overweight r g = 0.19, waist-to-hip ratio r g = 0.32), smoking ( r g = 0.54), cognitive ability (intelligence r g = −0.40), educational attainment (years of schooling r g = −0.46) and reproductive traits (age at first birth r g = −0.58, father’s age at death r g = −0.54). Our findings provide a starting point toward identifying critical biosocial risk mechanisms for the development of ASB.
Introductory Paragraph Recently, large-scale sex-stratified Genome-Wide Association Studies (GWASs) have been conducted that compare heritabilities and estimate global genetic correlations between males and females 1–5 . These studies have identified a small number of traits that show genetic sex differences on a global level. However, such global analyses cannot elucidate the local architecture of genetic sex differences 6 . To address this gap and gain insight into local (dis-)similarities in genetic signal between males and females, we estimated local genetic correlations and tested whether these differ from one, compared local heritability estimates, and tested for equality of local genetic effects in 157 quantitative traits. While the vast majority of loci do not significantly differ between males and females, almost every trait we studied had at least one locus that did. We show that these loci can highlight trait-relevant biology and may partly explain observed phenotypic sex differences.
Most neuropsychiatric disorders are highly polygenic, implicating hundreds to thousands of causal genetic variants that span much of the genome. This widespread polygenicity complicates biological understanding because no single variant can explain disease etiology. A strategy to advance biological insight is to seek convergent functions among the large set of variants and map them to a smaller set of disease-relevant genes and pathways. Accordingly, functional genomic resources that provide data on intermediate molecular phenotypes, such as gene-expression and methylation status, can be leveraged to functionally annotate variants and map them to genes. Such molecular quantitative trait locus mappings can be integrated with genome-wide association studies to make sense of the polygenic signal that underlies complex disease. Other resources that provide data on the 3-dimensional structure of chromatin and functional importance of specific genomic regions can be integrated similarly. In addition, mapped genes can then be tested for convergence in biological function, tissue, cell type, or developmental stage. In this review, we provide an overview of functional genomic resources and methods that can be used to interpret results from genome-wide association studies, and we discuss current challenges for biological understanding and future requirements to overcome them.
Most neuropsychiatric disorders are highly polygenic, implicating hundreds to thousands of causal genetic variants that span much of the genome. This widespread polygenicity complicates biological understanding because no single variant can explain disease etiology. A strategy to advance biological insight is to seek convergent functions among the large set of variants and map them to a smaller set of disease-relevant genes and pathways. Accordingly, functional genomic resources that provide data on intermediate molecular phenotypes, such as gene-expression and methylation status, can be leveraged to functionally annotate variants and map them to genes. Such molecular quantitative trait locus mappings can be integrated with genome-wide association studies to make sense of the polygenic signal that underlies complex disease. Other resources that provide data on the 3-dimensional structure of chromatin and functional importance of specific genomic regions can be integrated similarly. In addition, mapped genes can then be tested for convergence in biological function, tissue, cell type, or developmental stage. In this review, we provide an overview of functional genomic resources and methods that can be used to interpret results from genome-wide association studies, and we discuss current challenges for biological understanding and future requirements to overcome them.
Genome-wide association studies (GWAS) test hundreds of thousands of genetic variants across many genomes to find those statistically associated with a specific trait or disease. This methodology has generated a myriad of robust associations for a range of traits and diseases, and the number of associated variants is expected to grow steadily as GWAS sample sizes increase. GWAS results have a range of applications, such as gaining insight into a phenotype’s underlying biology, estimating its heritability, calculating genetic correlations, making clinical risk predictions, informing drug development programmes and inferring potential causal relationships between risk factors and health outcomes. In this Primer, we provide the reader with an introduction to GWAS, explaining their statistical basis and how they are conducted, describe state-of-the art approaches and discuss limitations and challenges, concluding with an overview of the current and future applications for GWAS results.