Polygenic risk scores, or PRS, have been widely used across many traits to estimate polygenic risk, pleiotropy, and disease prediction. While PRS has the potential to be simple and informative they are not generally interpretable and not directly comparable between studies. They depend on the specific approach used, the chosen parameters, the number of genetic variants included and the underlying population structure. One way to improve comparability is to place an individual’s score in the context of the PRS distribution of the ancestry-matched population. In this work, we present a method to estimate the parameters of PRS distributions in any population, using only publicly available summary data. It can be applied to quickly assess individual’s polygenic disease risk for any complex genetic disorder assuming that risk loci are shared across populations. We demonstrate the accuracy of this method through simulations and present population-specific PRS examples derived from genome-side association studies (GWAS) of two neurodegenerative diseases, Alzheimer’s disease (AD) and amyotrophic lateral sclerosis (ALS) using data from the 1000 Genomes Project.
Abstract Alzheimer’s disease pathology and cognitive outcomes frequently diverge, yet current single-axis definitions cannot identify resilient (high pathology, preserved cognition) and resistant (high risk, low pathology) subgroups reliably at scale, obscuring the mechanisms that uncouple pathological burden from cognitive decline. Here, we developed a multivariate blood-based framework integrating 19 molecular assays and six risk instruments in the Bio-Hermes-001 cohort (n=1,009). Unsupervised clustering identified resilient (n=91) and resistant (n=81) subgroups, together comprising 17% of the cohort, with distinct amyloid, tau, and neurodegeneration profiles. Amyloid-PET yielded convergent but only partially overlapping classifications. Proteomic, cytokine, and polygenic profiling further distinguished resistance through an APOE-centred genomic signature and resilience through neuroinflammatory markers associated with progression toward clinical Alzheimer’s disease. A four-biomarker panel (Aβ40, p-tau217, p-tau181, NfL) reproduced subgroup assignments with 83% accuracy. These findings support resilience and resistance as molecularly distinct subgroups and provide a scalable framework for pathology-informed stratification and mechanistic investigation.
Large-scale plasma proteomics can capture molecular changes across the Alzheimer's disease (AD) continuum and provide insight into biological mechanisms associated with AD pathology. We analysed the Bio-Hermes cohort (n = 961), with participants enrolled across 17 sites in the United States from April 2021 to November 2022. Participants were stratified by clinical status and amyloid PET scan-based Core1 biomarker status (CN Core1–, CN Core1+, MCI Core1+ and AD dementia Core1+). We performed differential abundance analyses across biologically defined contrasts, clustered proteins into co-expression networks, and evaluated protein panels to distinguish participants with biologically defined AD from amyloid-negative cognitively normal controls. We also used Mendelian randomization (MR) to assess genetic evidence for potential causal relationships with AD risk. The biologically defined contrast, Core1+ vs. CN Core1– , identified 69 differentially abundant proteins. Across AD stages, eight core proteins were consistently dysregulated from preclinical through prodromal and dementia phases, and three additional proteins emerged at MCI Core1+ and remained altered in AD dementia Core1+. We identified 29 co-expression modules, six of which varied significantly across the AD continuum. Among differential abundance proteins, ACHE ranked highest for distinguishing biologically defined AD from CN Core1–. Stage-specific protein panels improved the discriminatory performance for MCI Core1+ (AUC = 0.850) and AD dementia Core1+ (AUC = 0.856). MR provided genetic evidence consistent with an association between plasma ACHE abundance and AD risk. Plasma proteomics delineated a stage-spanning core signature across the AD continuum. These findings nominate co-expression modules and candidate proteins for further validation in early detection and AD screening.
Abstract Background The success of selecting high risk or early-stage Alzheimer’s disease individuals for the delivery of clinical trials depends on the design and the appropriate recruitment of participants. Polygenic risk scores (PRS) show potential for identifying individuals at risk for Alzheimer’s disease (AD). Our study comprehensively examines AD PRS utility using various methods and models. Methods We compared the PRS prediction accuracy in ADNI (N = 568) and BioFINDER (N = 766) cohorts using five disease risk modelling approaches, three PRS derivation methods, two AD genome-wide association study (GWAS) statistics and two sets of SNPs: the whole genome and microglia-selective regions only. Results The best prediction accuracy was achieved when modelling genetic risk by using two predictors: APOE and remaining PRS (AUC = 0.72–0.76). Microglial PRS showed comparable accuracy to the whole genome (AUC = 0.71–0.74). The individuals’ risk scores differed substantially, with the largest discrepancies (up to 70%) attributable to the GWAS statistics used. Conclusions Our work benchmarks the best PRS derivation and modelling strategies for AD genetic prediction.
Background Postpartum psychosis is the most severe postpartum mental illness, affecting 1-2 in every 1000 childbirths. Research into the causes and risk factors of postpartum psychosis has been impeded by confusion around its classification. Recent evidence suggests that the genetic architecture of postpartum psychosis differs from that of bipolar disorder. This work aimed to explore this further. Methods Cases were ascertained through the Bipolar Disorder Research Network (BDRN), who were assessed using semi-structured interviews and case note review. Postpartum psychosis was defined as a manic, mixed, or psychotic depression episode within 6 weeks of delivery. Healthy female controls were recruited through the national UK Blood Services, the 1958 British Birth Cohort (UK National Child Development Study) and the UK Household Longitudinal Study. The total sample size for the genome-wide association study was 772 cases and 8,537 controls. Heritability was estimated using LDAK with restricted maximum likelihood. Genetic correlations with schizophrenia and major depression were conducted using Linkage Disequilibrium Score Regression (LDSC). Muti-marker Analysis of GenoMic Annotation (MAGMA) through the online tool Functional Mapping and Annotation of Genome-Wide Association Studies (FUMA) was used to identify independent lead SNPs and conduct functional gene mapping. Results No genome-wide significant SNPs were identified, however, several were suggestive at a genome-wide significance level of p < 1 × 10-5. These mapped onto 13 genetic loci and 23 genes. No significant enrichment was found for any gene sets or specific tissues. SNP heritability for postpartum psychosis was estimated to be 0.54 (SE: 0.038). Genetic correlation with bipolar type-I was high (0.57, SE: 0.11, p = 5.18 × 10-8) but significantly different from 1 (p = 9.26 × 10-5). Genetic correlation with schizophrenia was similarly high (0.55, SE: 0.10, p = 1.84 × 10-8) and lower for bipolar type-II and major depression (0.31, SE: 0.17, p = 0.061; 0.34, SE: 0.15, p = 0.029). Discussion These results suggest substantial genetic overlap between postpartum psychosis and bipolar type-I and schizophrenia. They also provide evidence for the genetic architecture of postpartum psychosis being partially distinct from bipolar disorder. This supports previous literature and advocates for postpartum psychosis being included as its own distinct entity within diagnostic manuals.
Common forms of Alzheimer's disease (AD) are complex and polygenic. We have created a research resource that seeks to capture the extremes of polygenic risk in a collection of human induced pluripotent stem cell (iPSC) lines from over 100 donors: the IPMAR Resource (iPSC Platform to Model Alzheimer's Disease Risk). Donors were selected from a large UK cohort of 6,000+ research-diagnosed early or late-onset AD cases and elderly cognitively healthy controls, many of whom have lived through the age of risk for disease development (>85 years). We include iPSC with extremes of global AD polygenic risk (high-risk late-onset AD: 34; high-risk early-onset AD: 29; low-risk control: 27) as well as those reflecting complement pathway-specific genetic risk (high-risk AD: 9; low-risk controls: 10). All iPSC have associated clinical, longitudinal, and genetic datasets and will be available through collaboration or from cell (EBiSC) and data (DPUK) repositories.
Increasing evidence supports a role for deficient Wnt signaling in Alzheimer's disease (AD). Studies reveal that the secreted Wnt antagonist Dickkopf-3 (DKK3) colocalizes to amyloid plaques in AD patients. Here, we investigate the contribution of DKK3 to synapse integrity in healthy and AD brains. Our findings show that DKK3 expression is upregulated in the brains of AD subjects and that DKK3 protein levels increase at early stages in the disease. In hAPP-J20 and hAPPNL-G-F/NL-G-F mouse AD models, extracellular DKK3 levels are increased and DKK3 accumulates at dystrophic neuronal processes around plaques. Functionally, DKK3 triggers the loss of excitatory synapses through blockade of the Wnt/GSK3β signaling with a concomitant increase in inhibitory synapses via activation of the Wnt/JNK pathway. In contrast, DKK3 knockdown restores synapse number and memory in hAPP-J20 mice. Collectively, our findings identify DKK3 as a novel driver of synaptic defects and memory impairment in AD.
Frontotemporal dementia (FTD) is the second most common cause of early-onset dementia after Alzheimer disease (AD). Efforts in the field mainly focus on familial forms of disease (fFTDs), while studies of the genetic etiology of sporadic FTD (sFTD) have been less common. In the current work, we analyzed 4,685 sFTD cases and 15,308 controls looking for common genetic determinants for sFTD. We found a cluster of variants at the MAPT (rs199443; p = 2.5 × 10-12, OR = 1.27) and APOE (rs6857; p = 1.31 × 10-12, OR = 1.27) loci and a candidate locus on chromosome 3 (rs1009966; p = 2.41 × 10-8, OR = 1.16) in the intergenic region between RPSA and MOBP, contributing to increased risk for sFTD through effects on expression and/or splicing in brain cortex of functionally relevant in-cis genes at the MAPT and RPSA-MOBP loci. The association with the MAPT (H1c clade) and RPSA-MOBP loci may suggest common genetic pleiotropy across FTD and progressive supranuclear palsy (PSP) (MAPT and RPSA-MOBP loci) and across FTD, AD, Parkinson disease (PD), and cortico-basal degeneration (CBD) (MAPT locus). Our data also suggest population specificity of the risk signals, with MAPT and APOE loci associations mainly driven by Central/Nordic and Mediterranean Europeans, respectively. This study lays the foundations for future work aimed at further characterizing population-specific features of potential FTD-discriminant APOE haplotype(s) and the functional involvement and contribution of the MAPT H1c haplotype and RPSA-MOBP loci to pathogenesis of sporadic forms of FTD in brain cortex.
INTRODUCTION:The Dementias Platform UK (DPUK) Data Portal is a data repository bringing together a wide range of cohorts. Neurodegenerative dementias are a group of diseases with highly heterogeneous pathology and an overlapping genetic component that is poorly understood. The DPUK collection of independent cohorts can facilitate research in neurodegeneration by combining their genetic and phenotypic data. METHODS:For genetic data processing, pipelines were generated to perform quality control analysis, genetic imputation, and polygenic risk score (PRS) derivation with six genome-wide association studies of neurodegenerative diseases. Pipelines were applied to five cohorts. DISCUSSION:The data processing pipelines, research-ready imputed genetic data, and PRS scores are now available on the DPUK platform and can be accessed upon request though the DPUK application process. Harmonizing genome-wide data for multiple datasets increases scientific opportunity and allows the wider research community to access and process data at scale and pace.
Although there are several genome-wide association studies available which highlight genetic variants associated with Alzheimer's disease (AD), often the X chromosome is excluded from the analysis. We conducted an X-chromosome-wide association study (XWAS) in three independent studies with a pathologically confirmed phenotype (total 1970 cases and 1113 controls). The XWAS was performed in males and females separately, and these results were then meta-analysed. Four suggestively associated genes were identified which may be of potential interest for further study in AD, these are DDX53 (rs12006935, OR = 0.52, p = 6.9e-05), IL1RAPL1 (rs6628450, OR = 0.36, p = 4.2e-05; rs137983810, OR = 0.52, p = 0.0003), TBX22 (rs5913102, OR = 0.74, p = 0.0003) and SH3BGRL (rs186553004, OR = 0.35, p = 0.0005; rs113157993, OR = 0.52, p = 0.0003), which replicate across at least two studies. The SNP rs5913102 in TBX22 achieves chromosome-wide significance in meta-analysed data. DDX53 shows highest expression in astrocytes, IL1RAPL1 is most highly expressed in oligodendrocytes and neurons and SH3BGRL is most highly expressed in microglia. We have also identified SNPs in the NXF5 gene at chromosome-wide significance in females (rs5944989, OR = 0.62, p = 1.1e-05) but not in males (p = 0.83). The discovery of relevant AD associated genes on the X chromosome may identify AD risk differences and similarities based on sex and lead to the development of sex-stratified therapeutics.
DPUK is a data repository bringing together a wide range of population and clinical cohorts from the UK, Europe and South Korea. It enables data discovery, variable selection, and data access for multi-modal, multi-cohort analysis in a shared and secure environment. Neurodegenerative dementias are group of diseases with highly heterogeneous pathology and with an overlapping genetic component that is poorly understood. Combing genetic data from different studies is important due to the increase in power this provides, but this presents an additional challenge due to differences in genotyping arrays and lack of overlapping variants. We have curated and harmonized genome-wide data from 6 studies (Brains for Dementia Research, EPIC Norfolk, Generation Scotland, Airwave-chip 1, Airwave-chip 2, MRC National Survey of Health Development (NSHD)) within the DPUK platform. We created a pipeline that performs rigorous quality-control (QC) analysis and imputes genetic variance with the 1000 Genome reference panel using the minimac algorithm. Further, standard pruning and thresholding Polygenic Risk Score (PRS) have been generated with 5 summary statistics related to neurodegenerative diseases (clinical AD, clinical/proxy AD, FTD, ALS, PD) for all 6 studies separately and for the combined dataset. Research-ready imputed genetic data that have undergone QC and the PRS, based upon the latest neurodegenerative GWAS, are available for the 6 cohorts separately, and for the combined dataset of 60,670 individuals. Preparing multiple datasets to a common standard for research-readiness increases scientific opportunity and allows the wider research community to access and process data at scale and pace.
Background Metformin, a medication for type 2 diabetes, has been linked to many non-diabetes health benefits including increasing healthy lifespan. Previous work has only examined the benefits of metformin over periods of less than ten years, which may not be long enough to capture the true effect of this medication on longevity. Methods We searched medical records for Wales, UK, using the Secure Anonymised Information Linkage dataset for type 2 diabetes patients treated with metformin (N = 129,140) and sulphonylurea (N = 68,563). Non-diabetic controls were matched on sex, age, smoking, and history of cancer and cardiovascular disease. Survival analysis was performed to examine survival time after first treatment, using a range of simulated study periods. Findings Using the full twenty-year period, we found that type 2 diabetes patients treated with metformin had shorter survival time than matched controls, as did sulphonylurea patients. Metformin patients had better survival than sulphonylurea patients, controlling for age. Within the first three years, metformin therapy showed a benefit over matched controls, but this reversed after five years of treatment. Interpretation While metformin does appear to confer benefits to longevity in the short term, these initial benefits are outweighed by the effects of type 2 diabetes when patients are observed over a period of up to twenty years. Longer study periods are therefore recommended for studying longevity and healthy lifespan. Evidence before this study Work examining the non-diabetes outcomes of metformin therapy has suggested that there metformin has a beneficial effect on longevity and healthy lifespan. Both clinical trials and observational studies broadly support this hypothesis, but tend to be limited in the length of time over which they can study patients or participants. Added value of this study By using medical records we are able to study individuals with Type 2 diabetes over a period of two decades. We are also able to account for the effects of cancer, cardiovascular disease, hypertension, deprivation, and smoking on longevity and survival time following treatment. Implications of all the available evidence We confirm that there is an initial benefit to longevity of metformin therapy, but this benefit does not outweigh the negative effect on longevity of diabetes. Therefore, we suggest that longer study periods are required for inference to be made about longevity in future research.
Sixteen control subjects and six right brain-damaged patients with left hemiparesis (three showing signs of left unilateral neglect, three with no signs of neglect) performed a straight-ahead pointing task with their right hand while blindfolded. The aim was to test the hypothesis that the egocentric reference shows significant ipsilesional deviation in left neglect patients. We found no correlation between the position of the egocentric reference and the presence of neglect signs. Neglect patients, like non-neglect patients, showed leftware, rightward or no significant deviation when pointing straight ahead. Results are discussed with reference to egocentric hypotheses of neglect and experimental remission of neglect.
Introduction Both late-onset Alzheimer’s disease (AD) and ageing have a strong genetic component. In each case, many associated variants have been discovered, but how much missing heritability remains to be discovered is debated. Variability in the estimation of SNP-based heritability could explain the differences in reported heritability. Methods We compute heritability in five large independent cohorts (N = 7,396, 1,566, 803, 12,528 and 3,963) to determine whether a consensus for the AD heritability estimate can be reached. These cohorts vary by sample size, age of cases and controls and phenotype definition. We compute heritability a) for all SNPs, b) excluding APOE region, c) excluding both APOE and genome-wide association study hit regions, and d) SNPs overlapping a microglia gene-set. Results SNP-based heritability of late onset Alzheimer’s disease is between 38 and 66% when age and genetic disease architecture are correctly accounted for. The heritability estimates decrease by 12% [SD = 8%] on average when the APOE region is excluded and an additional 1% [SD = 3%] when genome-wide significant regions were removed. A microglia gene-set explains 69–84% of our estimates of SNP-based heritability using only 3% of total SNPs in all cohorts. Conclusion The heritability of neurodegenerative disorders cannot be represented as a single number, because it is dependent on the ages of cases and controls. Genome-wide association studies pick up a large proportion of total AD heritability when age and genetic architecture are correctly accounted for. Around 13% of SNP-based heritability can be explained by known genetic loci and the remaining heritability likely resides around microglial related genes.
Prediction models of Alzheimer’s disease using genetic information, such as polygenic risk scores, have been able to reach high levels of prediction accuracy. However, since confirmation of the disease is only possible post-mortem, these levels of prediction accuracy are typically only achievable in pathologically confirmed cohorts. In living individuals who have been clinically assessed for AD, prediction accuracy by genetics is still good but much lower. Biomarkers can indirectly assess AD pathologies, and so may be able to bridge the prediction accuracy gap. Blood plasma biomarkers may have additional clinical utility as they are cheaper and more accessible compared to traditional CSF or PET methods. We measured five blood plasma biomarkers known to be linked to AD pathologies (Aβ40, Aβ42, GFAP, NfL, P-tau181) in a cohort of AD cases (N=1439, mean age 68) and elderly screened controls (N=508, mean age 82). We also gathered information on APOE genotype, age at sample collection, sex, and age at onset and disease duration in cases. Linear regression models showed that all biomarker measurements were associated with age at onset in cases and most were associated with disease duration. Biomarkers were also associated with age at sample collection in both cases and controls, demonstrating their effectiveness for tracking neurological change over time. Using logistic regression, we found prediction accuracies for AD status for each biomarker individually of AUC=0.56-0.66 and by APOE and PRS AUC=0.73. A model combining all biomarkers had an AUC=0.75. The best prediction accuracy was achieved by combining all biomarkers with genetics and age at sample collection, which reached an AUC=0.81, and explained variance of R 2 =0.29. We found that blood plasma biomarkers predicted AD status and were associated with disease duration. Furthermore, biomarkers explain some variance not captured by genetic factors and therefore improve accuracy when combined in predictive models. Biomarkers also have the advantage of specificity over clinical assessments, which may confuse dementia subtypes due to phenotype similarity. Therefore, blood plasma biomarkers can be a useful tool for the assessment and prediction of AD on their own or in combination with genetic predictors.
The APOE-epsilon 4 allele is known to predispose to amyloid deposition and consequently is strongly associated with Alzheimer's disease (AD) risk. There is debate as to whether the APOE gene accounts for all genetic variation of the APOE locus. Another question which remains is whether APOE-epsilon 4 carriers have other genetic factors influencing the progression of amyloid positive individuals to AD. We conducted a genome-wide association study in a sample of 5,390 APOE-epsilon 4 homozygous (epsilon 4 epsilon 4) individuals (288 cases and 5102 controls) aged 65 or over in the UK Biobank. We found no significant associations of SNPs in the APOE locus with AD in the sample of 8 4 8 4 individuals. However, we identified a novel genome-wide significant locus associated to AD, mapping to DAB1 (rs112437613, OR = 2.28, CI = 1.73-3.01, p = 5.4 x 10(-9)). This identification of DAB1 led us to investigate other components of the DAB1-RELN pathway for association. Analysis of the DAB1-RELN pathway indicated that the pathway itself was associated with AD, therefore suggesting an epistatic interaction between the APOE locus and the DAB1-RELN pathway. (C) 2022 The Author(s). Published by Elsevier Inc.
Recent genome-wide studies have identified over 70 risk loci for late onset Alzheimer’s Disease (AD) (Kunkle et al. 2019, Bellenguez et el. 2021, Wightman et al. 2021). Analysis in this study focused on developing Machine Learning (ML) models to predict AD risk from genetic data. We compared the prediction accuracy of ML to polygenic risk scores (PRS) using SNPs in disease-associated biological pathways. We used the Genetic and Environmental Risk for Alzheimer’s Disease consortium dataset (Harold et al. 2009). SNPs were selected from the AD associated pathways in Kunkle et al. 2019. Two decision tree-based ML algorithms were used, Random Forests (RFs) and Gradient Boosting (GB). The prediction ability of these methods was compared to Polygenic Risk Score (PRS). RFs and GB models were developed using the Python library sklearn . These were trained and tested using 5-fold cross-validation. Clumping and thresholding (CT), as implemented in PLINK, was used to generate PRS, with logistic regression used for prediction. CT PRS was compared to the PRS generated by the PRS-CS method (Ge, T., Chen, CY., Ni, Y. et al). Initial results demonstrate that PRS, PRS-CS and ML perform similarly for pathway specific analysis (AUC∼69%) when pathways include APOE . Subsequent analyses will compare the performance of the methods when APOE SNPs are removed from the pathways, and in multivariate analyses modelling multiple pathway effects simultaneously.
Polygenic risk scores (PRSs) can boost risk-prediction in late-onset Alzheimer’s disease (LOAD) beyond apolipoprotein E ( APOE) but have not been leveraged to identify genetic resilience factors. Here, we sought to identify resilience-conferring common genetic variants in 1) unaffected individuals having high PRSs for LOAD, and 2) unaffected APOE - ε 4 carriers also having high PRSs for LOAD. We used genome-wide association study (GWAS) to contrast “resilient” unaffected individuals at the highest genetic risk for LOAD with LOAD cases at comparable risk. From GWAS results, we constructed polygenic resilience scores to aggregate the addictive contributions of risk-orthogonal common variants that promote resilience to LOAD. Replication of resilience scores was undertaken in eight independent studies. We successfully replicated two polygenic resilience scores that reduce genetic-risk penetrance for LOAD. We also showed that polygenic resilience scores positively correlate with polygenic risk scores in unaffected individuals, perhaps aiding in staving off disease. Our findings align with the hypothesis that a combination of risk-independent common variants mediates resilience to LOAD by moderating genetic disease risk.
Polygenic risk scores (PRS) have been widely adopted as a tool for measuring common variant liability and they have been shown to predict lifetime risk of Alzheimer’s disease (AD) development. However, the relationship between PRS and AD pathogenesis is largely unknown. To this end, we performed a differential gene-expression and associated disrupted biological pathway analyses of AD PRS vs. case/controls in human brain-derived cohort sample (cerebellum/temporal cortex; MayoRNAseq). The results highlighted already implicated mechanisms: immune and stress response, lipids, fatty acids and cholesterol metabolisms, endosome and cellular/neuronal death, being disrupted biological pathways in both case/controls and PRS, as well as previously less well characterised processes such as cellular structures, mitochondrial respiration and secretion. Despite heterogeneity in terms of differentially expressed genes in case/controls vs. PRS, there was a consensus of commonly disrupted biological mechanisms. Glia and microglia-related terms were also significantly disrupted, albeit not being the top disrupted Gene Ontology terms. GWAS implicated genes were significantly and in their majority, up-regulated in response to different PRS among the temporal cortex samples, suggesting potential common regulatory mechanisms. Tissue specificity in terms of disrupted biological pathways in temporal cortex vs. cerebellum was observed in relation to PRS, but limited tissue specificity when the datasets were analysed as case/controls. The largely common biological mechanisms between a case/control classification and in association with PRS suggests that PRS stratification can be used for studies where suitable case/control samples are not available or the selection of individuals with high and low PRS in clinical trials.