Obesity is a highly heritable trait, but rising obesity rates suggest environmental change is also of profound importance. We conducted a cross-cohort analysis to examine how associations between genetic risk for high BMI and observed BMI differed in four British birth cohorts born before and amidst the obesity epidemic (1946, 1958, 1970 and ~2001; N = 19,379). BMI (kg/m 2 ) was measured at multiple time points between ages 3 and 69 years. We used polygenic indices (PGI) derived from GWAS of adulthood and childhood BMI, respectively, with mixed effects models used to estimate associations with mean BMI and quantile regression used to assess associations across the distribution of BMI. We further used linear regression to estimate PGI-heritability (PGI-h 2 ; incremental variance explained by the PGI) and Genomic Relatedness Restricted Maximum Likelihood (GREML) to calculate SNP-heritability (SNP-h 2 ) by cohort and age. Adulthood BMI PGI was associated with BMI in all cohorts and ages but was more strongly associated with BMI in more recently born generations, e.g., at age 16y, a 1 SD increase in the adulthood PGI was associated with 0.46 kg/m 2 (0.37, 0.55) higher BMI in the 1946c and 0.90 kg/m 2 (0.83, 0.97) higher BMI in the 2001c. Cross-cohort differences widened with age and were larger at the upper end of the BMI distribution, indicating disproportionate increases in obesity in more recent generations for those with higher PGIs. Differences were also observed when using the childhood PGI. There were no clear, consistent differences in PGI-h 2 or SNP-h 2 , possibly due to limited statistical power, except that PGI-h 2 was highest in the most recently born cohort (2001c) when using the most predictive PGI for adulthood BMI. Findings highlight how the environment can modify genetic associations; genetic associations with BMI differed by birth cohort, age, and outcome centile.
Social scientists have long sought to investigate whether the predictors of educational attainment (EA) have changed across time. Here, we provide insights by incorporating genetic predictors of education in three nationally representative British birth cohorts born in 1946, 1958, and 1970. We investigated whether individual characteristics as proxied by polygenic indexes (PGIs) for EA and cognition have become more relevant to educational success over time and whether returns to genetic predisposition were moderated by early life socioeconomic background. We present three findings. First, associations between the EA PGI and attainment increased over time, with increasing incremental variance explained by the EA PGI. Second, associations between the cognition PGI and attainment were broadly consistent across cohorts, and there was no clear change in explained variance. Since the EA PGI captures multiple traits related to educational success, factors other than those related to cognition may have become more relevant over time. Third, we observed strong evidence of interaction: Associations between the EA PGI and EA were disproportionately larger among those from more advantaged socioeconomic backgrounds. The strength and pattern of associations varied when using EA PGIs that were less conservatively filtered for SNPs. Our findings suggest EA is influenced by social and genetic factors both independently and jointly. Genetic liability and social background could be considered as two forms of inherited advantage which synergistically influence education attainment.
Birth cohort studies have a rich history of contributing to science across disciplinary fields, notably health and social sciences. Here, we introduce a curated resource comprising genomic data from five British birth cohort studies: longitudinal studies with extensive data collected prospectively across life, each deliberately sampled to be nationally representative (born 1946 to 2001). These contain health and social data from birth to older age, enabling longitudinal and cross-cohort genetically informed research. The Millennium Cohort Study additionally includes data on parents and offspring, enabling within-family analyses. Across five cohorts born in 1946, 1958, 1970, 1989/90, and 2000/2002, 27,432 participants have harmonized, imputed, and quality-controlled genetic data from genotyping arrays covering 6.7 million common SNPs. The Millennium Cohort Study contains over 6,000 mother-offspring pairs and over 3,000 mother-father-offspring trios. Pseudonymized data are freely available to the global research community upon approval of a data access request (https://cls.ucl.ac.uk/data-access-training). ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement The Centre for Longitudinal Studies is funded by the Economic and Social Research Council (grant numbers ES/M001660/1 and ES/W013142/1). DB, LW and NMD are supported by the Medical Research Council (MR/V002147/1). NMD is supported via a Norwegian Research Council Grant (295989) and the UCL Division of Psychiatry (https://www.ucl.ac.uk/psychiatry/division-psychiatry). NC and AW are supported by the Medical Research Council (MR/Y014022/1). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Ethical approval was obtained in each study: 1946c (North Thames Multicentre Research Ethics Committee: reference 98/2/121 and 07/H1008/168), 1958c (South East Multi-centre Research Ethics Committee: ref 01/1/44), 1970c (South East Coast Brighton & Sussex: ref 15/LO/1446), 1989c (East of England Cambridge Central Research Ethics Committee: ref 22/EE/0052), 2001c (London-Central REC: 13/LO/1786). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Pseudonymised data are freely available to the research community upon approval of a data application request.
Lewy body (LB) diseases are an umbrella term encompassing a range of neurodegenerative conditions all characterized by the hallmark of intra-neuronal α-synuclein associated with the development of motor and cognitive dysfunction. In this study, we have conducted a large meta-analysis of DNA methylation across multiple cortical brain regions, in relation to increasing burden of LB pathology. Utilizing a combined dataset of 1239 samples across 855 unique donors, we identified a set of 30 false discovery rate (FDR) significant loci that are differentially methylated in association with LB pathology, the most significant of which were located in UBASH3B and PTAFR, as well as an intergenic locus. Ontological enrichment analysis of our meta-analysis results highlights several neurologically relevant traits, including synaptic, inflammatory and vascular alterations. We leverage our summary statistics to compare DNA methylation signatures between different neurodegenerative pathologies and highlight a shared epigenetic profile across LB diseases, Alzheimer's disease and Huntington's disease, although the top-ranked loci show disease specificity. Finally, utilizing summary statistics from previous large-scale genome-wide association studies we report FDR significant enrichment of DNA methylation differences with respect to increasing LB pathology in the SNCA genomic region, a gene previously associated with Parkinson's disease and dementia with Lewy bodies.
Major depression (MD) is a leading cause of global disease burden, and both experimental and population-based studies suggest that differences in DNA methylation may be associated with the condition. However, previous DNA methylation studies have, so far, not been widely replicated, suggesting a need for larger meta-analysis studies. Here we conducted a meta-analysis of methylome-wide association analysis for lifetime MD across 18 studies of 24,754 European-ancestry participants (5,443 MD cases) and an East Asian sample (243 cases, 1,846 controls). We identified 15 CpG sites associated with lifetime MD with methylome-wide significance. The methylation score created using the methylome-wide association analysis summary statistics was significantly associated with MD status in an out-of-sample classification analysis (area under the curve 0.53). Methylation score was also associated with five inflammatory markers, with the strongest association found with tumor necrosis factor beta. Mendelian randomization analysis revealed 23 CpG sites potentially causally linked to MD, with 7 replicated in an independent dataset. Our study provides evidence that variations in DNA methylation are associated with MD, and further evidence supporting involvement of the immune system.
BACKGROUND:Children with obesity are more likely to have parents with obesity than those without. Several environmental explanations have been proposed for this correlation, including foetal programming and parenting practices. However, body mass index (BMI) is a heritable trait; child-parent correlations may reflect direct inheritance of adiposity-related genes. There is some evidence that mothers' BMI associates with offspring BMI net of direct genetic inheritance, consistent with both intrauterine and parenting effects, but this requires replication. Here, we also investigate the role of fathers' BMI as well as offsprings' diet as a mediating factor. METHODS:We used Mendelian Randomization (MR) with genetic trio (mother-father-offspring) data from 2,630 families in the Millennium Cohort Study, a UK birth cohort study of individuals born in 2000/02, to examine the association between parental BMI (kg/m2) and offspring birthweight and BMI and diet measured at six-time points between ages 3y and 17y. Paternal and maternal BMI were instrumented with polygenic indices (PGI) for BMI conditioning upon offspring PGI. This allowed us to separate direct and indirect ("genetic nurture") genetic effects. We compared these results with associations obtained using standard multivariable regression techniques using phenotypic BMI data only. RESULTS:Mothers' and fathers' BMI were positively associated with offspring BMI to similar degrees. However, in MR analysis, associations between father's BMI and offspring BMI were close to the null. In contrast, mother's BMI was consistent in MR analysis with phenotypic associations. Maternal indirect genetic effects were between 25-50% the size of direct genetic effects. There was limited and inconsistent evidence of associations with offspring diet and some evidence that mothers', but not fathers', BMI was related to birthweight in both MR and multivariable regression models. CONCLUSIONS:Results suggest maternal BMI may be particularly important for offspring BMI: associations may arise due to both direct transmission of genetic effects and indirect (genetic nurture) effects. Associations of father's and offspring adiposity that do not account for direct genetic inheritance may yield biased estimates of paternal influence. Larger studies are required to confirm these findings.
Birth cohort studies involve repeated surveys of large numbers of individuals from birth and throughout their lives. They collect information useful for a wide range of life course research domains, and biological samples which can be used to derive data from an increasing collection of omic technologies. This rich source of longitudinal data, when combined with genomic data, offers the scientific community valuable insights ranging from population genetics to applications across the social sciences. Here we present quality-controlled whole exome sequencing data from three UK birth cohorts: the Avon Longitudinal Study of Parents and Children (8,436 children and 3,215 parents), the Millenium Cohort Study (7,667 children and 6,925 parents) and Born in Bradford (8,784 children and 2,875 parents). The overall objective of this coordinated effort is to make the resulting high-quality data widely accessible to the global research community in a timely manner. We describe how the datasets were generated and subjected to quality control at the sample, variant and genotype level. We then present some preliminary analyses to illustrate the quality of the datasets and probe potential sources of bias. We introduce measures of ultra-rare variant burden to the variables available for researchers working on these cohorts, and show that the exome-wide burden of deleterious protein-truncating variants, S het burden, is associated with educational attainment and cognitive test scores. The whole exome sequence data from these birth cohorts (CRAM & VCF files) are available through the European Genome-Phenome Archive, and here we provide guidance for their use.
Birth cohort studies involve repeated surveys of large numbers of individuals from birth and throughout their lives. They collect information useful for a wide range of life course research domains, and biological samples which can be used to derive data from an increasing collection of omic technologies. This rich source of longitudinal data, when combined with genomic data, offers the scientific community valuable insights ranging from population genetics to applications across the social sciences. Here we present quality-controlled whole exome sequencing data from three UK birth cohorts: the Avon Longitudinal Study of Parents and Children (8,436 children and 3,215 parents), the Millenium Cohort Study (7,667 children and 6,925 parents) and Born in Bradford (8,784 children and 2,875 parents). The overall objective of this coordinated effort is to make the resulting high-quality data widely accessible to the global research community in a timely manner. We describe how the datasets were generated and subjected to quality control at the sample, variant and genotype level. We then present some preliminary analyses to illustrate the quality of the datasets and probe potential sources of bias. We introduce measures of ultra-rare variant burden to the variables available for researchers working on these cohorts, and show that the exome-wide burden of deleterious protein-truncating variants, S het burden, is associated with educational attainment and cognitive test scores. The whole exome sequence data from these birth cohorts (CRAM & VCF files) are available through the European Genome-Phenome Archive, and here we provide guidance for their use.
ABSTRACT INTRODUCTION Given the established association between DNA methylation and the pathophysiology of dementia and its plausible role as a molecular mediator of lifestyle and environment, blood-derived DNA methylation data could enable early detection of dementia risk. METHODS In conjunction with an extensive array of machine learning techniques, we employed whole blood genome-wide DNA methylation data as a surrogate for 14 modifiable and non-modifiable factors in the assessment of dementia risk in two independent cohorts of Alzheimer’s disease (AD) and Parkinson’s disease (PD). RESULTS We established a multivariate methylation risk score (MMRS) to identify the status of mild cognitive impairment (MCI) cross-sectionally, independent of age and sex. We further demonstrated significant predictive capability of this score for the prospective onset of cognitive decline in AD and PD. DISCUSSION Our work shows the potential of employing blood-derived DNA methylation data in the assessment of dementia risk.
BACKGROUND: Schizophrenia is associated with increased risk of developing multiple aging-related diseases, including metabolic, respiratory, and cardiovascular diseases, and Alzheimer's and related dementias, leading to the hypothesis that schizophrenia is accompanied by accelerated biological aging. This has been difficult to test because there is no widely accepted measure of biological aging. Epigenetic clocks are promising algorithms that are used to calculate biological age on the basis of information from combined cytosine-phosphate-guanine sites (CpGs) across the genome, but they have yielded inconsistent and often negative results about the association between schizophrenia and accelerated aging. Here, we tested the schizophrenia-aging hypothesis using a DNA methylation measure that is uniquely designed to predict an individual's rate of aging. METHODS: We brought together 5 case-control datasets to calculate DunedinPACE (Pace of Aging Calculated from the Epigenome), a new measure trained on longitudinal data to detect differences between people in their pace of aging over time. Data were available from 1812 psychosis cases (schizophrenia or first-episode psychosis) and 1753 controls. Mean chronological age was 38.9 (SD = 13.6) years. RESULTS: We observed consistent associations across datasets between schizophrenia and accelerated aging as measured by DunedinPACE. These associations were not attributable to tobacco smoking or clozapine medication. CONCLUSIONS: Schizophrenia is accompanied by accelerated biological aging by midlife. This may explain the wideranging risk among people with schizophrenia for developing multiple different age-related physical diseases, including metabolic, respiratory, and cardiovascular diseases, and dementia. Measures of biological aging could prove valuable for assessing patients' risk for physical and cognitive decline and for evaluating intervention effectiveness.
The cortical epigenetic clock was developed in brain tissue as a biomarker of brain aging. As one way to identify mechanisms underlying aging, we conducted a GWAS of cortical age. We leveraged postmortem cortex tissue and genotyping array data from 694 participants of the Rush Memory and Aging Project and Religious Orders Study (ROSMAP; 11000,000 SNPs), and meta-analysed ROSMAP with 522 participants of Brains for Dementia Research (5,000,000 overlapping SNPs). We confirmed results using eQTL (cortical bulk and single nucleus gene expression), cortical protein levels (ROSMAP), and phenome-wide association studies (clinical/neuropathologic phenotypes, ROSMAP). In the meta-analysis, the strongest association was rs4244620 (p = 1.29 × 10-7), which also exhibited FDR-significant cis-eQTL effects for CD46 in bulk and single nucleus (microglia, astrocyte, oligodendrocyte, neuron) cortical gene expression. Additionally, rs4244620 was nominally associated with lower cognition, faster slopes of cognitive decline, and greater Parkinsonian signs (n ~ 1700 ROSMAP with SNP/phenotypic data; all p ≤ 0.04). In ROSMAP alone, the top SNP was rs4721030 (p = 8.64 × 10-8) annotated to TMEM106B and THSD7A. Further, in ROSMAP (n = 849), TMEM106B and THSD7A protein levels in cortex were related to many phenotypes, including greater AD pathology and lower cognition (all p ≤ 0.0007). Overall, we identified converging evidence of CD46 and possibly TMEM106B/THSD7A for potential roles in cortical epigenetic clock age.
Aims Epigenetic clocks are widely applied as surrogates for biological age in different tissues and/or diseases, including several neurodegenerative diseases. Despite white matter (WM) changes often being observed in neurodegenerative diseases, no study has investigated epigenetic ageing in white matter. Methods We analysed the performances of two DNA methylation-based clocks, DNAmClock Multi and DNAmClock Cortical , in post-mortem WM tissue from multiple subcortical regions and the cerebellum, and in oligodendrocyte-enriched nuclei. We also examined epigenetic ageing in control and multiple system atrophy (MSA) (WM and mixed WM and grey matter), as MSA is a neurodegenerative disease comprising pronounced WM changes and α-synuclein aggregates in oligodendrocytes. Results Estimated DNA methylation (DNAm) ages showed strong correlations with chronological ages, even in WM (e.g., DNAmClock Cortical , r = [0.80-0.97], p<0.05). However, performances and DNAm age estimates differed between clocks and brain regions. DNAmClock Multi significantly underestimated ages in all cohorts except in the MSA prefrontal cortex mixed tissue, whereas DNAmClock Cortica tended towards age overestimations. Pronounced age overestimations in the oligodendrocyte-enriched cohorts (e.g., oligodendrocyte-enriched nuclei, p=6.1×10 -5 ) suggested that this cell-type ages faster. Indeed, significant positive correlations were observed between estimated oligodendrocyte proportions and DNAm age acceleration estimated by DNAmClock Cortica (r>0.31, p<0.05), and similar trends with DNAmClock Multi . Although increased age acceleration was observed in MSA compared to controls, no significant differences were observed upon adjustment for possible confounders (e.g., cell-type proportions). Conclusions Our findings show that oligodendrocyte proportions positively influence epigenetic age acceleration across brain regions and highlight the need to further investigate this in ageing and neurodegeneration.
The majority of epigenetic epidemiology studies to date have generated genome-wide profiles from bulk tissues (e.g., whole blood) however these are vulnerable to confounding from variation in cellular composition. Proxies for cellular composition can be mathematically derived from the bulk tissue profiles using a deconvolution algorithm; however, there is no method to assess the validity of these estimates for a dataset where the true cellular proportions are unknown. In this study, we describe, validate and characterize a sample level accuracy metric for derived cellular heterogeneity variables. The CETYGO score captures the deviation between a sample's DNA methylation profile and its expected profile given the estimated cellular proportions and cell type reference profiles. We demonstrate that the CETYGO score consistently distinguishes inaccurate and incomplete deconvolutions when applied to reconstructed whole blood profiles. By applying our novel metric to >6,300 empirical whole blood profiles, we find that estimating accurate cellular composition is influenced by both technical and biological variation. In particular, we show that when using a common reference panel for whole blood, less accurate estimates are generated for females, neonates, older individuals and smokers. Our results highlight the utility of a metric to assess the accuracy of cellular deconvolution, and describe how it can enhance studies of DNA methylation that are reliant on statistical proxies for cellular heterogeneity. To facilitate incorporating our methodology into existing pipelines, we have made it freely available as an R package (https://github.com/ds420/CETYGO).
Inflammation and ageing-related DNA methylation patterns in the blood have been linked to a variety of morbidities, including cognitive decline and neurodegenerative disease. However, it is unclear how these blood-based patterns relate to patterns within the brain, and how each associates with central cellular profiles. In this study, we profiled DNA methylation in both the blood and in five post-mortem brain regions (BA17, BA20/21, BA24, BA46 and hippocampus) in 14 individuals from the Lothian Birth Cohort 1936. Microglial burdens were additionally quantified in the same brain regions. DNA methylation signatures of five epigenetic ageing biomarkers (‘epigenetic clocks’), and two inflammatory biomarkers (DNA methylation proxies for C-reactive protein and interleukin-6) were compared across tissues and regions. Divergent correlations between the inflammation and ageing signatures in the blood and brain were identified, depending on region assessed. Four out of the five assessed epigenetic age acceleration measures were found to be highest in the hippocampus (β range=0.83-1.14, p≤0.02). The inflammation-related DNA methylation signatures showed no clear variation across brain regions. Reactive microglial burdens were found to be highest in the hippocampus (β=1.32, p=5×10 -4 ); however, the only association identified between the blood- and brain-based methylation signatures and microglia was a significant positive association with acceleration of one epigenetic clock (termed DNAm PhenoAge) averaged over all five brain regions (β=0.40, p=0.002). This work highlights a potential vulnerability of the hippocampus to epigenetic ageing and provides preliminary evidence of a relationship between DNA methylation signatures in the brain and differences in microglial burdens.
Most epigenetic epidemiology to date has utilized microarrays to identify positions in the genome where variation in DNA methylation is associated with environmental exposures or disease. However, these profile less than 3% of DNA methylation sites in the human genome, potentially missing affected loci and preventing the discovery of disrupted biological pathways. Third generation sequencing technologies, including Nanopore sequencing, have the potential to revolutionize the generation of epigenetic data, not only by providing genuine genome-wide coverage but profiling epigenetic modifications direct from native DNA. Here we assess the viability of using Nanopore sequencing for epidemiology by performing a comparison with DNA methylation quantified using the most comprehensive microarray available, the Illumina EPIC array. We implemented a CRISPR-Cas9 targeted sequencing approach in concert with Nanopore sequencing to profile DNA methylation in three genomic regions to attempt to rediscover genomic positions that existing technologies have shown are differentially methylated in tobacco smokers. Using Nanopore sequencing reads, DNA methylation was quantified at 1779 CpGs across three regions, providing a finer resolution of DNA methylation patterns compared to the EPIC array. The correlation of estimated levels of DNA methylation between platforms was high. Furthermore, we identified 12 CpGs where hypomethylation was significantly associated with smoking status, including 10 within the AHRR gene. In summary, Nanopore sequencing is a valid option for identifying genomic loci where large differences in DNAm are associated with a phenotype and has the potential to advance our understanding of the role differential methylation plays in the etiology of complex disease.
Cognitive impairment is a debilitating symptom in Parkinson’s disease (PD). We aimed to establish an accurate multivariate machine learning (ML) model to predict cognitive outcome in newly diagnosed PD cases from the Parkinson’s Progression Markers Initiative (PPMI). Annual cognitive assessments over an 8-year time span were used to define two cognitive outcomes of (i) cognitive impairment, and (ii) dementia conversion. Selected baseline variables were organized into three subsets of clinical, biofluid and genetic/epigenetic measures and tested using four different ML algorithms. Irrespective of the ML algorithm used, the models consisting of the clinical variables performed best and showed better prediction of cognitive impairment outcome over dementia conversion. We observed a marginal improvement in the prediction performance when clinical, biofluid, and epigenetic/genetic variables were all included in one model. Several cerebrospinal fluid measures and an epigenetic marker showed high predictive weighting in multiple models when included alongside clinical variables.
Parkinson's disease (PD) and dementia with Lewy bodies (DLB) are closely related progressive disorders with no available disease-modifying therapy, neuropathologically characterized by intraneuronal aggregates of misfolded α-synuclein. To explore the role of DNA methylation changes in PD and DLB pathogenesis, we performed an epigenome-wide association study (EWAS) of 322 postmortem frontal cortex samples and replicated results in an independent set of 200 donors. We report novel differentially methylated replicating loci associated with Braak Lewy body stage near TMCC2, SFMBT2, AKAP6 and PHYHIP. Differentially methylated probes were independent of known PD genetic risk alleles. Meta-analysis provided suggestive evidence for a differentially methylated locus within the chromosomal region affected by the PD-associated 22q11.2 deletion. Our findings elucidate novel disease pathways in PD and DLB and generate hypotheses for future molecular studies of Lewy body pathology.
Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disease with an estimated heritability between 40 and 50%. DNA methylation patterns can serve as proxies of (past) exposures and disease progression, as well as providing a potential mechanism that mediates genetic or environmental risk. Here, we present a blood-based epigenome-wide association study meta-analysis in 9706 samples passing stringent quality control (6763 patients, 2943 controls). We identified a total of 45 differentially methylated positions (DMPs) annotated to 42 genes, which are enriched for pathways and traits related to metabolism, cholesterol biosynthesis, and immunity. We then tested 39 DNA methylation–based proxies of putative ALS risk factors and found that high-density lipoprotein cholesterol, body mass index, white blood cell proportions, and alcohol intake were independently associated with ALS. Integration of these results with our latest genome-wide association study showed that cholesterol biosynthesis was potentially causally related to ALS. Last, DNA methylation at several DMPs and blood cell proportion estimates derived from DNA methylation data were associated with survival rate in patients, suggesting that they might represent indicators of underlying disease processes potentially amenable to therapeutic interventions.
Cognitive impairment is a common and debilitating symptom in Parkinson’s disease (PD) with high variability in individual trajectory of decline. We sought to explore heterogeneity in the trajectory of individual cognitive change in a cohort of early stage PD patients and test association to cumulative genetic risk identified in large scale Genome Wide Association Studies (GWAS). Using longitudinal measures of the Montreal Cognitive Assessment (MoCA) we employed latent class mixed modelling (LCMM) to identify and investigate unknown populations in the Parkinson’s Progression Markers Inititative (PPMI) de-novo PD cohort. Tranformed MoCA scores were modelled as a quadratic function of years from baseline, controlling for age, gender and motor symptom severity. Optimal group number was identified and determined using standardly advised model fit metrics. Polygenic risk scores (PRS) for five GWAS were calculated using PRSice-2 applied to genotyping array data. Association of PRS with cognitive groups was tested using linear models and ANOVA tests. LCMM showed optimal fit statistics for three classes (lowest BIC, high entropy) and these groups were retained for further analysis The largest identified class (n = 240) on average, presented at baseline with higher MoCA scores and remained stable over time. The second class (n = 132) presented with lower MoCA scores and showed a slow declining trajectory whilst the smallest class (n = 13) presented with lower MoCA scores but declined at a rapid rate. Educational attainment and Alzheimer’s disease (AD) GWAS derived PRS were significantly associated with cognitive class and explained the highest amount of phenotypic variance. For PD case-control status, only the PD PRS was significantly associated with Parkinson’s status and explained a similar level of phenotypic variation. Latent class analysis may provide utility in subsetting longitudinal cognitive outcome groups for use in groupwise comparisons. Using this method we show evidence for association of educational attainment and AD cumulative genetic risk and worse cognitive outcomes in early PD.