Abstract Genetic variants can cause protein-coding mutations that result in disease. Variants are typically interpreted using the reference transcript for a gene. However, most human multi-exon genes have alternative isoforms. We show that, consistent with their reduced evolutionary constraint, coding exons in alternative isoforms harbour more population variants than exons of reference isoforms, and that these variants are more likely to cause nonsynonymous mutations. Common and rare disease-associated variants mapping to alternative transcripts can lead to amino acid substitutions predicted to be structurally damaging in the corresponding protein isoform. The alternative transcripts to which disease-associated variants map demonstrate high tissue-specificity, with many unannotated in reference human genomes, and only revealed by long-read RNA-sequencing. As an example, we report an unannotated, alternative transcript of the inflammasome regulator DPP9 that is lung epithelium-specific, that harbours a common genetic variant associated with severe COVID-19 and lung fibrosis. Using deep RNA sequencing of full-length transcript isoforms by targeted capture, we confirm the expression of the unannotated DPP9 isoform. The DPP9 isoform variant causes a p.Leu8Pro missense mutation in an alternative first exon, predicted to disrupt the encoded alpha helix, and we show that the variant alters DPP9 enzymatic activity. Our findings highlight the importance of considering alternative isoforms, their tissue-specific expression, and full-length transcripts in variant interpretation, with implications for uncovering underappreciated mechanisms of both common and rare disease.
Preterm birth is closely associated with immune dysregulation in early life and subsequent learning and psychiatric disorders, but methods for stratifying infants at risk remain elusive. Protein epigenetic Scores (EpiScores) are DNA methylation (DNAm)-based proxies of circulating proteins and can capture health-related exposures such as chronic inflammation. EpiScore of C-reactive protein (DNAm CRP) is associated with inflammatory burden in early life, atypical brain development following preterm birth and adult cognitive ability. To evaluate the utility of neonatal protein EpiScores for predicting childhood cognition, we examined associations of DNAm CRP and 42 other saliva-based EpiScores enriched for inflammatory proteins correlated with low gestational age, with cognition in a cohort of 231 children, including 154 preterm children assessed at 2 years and 127 preterm and term-born children assessed at 5 years. DNAm CRP was negatively associated with 5-year Mullen Scales of Early Learning Composite (ELC) (β = -0.273, p = 0.002). Association magnitudes were larger for children born earlier (DNAm CRP x gestational age, βinteraction = 0.181). DNAm CD209 was positively associated with 5-year ELC (β = 0.267, adjusted p < 0.005). Fourteen other EpiScores were nominally associated with either 2-year Bayley-III Cognitive composite or 5-year ELC (absolute β range 0.180 to 0.245, p < 0.05). For preterm children, associations of DNAm CCL18 with 2-year cognition (β = 0.182, p = 0.039) and of DNAm CRP (β = -0.318, p = 0.021) and DNAm CRTAM (β = -0.307, p = 0.008) with 5-year cognition remained significant after adjustment for inflammatory exposures. We demonstrate associations between a range of neonatal salivary EpiScores and childhood cognition, suggesting the clinical value of EpiScores as early life markers of cognitive ability in children at risk of impairment warrants further investigation.
The Genetics Core of the Edinburgh Clinical Research Facility provides DNA methylation analysis services using optimized automated protocols. This protocol details a high-throughput DNA bisulfite conversion method that uses Zymo EZ-96 DNA Methylation-Lightning MagPrep Kits and automated using the Integra VIAFLO 96 platform. The method is used as the first step in analysing DNA samples on Illumina Infinium Methylation Screening Arrays-48. This protocol allows 192 samples to be processed together.
Clonal hematopoiesis (CH) is characterized by expanding blood cell clones carrying somatic mutations in healthy aged individuals and is associated with various age-related diseases and all-cause mortality. While CH mutations affect diverse genes associated with myeloid malignancies, their mechanisms of expansion and disease associations remain poorly understood. We investigate the relationship between clonal fitness and clinical outcomes by integrating data from three longitudinal aging cohorts (n=713, observations=2,341). We demonstrate pathway-specific fitness advantage and clonal composition influence clonal dynamics. Further, the timing of mutation acquisition is necessary to determine the extent of clonal expansion reached during the host individual’s lifetime. We introduce MACS120, a metric combining mutation context, timing, and variant fitness to predict future clonal growth, outperforming traditional variant allele frequency measurements in predicting clinical outcomes. Our unified analytical framework enables standardized clonal dynamics inference across cohorts, advancing our ability to predict and potentially intervene in CH-related pathologies.
We have generated whole-blood DNA methylation profiles from 18,869 Generation Scotland Scottish Family Health Study (GS) participants, resulting in, at the time of writing, the largest single-cohort DNA methylation resource for basic biological and medical research: Methylation in Generation Scotland (MeGS). GS is a community- and family-based cohort, which recruited over 24,000 participants from Scotland between 2006 and 2011. Comprehensive phenotype information, including detailed data on cognitive function, personality traits, and mental health, is available for all participants. The majority (83%) have genome-wide SNP genotype data (Illumina HumanOmniExpressExome-8 array v1.0 and v1.2), and over 97% of GS participants have given consent for health record linkage and re-contact. At baseline, blood-based DNA methylation was characterised at ~850,000 sites across four batches using the Illumina EPICv1 array. MeGS participants were aged between 17 and 99 years at the time of enrolment to GS. Blood-based DNA methylation EPICv1 array profiles collected at a follow-up appointment that took place 4.3-12.2 years (mean=7.1 years) after baseline are also available for 796 MeGS participants. Access to MeGS for researchers in the UK and international collaborators is via application to the GS Access Committee (access@generationscotland.org). ### Competing Interest Statement Daniel L. McCartney is a part time employee of Optima Partners Ltd. Lee Murphy has received speaker and consultancy fees from Illumina. Andrew M. McIntosh has received research support from Eli Lilly, Janssen, and the Sackler Foundation. Andrew M. McIntosh has also received speaker fees from Illumina and Janssen and consulting fees. Riccardo E. Marioni has received speaker fees from Illumina. Cathie L. Sudlow is Director of the British Heart Foundation Data Science Centre and Chief Scientist of the multifunder institute, Health Data Research UK. ### Funding Statement MeGS was primarily funded through Wellcome Trust support (reference 104036/Z/14/Z, 220857/Z/20/Z). Additional funding came from: a NARSAD Young Investigator Grant from the Brain & Behavior Research Foundation (Ref: 27404; awardee: David M Howard); a JMAS SIM fellowship from the Royal College of Physicians of Edinburgh (Awardee: Heather C Whalley); and a NARSAD Independent Investigator Award from the Brain & Behavior Research Foundation (Ref: 21956; awardee: Kathryn L Evans). The Chief Scientist Office of the Scottish Government and the Scottish Funding Council (HR03006) provided core support for Generation Scotland: Scottish Family Health Study, alongside a grant from the Scottish Government Health Department, Chief Scientist Office (Number CZD/16/6). 'NextGenScot' is funded by the Wellcome Trust (ref 216767/Z/19/Z). PN is funded by BBSRC grant BBS/E/RL/230001A and acknowledges support from the MRC Human Genetics Unit program grant, U. MC\_UU\_00007/10, and grant MC\_PC\_U127592696. For the purpose of open access, the author has applied a Creative Commons Attribution (CC BY) licence to any Author Accepted Manuscript version arising from this submission. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: GS obtained ethical approval from the NHS Tayside Committee on Medical Research Ethics, on behalf of the National Health Service (reference: 05/S1401/89) and has Research Tissue Bank Status (reference: 20/ES/0021). All components of STRADL received formal, national ethical approval from the NHS Tayside committee on research ethics (reference 14/SS/0039). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Researchers wishing to access the MeGS resource and wider Generation Scotland study data can do so by submitting an access application form to access@generationscotland.org (contact person Dr D. McCartney). Access applications are subject to review through GS access processes, which ensure that all research using the resource aims to benefit the health and wellbeing of patients and the public. Approved projects are subject to a Data & Materials Transfer Agreement (DMTA) or commercial contract. Full information on the access procedure including application forms and DMTA templates is available at https://www.ed.ac.uk/generation-scotland/for-researchers/access. Data dictionaries describing the full GS resource are available online at https://datashare.ed.ac.uk/handle/10283/2988.
In susceptible patients, COVID-19 causes life-threatening disease driven by immune-mediated inflammatory lung injury. We have previously shown that multiple common host genetic variants are significantly associated with susceptibility to critical Covid-19, and in one case, we demonstrated that such variants can inform development of new, effective drug treatment. Here we report an association analysis of whole-genome sequences (WGS) from 11,423 cases from the GenOMICC study and 60,628 controls, together with meta-analyses with available genome-wide data. We identify a rare association signal at SLC50A1, primarily driven by a missense variant rs147850817 (1:155138217:G:T, Arg201Leu) that may interfere with transport function, and we identify four common association signals near ARF1, ZNF462, KLF13 and MVP genes. Finally, we build a WGS-derived polygenic risk score (PRS) for critical Covid-19, which offers only marginal improvement in risk estimation for the general population but may provide clinically-valuable discrimination for extreme susceptibility. ### Competing Interest Statement The authors have declared no competing interest. ### Clinical Protocols ### Funding Statement GenOMICC was funded by Sepsis Research (the Fiona Elizabeth Agnew Trust), the Intensive Care Society, a Wellcome Trust Senior Research Fellowship (J.K.Baillie, 223164/Z/21/Z), the Department of Health and Social Care (DHSC), Illumina, LifeArc, the Medical Research Council, UKRI, a BBSRC Institute Strategic Program Support Grant to the Roslin Institute (BBS/E/D/20002172, BBS/E/D/10002070 and BBS/E/D/30002275) and UKRI grants MC PC 20004, MC PC 19025, MC PC 1905, and MRNO2995X/1. ADB acknowledges funding from the Wellcome PhD training fellowship for clinicians (204979/Z/16/Z), the Edinburgh Clinical Academic Track (ECAT) programme. This research is supported in part by the Data and Connectivity National Core Study, led by Health Data Research UK in partnership with the Office for National Statistics and funded by UK Research and Innovation (grant ref MC PC 20029). This study owes a great deal to the National Institute for Healthcare Research Clinical Research Network (NIHR CRN) and the Chief Scientist's Office (Scotland), who facilitate recruitment into research studies in NHS hospitals, and to the global ISARIC and InFACT consortia. This work forms part of the translational research portfolio of the National Institute for Health and Care Research Barts Biomedical Research Centre. T.M. is supported by Cancer Research UK grant DRCRPG-May23/100002 to C. Siebold. Genomics England: This research was made possible through access to data in the National Genomic Research Library, which is managed by Genomics England Limited (a wholly owned company of the Department of Health and Social Care). The National Genomic Research Library (\url{https://www.genomicsengland.co.uk/research}) holds data provided by patients and collected by the NHS as part of their care and data collected as part of their participation in research. The National Genomic Research Library is funded by the National Institute for Health Research and NHS England. The Wellcome Trust, Cancer Research UK and the Medical Research Council have also funded research infrastructure. REACT: National Institute for Health and Care Research (NIHR) and UK Research and Innovation (UKRI) - REACT-Genomics England (REACT-GE) (MR/V030841/1) and REACT-Long COVID (REACT-LC) (COV-LT-0040). The REACT study was funded by the UK Department of Health and Social Care with supplemental funding from the Huo Family Foundation. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: GenOMICC was approved by the following research ethics committees: Scotland A Research Ethics Committee (15/SS/0110) and Coventry and Warwickshire Research Ethics Committee (England, Wales and Northern Ireland) (19/WM/0247). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All other data produced in the present study are available upon reasonable request to the authors
Genetic variants can cause protein-coding mutations that result in disease. Variants are typically interpreted using the reference transcript for a gene. However, most human multi-exon genes encode alternative isoforms. Here, we show that coding exons in alternative isoforms harbour more population variants than exons of reference isoforms, consistent with their reduced evolutionary constraint, and that these variants are more likely to cause nonsynonymous coding mutations. Common and rare disease-associated variants mapping to alternative transcripts can lead to amino acid substitutions predicted to be structurally damaging in the corresponding protein isoform. The alternative transcripts to which disease-associated variants map demonstrate high tissue-specific expression, with many unannotated in reference human genomes, revealed only by long-read RNA-sequencing. As an example, we report an unannotated alternative transcript of the inflammasome regulator DPP9 that is lung epithelium-specific and which harbours a common genetic variant associated with severe COVID-19 and lung fibrosis. The variant causes a p.Leu8Pro missense mutation in an alternative first exon, predicted to disrupt the encoded alpha helix. These findings highlight the importance of considering alternative isoforms, their tissue-specific expression, and full-length transcripts in variant interpretation, with implications for uncovering underappreciated mechanisms of both common and rare disease. ### Competing Interest Statement The authors have declared no competing interest. MRC Human Genetics Unit, MC\_UU\_00035/7, MC\_UU\_00035/9 Wellcome Trust, https://ror.org/029chgv08, 223164/Z/21/Z Chief Scientist Office, https://ror.org/01613vh25, PCL/20//02 Intensive Care Society, New Investigator Award European Union, https://ror.org/019w4f821, 101001169 Medical Research Council, Doctoral Training Programmes University of Sydney, https://ror.org/0384j8v12, Doctoral Training Programme
Preterm birth is closely associated with immune dysregulation in early life and subsequent learning and psychiatric disorders, but methods for stratifying infants at risk remain elusive. Epigenetic Scores (EpiScores) are relatively stable DNA methylation (DNAm)-based proxies of circulating proteins that can capture health-related exposures such as chronic inflammation. EpiScore of C-reactive protein (DNAm CRP) is associated with inflammatory burden in early life, atypical brain development following preterm birth (encephalopathy of prematurity), and adult cognitive ability. To evaluate the utility of EpiScores for predicting cognition in children born preterm, we investigated relations between 43 neonatal saliva-based EpiScores known to associate with low gestational age, and cognition assessed at 2 and 5 years of age in a cohort of 232 preterm and term-born children. DNAm CRP was negatively associated with 5-year Mullen Scales of Early Learning Composite (ELC) (β = -0.273, p = 0.002). Association magnitudes were larger for children born earlier (DNAm CRP x gestational age, βinteraction = 0.181). EpiScores of CRTAM, NCAM1 and SLITRK5 were also associated with 5-year ELC in the full cohort (absolute β range 0.219 to 0.267, Bonferroni-adjusted p-values <0.01). For preterm children, associations for DNAm CRP (β = -0.318, p = 0.021) and DNAm CRTAM (β = -0.307, p = 0.006) with 5-year ELC remained significant after adjustment for inflammatory exposures. We demonstrate associations between a range of neonatal salivary EpiScores and childhood cognition, suggesting the clinical value of EpiScores as early life markers of cognitive ability in children at risk of impairment warrants further investigation. ### Competing Interest Statement LM has received speaker and consultancy fees from Illumina. REM is a scientific advisor to the Epigenetic Clock Development Foundation and to Optima partners. All other authors declare that they have no competing interests. ### Funding Statement RS is supported by the Translational Neuroscience PhD Programme at the University of Edinburgh, funded by Wellcome (218493/Z/19/Z). SRC is supported by a Sir Henry Dale Fellowship jointly funded by the Wellcome Trust and the Royal Society (221890/Z/20/Z). The Theirworld Edinburgh Birth Cohort (TEBC) was initiated and is maintained by Theirworld (www.theirworld.org). This work was supported through the PRENCOG study (PReterm birth as a determinant of Neurodevelopment and COGnition in children), funded by a UKRI Medical Research Council Programme Grant, MR/X003434/1. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The UK National Research Ethics Service, South East Scotland Research Ethics Committee (11/55/0061, 13/SS/0143 and 16/SS/0154) gave ethical approval for this work. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Anonymised data including EpiScores used in these analyses are deposited in Edinburgh DataVault (https://doi.org/10.7488/e65499db-2263-4d3c-9335-55ae6d49af2b) (66) and are available to researchers under the terms of the Theirworld Edinburgh Birth Cohort Data Access and Collaboration policy available at https://reproductive-health.ed.ac.uk/theirworld-edinburgh-birth-cohort-tebc/for-researchers/data-access-and-collaboration.
DNA methylation offers an objective method to assess the impact of smoking. In this work, we conduct a Bayesian EWAS of smoking pack years (n = 17,865, ~850k sites, Illumina EPIC array) and extend it by analysing whole genome data of smokers and non-smokers from Generation Scotland (n = 46, ~4-21 million sites via TWIST and Oxford Nanopore sequencing). We develop mCigarette, an epigenetic biomarker of smoking, and test it in two British cohorts. Results of brain- and blood-based EWAS (nbrain=14, nblood = 882, >450k sites, Illumina arrays) reveal several loci with near-perfect discrimination of smoking status, but which do not overlap across tissues. Furthermore, we perform a GWAS of epigenetic smoking, identifying several smoking-related loci. Overall, we improve smoking-related biomarker accuracy and enhance the understanding of the effects of smoking by integrating DNA methylation data from multiple tissues and cohorts.
Tens of millions of people worldwide have inherited chromosomally integrated human herpesvirus 6 (iciHHV-6), yet we know little about the consequences. iciHHV-6-positive individuals inherit the genome of HHV-6A or HHV-6B in the germline, and viral genomes are therefore present in every nucleated cell. To investigate the epidemiology of iciHHV-6 in the UK, almost 32,000 individuals were screened from two volunteer research studies: the family-based Generation Scotland: Scottish Family Health Study (GS:SFHS) and the Breakthrough Generations Study (BGS). iciHHV-6 prevalence in GS:SFHS was, to our knowledge, higher than that in other large studies at 2.74% (647/23,637), with an iciHHV-6B prevalence of 2.55%. Scottish participants were more likely to be iciHHV-6B-positive than English (P < 0.001), and the BGS results suggested a north-south gradient of iciHHV-6B prevalence in mainland Britain. Disease association analysis confirmed the previously reported association with angina, with an odds ratio of 1.91 (95% confidence interval, 1.29, 2.82) following adjustment for known risk factors, providing compelling evidence that iciHHV-6 contributes to the risk of a common symptom. De novo integrations were not detected within GS:SFHS pedigrees; rather, our findings indicated that three viral lineages accounted for over 95% of iciHHV-6A-positive samples, and six viral lineages accounted for 90% of iciHHV-6B-positive samples in GS:SFHS. This study demonstrates that iciHHV-6 is common in the UK, shows significant regional heterogeneity in prevalence, is not entirely harmless, and is largely derived from a relatively small number of ancestral viral lineages.IMPORTANCEHuman herpesvirus 6 (HHV-6) has the unusual ability to integrate into the host chromosome telomeres. Most of the world's population is infected by HHV-6 in early childhood, but around 1% inherit the virus as a chromosomally integrated viral genome-referred to as inherited chromosomally integrated HHV-6 (iciHHV-6). Little is known about the consequences of iciHHV-6, which has the potential to cause disease through various mechanisms. Here, we have used large cohorts to study iciHHV-6 prevalence, lineages, and phenotypic associations. We replicate a previously reported association between iciHHV-6 and angina, suggesting that iciHHV-6 is not entirely benign. We show significant variation in iciHHV-6 prevalence within the UK with almost 3% of Scottish people carrying iciHHV-6. In the first detailed analysis of viral lineages at the population level, we show that 90% of iciHHV-6 is explained by nine ancestral viral lineages. These results have important implications for future disease association analyses.
BACKGROUND: Cardiovascular disease (CVD) is among the leading causes of death worldwide. The discovery of new omics biomarkers could help to improve risk stratification algorithms and expand our understanding of molecular pathways contributing to the disease. Here, ASSIGN—a cardiovascular risk prediction tool recommended for use in Scotland—was examined in tandem with epigenetic and proteomic features in risk prediction models in ≥12 657 participants from the Generation Scotland cohort. METHODS: Previously generated DNA methylation–derived epigenetic scores (EpiScores) for 109 protein levels were considered, in addition to both measured levels and an EpiScore for cTnI (cardiac troponin I). The associations between individual protein EpiScores and the CVD risk were examined using Cox regression (n cases ≥1274; n controls ≥11 383) and visualized in a tailored R application. Splitting the cohort into independent training (n=6880) and test (n=3659) subsets, a composite CVD EpiScore was then developed. RESULTS: Sixty-five protein EpiScores were associated with incident CVD independently of ASSIGN and the measured concentration of cTnI ( P <0.05), over a follow-up of up to 16 years of electronic health record linkage. The most significant EpiScores were for proteins involved in metabolic, immune response, and tissue development/regeneration pathways. A composite CVD EpiScore (based on 45 protein EpiScores) was a significant predictor of CVD risk independent of ASSIGN and the concentration of cTnI (hazard ratio, 1.32; P =3.7×10 − 3 ; 0.3% increase in C-statistic). CONCLUSIONS: EpiScores for circulating protein levels are associated with CVD risk independent of traditional risk factors and may increase our understanding of the etiology of the disease.
Abstract Background Epigenetic scores (EpiScores), reflecting DNA methylation (DNAm)-based surrogates for complex traits, have been developed for multiple circulating proteins. EpiScores for pro-inflammatory proteins, such as C-reactive protein (DNAm CRP), are associated with brain health and cognition in adults and with inflammatory comorbidities of preterm birth in neonates. Social disadvantage can become embedded in child development through inflammation, and deprivation is overrepresented in preterm infants. We tested the hypotheses that preterm birth and socioeconomic status (SES) are associated with alterations in a set of EpiScores enriched for inflammation-associated proteins. Results In total, 104 protein EpiScores were derived from saliva samples of 332 neonates born at gestational age (GA) 22.14 to 42.14 weeks. Saliva sampling was between 36.57 and 47.14 weeks. Forty-three (41%) EpiScores were associated with low GA at birth (standardised estimates |0.14 to 0.88|, Bonferroni-adjusted p-value < 8.3 × 10−3). These included EpiScores for chemokines, growth factors, proteins involved in neurogenesis and vascular development, cell membrane proteins and receptors, and other immune proteins. Three EpiScores were associated with SES, or the interaction between birth GA and SES: afamin, intercellular adhesion molecule 5, and hepatocyte growth factor-like protein (standardised estimates |0.06 to 0.13|, Bonferroni-adjusted p-value < 8.3 × 10−3). In a preterm subgroup (n = 217, median [range] GA 29.29 weeks [22.14 to 33.0 weeks]), SES–EpiScore associations did not remain statistically significant after adjustment for sepsis, bronchopulmonary dysplasia, necrotising enterocolitis, and histological chorioamnionitis. Conclusions Low birth GA is substantially associated with a set of EpiScores. The set was enriched for inflammatory proteins, providing new insights into immune dysregulation in preterm infants. SES had fewer associations with EpiScores; these tended to have small effect sizes and were not statistically significant after adjusting for inflammatory comorbidities. This suggests that inflammation is unlikely to be the primary axis through which SES becomes embedded in the development of preterm infants in the neonatal period. Graphical abstract
The FLG gene encodes the filaggrin protein which is essential for epidermal barrier formation and hydration. Mutations in the filaggrin gene are associated with a broad range of skin and allergic diseases. This protocol is used to genotype four mutations (R501X, 2282del4, R2447X and S3247X) within the FLG gene.
Introduction Preterm birth (PTB) is strongly associated with encephalopathy of prematurity (EoP) and neurocognitive impairment. The biological axes linking PTB with atypical brain development are uncertain. We aim to elucidate the roles of neuroendocrine stress activation and immune dysregulation in linking PTB with EoP. Methods and analysis PRENCOG (PREterm birth as a determinant of Neurodevelopment and COGnition in children: mechanisms and causal evidence) is an exposure-based cohort study at the University of Edinburgh. Three hundred mother–infant dyads comprising 200 preterm births (gestational age, GA <32 weeks, exposed) and 100 term births (GA >37 weeks, non-exposed), will be recruited between January 2023 and December 2027. We will collect parental and infant medical, demographic, socioeconomic characteristics and biological data which include placental tissue, umbilical cord blood, maternal and infant hair, infant saliva, infant dried blood spots, faecal material, and structural and diffusion MRI. Infant biosamples will be collected between birth and 44 weeks GA. EoP will be characterised by MRI using morphometric similarity networks (MSNs), hierarchical complexity (HC) and magnetisation transfer saturation imaging (MTsat). We will conduct: first, multivariable regressions and statistical association assessments to test how PTB-associated risk factors (PTB-RFs) relate to MSNs, HC and or MTsat; second, structural equation modelling to investigate neuroendocrine stress activation and immune dysregulation as mediators of PTB-RFs on features of EoP. PTB-RF selection will be informed by the variables that predict real-world educational outcomes, ascertained by linking the UK National Neonatal Research Database with the National Pupil Database. Ethics and dissemination A favourable ethical opinion has been given by the South East Scotland Research Ethics Committee 02 (23/SS/0067) and NHS Lothian Research and Development (2023/0150). Results will be reported to the Medical Research Council, in scientific media, via stakeholder partners and on a website in accessible language ( https://www.ed.ac.uk/centre-reproductive-health/prencog ).
The epigenome, including the methylation of cytosine bases at CG dinucleotides, is intrinsically linked to transcriptional regulation. The tight regulation of gene expression during skeletal development is essential, with ~1/500 individuals born with skeletal abnormalities. Furthermore, increasing evidence is emerging to link age-associated complex genetic musculoskeletal diseases, including osteoarthritis (OA), to developmental factors including joint shape. Multiple studies have shown a functional role for DNA methylation in the genetic mechanisms of OA risk using articular cartilage samples taken from aged patients. Despite this, our knowledge of temporal changes to the methylome during human cartilage development has been limited. We quantified DNA methylation at ~700,000 individual CpGs across the epigenome of developing human articular cartilage in 72 samples ranging from 7-21 post-conception weeks, a time period that includes cavitation of the developing knee joint. We identified significant changes in 8% of all CpGs, and >9400 developmental differentially methylated regions (dDMRs). The largest hypermethylated dDMRs mapped to transcriptional regulators of early skeletal patterning including MEIS1 and IRX1. Conversely, the largest hypomethylated dDMRs mapped to genes encoding extracellular matrix proteins including SPON2 and TNXB and were enriched in chondrocyte enhancers. Significant correlations were identified between the expression of these genes and methylation within the hypomethylated dDMRs. We further identified 811 CpGs at which significant dimorphism was present between the male and female samples, with the majority (68%) being hypermethylated in female samples. Following imputation, we captured the genotype of these samples at >5 million variants and performed epigenome-wide methylation quantitative trait locus (mQTL) analysis. Colocalization analysis identified 26 loci at which genetic variants exhibited shared impacts upon methylation and OA genetic risk. This included loci which have been previously reported to harbour OA-mQTLs (including GDF5 and ALDH1A2), yet the majority (73%) were novel (including those mapping to CHST3, FGF1 and TEAD1). To our knowledge, this is the first extensive study of DNA methylation across human articular cartilage development. We identify considerable methylomic plasticity within the development of knee cartilage and report active epigenomic mediators of OA risk operating in prenatal joint tissues.
Increasing evidence is emerging to link age-associated complex musculoskeletal diseases, including osteoarthritis (OA), to developmental factors. Multiple studies have shown a functional role for DNA methylation in the genetic mechanisms of OA risk using articular cartilage samples taken from aged individuals, yet knowledge of temporal changes to the methylome during human cartilage development is limited. We quantified DNA methylation at ∼700,000 individual CpGs across the epigenome of developing human chondrocytes in 72 samples ranging from 7 to 21 post-conception weeks. We identified significant changes in 3% of all CpGs and >8,200 developmental differentially methylated regions. We further identified 24 loci at which OA genetic variants colocalize with methylation quantitative trait loci. Through integrating developmental and mature human chondrocyte datasets, we find evidence for functional effects exerted solely in development or throughout the life course. This will have profound impacts on future approaches to translating genetic pathways for therapeutic intervention.
PREVENT is a multi-centre prospective cohort study in the UK and Ireland that aims to examine midlife risk factors for dementia and identify and describe the earliest indices of disease development. The PREVENT dementia programme is one of the original epidemiological initiatives targeting midlife as a critical window for intervention in neurodegenerative conditions. This paper provides an overview of the study protocol and presents the first summary results from the initial baseline data to describe the cohort. Participants in the PREVENT cohort provide demographic data, biological samples (blood, saliva, urine and optional cerebrospinal fluid), lifestyle and psychological questionnaires, undergo a comprehensive cognitive test battery and are imaged using multi-modal 3-T MRI scanning, with both structural and functional sequences. The PREVENT cohort governance structure is described, which includes a steering committee, a scientific advisory board and core patient and public involvement groups. A number of sub-studies that supplement the main PREVENT cohort are also described. The PREVENT cohort baseline data include 700 participants recruited between 2014 and 2020 across five sites in the UK and Ireland (Cambridge, Dublin, Edinburgh, London and Oxford). At baseline, participants had a mean age of 51.2 years (range 40-59, SD +/- 5.47), with the majority female (n = 433, 61.9%). There was a near equal distribution of participants with and without a parental history of dementia (51.4% versus 48.6%) and a relatively high prevalence of APOE epsilon 4 carriers (n = 264, 38.0%). Participants were highly educated (16.7 +/- 3.44 years of education), were mainly of European Ancestry (n = 672, 95.9%) and were cognitively healthy as measured by the Addenbrookes Cognitive Examination-III (total score 95.6 +/- 4.06). Mean white matter hyperintensity volume at recruitment was 2.26 +/- 2.77 ml (median = 1.39 ml), with hippocampal volume being 8.15 +/- 0.79 ml. There was good representation of known dementia risk factors in the cohort. The PREVENT cohort offers a novel data set to explore midlife risk factors and early signs of neurodegenerative disease. Data are available open access at no cost via the Alzheimer's Disease Data Initiative platform and Dementia Platforms UK platform pending approval of the data access request from the PREVENT steering group committee. PREVENT is a multi-centre prospective cohort study with 700 participants aged 40-59 from the UK and Ireland that aims to examine midlife risk factors for dementia. This paper provides an overview of the study protocol and presents the first summary results from the initial baseline data to describe the cohort. Graphical Abstract
Purpose (the aim of the study): The articulating human knee joint develops from a homogenous condensation of mesenchymal stem cells. This process requires the orchestrated expression of a multitude of genes. The transcriptome is primarily regulated through the spatiotemporal expression of transcription factors, yet epigenetic processes, including DNA methylation (DNAm) both underlie and reinforce developmental plasticity.
The prevalence of clonal haematopoiesis of indeterminate potential (CHIP) in healthy individuals increases rapidly from age 60 onwards and has been associated with increased risk for malignancy, heart disease and ischemic stroke. CHIP is driven by somatic mutations in stem cells that are also drivers of myeloid malignancies. We previously set out to quantify the fitness effects of CHIP drivers over a 12-year timespan in older age, using longitudinal error-corrected sequencing data. We now extended our longitudinal data to include participants from two other longitudinal cohorts of aging. We show that fitness is a better predictor of all-cause mortality compared to clone size. Moreover, we generated a tool to allow for deconvolution of clonal complexities in the blood of aged individual, highlighting clonal composition and co-occurrence of mutations. We identified commonly co-occurring mutations with older age and will discuss consequences of those co-occurring mutations. In summary, longitudinal cohorts allow identification of clonal evolution in individuals with advanced age.
Background Epigenetic scores (EpiScores) can provide biomarkers of lifestyle and disease risk. Projecting new datasets onto a reference panel is challenging due to separation of technical and biological variation with array data. Normalisation can standardise data distributions but may also remove population-level biological variation. Results We compare two birth cohorts (Lothian Birth Cohorts of 1921 and 1936 — n LBC1921 = 387 and n LBC1936 = 498) with blood-based DNA methylation assessed at the same chronological age (79 years) and processed in the same lab but in different years and experimental batches. We examine the effect of 16 normalisation methods on a novel BMI EpiScore (trained in an external cohort, n = 18,413), and Horvath’s pan-tissue DNA methylation age, when the cohorts are normalised separately and together. The BMI EpiScore explains a maximum variance of R 2 =24.5% in BMI in LBC1936 (SWAN normalisation). Although there are cross-cohort R 2 differences, the normalisation method makes a minimal difference to within-cohort estimates. Conversely, a range of absolute differences are seen for individual-level EpiScore estimates for BMI and age when cohorts are normalised separately versus together. While within-array methods result in identical EpiScores whether a cohort is normalised on its own or together with the second dataset, a range of differences is observed for between-array methods. Conclusions Normalisation methods returning similar EpiScores, whether cohorts are analysed separately or together, will minimise technical variation when projecting new data onto a reference panel. These methods are important for cases where raw data is unavailable and joint normalisation of cohorts is computationally expensive.