BACKGROUND: Polygenic scores (PGSs) are weighted sum scores of trait-associated alleles from up to millions of SNPs. As PGS research pivots to translation into health care settings a key issue for laboratories providing PGS is demonstration of analytical validity of PGS. METHODS: We report data from 6 individuals who have been genotyped multiple times using the same and different technologies. These data were generated as part of standard experimental design for quality control purposes in two research settings over many studies and over many years. Using this opportunistic design of technical variability, we provide an empirical evaluation of technical reproducibility of PGS from 115 traits of different genetic architectures. RESULTS: Given a predefined set of SNP weights variability in PGS can reflect only SNP missingness or incorrect genotype call. We find very high reproducibility of SNP genotypes. In particular, the technical reproducibility of PGS generated from the same array technology and processed through the same quality control and imputation pipeline is very high. However, impact of missing SNPs varies between traits depending on the SNP’s weight for a trait. We provide a PGS quality score statistic (PGS:QS) that can be reported for each trait-specific score for an individual, to provide a quantitative assessment of the proportion of variation of the score that is captured by the SNPs genotyped/imputed for the individual. We provide an algorithm (PGS-impute) that updates the SNP weights of the scoring algorithm to the SNPs available for an individual, improving PGS accuracy. CONCLUSIONS: While validity of directly measured genotypes (whether from microarray or whole genome sequencing) is well-established, objective approaches to evaluate analytical reproducibility of PGS post-genotyping pipeline have been lacking. Here, we provide empirical data and an analysis framework which can be used by PGS providers to support understanding of analytical reproducibility and robustness.
Background: Endometriosis is a complex condition with substantial diagnostic delays and health burden. This study aimed to develop and evaluate a polygenic risk score (PRS) for endometriosis and compare its predictive utility with, and alongside, family history (FH). Methods: Genetic effect estimates were derived from a large genome-wide association study meta-analysis and applied to an independent Australian cohort (n≈3,400) with linked survey and health data. PRS was calculated using clumping/thresholding and Bayesian (SBayesRC) approaches. Predictive performance was assessed using logistic regression and area under the receiver operating characteristic curve (AUC). Findings: PRS was associated with endometriosis risk, with the best-performing model achieving modest discrimination (AUC≈0.61). Individuals in the highest PRS decile had increased risk compared with the lowest decile. FH was a stronger predictor, with 3–5-fold higher odds of disease among those with an affected relative. PRS and FH were weakly correlated and independently associated with endometriosis. Combining them improved prediction (AUC up to ~0.72). Interpretation: PRS and FH capture complementary aspects of endometriosis risk. While FH remains a robust and clinically accessible predictor, PRS provides additional genetic information that enhances risk stratification. Integrating these measures may support improved risk assessment and earlier diagnosis.
Objectives: Sleep and circadian rhythm disturbances (SCRDs) are proposed to be pathophysiological mechanisms underlying some cases of depressive and bipolar (mood) disorders. An unresolved clinical question is whether sleep and circadian based therapies are effective antidepressants for young adults (18-30-years) with a mood disorder. Method & analysis: MELODY (Melatonin for Depression in Youth) is an investigator-initiated, single-centre, randomised, placebo-controlled, phase 3 clinical trial. The trial is testing whether 12 weeks of adjunctive melatonin or digital cognitive behavioural therapy for insomnia (dCBT-I) are more effective than pill placebo at reducing depressive symptoms in 660 young people aged 18-30-years with a Structured Clinical Interview for DSM-5 (SCID-5) diagnosis of Major Depressive Disorder or Bipolar Disorder type II, moderate-to-severe depressive symptoms, and significant sleep or sleep-wake complaints. The week 12 primary outcome is depressive symptoms (Quick Inventory of Depressive Symptomatology, Adolescent version). Secondary outcomes include partial remission of the Major Depressive Episode (SCID-5) and change in other mental health symptoms (e.g., anxiety), sleep (e.g., insomnia); a subset of participants will undergo in-home assessment of sleep physiology (Sleep ProfilerTM) and in-lab assessment of circadian rhythms (dim-light melatonin onset) to examine biological change. A 6-month follow-up will explore durability of clinical or sleep effects. Mediation analyses will test whether sleep or circadian rhythm changes play a causal role in antidepressant effects. Ethics & dissemination: MELODY has been reviewed and approved by the Human Research Ethics Committee (HREC) of the Sydney Local Health District (HREC Approval Number: X23-0450, Protocol version: 1.4, 7/7/25). The findings of the MELODY trial will be disseminated into the scientific and clinical communities via refereed publications, talks, and other professional and media outlets. The Brain and Mind Centres Lived Experience Working Group (LEWG) will contribute to dissemination of MELODYs findings via youth-friendly methods (e.g., social media videos, explainers). Trial registration number: ACTRN12624000017527 ### Competing Interest Statement The views and opinions expressed in this article are those of the authors and should not be construed to represent the views of any of the sponsoring organisations, agencies, or US government. EMS is a Principal Research Fellow at the Brain and Mind Centre, The University of Sydney. She is Discipline Leader of Adult Mental Health, School of Medicine, University of Notre Dame, and a Consultant Psychiatrist. She was the Medical Director, Young Adult Mental Health Unit, St Vincents Hospital Darlinghurst until January 2021. She has received honoraria for educational seminars related to the clinical management of depressive disorders supported by Servier, Janssen and Eli-Lilly Pharmaceuticals. She has participated in a national advisory board for the antidepressant compound Pristiq, manufactured by Pfizer. She was the national coordinator of an antidepressant trial sponsored by Servier. IBH is the Co-Director, Health and Policy at the Brain and Mind Centre (BMC) University of Sydney, Australia. The BMC operates an early-intervention youth services at Camperdown under contract to headspace. He has previously led community-based and pharmaceutical industry-supported (Wyeth, Eli Lily, Servier, Pfizer, AstraZeneca) projects focused on the identification and better management of anxiety and depression and investigator-initiated studies of agomelatine. He is the Chief Scientific Advisor to, and a 3.2% equity shareholder in, InnoWell Pty Ltd. InnoWell was formed by the University of Sydney (45% equity) and PwC (Australia; 45% equity) to lead transformation of mental health services internationally through the use of innovative technologies. ### Clinical Trial Australian New Zealand Clinical Trials Registry: ACTRN12624000017527 ### Funding Statement This work was supported by the Wellcome Trust [227089/Z/23/Z]. The funder has no role in collection, management, analysis, and interpretation of data; writing of the report; and the decision to submit the report for publication. The funder mandated the use of some assessments as part of their common metrics scheme. Barring this mandate, the funder has no ultimate authority in any of the other activities. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: MELODY has been reviewed and approved by the Human Research Ethics Committee (HREC) of the Sydney Local Health District (HREC Approval Number: X23-0450, Protocol version: 1.4, 7/7/25). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The datasets generated and/or analysed during this study will be accessible upon request to Prof Ian Hickie (ian.hickie{at}sydney.edu.au). The data provided will be fully anonymised, with all participant-identifiable information removed. Access to the full dataset will be granted only after the formal reporting of study findings in a peer-reviewed scientific journal. Datasets will be available exclusively to bona fide scientific researchers. Requests must be submitted in writing to the Principal Investigator, including details about the investigators background and the intended purpose of the data. Each request will be evaluated based on the proposed analyses, with potential uses likely to include meta-analyses, for example. The studys Participant Information Sheet and consent forms explicitly mentioned the availability of anonymised data, a process approved by the Sydney Local Health District HREC.
Amyotrophic lateral sclerosis (ALS) is a neurogenerative disease resulting from progressive degeneration of motor neurons leading to systemic consequences. Despite being the most common motor neuron disease, with increasing global prevalence, limited treatment options exist. Emerging evidence from genetic studies and pathology analyses implicates RNA dysregulation in ALS pathogenesis, however, deep, comprehensive RNA sequencing studies have not been carried out. Here, we analysed >240 ALS and control whole blood transcriptome samples. Cross-sectional (Ncases=121, Ncontrols=53) and longitudinal (Nobservations=103) cohorts supported complementary expression analyses of disease mechanisms across disease stages. Both short (N=241) and long-read (N=16) technologies were utilised to discover splicing changes. Total RNA was extracted from PAXgene whole blood RNA tubes before libraries (Illumina Stranded Total RNA RiboZero Plus) were prepared and sequenced (~50M PE reads per sample). Long-read sequencing was performed using the Mas-Seq protocol with Kinnex full-length RNA prep kit and sequenced (PacBio Revio platform, 10M reads per sample) for full-length transcripts. Case-control cohort analyses identified 50 significantly differentially expressed genes, with pathway analyses implicating RNA processing and immune system regulation. Findings were corroborated using existing ALS RNAseq datasets from blood (correlation >0.4), iPSC-MN and post-mortem tissues. Alternative splicing (AS) analyses (LeafCutter) identified 62 clusters. Within-case analyses involved ALS cases with multiple (2-4) visits, detected 144 genes associated with disability progression over time. The long-read sequencing (Ncases=8, Ncontrols=8) provided novel discovery insights, in particular in the HLA region. This comprehensive blood-based transcriptomic dataset reveals both known and novel disease mechanisms in ALS, offering valuable insights that could inform future research and therapeutic development. The results of this study may inform and refine the prioritization of candidate genes and loci in future ALS research. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement We gratefully acknowledge all participants and their families for generously contributing to this research. We also thank the clinical teams and staff involved in supporting sample and data collection who made this project possible. The data generated in this project was funded by an IMPACT grant from FightMND (2022, to FCG). Additional project funding and support were provided by a Daniel McLoone MND Research Grant from Motor Neurone Disease Australia (IG2312, to FCG), (Australian) National Health and Medical Research Council (grants 1078901, 1113400, 1087889) to NRW and the MND Australia Ice Bucket Challenge Grant (to NRW). FCG was funded by a Scott Sullivan Fellowship (MND and Me/MNDRA). We thank Kelly Williams and Natalie Grima for their helpful responses re: GSE234297. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Metro North Health Human Research Ethics Committee EC00172 of the Royal Brisbane and Women's hospital gave ethical approval (2006/047) for this work. The University of Queensland Human Research Ethics Committee gave ethical approval for this work (2021/HE002682). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The RNA-seq blood data generated here is available from the author upon reasonable request; an online submission for upload is in progress.
Cell-free DNA (cfDNA), derived from dying cells, has demonstrated utility across multiple clinical applications. However, its potential in neurodegenerative diseases remains underexplored, with most existing cfDNA technologies tailored to specific disease contexts like cancer or non-invasive prenatal screening. To address this gap, we developed a novel approach to characterize epigenetic cfDNA profiles by identifying key regions of DNA methylation that reveal the tissues origins undergoing apoptosis or necrosis. We evaluated this method in the largest cfDNA study of amyotrophic lateral sclerosis (ALS) and other neurological diseases (OND) to date, encompassing two independent cohorts (n = 192) from Australia (UQ Ncases = 48, Ncontrols = 32, NOND = 15) and the USA, (UCSF Ncases = 50, Ncontrols = 45)). Our approach accurately distinguished ALS patients from controls (UQ AUC = 0.82, UCSF AUC = 0.99) and from individuals with other neurological diseases (AUC = 0.91). It also identified an asymptomatic carrier of a pathogenic C9orf72 variant, and strongly correlated with ALS disease progression measures (Pearson’s R = 0.66, p = 3.71 × 10⁻⁹). We identified DNA methylation signals from multiple tissue types in ALS cfDNA, highlighting diverse tissue involvement in ALS pathology. These findings promote epigenetic cfDNA analysis as a powerful tool for advancing our understanding of neurodegenerative disease.
Purpose Children with neurodevelopmental disorders (NDDs) such as autism spectrum disorder (ASD) and attention-deficit hyperactivity disorder (ADHD) face a range of challenges which impact their daily functioning and that of their family. NDDs are often associated with significant mental health problems which can influence the course. The Improving Outcomes in Mental Health cohort described in this article aims to investigate the risk factors for the persistence and severity of mental health problems in children with NDDs. Participants A total of 1084 families (primary caregivers and children) were recruited from the Child Development Program at the Children’s Health Queensland Hospital and Health Service in Brisbane, Australia. 1471 caregivers (female n=1036) participated in the study, which included 382 families with 2 or more caregivers participating. The children were predominantly male (71%), with the average age of all children 5.6 years. Findings to date The most prevalent child clinical diagnoses were ASD and ADHD, with half of children receiving more than one diagnosis. Caregiver reports indicated that children were experiencing clinical levels of depression (30.8%) and anxiety (27.6%). Approximately 39% of caregivers scored in the subclinical or clinical range for at least one Diagnostic and Statistical Manual of Mental Disorders measure, the majority reporting depressive problems. Future plans Future plans for this data set include analysis of environmental variables such as family structure, income, school achievements and leisure activities as risk factors for the persistence of mental health problems in children with NDDs. Genetic data will be used to provide insights into the heritability of mental illness and improve prediction.
Globally, over 65 million individuals are estimated to suffer from post-acute sequelae of COVID-19 (PASC). A large number of individuals living with PASC experience cardiovascular symptoms (i.e. chest pain and heart palpitations) (PASC-CVS). The role of chronic inflammation in these symptoms, in particular in individuals with symptoms persisting for >1 year after SARS-CoV-2 infection, remains to be clearly defined. In this cross-sectional study, blood samples were obtained from three different sites in Australia from individuals with i) a resolved SARS-CoV-2 infection (and no persistent symptoms i.e. ‘Recovered’), ii) individuals with prolonged PASC-CVS and iii) SARS-CoV-2 negative individuals. Individuals with PASC-CVS, relative to Recovered individuals, had a blood transcriptomic signature associated with inflammation. This was accompanied by elevated levels of pro-inflammatory cytokines (IL-12, IL-1β, MCP-1 and IL-6) at approximately 18 months post-infection. These cytokines were present in trace amounts, such that they could only be detected with the use of novel nanotechnology. Importantly, these trace-level cytokines had a direct effect on the functionality of pluripotent stem cell derived cardiomyocytes in vitro . This effect was not observed in the presence of dexamethasone. Plasma proteomics demonstrated further differences between PASC-CVS and Recovered patients at approximately 18 months post-infection including enrichment of complement and coagulation associated proteins in those with prolonged cardiovascular symptoms. Together, these data provide a new insight into the role of chronic inflammation in PASC-CVS and present nanotechnology as a possible novel diagnostic approach for the condition. ### Competing Interest Statement K.R.S. is a consultant for Sanofi, Pfizer, Roche and NovoNordisk. The opinions and data presented in this manuscript are of the authors and are independent of these relationships.
Cell-free DNA (cfDNA) is increasingly recognized as a promising biomarker candidate for disease monitoring. However, its utility in neurodegenerative diseases, like amyotrophic lateral sclerosis (ALS), remains underexplored. Existing biomarker discovery approaches are tailored to a specific disease context or are too expensive to be clinically practical. Here, we address these challenges through a new approach combining advances in molecular and computational technologies. First, we develop statistical tools to select tissue-informative DNA methylation sites relevant to a disease process of interest. We then employ a capture protocol to select these sites and perform targeted methylation sequencing. Multi-modal information about the DNA methylation patterns are then utilized in machine learning algorithms trained to predict disease status and disease progression. We applied our method to two independent cohorts of ALS patients and controls (n=192). Overall, we found that the targeted sites accurately predicted ALS status and replicated between cohorts. Additionally, we identified epigenetic features associated with ALS phenotypes, including disease severity. These findings highlight the potential of cfDNA as a non-invasive biomarker for ALS.
An estimated 65 million people globally suffer from post-acute sequelae of COVID-19 (PASC), with many experiencing cardiovascular symptoms (PASC-CVS) like chest pain and heart palpitations. This study examines the role of chronic inflammation in PASC-CVS, particularly in individuals with symptoms persisting over a year after infection. Blood samples from three groups—recovered individuals, those with prolonged PASC-CVS and SARS-CoV-2-negative individuals—revealed that those with PASC-CVS had a blood signature linked to inflammation. Trace-level pro-inflammatory cytokines were detected in the plasma from donors with PASC-CVS 18 months post infection using nanotechnology. Importantly, these trace-level cytokines affected the function of primary human cardiomyocytes. Plasma proteomics also demonstrated higher levels of complement and coagulation proteins in the plasma from patients with PASC-CVS. This study highlights chronic inflammation’s role in the symptoms of PASC-CVS. Sinclair et al. explore the contribution of chronic inflammation to cardiovascular symptoms associated with post-acute sequelae of SARS-CoV-2 infection (PASC-CVS). The authors identify trace levels of inflammatory cytokines in individuals with PASC-CVS that impair the function of cardiomyocytes derived from induced pluripotent stem cells.
PURPOSE:Children with neurodevelopmental disorders (NDDs) such as autism spectrum disorder (ASD) and attention-deficit hyperactivity disorder (ADHD) face a range of challenges which impact their daily functioning and that of their family. NDDs are often associated with significant mental health problems which can influence the course. The Improving Outcomes in Mental Health cohort described in this article aims to investigate the risk factors for the persistence and severity of mental health problems in children with NDDs. PARTICIPANTS:A total of 1084 families (primary caregivers and children) were recruited from the Child Development Program at the Children's Health Queensland Hospital and Health Service in Brisbane, Australia. 1471 caregivers (female n=1036) participated in the study, which included 382 families with 2 or more caregivers participating. The children were predominantly male (71%), with the average age of all children 5.6 years. FINDINGS TO DATE:The most prevalent child clinical diagnoses were ASD and ADHD, with half of children receiving more than one diagnosis. Caregiver reports indicated that children were experiencing clinical levels of depression (30.8%) and anxiety (27.6%). Approximately 39% of caregivers scored in the subclinical or clinical range for at least one Diagnostic and Statistical Manual of Mental Disorders measure, the majority reporting depressive problems. FUTURE PLANS:Future plans for this data set include analysis of environmental variables such as family structure, income, school achievements and leisure activities as risk factors for the persistence of mental health problems in children with NDDs. Genetic data will be used to provide insights into the heritability of mental illness and improve prediction.
Autism omics research has historically been reductionist and diagnosis centric, with little attention paid to common co-occurring conditions (for example, sleep and feeding disorders) and the complex interplay between molecular profiles and neurodevelopment, genetics, environmental factors and health. Here we explored the plasma lipidome (783 lipid species) in 765 children (485 diagnosed with autism spectrum disorder (ASD)) within the Australian Autism Biobank. We identified lipids associated with ASD diagnosis ( n = 8), sleep disturbances ( n = 20) and cognitive function ( n = 8) and found that long-chain polyunsaturated fatty acids may causally contribute to sleep disturbances mediated by the FADS gene cluster. We explored the interplay of environmental factors with neurodevelopment and the lipidome, finding that sleep disturbances and unhealthy diet have a convergent lipidome profile (with potential mediation by the microbiome) that is also independently associated with poorer adaptive function. In contrast, ASD lipidome differences were accounted for by dietary differences and sleep disturbances. We identified a large chr19p13.2 copy number variant genetic deletion spanning the LDLR gene and two high-confidence ASD genes ( ELAVL3 and SMARCA4 ) in one child with an ASD diagnosis and widespread low-density lipoprotein-related lipidome derangements. Lipidomics captures the complexity of neurodevelopment, as well as the biological effects of conditions that commonly affect quality of life among autistic people.
Background Amyotrophic lateral sclerosis (ALS), the most predominant form of Motor Neuron Disease (MND), is a progressive and fatal neurodegenerative condition that spreads throughout the neuromotor system by afflicting upper and lower motor neurons. Lower motor neurons project from the central nervous system and innervate muscle fibres at motor endplates, which degrade over the course of the disease leading to muscle weakness. The direction of neurodegeration from or to the point of neuromuscular junctions and the role of muscle itself in pathogenesis has continued to be a topic of debate in ALS research. Methods To assess the variation in gene expression between affected and nonaffected muscle tissue that might lead to this local degeneration of motor units, we generated RNA-seq skeletal muscle transcriptomes from 28 MND cases and 18 healthy controls and conducted differential expression analyses on gene-level counts, as well as an isoform switching analysis on isoform-level counts. Results We identified 52 differentially-expressed genes (Benjamini-Hochberg-adjusted p < 0.05) within this comparison, including 38 protein coding, 9 long non-coding RNA, and 5 pseudogenes. Of protein-coding genes, 31 were upregulated in cases including with notable genes including the collagenic COL25A1 ( p = 3.1 × 10−10), SAA1 which is released in response to tissue injury ( p = 3.6 × 10−5) as well as others of the SAA family, and the actin-encoding ACTC1 ( p = 2.3 × 10−5). Additionally, we identified 17 genes which exhibited a functional isoform switch with likely functional consequences between cases and controls. Conclusions Our analyses provide evidence of increased tissue generation in MND cases, which likely serve to compensate for the degeneration of motor units and skeletal muscle. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This research was supported through funding from the Scott Sullivan MND Research Fellowship to S.T.N. (2015-2020; MND and Me Foundation and RBWH Foundation), FightMND Mid-Career Fellowship (to S.T.N.), MNDRA Charcot Grant (GIA 1701 to S.T.N. and F.J.S.), the National Health and Medical Research Council Australia (1101085 to S.T.N. and F.J.S., 1113400 & 1173790 to N.R.W.) and the Australian Research Council (Future Fellowship 200100837 to A.F.M.). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: All work performed in this study was approved by the Royal Brisbane and Women's Hospital and University of Queensland human research ethics committees. All participants provided written informed consent. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The data that support the findings of this study are available on request from the corresponding author. The raw transcript data are not publicly available due to ethical restrictions; however, gene and transcript counts have been made available. The datasets generated during and/or analysed during the current study are available in the University of Queensland data collection repository. * ALS-FRS-R : Revised Amyotrophic Lateral Sclerosis Functional Rating Scale ALS : Amyotrophic lateral sclerosis EM : Expectation maximalization GIF : Surrogate variable analysis lncRNA : Long non-coding RNA MND : Motor neurone disease NMJ : Neuromuscular junction ORF : Open reading frame PC : Protein coding PCA : Principal component analysis PP : Processed pseudogene RIN : RNA integrity number TUP : Transcribed unprocessed pseudogene. UP : Unprocessed pseudogene
Individuals encounter varying environmental exposures throughout their lifetimes. Some exposures such as smoking are readily observed and have high personal recall; others are more indirect or sporadic and might only be inferred from long occupational histories or lifestyles. We evaluated the utility of using lifetime-long self-reported exposures for identifying differential methylation in an amyotrophic lateral sclerosis cases-control cohort of 855 individuals. Individuals submitted paper-based surveys on exposure and occupational histories as well as whole blood samples. Genome-wide DNA methylation levels were quantified using the Illumina Infinium Human Methylation450 array. We analyzed 15 environmental exposures using the OSCA software linear and MOA models, where we regressed exposures individually by methylation adjusted for batch effects and disease status as well as predicted scores for age, sex, cell count, and smoking status. We also regressed on the first principal components on clustered environmental exposures to detect DNA methylation changes associated with a more generalised definition of environmental exposure. Five DNA methylation probes across three environmental exposures (cadmium, mercury and metalwork) were significantly associated using the MOA models and seven through the linear models, with one additionally across a principal component representing chemical exposures. Methylome-wide significance for four of these markers was driven by extreme hyper/hypo-methylation in small numbers of individuals. The results indicate the potential for using self-reported exposure histories in detecting DNA methylation changes in response to the environment, but also highlight the confounded nature of environmental exposure in cohort studies.
Background Amyotrophic lateral sclerosis (ALS) is a complex, late-onset, neurodegenerative disease with a genetic contribution to disease liability. Genome-wide association studies (GWAS) have identified ten risk loci to date, including the TNIP1 / GPX3 locus on chromosome five. Given association analysis data alone cannot determine the most plausible risk gene for this locus, we undertook a comprehensive suite of in silico, in vivo and in vitro studies to address this. Methods The Functional Mapping and Annotation (FUMA) pipeline and five tools (conditional and joint analysis (GCTA-COJO), Stratified Linkage Disequilibrium Score Regression (S-LDSC), Polygenic Priority Scoring (PoPS), Summary-based Mendelian Randomisation (SMR-HEIDI) and transcriptome-wide association study (TWAS) analyses) were used to perform bioinformatic integration of GWAS data ( N cases = 20,806, N controls = 59,804) with ‘omics reference datasets including the blood (eQTLgen consortium N = 31,684) and brain ( N = 2581). This was followed up by specific expression studies in ALS case-control cohorts (microarray N total = 942, protein N total = 300) and gene knockdown (KD) studies of human neuronal iPSC cells and zebrafish-morpholinos (MO). Results SMR analyses implicated both TNIP1 and GPX3 ( p < 1.15 × 10 −6 ), but there was no simple SNP/expression relationship. Integrating multiple datasets using PoPS supported GPX3 but not TNIP1 . In vivo expression analyses from blood in ALS cases identified that lower GPX3 expression correlated with a more progressed disease (ALS functional rating score, p = 5.5 × 10 −3 , adjusted R 2 = 0.042, B effect = 27.4 ± 13.3 ng/ml/ALSFRS unit) with microarray and protein data suggesting lower expression with risk allele (recessive model p = 0.06, p = 0.02 respectively). Validation in vivo indicated gpx3 KD caused significant motor deficits in zebrafish-MO (mean difference vs. control ± 95% CI, vs. control, swim distance = 112 ± 28 mm, time = 1.29 ± 0.59 s, speed = 32.0 ± 2.53 mm/s, respectively, p for all < 0.0001), which were rescued with gpx3 expression, with no phenotype identified with tnip1 KD or gpx3 overexpression. Conclusions These results support GPX3 as a lead ALS risk gene in this locus, with more data needed to confirm/reject a role for TNIP1 . This has implications for understanding disease mechanisms ( GPX3 acts in the same pathway as SOD1 , a well-established ALS-associated gene) and identifying new therapeutic approaches. Few previous examples of in-depth investigations of risk loci in ALS exist and a similar approach could be applied to investigate future expected GWAS findings.
Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disease with an estimated heritability between 40 and 50%. DNA methylation patterns can serve as proxies of (past) exposures and disease progression, as well as providing a potential mechanism that mediates genetic or environmental risk. Here, we present a blood-based epigenome-wide association study meta-analysis in 9706 samples passing stringent quality control (6763 patients, 2943 controls). We identified a total of 45 differentially methylated positions (DMPs) annotated to 42 genes, which are enriched for pathways and traits related to metabolism, cholesterol biosynthesis, and immunity. We then tested 39 DNA methylation–based proxies of putative ALS risk factors and found that high-density lipoprotein cholesterol, body mass index, white blood cell proportions, and alcohol intake were independently associated with ALS. Integration of these results with our latest genome-wide association study showed that cholesterol biosynthesis was potentially causally related to ALS. Last, DNA methylation at several DMPs and blood cell proportion estimates derived from DNA methylation data were associated with survival rate in patients, suggesting that they might represent indicators of underlying disease processes potentially amenable to therapeutic interventions.
Additional file 1: Supplementary Tables. Table S1: Significant risk loci detected from ALS case-control GWAS. Table S2: Lead SNPs identified from independent significant SNPs of ALS GWAS. Table S3: Independent significant SNPs at r2 < 0:6 identified from ALS GWAS. Table S4: FUMA identifies 92 genes (ALS GWAS). Table S5: Prioritized genes from ALS GWAS by functional mapping. Table S6: FUMA annotation pathway categories. Table S7: S-LDSC annotation categories enriched in ALS GWAS. Table S8: S-LDSC cell-type categories enriched in ALS GWAS. Table S9: GCTA-COJO analyses. Table S10: GCTA-COJO analysis conditioned based on lead SNP (rs10463311) and different co-linearity thresholds. Table S11: FUMA chromatin interactions. Table S12: SMR results for GPX3, TNIP1 and C9orf72. Table S13: SMR significant genes. Table S14: Brain eQTL SMR results to follow-up significant genes in blood eQTL data. Table S15: The top GWAS and SMR chromosome 5 SNPs in brain expression datasets. Table S16: Significantly associated genes using TWAS models elastic net and CONTENT models. Table S17: Significantly associated genes using TWAS models elastic net and CONTENT models for each identified tissue. Table S18: TWAS CONTENT full model with significantly expressed GPX3 or TNIP1 in relevant tissue types. Table S19: Gene correlation in 48 GTEx tissues. Table S20: Gene correlation of 16 ALS genes in 45 GTEx tissues. Table S21: Top two prioritized candidates per chromosome in ALS GWAS using PoPS. Table S22: Platforms used for microarray ALS and control samples and genotype counts. Table S23: Morpholino (MO) abnormalities after injection of GPX3-knock-down mRNA were no different from controls or non-injected MO. Table S24: No severe motor abnormalities in zebrafish morpholinos (MO) with increasing GPX3 mRNA.
Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disease with a lifetime risk of one in 350 people and an unmet need for disease-modifying therapies. We conducted a cross-ancestry genome-wide association study (GWAS) including 29,612 patients with ALS and 122,656 controls, which identified 15 risk loci. When combined with 8,953 individuals with whole-genome sequencing (6,538 patients, 2,415 controls) and a large cortex-derived expression quantitative trait locus (eQTL) dataset (MetaBrain), analyses revealed locus-specific genetic architectures in which we prioritized genes either through rare variants, short tandem repeats or regulatory effects. ALS-associated risk loci were shared with multiple traits within the neurodegenerative spectrum but with distinct enrichment patterns across brain regions and cell types. Of the environmental and lifestyle risk factors obtained from the literature, Mendelian randomization analyses indicated a causal role for high cholesterol levels. The combination of all ALS-associated signals reveals a role for perturbations in vesicle-mediated transport and autophagy and provides evidence for cell-autonomous disease initiation in glutamatergic neurons.
Background The schizophrenia polygenic risk score (SCZ-PRS) is an emerging tool in psychiatry. Aims We aimed to evaluate the utility of SCZ-PRS in a young, transdiagnostic, clinical cohort. Method SCZ-PRSs were calculated for young people who presented to early-intervention youth mental health clinics, including 158 patients of European ancestry, 113 of whom had longitudinal outcome data. We examined associations between SCZ-PRS and diagnosis, clinical stage and functioning at initial assessment, and new-onset psychotic disorder, clinical stage transition and functional course over time in contact with services. Results Compared with a control group, patients had elevated PRSs for schizophrenia, bipolar disorder and depression, but not for any non-psychiatric phenotype (for example cardiovascular disease). Higher SCZ-PRSs were elevated in participants with psychotic, bipolar, depressive, anxiety and other disorders. At initial assessment, overall SCZ-PRSs were associated with psychotic disorder (odds ratio (OR) per s.d. increase in SCZ-PRS was 1.68, 95% CI 1.08–2.59, P = 0.020), but not assignment as clinical stage 2+ (i.e. discrete, persistent or recurrent disorder) (OR = 0.90, 95% CI 0.64–1.26, P = 0.53) or functioning (R = 0.03, P = 0.76). Longitudinally, overall SCZ-PRSs were not significantly associated with new-onset psychotic disorder (OR = 0.84, 95% CI 0.34–2.03, P = 0.69), clinical stage transition (OR = 1.02, 95% CI 0.70–1.48, P = 0.92) or persistent functional impairment (OR = 0.84, 95% CI 0.52–1.38, P = 0.50). Conclusions In this preliminary study, SCZ-PRSs were associated with psychotic disorder at initial assessment in a young, transdiagnostic, clinical cohort accessing early-intervention services. Larger clinical studies are needed to further evaluate the clinical utility of SCZ-PRSs, especially among individuals with high SCZ-PRS burden.
Background Autism spectrum disorder (ASD) is a complex neurodevelopmental condition whose biological basis is yet to be elucidated. The Australian Autism Biobank (AAB) is an initiative of the Cooperative Research Centre for Living with Autism (Autism CRC) to establish an Australian resource of biospecimens, phenotypes and genomic data for research on autism. Methods Genome-wide single-nucleotide polymorphism genotypes were available for 2,477 individuals (after quality control) from 546 families (436 complete), including 886 participants aged 2 to 17 years with diagnosed ( n = 871) or suspected ( n = 15) ASD, 218 siblings without ASD, 1,256 parents, and 117 unrelated children without an ASD diagnosis. The genetic data were used to confirm familial relationships and assign ancestry, which was majority European ( n = 1,964 European individuals). We generated polygenic scores (PGS) for ASD, IQ, chronotype and height in the subset of Europeans, and in 3,490 unrelated ancestry-matched participants from the UK Biobank. We tested for group differences for each PGS, and performed prediction analyses for related phenotypes in the AAB. We called copy-number variants (CNVs) in all participants, and intersected these with high-confidence ASD- and intellectual disability (ID)-associated CNVs and genes from the public domain. Results The ASD ( p = 6.1e−13), sibling ( p = 4.9e−3) and unrelated ( p = 3.0e−3) groups had significantly higher ASD PGS than UK Biobank controls, whereas this was not the case for height—a control trait. The IQ PGS was a significant predictor of measured IQ in undiagnosed children ( r = 0.24, p = 2.1e−3) and parents ( r = 0.17, p = 8.0e−7; 4.0% of variance), but not the ASD group. Chronotype PGS predicted sleep disturbances within the ASD group ( r = 0.13, p = 1.9e−3; 1.3% of variance). In the CNV analysis, we identified 13 individuals with CNVs overlapping ASD/ID-associated CNVs, and 12 with CNVs overlapping ASD/ID/developmental delay-associated genes identified on the basis of de novo variants. Limitations This dataset is modest in size, and the publicly-available genome-wide-association-study (GWAS) summary statistics used to calculate PGS for ASD and other traits are relatively underpowered. Conclusions We report on common genetic variation and rare CNVs within the AAB. Prediction analyses using currently available GWAS summary statistics are largely consistent with expected relationships based on published studies. As the size of publicly-available GWAS summary statistics grows, the phenotypic depth of the AAB dataset will provide many opportunities for analyses of autism profiles and co-occurring conditions, including when integrated with other omics datasets generated from AAB biospecimens (blood, urine, stool, hair).
Amyotrophic Lateral Sclerosis (ALS) is recognised to be a complex neurodegenerative disease involving both genetic and non-genetic risk factors. The underlying causes and risk factors for the majority of cases remain unknown; however, ever-larger genetic data studies and methodologies promise an enhanced understanding. Recent analyses using published summary statistics from the largest ALS genome-wide association study (GWAS) (20,806 ALS cases and 59,804 healthy controls) identified that schizophrenia (SCZ), cognitive performance (CP) and educational attainment (EA) related traits were genetically correlated with ALS. To provide additional evidence for these correlations, we built single and multi-trait genetic predictors using GWAS summary statistics for ALS and these traits, (SCZ, CP, EA) in an independent Australian cohort (846 ALS cases and 665 healthy controls). We compared methods for generating the risk predictors and found that the combination of traits improved the prediction (Nagelkerke-R 2 ) of the case–control logistic regression. The combination of ALS, SCZ, CP, and EA, using the SBayesR predictor method gave the highest prediction (Nagelkerke-R 2 ) of 0.027 ( P value = 4.6 × 10 −8 ), with the odds-ratio for estimated disease risk between the highest and lowest deciles of individuals being 3.15 (95% CI 1.96–5.05). These results support the genetic correlation between ALS, SCZ, CP and EA providing a better understanding of the complexity of ALS.