A single gene can encode multiple versions of a protein, dubbed isoforms, with varying functionality. Cellular control of isoform abundances is critical for multiple aspects of biology and is only partially regulated by transcript levels. While long-read sequencing facilitates transcript quantification, quantifying the resulting protein isoforms on a large scale is a major challenge, complicating biological interpretation of transcript alterations. Standard "bottom up" mass spectrometry can assess only short portions of isoforms called peptides, and these peptides often map onto more than one isoform. We introduce PAQu, a novel Bayesian method that leverages multiomic information from the peptidome and transcriptome to provide accurate estimates of isoform abundance even when peptide mapping is ambiguous. PAQu offers several advantages over existing methods in a unified framework. It provides uncertainty quantification, integrates multiomic information for improved accuracy, and provides a rigorous framework for hypothesis testing. Extensive simulations show that PAQu consistently outperforms competing methods in detecting differentially expressed protein isoforms and estimating their abundances. We use PAQu to investigate differences in isoform abundance levels between people with schizophrenia and control subjects, confirming a long held hypothesis that levels of the C4A isoform of Complement Component 4 are increased in schizophrenia while C4B is not. These results demonstrate that PAQu can identify significant variations in isoform abundance levels not previously possible.
Obsessive-compulsive disorder (OCD) is a chronic psychiatric illness associated with altered function in cortico-striatal-thalamo-cortical (CSTC) circuits. In this pilot study, we examined differential RNA expression in the thalamus using postmortem human brain tissue samples from 11 subjects with OCD and 10 unaffected subjects. We individually dissected the mediodorsal magnocellular, mediodorsal parvocellular, and ventral anterior nuclei, which participate in orbitofrontal and anterior cingulate CSTC circuits most frequently associated with OCD, and the posterior ventrolateral nucleus, which participates in premotor and motor circuits that are increasingly implicated in OCD. Preselected GABAergic, glutamatergic and ion channel genes were analyzed via qPCR. Two genes required for GABA synthesis and release, GAD1 and SLC32A1, were found to be downregulated in OCD subjects across all nuclei, and potassium channel KCNN3 was upregulated. In parallel, we performed an exploratory total RNAseq differential expression analysis. We identified few (12-52) differentially expressed genes (DEGs) in each nucleus, and only one DEG in a pooled analysis of all nuclei. No DEGs were significant after correction for multiple comparisons. Investigation by model selection indicated that OCD diagnosis was not a useful factor in modelling gene expression in our dataset. OCD was also not associated with any modules of co-expressed genes identified using weighted gene correlation network analysis. Overall, we found minimal evidence of differential RNA expression in these thalamic nuclei in OCD. These findings contrast with our previous work including many of the same subjects where we found widespread differential mRNA expression in the orbitofrontal cortex and striatum in OCD.
Autism spectrum disorder (ASD) is estimated to be up to four times as common in males as in females, yet the causes of this prevalence difference are not well established. One possible driver is genetic variation on the X chromosome, as it contains genes capable of contributing to ASD (e.g., PTCHD1, MECP2) and is known to play a role in genetic disorders with differential sex prevalence (e.g., color blindness). However, a lack of power compared to the autosomes combined with the complexities of modeling its biology have led to the X being largely overlooked in sequencing studies. Here, we develop quantitative X-linked TADA, a new model designed specifically for application to this chromosome, and use it to analyze rare variation from 50,663 individuals with ASD (and 136,670 individuals total). We find 9 genes on the X associated with ASD at a false discovery rate (FDR) < 0.05 and an additional 9 genes at FDR < 0.2, with many of these previously identified as involved in specific neurodevelopmental disorders. Point estimates of the liability conferred by de novo variants on the X are similar in females and males, with both sexes' estimates elevated >20% above the corresponding autosomal values. We also develop a general theory of how X-linked variation of any additive or non-additive effect influences liability and describe its implications for prevalence. Using this theory and our empirical results, we show how genetic variation on the X could contribute to the sex-differential prevalence of ASD.
Motivation Gene-damaging mutations are highly informative for studies seeking to discover genes underlying developmental disorders. Traditionally, these de novo variants are recognized by evaluating high-quality DNA sequence from affected offspring and parents. However, when parental sequence is unavailable, methods are required to infer de novo status and use this inference for association studies.Results We use data from autism spectrum disorder to illustrate and evaluate methods. Separating de novo from rare inherited variants is challenging because the latter are far more common. Using a classifier for unbalanced data and variants of known inheritance class, we build an inheritance model and then a de novo score for variants when parental data are missing. Next, we propose a new Random Draw (RD) model to use this score for gene discovery. Built into an existing inferential framework, RD produces a more powerful gene-based association test and controls the false discovery rate.Availability and implementation Codes are available at Github (https://github.com/HaeunM/TADA-RD) and Zenodo (DOI: https://doi.org/10.5281/zenodo.18531769).
Autism spectrum disorder is a heritable neurodevelopmental condition affecting approximately 3% of children that presents with core behavioral features and a range of possible comorbidities, including intellectual disability. While common variants contribute substantially to autism liability, the discovery of specific autism-associated genes has largely been driven by studies of rare and de novo variants. Many of these genes are also linked with broadly defined developmental disorders, but their involvement in other conditions has not been mapped at scale. Here, we analyze autosomal rare coding variation from 62,429 individuals with autism from research and clinical cohorts to identify 253 autism-associated genes at an estimated false discovery rate < 0.001. We cluster them based on association evidence from large-scale studies of developmental disorders, schizophrenia, bipolar disorder, and epilepsy, generating six clusters of genes with differing biological pathway enrichments and patterns of comorbidities. Investigating rare variant associations in the population using the UK Biobank and All of Us, we identify autism-associated genes displaying pleiotropy across physiological systems. In addition, we report 497 genes impacting development in a meta-analysis with 26,109 published developmental disorders samples. Collectively drawing upon data from over 1.5 million individuals, our study finds that rare variants across hundreds of genes contribute to autism with variable phenotypic outcomes.
The past decade has seen remarkable progress in identifying genes that, when impacted by deleterious coding variation, confer high likelihood for autism spectrum disorder (ASD), intellectual disability and other associated developmental disorders. However, most underlying gene discovery efforts have focused on individuals of European ancestry, limiting insights into genetic liability across diverse populations. To help address this, the Genomics of Autism in Latin American Ancestries (GALA) Consortium was formed, presenting here the largest sequencing study of autism in Latin American individuals (n > 15,000, including 4,717 participants with an ASD diagnosis). We identified 35 genome-wide significant (false discovery rate < 0.05) autism-associated genes, with substantial overlap with findings from European cohorts, and highly constrained genes showing consistent signal across populations. The results provide support for emerging (for example, MARK2, YWHAG, PACS1, RERE, SPEN, GSE1, GLS, TNPO3 and ANKRD17) and established autism genes and for the utility of genetic testing approaches for deleterious variants in individuals from diverse backgrounds; the results also demonstrate the ongoing need for more inclusive genetic research and testing. We conclude that the biology of autism is consistent across populations, with no detectable influence of ancestry.
PACS1 syndrome is a neurodevelopmental disorder (NDD) resulting from a unique de novo p.R203W variant in Phosphofurin Acidic Cluster Sorting protein 1 (PACS1). PACS1 encodes a multifunctional sorting protein required for localizing furin to the trans-Golgi network. Although few studies have started to investigate the impact of the PACS1 p.R203W variant, the mechanisms by which the variant affects neurodevelopment are still poorly understood. In recent years, autism spectrum disorder (ASD) patient-derived brain organoids have been increasingly used to identify pathogenic mechanisms and possible therapeutic targets. While most of these studies evaluate the mechanisms by which ASD-risk genes affect the transcriptome, studies considering the proteome are limited. Here, we examine the effect of PACS1 p.R203W on the proteomic landscape of brain organoids using tandem mass tag (TMT) mass-spectrometry. Time series analysis between PACS1(+/+) and PACS1(+/R203W) organoids uncovered several proteins with dysregulated abundance or phosphorylation status, including known PACS1 interactors. Although we observed low overlap between proteins with altered expression and phosphorylation, the resulting dysregulated processes converged. The presence of the PACS1 p.R203W variant accelerated the emergence of proteins related to synaptogenesis and impaired vesicle loading and recycling. The earlier presence of these proteins and their related processes could lead to defective and/or incomplete synaptic function. Key dysregulated proteins observed in PACS1(+/R203W) organoids have been associated with several neurological diseases, and many are classified as NDD-causative and ASD-risk genes. Our results highlight that proteomic analyses not only enhance our understanding of general NDD mechanisms by complementing transcriptomic studies, but could also uncover additional targets, and therefore facilitate therapy development.
Elevated maternal pre-pregnancy body mass index (BMI) has been suggested to increase risk of offspring autism spectrum disorder (ASD) but evidence is mixed across heterogeneous studies and robust estimates spanning the full BMI range are lacking. This study examined the association between maternal BMI and offspring ASD in a harmonized, two-nation study and across the full BMI range. We included all singleton children born in Denmark 2004–2018 and Sweden 1998–2019 to parents of Nordic origin (n = 2,072,445), with follow-up from age 2 until 31 December 2021, or 2022, respectively. Maternal BMI recorded at the first antenatal visit was obtained from the Swedish and Danish Medical Birth Registers and was analyzed as a continuous variable and in World Health Organization-defined categories of underweight (BMI < 18.5), normal weight (18.5–24.9), overweight (25–29.9), obese class I (30–34.9), and obese class II–III (≥ 35). The relative risk of ASD was estimated as hazard ratios (HR) from Cox regression models, adjusted for birth year and parental age, educational level, income, and psychiatric history at time of childbirth, using data from national health and population registers. Both country-specific and pooled analyses were conducted. Subgroup and sensitivity analyses, including a sibling comparison, were performed to address the specificity and robustness of findings. A total of 58,416 (2.8
The past decade has seen remarkable progress in identifying genes that, when impacted by deleterious coding variation, confer high risk for autism spectrum disorder (ASD), intellectual disability, and other developmental disorders. However, most underlying gene discovery efforts have focused on individuals of European ancestry, limiting insights into genetic risks across diverse populations. To help address this, the Genomics of Autism in Latin American Ancestries Consortium (GALA) was formed, presenting here the largest sequencing study of ASD in Latin American individuals (n>15,000). We identified 35 genome-wide significant (FDR < 0.05) ASD risk genes, with substantial overlap with findings from European cohorts, and highly constrained genes showing consistent signal across populations. The results provide support for emerging (e.g., MARK2, YWHAG, PACS1, RERE, SPEN, GSE1, GLS, TNPO3, ANKRD17) and established ASD genes, and for the utility of genetic testing approaches for deleterious variants in diverse populations, while also demonstrating the ongoing need for more inclusive genetic research and testing. We conclude that the biology of ASD is universal and not impacted to any detectable degree by ancestry.
INTRODUCTION Individuals with Alzheimer's disease (AD) commonly experience neuropsychiatric symptoms of psychosis (AD+P) and/or affective disturbance (depression, anxiety, and/or irritability, AD+A). This study's goal was to identify the genetic architecture of AD+P and AD+A, as well as their genetically correlated phenotypes. METHOD SGenome-wide association meta-analysis of 9988 AD participants from six source studies with participants characterized for AD+P AD+A, and a joint phenotype (AD+A+P). RESULTS AD+P and AD+A were genetically correlated. However, AD+P and AD+A diverged in their genetic correlations with psychiatric phenotypes in individuals without AD. AD+P was negatively genetically correlated with bipolar disorder and positively with depressive symptoms. AD+A was positively correlated with anxiety disorder and more strongly correlated than AD+P with depressive symptoms. AD+P and AD+A+P had significant estimated heritability, whereas AD+A did not. Examination of the loci most strongly associated with the three phenotypes revealed overlapping and unique associations. DISCUSSION AD+P, AD+A, and AD+A+P have both shared and divergent genetic associations pointing to the importance of incorporating genetic insights into future treatment development.
Polygenic scores (PGSs) are quantitative metrics for predicting phenotypic values, such as human height or disease status. Some PGS methods require only summary statistics of a relevant genome-wide association study (GWAS) for their score. One such method is Lassosum, which inherits the model selection advantages of Lasso to select a meaningful subset of the GWAS single-nucleotide polymorphisms as predictors from their association statistics. However, even efficient scores like Lassosum, when derived from European-based GWASs, are poor predictors of phenotype for subjects of non-European ancestry; that is, they have limited portability to other ancestries. To increase the portability of Lassosum, when GWAS information and estimates of linkage disequilibrium are available for both ancestries, we propose Joint-Lassosum (JLS). In the simulation settings we explore, JLS provides more accurate PGSs compared to other methods, especially when measured in terms of fairness. In analyses of UK Biobank data, JLS was computationally more efficient but slightly less accurate than a Bayesian comparator, SDPRX. Like all PGS methods, JLS requires selection of predictors, which are determined by data-driven tuning parameters. We describe a new approach to selecting tuning parameters and note its relevance for model selection for any PGS. We also draw connections to the literature on algorithmic fairness and discuss how JLS can help mitigate fairness-related harms that might result from the use of PGSs in clinical settings. While no PGS method is likely to be universally portable, due to the diversity of human populations and unequal information content of GWASs for different ancestries, JLS is an effective approach for enhancing portability and reducing predictive bias.
BACKGROUND:Risk for Tourette disorder, and chronic motor or vocal tic disorders (referenced here inclusively as CTD), arise from a combination of genetic and environmental factors. While multiple studies have demonstrated the importance of direct additive genetic variation for CTD risk, little is known about the role of cross-generational transmission of genetic risk, such as maternal effect, which is not transmitted via the inherited parental genomes. Here, we partition sources of variation on CTD risk into direct additive genetic effect (narrow-sense heritability) and maternal effect. METHODS:The study population consists of 2 522 677 individuals from the Swedish Medical Birth Register, who were born in Sweden between 1 January 1973 and 31 December 2000, and followed for a diagnosis of CTD through 31 December, 2013. We used generalised linear mixed models to partition the liability of CTD into: direct additive genetic effect, genetic maternal effect and environmental maternal effect. RESULTS:We identified 6227 (0.2%) individuals in the birth cohort with a CTD diagnosis. A study of half-siblings showed that maternal half-siblings had twice higher risk of developing a CTD compared with paternal ones. We estimated 60.7% direct additive genetic effect (95% credible interval, 58.5% to 62.4%), 4.8% genetic maternal effect (95% credible interval, 4.4% to 5.1%) and 0.5% environmental maternal effect (95% credible interval, 0.2% to 7%). CONCLUSIONS:Our results demonstrate genetic maternal effect contributes to the risk of CTD. Failure to account for maternal effect results in an incomplete understanding of the genetic risk architecture of CTD, as the risk for CTD is impacted by maternal effect which is above and beyond the risk from transmitted genetic effect.
The genetic architectures underlying symptoms of conduct problems and depression have largely been examined separately and without incorporating temperament, despite evidence for their genetic overlap. We examined how symptoms and temperament dimensions were transmitted together in families to identify highly heritable composite phenotypes, and how these composite phenotypes predicted alcohol outcomes in young adulthood. Participants (N = 486) were drawn from the third generation of families oversampled for alcohol use disorder in the first generation. Conduct problems, depression, and temperament were reported at 11–19 years old and alcohol outcomes at 18–26 years old. Using principal components of heritability analysis, we found seven highly heritable composite phenotypes, five of which predicted alcohol outcomes: three characterized by co-occurring conduct problems and depression and two by conduct problems. Novel composite phenotypes that were characterized by both conduct problems and depression showed different types of symptoms, temperament features, and genetic underpinnings. Children manifesting differing composite phenotypes might benefit from distinct treatments based on their unique etiologies.
Molecular mechanisms of neuropsychiatric disorders are challenging to study in human brain. For decades, the preferred model has been to study postmortem human brain samples despite the limitations they entail. A recent study generated RNA sequencing data from biopsies of prefrontal cortex from living patients with Parkinson's Disease and compared gene expression to postmortem tissue samples, from which they found vast differences between the two. This led the authors to question the utility of postmortem human brain studies. Through re-analysis of the same data, we unexpectedly found that the living brain tissue samples were of much lower quality than the postmortem samples across multiple standard metrics. We also performed simulations that illustrate the effects of ignoring RNA degradation in differential gene expression analyses, showing the effects can be substantial and of similar magnitude to what the authors find. For these reasons, we believe the authors' conclusions are unjustified. To the contrary, while opportunities to study gene expression in the living brain are welcome, evidence that this eclipses the value of postmortem analyses is not apparent.
Some individuals with autism spectrum disorder (ASD) carry functional mutations rarely observed in the general population. We explored the genes disrupted by these variants from joint analysis of protein-truncating variants (PTVs), missense variants and copy number variants (CNVs) in a cohort of 63,237 individuals. We discovered 72 genes associated with ASD at false discovery rate (FDR) ≤ 0.001 (185 at FDR ≤ 0.05). De novo PTVs, damaging missense variants and CNVs represented 57.5%, 21.1% and 8.44% of association evidence, while CNVs conferred greatest relative risk. Meta-analysis with cohorts ascertained for developmental delay (DD) (n = 91,605) yielded 373 genes associated with ASD/DD at FDR ≤ 0.001 (664 at FDR ≤ 0.05), some of which differed in relative frequency of mutation between ASD and DD cohorts. The DD-associated genes were enriched in transcriptomes of progenitor and immature neuronal cells, whereas genes showing stronger evidence in ASD were more enriched in maturing neurons and overlapped with schizophrenia-associated genes, emphasizing that these neuropsychiatric disorders may share common pathways to risk.
DNA methylation (DNAm), the addition of a methyl group to a cytosine in DNA, plays an important role in the regulation of gene expression. Single-nucleotide polymorphisms (SNPs) associated with schizophrenia (SZ) by genome-wide association studies (GWAS) often influence local DNAm levels. Thus, DNAm alterations, acting through effects on gene expression, represent one potential mechanism by which SZ-associated SNPs confer risk. In this study, we investigated genome-wide DNAm in postmortem superior temporal gyrus from 44 subjects with SZ and 44 non-psychiatric comparison subjects using Illumina Infinium MethylationEPIC BeadChip microarrays, and extracted cell-type-specific methylation signals by applying tensor composition analysis. We identified SZ-associated differential methylation at 242 sites, and 44 regions containing two or more sites (FDR cutoff of q = 0.1) and determined a subset of these were cell-type specific. We found mitotic arrest deficient 1-like 1 ( MAD1L1 ), a gene within an established GWAS risk locus, harbored robust SZ-associated differential methylation. We investigated the potential role of MAD1L1 DNAm in conferring SZ risk by assessing for colocalization among quantitative trait loci for methylation and gene transcripts (mQTLs and tQTLs) in brain tissue and GWAS signal at the locus using multiple-trait-colocalization analysis. We found that mQTLs and tQTLs colocalized with the GWAS signal (posterior probability >0.8). Our findings suggest that alterations in MAD1L1 methylation and transcription may mediate risk for SZ at the MAD1L1 -containing locus. Future studies to identify how SZ-associated differential methylation affects MAD1L1 biological function are indicated.