Rare-variant analysis is commonly used in whole-exome or genome sequencing studies. Compared to common variants, rare variants tend to have larger effect sizes and often directly point out causal genes. These potential benefits make association analysis with rare variants a priority for human genetics researchers. To improve the power of such studies, numerous methods have been developed to aggregate information of all variants of a gene. However, these gene-based methods often make unrealistic assumptions, e.g., the commonly used burden test effectively assumes that all variants chosen in the analysis have the same effects. In practice, current methods are often underpowered. We propose a Bayesian method: mixture-model-based rare-variant analysis on genes (MIRAGE). MIRAGE analyzes summary statistics (i.e., variant counts from inherited variants in trio sequencing or from ancestry-matched case-control studies). MIRAGE captures the heterogeneity of variant effects by treating all variants of a gene as a mixture of risk and non-risk variants and uses external information of variants to model the prior probabilities of being risk variants. We demonstrate, in both simulations and analysis of an exome-sequencing dataset of autism, that MIRAGE significantly outperforms current methods for rare-variant analysis. The top genes identified by MIRAGE are highly enriched with known or plausible autism-risk genes.
Autism Spectrum Disorder (ASD) arises from complex genetic and environmental factors, with inherited genetic variation playing a substantial role. This study introduces a novel approach to uncover moderate effect size (MES) genes in ASD, which individually do not meet the ASD liability threshold but collectively contribute when paired with specific other MES genes. Analyzing 10,795 families from the SPARK dataset, we identified 97 MES genes forming 50 significant gene pairs, demonstrating a substantial association with ASD when considered in tandem, but not individually. Our method leverages familial inheritance patterns and statistical analyses, refined by comparisons against control cohorts, to elucidate these gene pairs' contribution to ASD liability. Furthermore, expression profile analyses of these genes in brain tissues underscore their relevance to ASD pathology. This study underscores the complexity of ASD's genetic landscape, suggesting that gene combinations, beyond high impact single-gene mutations, significantly contribute to the disorder's etiology and heterogeneity. Our findings pave the way for new avenues in understanding ASD's genetic underpinnings and developing targeted therapeutic strategies.
De novo mutations in protein-coding regions are strongly associated with autism, and family-based sequencing studies have identified numerous genes that harbor excess mutations in probands. However, the aggregate contribution of this class of variation to autism remains unclear. Here, we model the distribution of de novo autosomal coding variant effect sizes in 38,680 autism trios to estimate fundamental features of de novo genetic architecture. We find that damaging de novo single-nucleotide variants and frameshift indels explain 3.4% (95% CI: 2.1% - 4.7%) of autism variance on the observed scale. Approximately 7.0% (95% CI: 5.6% - 8.4%) of cases carry a large-effect mutation (rate ratio > 5), and most such mutations are incompletely penetrant. Although hundreds of genes make some nonzero contribution, 50% of mutational variance on the autosomes is explained by just 15 genes. De novo enrichments vary across cohorts with different ascertainment strategies; making projections for future trio studies, we show that many large-effect genes remain to be found. ### Competing Interest Statement K.J.K. is a member of the scientific advisory board of Nurture Genomics. M.E.T. has received research and/or financial support from Illumina Inc, Microsoft Inc, Pacific Biosciences, Ionis Pharmaceuticals, Levo Therapeutics, BridgeBio, and First Genomic Insights. Z.Z., M.M.M., and P.K. are or were employees of and may be shareholders of GeneDx, LLC. The other authors declare no conflicts of interest. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: All samples analyzed in this manuscript are described in detail in the flagship consortium preprint: Satterstrom, F.K. et al. (2026) Rare variation illuminates the distinct and pleiotropic genetic architecture of autism across neuropsychiatric traits, medRxiv [Preprint]. The data had been de-identified before use in the study. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Code for running burdenMLE-DN is available at https://github.com/ajaynadig/burdenMLE-DN, including a wiki and tutorial.
Abstract Rare variant association studies (RVAS) have identified hundreds of genes contributing to human disease, yet gene-level signals provide limited insight into the molecular mechanisms underlying pathogenicity. Missense variants, which can be mapped onto three-dimensional protein structures, offer an opportunity to gain novel mechanistic insights. Here, we develop a scalable framework for systematically mapping case and control variants onto protein structures and identifying spatially localized regions enriched for case variants. Our framework builds on the 3D Neighborhood Test (3DNT), which we recently introduced in a single-gene analysis of ATP2B2 , and enables the genome-wide analysis of rare coding variation beyond standard gene-level approaches. We applied 3DNT across multiple large-scale datasets, including Mendelian disease variants from ClinVar, de novo mutations from 37,486 autism spectrum disorder (ASD) probands, and case-control exome sequencing cohorts for epilepsy and schizophrenia. We identified significant clusters in 872 genes for Mendelian disease, in 70 genes for autism, in one gene for epilepsy, and in three genes for schizophrenia. These clusters are strongly enriched for known functional sites and provide insight into both known and previously unrecognized disease genes. Our results demonstrate that scalably integrating RVAS data with protein structure predictions localizes disease-associated variation to specific functional regions and reveals a layer of disease biology that is largely invisible to standard analyses.
Autism spectrum disorder (ASD) is estimated to be up to four times as common in males as in females, yet the causes of this prevalence difference are not well established. One possible driver is genetic variation on the X chromosome, as it contains genes capable of contributing to ASD (e.g., PTCHD1, MECP2) and is known to play a role in genetic disorders with differential sex prevalence (e.g., color blindness). However, a lack of power compared to the autosomes combined with the complexities of modeling its biology have led to the X being largely overlooked in sequencing studies. Here, we develop quantitative X-linked TADA, a new model designed specifically for application to this chromosome, and use it to analyze rare variation from 50,663 individuals with ASD (and 136,670 individuals total). We find 9 genes on the X associated with ASD at a false discovery rate (FDR) < 0.05 and an additional 9 genes at FDR < 0.2, with many of these previously identified as involved in specific neurodevelopmental disorders. Point estimates of the liability conferred by de novo variants on the X are similar in females and males, with both sexes' estimates elevated >20% above the corresponding autosomal values. We also develop a general theory of how X-linked variation of any additive or non-additive effect influences liability and describe its implications for prevalence. Using this theory and our empirical results, we show how genetic variation on the X could contribute to the sex-differential prevalence of ASD.
Autism spectrum disorder is a heritable neurodevelopmental condition affecting approximately 3% of children that presents with core behavioral features and a range of possible comorbidities, including intellectual disability. While common variants contribute substantially to autism liability, the discovery of specific autism-associated genes has largely been driven by studies of rare and de novo variants. Many of these genes are also linked with broadly defined developmental disorders, but their involvement in other conditions has not been mapped at scale. Here, we analyze autosomal rare coding variation from 62,429 individuals with autism from research and clinical cohorts to identify 253 autism-associated genes at an estimated false discovery rate < 0.001. We cluster them based on association evidence from large-scale studies of developmental disorders, schizophrenia, bipolar disorder, and epilepsy, generating six clusters of genes with differing biological pathway enrichments and patterns of comorbidities. Investigating rare variant associations in the population using the UK Biobank and All of Us, we identify autism-associated genes displaying pleiotropy across physiological systems. In addition, we report 497 genes impacting development in a meta-analysis with 26,109 published developmental disorders samples. Collectively drawing upon data from over 1.5 million individuals, our study finds that rare variants across hundreds of genes contribute to autism with variable phenotypic outcomes.
The past decade has seen remarkable progress in identifying genes that, when impacted by deleterious coding variation, confer high likelihood for autism spectrum disorder (ASD), intellectual disability and other associated developmental disorders. However, most underlying gene discovery efforts have focused on individuals of European ancestry, limiting insights into genetic liability across diverse populations. To help address this, the Genomics of Autism in Latin American Ancestries (GALA) Consortium was formed, presenting here the largest sequencing study of autism in Latin American individuals (n > 15,000, including 4,717 participants with an ASD diagnosis). We identified 35 genome-wide significant (false discovery rate < 0.05) autism-associated genes, with substantial overlap with findings from European cohorts, and highly constrained genes showing consistent signal across populations. The results provide support for emerging (for example, MARK2, YWHAG, PACS1, RERE, SPEN, GSE1, GLS, TNPO3 and ANKRD17) and established autism genes and for the utility of genetic testing approaches for deleterious variants in individuals from diverse backgrounds; the results also demonstrate the ongoing need for more inclusive genetic research and testing. We conclude that the biology of autism is consistent across populations, with no detectable influence of ancestry.
De novo variants are a leading cause of neurodevelopmental disorders (NDDs), but because every monogenic NDD is different and usually extremely rare, it remains a major challenge to understand the complete phenotype and genotype spectrum of any morbid gene. According to OMIM, heterozygous variants in KDM6B cause "neurodevelopmental disorder with coarse facies and mild distal skeletal abnormalities."Here, by examining the molecular and clinical spectrum of 85 reported individuals with mostly de novo (likely) pathogenic KDM6B variants, we demonstrate that this description is inaccurate and potentially misleading. Cognitive deficits are seen consistently in all individuals, but the overall phenotype is highly variable. Notably, coarse facies and distal skeletal anomalies, as defined by OMIM, are rare in this expanded cohort while other features are unexpectedly common (e.g., hypotonia, psychosis, etc.). Using 3D protein structure analysis and an innovative dual Drosophila gain-of-function assay, we demonstrated a disruptive effect of 11 missense/in-frame indels located in or near the enzymatic JmJC or Zn-containing domain of KDM6B. Consistent with the role of KDM6B in human cognition, we demonstrated a role for the Drosophila KDM6B ortholog in memory and behavior. Taken together, we accurately define the broad clinical spectrum of the KDM6B-related NDD, introduce an innovative functional testing paradigm for the assessment of KDM6B variants, and demonstrate a conserved role for KDM6B in cognition and behavior. Our study demonstrates the critical importance of international collaboration, sharing of clinical data, and rigorous functional analysis of genetic variants to ensure correct disease diagnosis for rare disorders.
Autism is four times more prevalent in males than females. To study whether this reflects a difference in genetic predisposition attributed to autosomal rare variants, we evaluated sex differences in effect size of damaging protein-truncating and missense variants on autism predisposition in 47,061 autistic individuals using a liability model with differing thresholds. Given the sex differences in the rates of cognitive impairment among autistic individuals, we also compared effect sizes of rare variants between individuals with and without cognitive impairment or motor delay. Although these variants mediated different likelihoods of autism with versus without cognitive or motor difficulties, their effect sizes on the liability scale did not differ significantly by sex exome wide or in genes sex-differentially expressed in the cortex. De novo mutations were enriched in genes with male-biased expression in the adult cortex, but these genes did not show a significant sex difference on the liability scale, nor did the liability conferred by these genes differ significantly from other genes with similar loss-of-function intolerance and sex-averaged cortical expression. Exome-wide female bias in de novo protein-truncating mutation rates on the observed scale was driven by high-confidence and syndromic autism-predisposition genes. In summary, autosomal rare and damaging coding variants confer similar liability for autism in females and males.
Attention deficit hyperactivity disorder (ADHD) is a childhood-onset neurodevelopmental disorder with a large genetic component1. It affects around 5% of children and 2.5% of adults2, and is associated with several severe outcomes3-11. Common genetic variants associated with the disorder have been identified12,13, but the role of rare variants in ADHD is mostly unknown. Here, by analysing rare coding variants in exome-sequencing data from 8,895 individuals with ADHD and 53,780 control individuals, we identify three genes (MAP1A, ANO8 and ANK2; P < 3.07 × 10-6; odds ratios 5.55-15.13) that are implicated in ADHD. The protein-protein interaction networks of these three genes were enriched for rare-variant risk genes of other neurodevelopmental disorders, and for genes involved in cytoskeleton organization, synapse function and RNA processing. Top associated rare-variant risk genes showed increased expression across pre- and postnatal brain developmental stages and in several neuronal cell types, including GABAergic (γ-aminobutyric-acid-producing) and dopaminergic neurons. Deleterious variants were associated with lower socioeconomic status and lower levels of education in individuals with ADHD, and a decrease of 2.25 intelligence quotient (IQ) points per rare deleterious variant in a sample of adults with ADHD (n = 962). Individuals with ADHD and intellectual disability showed an increased load of rare variants overall, whereas other psychiatric comorbidities had an increased load only for specific gene sets associated with those comorbidities. This suggests that psychiatric comorbidity in ADHD is driven mainly by rare variants in specific genes, rather than by a general increased load across constrained genes.
Background Autism spectrum disorder (ASD) and attention deficit hyperactivity disorder (ADHD) are heterogeneous neurodevelopmental disorders with high heritability and frequent co-occurrence. Our previous work on the initial iPSYCH exomes (Satterstrom et al., 2019) suggested a similar burden of rare protein-truncating variants (PTVs) across ASD and ADHD and identified MAP1A as a shared risk gene implicated by rare PTVs in both disorders. To build upon these findings, we aimed to 1) expand our gene discovery analysis using an updated dataset with nearly twice the sample size from the latest iPSYCH exomes, 2) quantify the burden heritability attributable to rare coding variants in ASD and ADHD, and 3) evaluate the burden genetic correlation between two disorders. Methods We analyzed exomes of 25,208 individuals from iPSYCH, comprising 7,119 diagnosed with ASD alone (ASD-only), 5,598 with ADHD alone (ADHD-only), 3,794 diagnosed with both conditions (ASD+ADHD), and 8,697 controls. Multivariate Poisson regression models were applied to systematically assess rare variant burdens in various gene sets across the three case groups, further stratifying by the presence or absence of intellectual disability (ID). We employed c-alpha tests to compare the distribution of rare deleterious variants between ASD-only and ADHD-only. We performed burden heritability regression analyses to estimate the burden heritability of ASD and ADHD, respectively, and to measure their burden genetic correlation. For gene discovery, we combined individuals diagnosed with ASD and/or ADHD into a single case group, included non-psychiatric non-Finnish European exome subset of gnomAD as external controls, and applied Fisher’s exact test to identify genes reaching exome-wide significance. Results Consistent with our previous findings, all three case groups demonstrated comparable elevated burdens of class I variants - including rare PTVs and highly deleterious missense variants (AlphaMissense ≥ 0.98 and MPC ≥ 2) in constrained genes compared to controls (ASD-only: OR = 1.49, 95% CI [1.40, 1.58]; ADHD-only: OR = 1.40, 95% CI [1.31, 1.50]; ASD+ADHD: OR = 1.46, 95% CI [1.35, 1.57]). C-alpha tests indicated no significant differences in the distribution of these variants between ASD-only and ADHD-only (P= 0.40), whereas significant differences were observed when comparing each group to controls. Burden heritability estimates of class I variants were 1.8% (s.e. = 0.4%) for ASD and 3.2% (s.e. = 0.7%) for ADHD on the liability scale. The burden genetic correlation between the two disorders was 0.46 (s.e. = 0.17), aligning closely with previously reported common-variant genetic correlation (0.42, s.e. = 0.05; Demontis et al., 2023). In gene discovery, we identified eight exome-wide significant genes associated with both disorders, including MAP1A (the first cross-disorder gene previously identified) and seven new risk genes: five previously implicated in ASD, developmental delay, and neurodevelopmental disorders; one strong candidate gene for ASD; and one novel gene not previously linked to either disorder. Discussion Our findings underscore a substantial shared genetic architecture involving rare coding variants between ASD and ADHD, reinforcing and expanding on earlier research. Moving forward, we aim to explore the distinct genetic risks specific to each disorder and to conduct sex-stratified analyses to uncover potential sex-specific genetic differences. The results will be presented at the conference.
Rare-variant analysis is commonly used in whole-exome or genome sequencing studies. Compared to common variants, rare variants tend to have larger effect sizes and often directly point out causal genes. These potential benefits make association analysis with rare variants a priority for human genetics researchers. To improve the power of such studies, numerous methods have been developed to aggregate information of all variants of a gene. However, these gene-based methods often make unrealistic assumptions, e.g., the commonly used burden test effectively assumes that all variants chosen in the analysis have the same effects. In practice, current methods are often underpowered. We propose a Bayesian method: mixture-model-based rare-variant analysis on genes (MIRAGE). MIRAGE analyzes summary statistics (i.e., variant counts from inherited variants in trio sequencing or from ancestry-matched case-control studies). MIRAGE captures the heterogeneity of variant effects by treating all variants of a gene as a mixture of risk and non-risk variants and uses external information of variants to model the prior probabilities of being risk variants. We demonstrate, in both simulations and analysis of an exome-sequencing dataset of autism, that MIRAGE significantly outperforms current methods for rare-variant analysis. The top genes identified by MIRAGE are highly enriched with known or plausible autism-risk genes.
The past decade has seen remarkable progress in identifying genes that, when impacted by deleterious coding variation, confer high risk for autism spectrum disorder (ASD), intellectual disability, and other developmental disorders. However, most underlying gene discovery efforts have focused on individuals of European ancestry, limiting insights into genetic risks across diverse populations. To help address this, the Genomics of Autism in Latin American Ancestries Consortium (GALA) was formed, presenting here the largest sequencing study of ASD in Latin American individuals (n>15,000). We identified 35 genome-wide significant (FDR < 0.05) ASD risk genes, with substantial overlap with findings from European cohorts, and highly constrained genes showing consistent signal across populations. The results provide support for emerging (e.g., MARK2, YWHAG, PACS1, RERE, SPEN, GSE1, GLS, TNPO3, ANKRD17) and established ASD genes, and for the utility of genetic testing approaches for deleterious variants in diverse populations, while also demonstrating the ongoing need for more inclusive genetic research and testing. We conclude that the biology of ASD is universal and not impacted to any detectable degree by ancestry.
Polygenic association studies implicate numerous genes in neuropsychiatric disorders, but linkage disequilibrium (LD) and cellular heterogeneity hinder mechanistic interpretation. Here, we integrate single-nucleus RNA-seq from human neurons with network inference and polygenic signal weighting to resolve pathway-level drivers. Neuron-resolved gene co-expression networks constructed across brain regions are reweighted by GWAS-derived polygenic signal (LD-aware heritability enrichment), prioritizing modules that disproportionately contribute to liability. Using this framework, we highlight the dysregulation of Ca2+ homeostasis as an etiological driver of neuropsychiatric disorders, and even relative to other neuronal gene sets, Ca2+ homeostasis exhibits the greatest concentration of rare variant signal. Furthermore, we find that a critical component of this molecular system, the P-type calcium ATPase ATP2B2 , exhibits marked expression deficits in both nuclear transcriptomic and synaptic proteomic datasets derived from the dorsolateral prefrontal cortices of individuals with schizophrenia. To connect sequence variation to structure and mechanism, we developed a residue-centric three-dimensional neighborhood analysis that integrates case–control missense variation with AlphaFold3 structural models to localize mutational hotspots of biological significance for downstream mechanistic interrogation. This approach identified an enrichment of deleterious missense variants - implicated across multiple neuropsychiatric disorders - that changed protein residues in close spatial proximity to both the Ca2+ permeation tunnel and the ATP:Mg2+ coordination site of ATP2B2. Cellular and biochemical analyses of the canonical Ca2+ binding site revealed clear loss-of-function effects, corroborating the earlier functional genomics evidence, and establishing a distinct molecular mechanism that converges on impaired Ca2+ extrusion, likely perturbing pre-and post-synaptic Ca2+ homeostatic equilibrium in excitatory neurons. Altogether, our study makes a significant contribution by linking genetic risk to neuronal dysfunction through a critical calcium signaling axis, offering mechanistic insight into the pathogenesis of neuropsychiatric disorders. In parallel, we develop a residue-centered 3D neighborhood framework that couples case–control genetics with structural models to discover pathogenic hotspots, generalizable across the proteome to any protein structure. ### Competing Interest Statement The authors have declared no competing interest.
Autism spectrum disorder (ASD) and attention deficit hyperactivity disorder (ADHD) are heterogeneous neurodevelopmental disorders with high heritability and frequent co-occurrence. Our previous work on the first phase of iPSYCH exomes (Satterstrom et al., 2019) suggested a similar burden of rare protein-truncating variants (PTVs) across ASD and ADHD and identified MAP1A as a shared risk gene implicated by rare PTVs in both disorders. This study aims to 1) extend these findings, employing a significantly larger iPSYCH exome dataset for gene discovery, 2) estimate the burden heritability explained by rare coding variants in ASD and ADHD, and 3) assess the burden genetic correlation between the two disorders.We analyzed exomes of 25,208 individuals from iPSYCH, encompassing 7,119 individuals diagnosed with ASD alone (ASD-only), 5,598 with ADHD alone (ADHD-only), 3,794 with both ASD and ADHD (ASD+ADHD), and 8,697 controls. We used multivariate Poisson regression models to systematically evaluate rare variant burdens in different gene sets across the three case groups and controls, stratified further by the presence or absence of intellectual disability (ID). The gene sets included all genes, genes intolerant to loss-of-function variants (pLI > 0.9), and gene sets associated with different disorders including ID, ASD, ADHD, schizophrenia, and a broader group of neurodevelopmental disorders. We applied c-alpha tests to assess whether the distribution of rare deleterious variants differs between ASD and ADHD. We employed burden heritability regression to estimate the burden heritability of ASD and ADHD, respectively, and the burden genetic correlation between the two disorders. For gene discovery, we combined individuals diagnosed with ASD and/or ADHD into a single case group and applied TADA+ to integrate with family data and Swedish PAGES case-control data from a recent large-scale ASD rare variant study (Fu et al., 2022).We observed similar burdens of class I variants including rare PTVs and rare deleterious missense variants (MPC > 3) in constrained genes across the three case groups, while they all showed a significant excess compared to controls: OR = 1.35, 95% CI = [1.26, 1.45] for ASD-only; OR = 1.35, CI = [1.25, 1.45] for ADHD-only; and OR = 1.39, CI = [1.28, 1.52] for ASD+ADHD. The c-alpha tests indicated no significant differences in the distribution of class I variants in constrained genes between ASD-only and ADHD-only groups (P= 0.39) while, when comparing the case groups to controls, significant differences were observed. The burden heritability of class I variants on the liability scale was estimated to 1.87% (SE = 0.51%) for ASD and 2.42% (s.e. = 0.72%) for ADHD. The class I variant burden genetic correlation between ASD and ADHD was 0.31 (s.e. = 0.26), which approximates the point estimate of their common-variant genetic correlation of 0.42 (s.e. = 0.05) (Demontis et al., 2023).Our findings suggest substantial sharing of rare variant risk between ASD and ADHD, reinforcing the results of our earlier work (Satterstrom et al., 2019). This motivated us to merge individuals diagnosed with ASD and/or ADHD into a single group to enhance the discovery of rare variant risk genes shared between the disorders. This gene discovery analysis is ongoing, and the results will be presented at the conference.
AbstractAutism is four times more prevalent in males than females. To study whether this reflects a difference in genetic predisposition attributed to autosomal rare variants, we evaluated the sex differences in effect size of damaging protein-truncating and missense variants on autism predisposition in 47,061 autistic individuals, then compared effect sizes between individuals with and without cognitive impairment or motor delay. Although these variants mediated differential likelihood of autism with versus without motor or cognitive impairment, their effect sizes on the liability scale did not differ significantly by sex exome-wide or in genes sex-differentially expressed in the cortex. Although de novo mutations were enriched in genes with male-biased expression in the fetal cortex, the liability they conferred did not differ significantly from other genes with similar loss-of-function intolerance and sex-averaged cortical expression. In summary, autosomal rare coding variants confer similar liability for autism in females and males.
The fields of autism and neurodevelopmental disorder (NDD) genetics are rapidly advancing. Catalyzed by the power of large cohorts and integration of all classes of de novo and inherited protein-coding variation, dozens of genes have emerged to harbor variants that confer high relative risk for autism, and hundreds of genes have been associated with NDDs more broadly. Through examination of protein-truncating variants (PTVs), predicted damaging missense variation, and copy number variants (CNVs), our prior analyses have begun to map the allelic diversity of perturbations within 72 autism-associated genes and 373 genes associated with NDDs, finding intriguing evidence of genes with significantly higher mutation rates and differences in the distribution of clinical phenotypes in autism compared to NDD (Fu et al., 2022; Satterstrom et al., 2020). Despite this progress, cohort sizes remain insufficient for disentangling the shared and distinct genetic architectures of autism, NDDs, and other neuropsychiatric conditions, as well as associating genes with more subtle impacts on neurodevelopment.To advance these boundaries, we present the largest to-date study of rare coding variants, consisting of 62,013 autistic individuals, including 38,088 probands and 9,567 unaffected siblings from complete trio and quartet families, respectively, and 23,925 additional autism cases without parental information contrasted against 26,931 controls. By aggregating across the Autism Sequencing Consortium (ASC), the Simons Simplex Collection (SSC), the Simons Foundation Powering Autism Research (SPARK), and individuals from a leading diagnostic laboratory (GeneDx), this dataset totals almost 200,000 individuals, nearly a three-fold increase over prior studies. When we stratified the clinically-referred GeneDx autistic probands by co-occurring DD/ID status, we found synonymous, missense, and PTV de novo mutation rates in autism probands without DD/ID from GeneDx that were nearly identical to individuals ascertained for a diagnosis of autism in the ASC, SSC, and SPARK research studies (0.296 vs 0.294, 0.767 vs 0.763, and 0.141 vs 0.145 respectively), while GeneDx autism probands with DD/ID exhibited mutation rates similar to those observed in previous research studies of DD.Further analyses of these data solidified previous observations of significant enrichment of de novo PTVs among autism probands of 3x compared to siblings among the genes most intolerant to PTVs in the human genome (i.e., lowest decile of LOEUF from gnomAD). We have also incorporated Alpha Missense (AM) pathogenicity estimates to complement our prior MPC scores for predicting damaging missense variation and identifying de novo missense variants acting with effect sizes comparable to de novo PTVs in constrained genes, with analysis of regional missense constraint within genes ongoing. We further leveraged the TADA Bayesian statistical method to jointly model these data in a single unified framework, leveraging genetic information across rare PTVs, damaging missense variants, and CNVs. This approach discovered hundreds of genes associated with autism, where we observe a steadily increasing contribution of variant classes other than de novo PTVs in newly associated genes. Analyses are ongoing to understand the gene networks, developmental timing, and biological functions by which these genes exert their influence on phenotypic manifestations of autism and related neuropsychiatric disorders.
Autism is highly heritable and has been associated with multiple classes of genetic variation. Common genetic variation contributes substantially to autism. Previously, with 18,381 autistic individuals and 27,969 non-autistic individuals, five genome-wide significant loci were identified. Now with 38,717 autistic individuals and 232,725 non-autistic individuals, we report an updated genome-wide association study (GWAS) of autism with 12 genome-wide significant loci. We observe a moderate genetic correlation (0.675, SE=0.0434) between Europe-based (Nautistic=22,643; Nnon-autistic=204,389) and United States-based (Nautistic =16,074; Nnon-autistic=28,346) autism cohorts, which contributes to the decline of the estimated single nucleotide polymorphism (SNP) heritability (from 0.118 (SE=0.010) to 0.068 (SE=0.003)). The genetic correlation between autism with intellectual disability (ID) (Nautistic=6,590; Nnon-autistic= 43,071; h2=0.062; SE=0.012) and autism without ID (Nautistic=23,173; Nnon-autistic= 204,679; h2=0.089; SE=0.005) is 0.658 (SE=0.086). In the United States family-based cohorts, the genetic correlation between autism with ID (Nfamily=3,993; h2=0.159; SE=0.033) and autism without ID (Nfamily=4,357; h2=0.171; SE=0.031) is 0.812 (SE=0.157). Autism without ID was positively genetically correlated with educational attainment (0.163; P=4.84 × 10-11) and intelligence (0.233; P=1.95 × 10-11). Autism with ID genetically correlated with neither educational attainment (0.036; P=0.409) nor intelligence (-0.072; P=0.235). As ID alone is negatively genetically correlated with intelligence, the lack of correlation between autism with ID and intelligence strongly suggests that autism with ID is genetically different from ID alone. This difference has implications for both research and clinical nosology. Rare and de novo variants contribute substantially to autism in some individuals. Through rare variant analyses, 72 genes have been associated with autism at a genome-wide significant level to date. While de novo protein truncating variants (PTVs) and copy number deletions have been associated with autism, we report preliminary findings that the burden of inherited PTVs and copy number deletions among autistic individuals was elevated compared to their non-autistic siblings (P=4.00 × 10-5). Integration of multiple genetic factors will help us better understand the etiology of autism.
Variant scoring methods (VSMs) aid in the interpretation of coding mutations and their potential impact on health, but their evaluation in the context of human genetics applications remains inconsistent. Here, we describe GeneticsGym, a systematic approach to evaluating the real-world impact of VSMs on human genetic analysis. We show that the relative performance of VSMs varies across regimes of natural selection, and that both variant-to-gene and gene-to-disease components contribute. ### Competing Interest Statement K.J.K. is a consultant for Tome Biosciences, AlloDx, and Vor Biosciences, and a member of the scientific advisory board of Nurture Genomics. M.J.D is a founder of Maze Therapeutics.
Abstract Background Whole-genome sequencing (WGS) analyses have found higher genetic burden in autistic females compared to males, supporting higher liability threshold in females. However, genomic evidence of sex differences has been limited to European ancestry to date and little is known about how genetic variation leads to autism-related traits within families across sex. Methods To address this gap, we present WGS data of Korean autism families (n = 2255) and a Korean general population sample (n = 2500), the largest WGS data of East Asian ancestry. We analyzed sex differences in genetic burden and compared with cohorts of European ancestry (n = 15,839). Further, with extensively collected family-wise Korean autism phenotype data (n = 3730), we investigated sex differences in phenotypic scores and gene-phenotype associations within family. Results We observed robust female enrichment of de novo protein-truncating variants in autistic individuals across cohorts. However, sex differences in polygenic burden varied across cohorts and we found that the differential proportion of comorbid intellectual disability and severe autism symptoms mainly drove these variations. In siblings, males of autistic females exhibited the most severe social communication deficits. Female siblings exhibited lower phenotypic severity despite the higher polygenic burden than male siblings. Mothers also showed higher tolerance for polygenic burden than fathers, supporting higher liability threshold in females. Conclusions Our findings indicate that genetic liability in autism is both sex- and phenotype-dependent, expanding the current understanding of autism’s genetic complexity. Our work further suggests that family-based assessments of sex differences can help unravel underlying sex-differential liability in autism.