Accurate estimates of allele frequencies aid in genetic discovery, including rare disease diagnosis, common disease investigations, and population genetics. Here, we present the Genome Aggregation Database version 4 (gnomAD v4), including 730,947 with exome sequences, a fivefold increase over previous releases. We demonstrate that statistical power to detect strong selective constraint continues to increase with sample size. We develop a new loss-of-function annotation pipeline, which learns genomic features predictive of nonsense-mediated decay and splicing effects from selection signals, achieving 90% precision for distinguishing likely true versus false positive loss-of-function variants. This improved pipeline, along with incorporation of highly deleterious missense variants into measures of loss-of-function intolerance, improves disease gene detection, particularly for short genes and those with gain-of-function mechanisms. To improve disease gene prediction, we systematically extract gene-disease associations from biomedical literature, map these to gene-level biological features, and integrate both with refined constraint metrics within a Bayesian framework, yielding state-of-the-art prediction of gene-disease relevance. We highlight genes under strong constraint but with limited clinical characterization, which are enriched in embryonic lethal and fertility phenotypes, thus prioritizing previously under-characterized disease genes. Together, these advances establish a unified framework for accelerating gene discovery and improving rare disease diagnosis.
The NeuroDev study, conducted in Kenya and South Africa, is a large-scale clinical, genetic, and epidemiologic characterization of neurodevelopmental disorders (NDDs) on the African continent. NeuroDev assessments capture birth, demographic, and developmental history; cognitive and behavioral outcomes; and physical health variables. DNA samples are collected for exome sequencing and clinical genetic analysis. This paper presents novel data from 521 children with NDDs, 739 of those children's parents, and 255 unrelated, typically-developing children. The analyses offer unique genetic and phenotypic characterizations of NDDs in two African countries and underscore the importance of including underrepresented populations in NDD research. Ultimately, 107 children with NDDs from the NeuroDev cohort (22.1%) had likely pathogenic or pathogenic variants in established NDD genes. High rates of genetic diagnosis were associated with high rates of environmental risk factors for NDDs. All data, materials, and measures generated from this study are publicly available through the US National Institute of Mental Health.
Reanalysis of genomic data in rare disease is highly effective in increasing diagnostic yields but remains limited by manual approaches. Automation and optimization for high specificity will be necessary to ensure scalability, adoption and sustainability of iterative reanalysis. We developed Talos, an open-source tool that automates variant prioritization by integrating dynamically updated gene-disease and variant-level evidence with inheritance-aware filtering and validated its performance using data from 1,089 individuals with rare disease. Trio-based analysis identified 90% of known diagnoses, returning 1.3 variants per case on average. Variant burden reduced to one variant per 200 cases on iterative monthly reanalysis. Application to an unselected cohort of 4,735 undiagnosed individuals identified 241 diagnoses (5.1% yield): 78 (32%) due to new gene-disease relationships, 54 (22%) due to new variant-level evidence and 109 (45%) due to improved analysis strategies. Our automated, iterative reanalysis model demonstrates the feasibility of delivering frequent, systematic reanalysis at scale.
Genomic sequencing is now widely accessible for genetic diagnostics and is emerging as a component of newborn screening. This technological development generates the need to characterize incoming mutations, create comprehensive datasets of genes causing rare Mendelian disorders, and identify pathogenic variants. Large-scale exome sequencing datasets such as Genome Aggregation Database (gnomAD) have been assembled to help address these challenges. The recent release of gnomAD (v4; n = 730,947) uncovers millions of rare coding variants, many of which have arisen more than once by independent recurrent mutations in the rapidly growing recent human population. Here, we use newly developed theoretical understanding of sampling properties of rare variants to estimate key population genetics parameters of practical importance to human genetics such as demography history, mutation rate, and selection. Solely relying on population data, our method Population Inferred Estimates of Selection (PIES) identifies novel genes with loss-of-function mutational hotspots likely due to selection in spermatogonia. PIES efficiently estimates selection coefficients for heterozygous loss-of-function variants. Combining population genetics inference with variant effect predictors, PIES predicts pathogenic missense mutations and improves variant prioritization for genetic diagnostics and newborn screening.
Autism spectrum disorder is a heritable neurodevelopmental condition affecting approximately 3% of children that presents with core behavioral features and a range of possible comorbidities, including intellectual disability. While common variants contribute substantially to autism liability, the discovery of specific autism-associated genes has largely been driven by studies of rare and de novo variants. Many of these genes are also linked with broadly defined developmental disorders, but their involvement in other conditions has not been mapped at scale. Here, we analyze autosomal rare coding variation from 62,429 individuals with autism from research and clinical cohorts to identify 253 autism-associated genes at an estimated false discovery rate < 0.001. We cluster them based on association evidence from large-scale studies of developmental disorders, schizophrenia, bipolar disorder, and epilepsy, generating six clusters of genes with differing biological pathway enrichments and patterns of comorbidities. Investigating rare variant associations in the population using the UK Biobank and All of Us, we identify autism-associated genes displaying pleiotropy across physiological systems. In addition, we report 497 genes impacting development in a meta-analysis with 26,109 published developmental disorders samples. Collectively drawing upon data from over 1.5 million individuals, our study finds that rare variants across hundreds of genes contribute to autism with variable phenotypic outcomes.
Reanalysis of genomic data in rare disease is highly effective in increasing diagnostic yields but remains limited by manual approaches. Automation and optimization for high specificity will be necessary to ensure scalability, adoption and sustainability of iterative reanalysis. We developed a publicly available automated tool, Talos, and validated its performance using data from 1,089 individuals with rare genetic disease. Trio-based analysis identified 86% of known in-scope diagnoses, returning one variant per case on average. Variant burden reduced to one variant per 200 cases on iterative monthly reanalysis cycles. Application to an unselected cohort of 4,735 undiagnosed individuals identified 248 diagnoses (5.2% yield): 73 (29%) due to new gene-disease relationships, 56 (23%) due to new variant-level evidence, and 119 (48%) due to improved filtering and analysis strategies. Our automated, iterative reanalysis model, applied to thousands of rare disease patients, demonstrates the feasibility of delivering frequent, systematic reanalysis at scale.
Rare diseases are collectively common, affecting approximately 1 in 20 individuals worldwide. In recent years, rapid progress has been made in rare disease diagnostics due to advances in next-generation sequencing, development of new computational and functional genomics approaches to prioritize genes and variants and increased global sharing of clinical and genetic data. However, more than half of individuals suspected to have a rare disease lack a genetic diagnosis. The Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) Consortium was initiated to study thousands of challenging rare disease cases and families and apply, standardize and evaluate emerging genomics technologies and analytics to accelerate their adoption in clinical practice. Furthermore, all data generated, currently representing over 7,500 individuals from over 3,000 families, are rapidly made available to researchers worldwide through the Analysis, Visualization and Informatics Lab-space (AnVIL) to catalyse global efforts to develop approaches for genetic diagnoses in rare diseases. Most of these families have undergone previous clinical genetic testing but remained unsolved, with most being exome-negative. Here we describe the collaborative research framework, datasets and discoveries comprising GREGoR that will provide foundational resources and substrates for the future of rare disease genomics.
Recent work has revealed an important role for rare, incompletely penetrant inherited coding variants in neurodevelopmental disorders (NDDs). Additionally, we have previously shown that common variants contribute to risk for rare NDDs. Here, we investigate whether common variants exert their effects by modifying gene expression, using multi-cis-expression quantitative trait loci (cis-eQTL) prediction models. We first performed a transcriptome-wide association study for NDDs using 6987 probands from the Deciphering Developmental Disorders (DDD) study and 9720 controls, and found one gene, RAB2A, that passed multiple testing correction (p = 6.7 x 10(-7)). We then investigated whether cis-eQTLs modify the penetrance of putatively damaging, rare coding variants inherited by NDD probands from their unaffected parents in a set of 1700 trios. We found no evidence that unaffected parents transmitting putatively damaging coding variants had higher genetically-predicted expression of the variant-harboring gene than their child. In probands carrying putatively damaging variants in constrained genes, the genetically-predicted expression of these genes in blood was lower than in controls (p = 2.7 x 10(-3)). However, results for proband-control comparisons were inconsistent across different sets of genes, variant filters and tissues. We find limited evidence that common cis-eQTLs modify penetrance of rare coding variants in a large cohort of NDD probands.
Autosomal recessive coding variants are well-known causes of rare disorders. We quantified the contribution of these variants to developmental disorders in a large, ancestrally diverse cohort comprising 29,745 trios, of whom 20.4% had genetically inferred non-European ancestries. The estimated fraction of patients attributable to exome-wide autosomal recessive coding variants ranged from ~2-19% across genetically inferred ancestry groups and was significantly correlated with average autozygosity. Established autosomal recessive developmental disorder-associated (ARDD) genes explained 84.0% of the total autosomal recessive coding burden, and 34.4% of the burden in these established genes was explained by variants not already reported as pathogenic in ClinVar. Statistical analyses identified two novel ARDD genes: KBTBD2 and ZDHHC16. This study expands our understanding of the genetic architecture of developmental disorders across diverse genetically inferred ancestry groups and suggests that improving strategies for interpreting missense variants in known ARDD genes may help diagnose more patients than discovering the remaining genes.
The Genome Aggregation Database (gnomAD) is a large-scale sequencing dataset that aggregated genetic variation from over 800,000 individuals. Leveraging population databases such as the gnomAD for estimating the genetic prevalence of monogenic disorders has been demonstrated for autosomal recessive disorders. However, dominant, and X-linked monogenic disorders, particularly those associated with severe or early onset phenotypes, represent a challenge as gnomAD is expected to be depleted for variants impacting reproductive fitness.
Background Although rare neurodevelopmental conditions have a large Mendelian component, common genetic variants also contribute to risk. However, little is known about how polygenic risk is distributed among patients with these conditions and their parents, how it interacts with rare variants, and whether parents’ polygenic background contributes to their children's risk beyond the direct effect of variants transmitted to the child (i.e. via indirect genetic effects potentially mediated through the prenatal environment or ‘genetic nurture’). Methods We used genetic data from 11,573 patients with neurodevelopmental conditions from the Deciphering Developmental Disorders study and the 100,000 Genomes project, 9,128 of their parents and 26,869 controls. We conducted genome-wide association studies and estimated heritability of neurodevelopmental conditions and genetic correlations with other brain-related traits. We calculated polygenic scores (PGS) for neurodevelopmental conditions and related traits and compared average PGS between patient with and without a monogenic diagnosis. We assessed effects of direct genetic effects from transmitted alleles on risk and that from parental non-transmitted alleles using trio-based regression models. Results Common variants explained ∼10% of variance in overall risk. We observed significant negative genetic correlations between neurodevelopmental conditions and educational attainment, cognitive performance, and the non-cognitive component of educational attainment, and significant positive genetic correlations with schizophrenia and ADHD. Patients with a monogenic diagnosis had significantly less polygenic risk than those without (p < 1.2 × 10-3), supporting a liability threshold model. Genetically undiagnosed patients and diagnosed patients with affected parents had significantly more polygenic risk than controls (p < 9.3 × 10-5).In trio-based analyses, using a PGS for neurodevelopmental conditions, the transmitted but not the non-transmitted parental alleles were associated with risk, suggesting a direct genetic effect. In contrast, we observed no direct genetic effect of PGSs for educational attainment and cognitive performance but saw a significant (p < 1.1 × 10-7) correlation between the child's risk and non-transmitted parental alleles, potentially due to indirect genetic effects and/or parental assortment. Observing a significant negative genetic correlation between educational attainment and premature delivery (rg=-0.30, p=2.3 × 10-10), we hypothesized that the latter might mediate the effect of non-transmitted alleles associated with lower educational attainment in the mother. Although we found no direct evidence of this, we found some evidence for a significant bi-directional causal relationship between lower educational attainment and premature delivery using Mendelian randomization. As expected under parental assortment, common variant predisposition for neurodevelopmental conditions was correlated with the rare variant component of risk. Discussion Our findings suggest that future studies should investigate the possible role and nature of indirect genetic effects on rare neurodevelopmental conditions. The correlation between non-transmitted common alleles and neurodevelopmental conditions may partly reflect transmitted rare variants, and future studies should consider the contribution of common and rare variants simultaneously when studying cognition-related phenotypes. Disclosure Nothing to disclose.
Although rare neurodevelopmental conditions have a large Mendelian component(1), common genetic variants also contribute to risk(2,3). However, little is known about how this polygenic risk is distributed among patients with these conditions and their parents nor its interplay with rare variants. It is also unclear whether polygenic background affects risk directly through alleles transmitted from parents to children, or whether indirect genetic effects mediated through the family environment(4) also play a role. Here we addressed these questions using genetic data from 11,573 patients with rare neurodevelopmental conditions, 9,128 of their parents and 26,869 controls. Common variants explained around 10% of variance in risk. Patients with a monogenic diagnosis had significantly less polygenic risk than those without, supporting a liability threshold model(5). A polygenic score for neurodevelopmental conditions showed only a direct genetic effect. By contrast, polygenic scores for educational attainment and cognitive performance showed no direct genetic effect, but the non-transmitted alleles in the parents were correlated with the child's risk, potentially due to indirect genetic effects and/or parental assortment for these traits(4). Indeed, as expected under parental assortment, we show that common variant predisposition for neurodevelopmental conditions is correlated with the rare variant component of risk. These findings indicate that future studies should investigate the possible role and nature of indirect genetic effects on rare neurodevelopmental conditions, and consider the contribution of common and rare variants simultaneously when studying cognition-related phenotypes.
Background The degree of gene and sequence preservation across species provides valuable insights into the relative necessity of genes from the perspective of natural selection. Here, we developed novel interspecies metrics across 462 mammalian species, GISMO (Gene identity score of mammalian orthologs) and GISMO-mis (GISMO-missense), to quantify gene loss traversing millions of years of evolution. GISMO is a measure of gene loss across mammals weighed by evolutionary distance relative to humans, whereas GISMO-mis quantifies the ratio of missense to synonymous variants across mammalian species for a given gene. Rationale Despite large sample sizes, current human constraint metrics are still not well calibrated for short genes. Traversing over 100 million years of evolution across hundreds of mammals can identify the most essential genes and improve gene-disease association. Beyond human genetics, these metrics provide measures of gene constraint to further enable mammalian genetics research. Results Our analyses showed that both metrics are strongly correlated with measures of human gene constraint for loss-of-function, missense, and copy number dosage derived from upwards of a million human samples, which highlight the power of interspecies constraint. Importantly, neither GISMO nor GISMO-mis are strongly correlated with coding sequence length. Therefore both metrics can identify novel constrained genes that were too small for existing human constraint metrics to capture. We also found that GISMO scores capture rare variant association signals across a range of phenotypes associated with decreased fecundity, such as schizophrenia, autism, and neurodevelopmental disorders. Moreover, common variant heritability of disease traits are highly enriched in the most constrained deciles of both metrics, further underscoring the biological relevance of these metrics in identifying functionally important genes. We further showed that both scores have the lowest duplication and deletion rate in the most constrained deciles for copy number variants in the UK Biobank, suggesting that it may be an important metric for dosage sensitivity. We additionally demonstrate that GISMO can improve prioritization of recessive disorder genes and captures homozygous selection. Conclusions Overall, we demonstrate that the most constrained genes for gene loss and missense variation capture the largest fraction of heritability, GISMO can help prioritize recessive disorder genes, and identify the most conserved genes across the mammalian tree. ![Figure][1] ### Competing Interest Statement The authors have declared no competing interest. [1]: pending:yes
While the role of de novo and recessively-inherited coding variation in risk for rare developmental disorders (DDs) has been well established, the contribution of damaging variation dominantly-inherited from parents is less explored. Here, we investigated the contribution of rare coding variants to DDs by analyzing 13,452 individuals with DDs, 18,613 of their family members, and 3,943 controls using a combination of family-based and case/control analyses. In line with previous studies of other neuropsychiatric traits, we found a significant burden of rare (allele frequency < 1x10-5) predicted loss-of-function (pLoF) and damaging missense variants, the vast majority of which are inherited from apparently unaffected parents. These predominantly inherited burdens are strongest in DD-associated genes or those intolerant of pLoF variation in the general population, however we estimate that ~10% of the excess of these variants in DD cases is found within the DD-associated genes, implying many more risk loci are yet to be identified. We found similar, but attenuated, burdens when comparing the unaffected parents of individuals with DDs to controls, indicating that parents have elevated risk of DDs due to these rare variants, which are overtransmitted to their affected children. We estimate that 6-8.5% of the population attributable risk for DDs are due to rare pLoF variants in those genes intolerant of pLoF variation in the general population. Finally, we apply a Bayesian framework to combine evidence from these analyses of rare, mostly-inherited variants with prior de novo mutation burden analyses to highlight an additional 25 candidate DD-associated genes for further follow up. ### Competing Interest Statement K.E.S. has received support from Microsoft for work related to rare disease diagnostics. E.J.G. is an employee of and holds shares in Insmed Incorporated. M.E.H. is a co-founder of, consultant to and holds shares in Congenica, a genetics diagnostic company. The remaining authors declare no competing interests. ### Funding Statement This study makes use of DECIPHER, which is funded by the Wellcome Trust. This research was funded in part by Wellcome (grant no. 220540/Z/20/A, Wellcome Sanger Institute Quinquennial Review 2021 - 2026). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The study was approved by the UK Research Ethics Committee (10/H0305/83 granted by the Cambridge South Research Ethics Committee, and GEN/284/12 granted by the Republic of Ireland Research Ethics Committee) I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Sequence and variant-level data and phenotype data from the DDD study data are available on the European Genome-phenome Archive (EGA; https://www.ebi.ac.uk/ega/) with study ID EGAS00001000775. Exome sequencing for the INTERVAL cohort is also available on EGA with study ID EGAD00001002221. Previously described databases were from the Genome Aggregation Database (gnomAD v2.1.1; https://gnomad.broadinstitute.org/downloads) and the Developmental Disorders Genotype-Phenotype Database (DDG2P; https://www.ebi.ac.uk/gene2phenotype/downloads).
BackgroundOne of the major hurdles in clinical genetics is interpreting the clinical consequences associated with germline missense variants in humans. Recent significant advances have leveraged natural variation observed in large-scale human populations to uncover genes or genomic regions that show a depletion of natural variation, indicative of selection pressure. We refer to this as "genetic constraint". Although existing genetic constraint metrics have been demonstrated to be successful in prioritising genes or genomic regions associated with diseases, their spatial resolution is limited in distinguishing pathogenic variants from benign variants within genes.MethodsWe aim to identify missense variants that are significantly depleted in the general human population. Given the size of currently available human populations with exome or genome sequencing data, it is not possible to directly detect depletion of individual missense variants, since the average expected number of observations of a variant at most positions is less than one. We instead focus on protein domains, grouping homologous variants with similar functional impacts to examine the depletion of natural variations within these comparable sets. To accomplish this, we develop the Homologous Missense Constraint (HMC) score. We utilise the Genome Aggregation Database (gnomAD) 125 K exome sequencing data and evaluate genetic constraint at quasi amino-acid resolution by combining signals across protein homologues.ResultsWe identify one million possible missense variants under strong negative selection within protein domains. Though our approach annotates only protein domains, it nonetheless allows us to assess 22% of the exome confidently. It precisely distinguishes pathogenic variants from benign variants for both early-onset and adult-onset disorders. It outperforms existing constraint metrics and pathogenicity meta-predictors in prioritising de novo mutations from probands with developmental disorders (DD). It is also methodologically independent of these, adding power to predict variant pathogenicity when used in combination. We demonstrate utility for gene discovery by identifying seven genes newly significantly associated with DD that could act through an altered-function mechanism.ConclusionsGrouping variants of comparable functional impacts is effective in evaluating their genetic constraint. HMC is a novel and accurate predictor of missense consequence for improved variant interpretation.
The fields of autism and neurodevelopmental disorder (NDD) genetics are rapidly advancing. Catalyzed by the power of large cohorts and integration of all classes of de novo and inherited protein-coding variation, dozens of genes have emerged to harbor variants that confer high relative risk for autism, and hundreds of genes have been associated with NDDs more broadly. Through examination of protein-truncating variants (PTVs), predicted damaging missense variation, and copy number variants (CNVs), our prior analyses have begun to map the allelic diversity of perturbations within 72 autism-associated genes and 373 genes associated with NDDs, finding intriguing evidence of genes with significantly higher mutation rates and differences in the distribution of clinical phenotypes in autism compared to NDD (Fu et al., 2022; Satterstrom et al., 2020). Despite this progress, cohort sizes remain insufficient for disentangling the shared and distinct genetic architectures of autism, NDDs, and other neuropsychiatric conditions, as well as associating genes with more subtle impacts on neurodevelopment.To advance these boundaries, we present the largest to-date study of rare coding variants, consisting of 62,013 autistic individuals, including 38,088 probands and 9,567 unaffected siblings from complete trio and quartet families, respectively, and 23,925 additional autism cases without parental information contrasted against 26,931 controls. By aggregating across the Autism Sequencing Consortium (ASC), the Simons Simplex Collection (SSC), the Simons Foundation Powering Autism Research (SPARK), and individuals from a leading diagnostic laboratory (GeneDx), this dataset totals almost 200,000 individuals, nearly a three-fold increase over prior studies. When we stratified the clinically-referred GeneDx autistic probands by co-occurring DD/ID status, we found synonymous, missense, and PTV de novo mutation rates in autism probands without DD/ID from GeneDx that were nearly identical to individuals ascertained for a diagnosis of autism in the ASC, SSC, and SPARK research studies (0.296 vs 0.294, 0.767 vs 0.763, and 0.141 vs 0.145 respectively), while GeneDx autism probands with DD/ID exhibited mutation rates similar to those observed in previous research studies of DD.Further analyses of these data solidified previous observations of significant enrichment of de novo PTVs among autism probands of 3x compared to siblings among the genes most intolerant to PTVs in the human genome (i.e., lowest decile of LOEUF from gnomAD). We have also incorporated Alpha Missense (AM) pathogenicity estimates to complement our prior MPC scores for predicting damaging missense variation and identifying de novo missense variants acting with effect sizes comparable to de novo PTVs in constrained genes, with analysis of regional missense constraint within genes ongoing. We further leveraged the TADA Bayesian statistical method to jointly model these data in a single unified framework, leveraging genetic information across rare PTVs, damaging missense variants, and CNVs. This approach discovered hundreds of genes associated with autism, where we observe a steadily increasing contribution of variant classes other than de novo PTVs in newly associated genes. Analyses are ongoing to understand the gene networks, developmental timing, and biological functions by which these genes exert their influence on phenotypic manifestations of autism and related neuropsychiatric disorders.
Autism is highly heritable and has been associated with multiple classes of genetic variation. Common genetic variation contributes substantially to autism. Previously, with 18,381 autistic individuals and 27,969 non-autistic individuals, five genome-wide significant loci were identified. Now with 38,717 autistic individuals and 232,725 non-autistic individuals, we report an updated genome-wide association study (GWAS) of autism with 12 genome-wide significant loci. We observe a moderate genetic correlation (0.675, SE=0.0434) between Europe-based (Nautistic=22,643; Nnon-autistic=204,389) and United States-based (Nautistic =16,074; Nnon-autistic=28,346) autism cohorts, which contributes to the decline of the estimated single nucleotide polymorphism (SNP) heritability (from 0.118 (SE=0.010) to 0.068 (SE=0.003)). The genetic correlation between autism with intellectual disability (ID) (Nautistic=6,590; Nnon-autistic= 43,071; h2=0.062; SE=0.012) and autism without ID (Nautistic=23,173; Nnon-autistic= 204,679; h2=0.089; SE=0.005) is 0.658 (SE=0.086). In the United States family-based cohorts, the genetic correlation between autism with ID (Nfamily=3,993; h2=0.159; SE=0.033) and autism without ID (Nfamily=4,357; h2=0.171; SE=0.031) is 0.812 (SE=0.157). Autism without ID was positively genetically correlated with educational attainment (0.163; P=4.84 × 10-11) and intelligence (0.233; P=1.95 × 10-11). Autism with ID genetically correlated with neither educational attainment (0.036; P=0.409) nor intelligence (-0.072; P=0.235). As ID alone is negatively genetically correlated with intelligence, the lack of correlation between autism with ID and intelligence strongly suggests that autism with ID is genetically different from ID alone. This difference has implications for both research and clinical nosology. Rare and de novo variants contribute substantially to autism in some individuals. Through rare variant analyses, 72 genes have been associated with autism at a genome-wide significant level to date. While de novo protein truncating variants (PTVs) and copy number deletions have been associated with autism, we report preliminary findings that the burden of inherited PTVs and copy number deletions among autistic individuals was elevated compared to their non-autistic siblings (P=4.00 × 10-5). Integration of multiple genetic factors will help us better understand the etiology of autism.