Many African great ape chromosomes possess large subterminal heterochromatic caps at their telomeres that are conspicuously absent from the human lineage. Leveraging the complete sequences of great ape genomes, we characterize the organization of subterminal caps and reconstruct the evolutionary history of these regions in chimpanzees and gorillas. Detailed analyses of the composition of the associated terminal 32 bp satellite array from chimpanzee (termed pCht) and intervening segmental duplication (SD) spacers confirm two independent origins in the Pan and gorilla lineages. In chimpanzee and bonobo, we estimate these structures emerged ∼7.7 million years ago (MYA) in contrast to gorilla, in which they expanded more recently, ∼5.0 MYA, and now make up 8.5% of the total gorilla genome. In both lineages, the SD spacers punctuating the pCht heterochromatic satellite arrays correspond to pockets of decreased methylation, although in gorilla such regions are significantly less methylated (P < 2.2 × 10-16) than in chimpanzee or bonobo. Allelic pairs of subterminal caps show a higher degree of sequence divergence than euchromatic sequences, with bonobo showing less divergent haplotypes and less differentially methylated spacers. In contrast, we identify virtually identical subterminal caps mapping to nonhomologous chromosomes within a species, suggesting ectopic recombination potentially mediated by SD spacers. We find that the transition regions from heterochromatic subterminal caps to euchromatin are enriched for structural variant insertions and lineage-specific duplicated genes. Our findings suggest independent evolution of subterminal caps converging on a common genetic and epigenetic structure that promoted ectopic exchange as well as the emergence of novel genes at transition regions between euchromatin and heterochromatin.
Long-read sequencing (LRS) and diploid genome assembly have enabled nearly complete structural variant (SV) discovery. Using 293 nearly complete genomes, we characterize the full spectrum of genetic variation and show that while 99% of the variants between any two genomes are single base-pair substitutions, 88% of the euchromatic variant base pairs are SVs, including insertions, deletions, duplications, and inversions. We identify 24 gene-rich regions subject to megabase-scale variation, 2,293 potentially unstable tandem repeats, and 890 novel expression quantitative trait loci associated with SVs in humans. Expanding to 1,218 LRS samples from the 1000 Genomes Project and applying a newly developed cross-platform breakpoint evaluation tool, BoostSV, we construct a nonredundant callset comprising 614,522 SVs. We demonstrate the utility of this population-level SV reference callset by filtering >99% of the common variation from 44 unsolved LRS probands from the Undiagnosed Diseases Network to discover likely disease-causing SVs. Second, we genotype 1,053 high-impact biallelic SVs from the pangenome callset in 232,090 samples from All of Us and discover 105 SVs with significant associations, including 26% where the SV is the lead variant. This publicly available pangenome SV resource will drive new disease associations and further our understanding of the missing heritability of human genetic disease.
The human Y chromosome is among the most structurally dynamic chromosomes in the human genome, yet much of its diversity remains unresolved because of extensive palindromes, ampliconic gene families, satellite-rich heterochromatin and large segmental duplications. What remained unclear was how these diverse forms of variation fit together across the full chromosome, how often similar structures recur in different lineages, and which aspects of organization remain constrained despite rapid sequence turnover. Here, we generated and analyzed 142 nearly complete human Y chromosome assemblies from 17 major haplogroups spanning approximately 180,000 years of evolution, creating a population-scale resource for studying Y chromosome biology and diversity. These assemblies show that structural change on the Y chromosome is recurrent but constrained, even in its most repetitive regions. In the fertility-associated azoospermia factor c (AZFc) region, recurrent inversions, deletions, and complex rearrangements generate a limited repertoire of structural haplotypes. Multicopy ampliconic gene families follow distinct evolutionary paths: DAZ paralogues differ in structural constraint, RBMY evolves within a modular array, and TSPY copy number varies mainly through local expansion and contraction. The centromere and Yq12 heterochromatin vary greatly in size but retain a stable higher-order organization, including a single hypomethylated centromeric core and conserved Yq12 repeat composition and orientation. Methylation across palindromic and ampliconic regions is likewise structured by repeat class, copy order and local architecture. Together, these results provide a population-scale resource for the human Y chromosome and show that its rapid structural evolution is repeatedly funneled into a limited set of architectural outcomes.
Genetic introgression from Neanderthals and Denisovans shaped modern human genomes; however, introgressed structural variants (SVs ≥ 50 base pairs) remain challenging to discover. We integrated high-quality phased assemblies from four new Papua New Guinea (PNG) haploid genomes with 94 published assemblies of diverse ancestry to infer an introgressed SV map. Introgressed SVs are enriched in genes (47%), including critical genomic disorder regions, and are most abundant in PNG genomes. We identified 11 centromeres likely derived from archaic hominins, adding unexplored diversity to centromere genomics. Pangenome genotyping of these 98 assemblies across 1363 samples revealed 16 adaptive SVs, many associated with immune-related genes and expression, in the PNG genomes. We hypothesize that archaic SVs contributed to reproductive success, underscoring introgression as a major force in human adaptive evolution.
NOTCH2NL (NOTCH2-N-terminus-like) genes arose from ape-specific chromosome 1 segmental duplications implicated in human brain cortical expansion, including an incomplete NOTCH2 gene. Genetic characterization of these loci and their regulation is complicated because they are embedded in large, nearly identical duplications that predispose to recurrent microdeletion syndromes. Using near-complete long-read assemblies generated from 70 human and 12 ape haploid genomes, we show independent recurrent duplication among apes with protein-coding copies emerging in humans 2.2-3.7 million years ago. We distinguish NOTCH2NL paralogs present in every human haplotype (NOTCH2NLA) from copy-number-variable ones. We also characterize large-scale structural variation, including gene conversion, for 28% of haplotypes, leading to a previously undescribed paralog, NOTCH2tv. Finally, we apply Fiber-seq and long-read transcript sequencing to human dorsal forebrain organoids to characterize the regulatory landscape and find that the most fixed paralogs, NOTCH2 and NOTCH2NLA, harbor the greatest number of paralog-specific elements potentially driving their regulation.
Abstract Long-read sequencing improves sensitivity to discover variation in complex repetitive regions, assign parent-of-origin, and distinguish germline from postzygotic mutations. We applied Illumina, Oxford Nanopore Technologies, and PacBio sequencing to discover and validate de novo mutations in 73 children from 42 autism families (157 individuals). We assay 2.77 Gbp of the human genome, yielding on average 95 de novo mutations per transmission (87.5 single-nucleotide substitutions, 7.8 indels), with no significant difference in mutation rate or profile between probands and their unaffected siblings. Long reads increase de novo mutation discovery by 20-40% and double the mutations classified as early embryonic. The germline mutation rate is 1.30×10−8 substitutions/base pair/generation; the postzygotic rate is 0.23×10−8. These rates are significantly increased in repetitive DNA, where segmental duplication mutability is dependent on length and percent identity. Here, we show that enrichment in repeats occurs predominantly postzygotically, likely resulting from faulty DNA repair and interlocus gene conversion.
Complete, haplotype-resolved genome assemblies have provided unprecedented insight into the evolution of structurally complex, rapidly evolving regions of human genomes; however, population-scale pangenome resources of our closest relatives, chimpanzees and bonobos (genus, Pan), are necessary to ascertain the origins and evolutionary context of these loci. Here, we sequence and assemble 58 haplotypes from four distinct Pan clades to high contiguity (median contig NG50=54 Mb), including eight near-T2T genomes. These genomes reveal previously intractable genetic variation increasing estimates of genome-wide genetic diversity 6-37% across populations compared to short-read estimates. We identify recurrent structural polymorphisms across species impacting genes associated with immune response and host-pathogen interaction and find that structural variants (SVs) are 170- to 260-fold more likely than single nucleotide variants (SNVs) to exhibit high-impact effects across species. Contrasting SV patterns across primates we find that transposable element mutation rates differ by as much as threefold between species. We show that human disease-associated short tandem repeat (TR) loci have uniquely expanded in humans sensitizing our species to these TR-expansion disorders. Physically phased haplotypes enable reconstruction of genome-wide genealogical histories, uncovering ancient, functional genetic variation maintained by balancing selection, as well as signatures of recent adaptation in chimpanzee subspecies. Several malaria-associated loci exhibit ancient structural polymorphism, including the African great ape-specific glycophorin (GYP) gene expansion. We characterize the sequence, structure, and composition of diverse glycophorin haplotypes in humans and chimpanzees. We identify independent malaria-protective GYPA-B fusion events in humans and novel chimpanzee glycophorin genes resulting from both ancient and recent fusion events demonstrating parallel adaptations to pathogen resistance across hominins. Together, our resource highlights the critical importance of nonhuman primate population-scale pangenomics for understanding the evolution of complex genome structures and the biodiversity of our endangered closest living relatives.
IntroductionDYRK1A, a protein kinase located on human chromosome 21, plays a role in postembryonic neuronal development and degeneration. Alterations to DYRK1A have been consistently associated with cognitive functioning and neurodevelopmental disorders (e.g., autism, intellectual disability). However, the broader cognitive and behavioral phenotype of DYRK1A syndrome requires further characterization. Specifically, executive functioning, or cognitive processes that are necessary for goal-directed behavior, has not yet been characterized in this population.MethodsIndividuals with DYRK1A variants (n = 29; ages 4 to 21 years) were assessed with a standardized protocol with multiple measures of executive functioning: Delis-Kaplan Executive Function Schedule, and chronologically age-appropriate caregiver-report forms of the Behavior Rating Inventory of Executive Function (BRIEF) and Achenbach System of Empirically Based Assessment (ASEBA). We first examined the feasibility and appropriateness of established executive functioning measures among participants with DYRK1A syndrome to inform selection of executive functioning tools in future research. We then characterized executive functioning among the group, including associations with other phenotypic features.ResultsNeurocognitive assessments of executive functioning were deemed infeasible due to cognitive and verbal functioning. Caregiver-report revealed elevated executive functioning concerns related to self-monitoring, working memory, and planning/organization on the BRIEF, and attention and ADHD on the CBCL. Only two participants had existing ADHD diagnoses; however, 5 participants (out of 10 participants with data) exceeded the cutoff on the BRIEF, 13 individuals (out of 27 with data) exceeded the cutoff on the ASEBA ADHD subscale, and 18 exceeded the cutoff on the ASEBA attention subscale. There was concordance between ADHD diagnosis and the ASEBA, but not BRIEF. Executive functioning was correlated with nonverbal IQ and autism traits.DiscussionObjective measures of executive functioning are needed for individuals with intellectual disability who are nonverbal and/or have motor limitations. Diagnostic overshadowing, or the tendency to attribute all problems to intellectual disability and to leave other co-existing conditions, such as executive functioning challenges or ADHD, undiagnosed, is common. Phenotypic characterization of executive functioning is therefore important for our understanding of DYRK1A syndrome and for ensuring that caregivers’ concerns are addressed, and individuals receive the clinical services that best meet their needs.
The crab-eating macaques (Macaca fascicularis) and rhesus macaques (Macacamulatta) are pivotal in biomedical and evolutionary research1, 2-3. However, their genomic complexity and interspecies genetic differences remain unclear4. Here, we present a complete genome assembly of a crab-eating macaque, revealing 46% fewer segmental duplications and 3.83 times longer centromeres than those of humans5,6. We also characterize 93 large-scale genomic differences between macaques and humans at a single-base-pair resolution, highlighting their impact on gene regulation in primate evolution. Using ten long-read macaque genomes, hundreds of short-read macaque genomes and full-length transcriptome data, we identified roughly 2 Mbp of fixed-genetic variants, roughly 240 Mbp of complex loci, 16.76 Mbp genetic differentiation regions and 110 alternative splice events, potentially associated with various phenotypic differences between the two macaque species. In summary, the integrated genetic analysis enhances understanding of lineage-specific phenotypes, adaptation and primate evolution, thereby improving their biomedical applications in human disease research.
Mutations in ADNP (Activity-Dependent Neuroprotective Protein) are among the most frequent monogenic causes of autism spectrum disorder (ASD) and lead to Helsmoortel-Van der Aa syndrome (HVDAS). Yet how ADNP dysfunction leads to HVDAS is unclear. We employed patient-derived induced pluripotent stem cells, cortical organoids and ADNP KO human neural stem cells (hNSCs) to clarify the cellular and molecular mechanism of HVDAS onset. We purified an ADNP-KDM1A-GTF2I (AKG) protein complex from hNSCs and show that it targets transposable elements (TEs) to repress nearby gene transcription. Upon ADNP KO, KDM1A binding is lost at promoters targeted by AKG, pointing to ADNP as the anchoring subunit of the AKG complex. HVDAS cortical organoids show impaired progenitor proliferation and accelerated neuronal differentiation, coupled with a sustained upregulation of neurogenesis transcriptional programs, including key transcription factors normally repressed by AKG. This work suggests that the AKG complex acts as the relevant ADNP unit in the molecular onset of HVDAS. ### Competing Interest Statement E.E.E. is a scientific advisory board (SAB) member of Variant Bio, Inc. The other authors declare no competing interests.
Rare diseases are collectively common, affecting approximately 1 in 20 individuals worldwide. In recent years, rapid progress has been made in rare disease diagnostics due to advances in next-generation sequencing, development of new computational and functional genomics approaches to prioritize genes and variants and increased global sharing of clinical and genetic data. However, more than half of individuals suspected to have a rare disease lack a genetic diagnosis. The Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) Consortium was initiated to study thousands of challenging rare disease cases and families and apply, standardize and evaluate emerging genomics technologies and analytics to accelerate their adoption in clinical practice. Furthermore, all data generated, currently representing over 7,500 individuals from over 3,000 families, are rapidly made available to researchers worldwide through the Analysis, Visualization and Informatics Lab-space (AnVIL) to catalyse global efforts to develop approaches for genetic diagnoses in rare diseases. Most of these families have undergone previous clinical genetic testing but remained unsolved, with most being exome-negative. Here we describe the collaborative research framework, datasets and discoveries comprising GREGoR that will provide foundational resources and substrates for the future of rare disease genomics.
Understanding the human de novo mutation (DNM) rate requires complete sequence information1. Here using five complementary short-read and long-read sequencing technologies, we phased and assembled more than 95% of each diploid human genome in a four-generation, twenty-eight-member family (CEPH 1463). We estimate 98-206 DNMs per transmission, including 74.5 de novo single-nucleotide variants, 7.4 non-tandem repeat indels, 65.3 de novo indels or structural variants originating from tandem repeats, and 4.4 centromeric DNMs. Among male individuals, we find 12.4 de novo Y chromosome events per generation. Short tandem repeats and variable-number tandem repeats are the most mutable, with 32 loci exhibiting recurrent mutation through the generations. We accurately assemble 288 centromeres and six Y chromosomes across the generations and demonstrate that the DNM rate varies by an order of magnitude depending on repeat content, length and sequence identity. We show a strong paternal bias (75-81%) for all forms of germline DNM, yet we estimate that 16% of de novo single-nucleotide variants are postzygotic in origin with no paternal bias, including early germline mosaic mutations. We place all this variation in the context of a high-resolution recombination map (~3.4 kb breakpoint resolution) and find no correlation between meiotic crossover and de novo structural variants. These near-telomere-to-telomere familial genomes provide a truth set to understand the most fundamental processes underlying human genetic variation.
Autism is highly heritable and diagnosed more frequently in males than females. To identify neurodevelopmental processes that might present sex-biased vulnerability, we generated transcriptomic and epigenomic profiles of cell types present in the prenatally developing human cerebral cortex of 27 males and 21 females. By intersecting sex-biased molecular signatures and genes with de novo mutations in male and female autistic probands, we reveal two points of vulnerability contributing to the sex-biased penetrance in neurodevelopmental disorders (NDDs). First, we show that NDD risk genes are biased towards higher expression in females, identifying the NDD gene MEF2C as a critical transcription factor for female-biased expression. Second, we identify a significant contribution of X chromosome genes to NDD pathobiology. We construct a gene regulatory map of X-linked risk genes to enable functional studies of genetic variants that likely disrupt gene expression in the developing brains of autistic males. Together, these results point towards an outsized contribution of the X-chromosome to both the origin of sex differences in the developing human cortex and NDD vulnerability. We propose a model where female-biased vulnerability is driven by coding variation within genes while male-biased vulnerability is driven by noncoding variation in regulatory elements that affect gene expression.
The attachment of the kinetochore to the centromere is essential for genome maintenance, yet the highly repetitive nature of satellite regional centromeres limits our understanding of their chromatin organization. We demonstrate that single-molecule chromatin fiber sequencing (Fiber-seq) can uniquely co-resolve kinetochore and surrounding chromatin architectures along point centromeres, revealing largely homogeneous single-molecule kinetochore occupancy. In contrast, the application of Fiber-seq to regional centromeres exposed marked per-molecule heterogeneity in their chromatin organization. Regional centromere cores uniquely contain a dichotomous chromatin organization (dichromatin) composed of compacted nucleosome arrays punctuated with highly accessible chromatin patches. CENP-B occupancy phases dichromatin to the underlying alpha-satellite repeat within centromere cores but is not necessary for dichromatin formation. Centromere core dichromatin is conserved between humans and primates, including along regional centromeres lacking satellite repeats. Overall, the chromatin organization of regional centromeres is defined by marked per-molecule heterogeneity, buffering kinetochore attachment against sequence and structural variability within regional centromeres.
The most common genomic disorder, chromosome 22q11.2 microdeletion syndrome (22q11.2DS), is mediated by highly identical and polymorphic segmental duplications (SDs) known as low copy repeats (LCRs; regions A-D) that have been challenging to sequence and characterize. Here, we report the sequence-resolved genomic architecture of 135 chromosome 22q11.2 haplotypes from diverse 1000 Genomes Project samples. We find that more than 90% of the copy number variation is polarized to the most proximal LCR region A (LCRA) where 50 distinct structural configurations are observed (~189 kbp to ~2.15 Mbp or 11-fold length variation). A higher-order SD cassette structure of 105 kbp in length, flanked by 25 kbp long inverted repeats, drives this variation and emerged in the human-chimpanzee ancestral lineage later expanding in humans ~1.0 [0.8-1.2] million years ago. African LCRA haplotypes are significantly longer (p=0.0047) when compared to non-Africans yet are predicted to be more protected against recurrent microdeletions (p=0.00053) due to a preponderance of flanking SDs in an inverted orientation. Conversely, we identified nine distinct inversion polymorphisms, including five recurrent ~2.28 Mbp inversions extending across the critical region (LCRA-D) and four smaller inversions (two LCRA-B, one LCRC-D, and one LCRB-D); 7/9 of these events were identified in haplotypes of African and admixed American ancestry. Finally, we sequence and assemble four families and show that LCRA-D deletion breakpoints map to the 105 kbp repeat unit while inversion breakpoints associate with the 25 kbp repeats adjacent to palindromic AT-rich regions. In one family, we observe evidence of more complex unequal crossover events associated with gene conversion and multiple breakpoints. Our findings suggest that specific haplotype configurations are protective and susceptible to chromosome 22q11.2DS while recurrent large-scale inversions help to explain why this syndrome is less prevalent among individuals of African descent.
All great apes differ karyotypically from humans due to the fusion of chromosomes 2a and 2b, resulting in human chromosome 2. Here, we show that the fusion was associated with multiple pericentric inversions, segmental duplications (SDs), and the turnover of subterminal repetitive DNA. We characterized the fusion site at the single-base-pair resolution and identified three distinct SDs that originated more than 5 million years ago. These three distinct SDs were differentially distributed among African great apes as a result of incomplete lineage sorting (ILS) and lineage-specific duplication. One of these SDs shares homology to a hypomethylated SD spacer sequence present in the subterminal heterochromatin of Pan but is completely absent subtelomerically in both humans and orangutans. CRISPR-Cas9-mediated depletion of the fusion site in human neural progenitor cells alters the expression of genes, indicating a potential regulatory consequence to this human-specific karyotypic change. Overall, this study offers insights into how complex regions subject to ILS may contribute to speciation.
Down syndrome is the most common form of human intellectual disability caused by precocious segregation and nondisjunction of chromosome 21. Differences in centromere structure have been hypothesized to play a potential role in this process in addition to the well-established risk of advancing maternal age. Using long-read sequencing, we completely sequenced and assembled the centromeres from a parent-child trio where Trisomy 21 arose in the child as a result of a meiosis I error. The proband carries three distinct chromosome 21 centromere haplotypes that vary by 11-fold in length--both the largest (H1) and smallest (H2) originating from the mother. The longest H1 allele harbors a less clearly defined centromere dip region (CDR) as defined by CpG methylation and a significantly reduced signal by CENP-A chromatin immunoprecipitation sequencing when compared to H2 or paternal H3 centromeres. These epigenetic signatures suggest less competent kinetochore attachment for the maternally transmitted H1. Analysis of H1 in the mother indicates that the reduced CENP-A ChIP-seq signal, but not the CDR profile, pre-existed the meiotic nondisjunction event. A comparison of the three proband centromeres to a population sampling of 35 completely sequenced chromosome 21 centromeres shows that H2 is the smallest centromere sequenced to date and all three haplotypes (H1-H3) share a common origin of ~15 thousand years ago. These results suggest that recent asymmetry in size and epigenetic differences of chromosome 21 centromeres may contribute to nondisjunction risk.