Mitochondrial sequences have integrated into the nuclear genome since the origin of eukaryotes. Recent insertions that retain homology to extant mitochondrial DNA (mtDNA), termed NUMTs, confound mtDNA sequence analysis. Here, we use great ape Telomere-to-Telomere (T2T) genomes to study NUMTs in bonobo, chimpanzee, human, gorilla, and Bornean and Sumatran orangutans. A phylogeny based on shared and lineage-specific NUMTs accurately recapitulates the great ape species tree topology. NUMTs are enriched at nonfunctional nonrepetitive regions of the nuclear genome, and depleted within enhancers and coding sequences, suggesting negative selection. We validate the presence of a 76-kilobase-long heterozygous NUMT in chimpanzee, which is larger than any other NUMT observed in great apes, and find that dozens of NUMTs on the Pan Y Chromosome expanded together with palindromes. Finally, by analyzing intraspecific variation, we confirm that the vast majority of species-specific NUMTs identified in T2T assemblies are fixed or present at high frequencies in each species. Our study highlights NUMTs as a dynamic evolutionary force contributing to shaping ape genomes, and is valuable for characterizing mtDNA in great apes.
Oocytes are densely packed with mitochondria, the energy-producing organelles that contain their own genome, mitochondrial DNA (mtDNA). Each cell contains multiple copies of mtDNA, with copy number varying among tissue types. Oocytes possess the highest mtDNA copy number, containing hundreds of thousands of mtDNA molecules per cell. Because mitochondria are inherited exclusively through the maternal lineage, accurate detection of mtDNA variants is essential for studies of inheritance, aging, and disease. The presence of multiple mtDNA copies allows wild-type and mutant molecules to coexist within the same cell, a condition known as heteroplasmy, in which low-frequency and de novo variants may occur at frequencies below 1%. Conventional next-generation sequencing (NGS) lacks sufficient accuracy to reliably distinguish these rare variants from errors introduced during library preparation and sequencing. Here, we present a protocol for enriching mtDNA from single human oocytes using Exonuclease V to remove linear DNA, followed by duplex sequencing library preparation for highly accurate mtDNA analysis. This workflow enables error-corrected sequencing of individual oocytes, facilitating reliable detection of low-frequency mtDNA variants and analysis of heteroplasmy and de novo mutagenesis. The protocol provides a reproducible approach for investigating mitochondrial genome variation in single oocytes using Illumina-compatible sequencing platforms.
Bird genomes are the smallest among amniotes; however, they remain challenging to assemble due to their structural complexity. This study presents a fully phased diploid telomere-to-telomere reference genome for the zebra finch (Taeniopygia guttata), a model organism for neuroscience and evolutionary genomics. Combining sequencing strategies allowed the closing of nearly all gaps, adding ∼90 Mbp of sequence (7.8%). The assembly contains gapless sequences for all microchromosomes, including the elusive dot chromosomes with their distinctive architecture. The genome was comprehensively annotated for genes, repeats, and structural variants. Complete centromeres were identified, along with candidate DNA-binding centromere protein B (CENP-B) homologs, suggesting that birds possess a kinetochore-associated CENP system similar to that of mammals. Relative to the previous reference generated by the Vertebrate Genomes Project, 2,710 (8.13%) previously unassembled or unannotated genes were identified. This complete genome of a songbird serves as a public reference and illuminates avian genome architecture and function.
Non-canonical (non-B) DNA structures-e.g., bent DNA, hairpins, G-quadruplexes (G4s), Z-DNA, etc.-which form at certain sequence motifs (e.g., A-phased repeats, inverted repeats, etc.), have emerged as important regulators of cellular processes and drivers of genome evolution. Yet, they have been understudied due to their repetitive nature and potentially inaccurate sequences generated with short-read technologies. Here we comprehensively characterize such motifs in the long-read telomere-to-telomere (T2T) genomes of human, bonobo, chimpanzee, gorilla, Bornean orangutan, Sumatran orangutan, and siamang. Non-B DNA motifs are enriched at the genomic regions added to T2T assemblies, and occupy 9-15%, 9-11%, and 12-38% of autosomes, and chromosomes X and Y, respectively. G4s and Z-DNA are enriched at promoters and enhancers, as well as at origins of replication. Repetitive sequences harbor more non-B DNA motifs than non-repetitive sequences, especially in the short arms of acrocentric chromosomes. Most centromeres and/or their flanking regions are enriched in at least one non-B DNA motif type, consistent with a potential role of non-B structures in determining centromeres. Our results highlight the uneven distribution of predicted non-B DNA structures across ape genomes and suggest their novel functions in previously inaccessible genomic regions.
The most dynamic and repetitive regions of great ape genomes have traditionally been excluded from comparative studies 1–3 . Consequently, our understanding of the evolution of our species is incomplete. Here we present haplotype-resolved reference genomes and comparative analyses of six ape species: chimpanzee, bonobo, gorilla, Bornean orangutan, Sumatran orangutan and siamang. We achieve chromosome-level contiguity with substantial sequence accuracy (<1 error in 2.7 megabases) and completely sequence 215 gapless chromosomes telomere-to-telomere. We resolve challenging regions, such as the major histocompatibility complex and immunoglobulin loci, to provide in-depth evolutionary insights. Comparative analyses enabled investigations of the evolution and diversity of regions previously uncharacterized or incompletely studied without bias from mapping to the human reference genome. Such regions include newly minted gene families in lineage-specific segmental duplications, centromeric DNA, acrocentric chromosomes and subterminal heterochromatin. This resource serves as a comprehensive baseline for future evolutionary studies of humans and our closest living ape relatives.
BACKGROUND:G-quadruplexes (G4s) are non-canonical DNA structures that can form at approximately 1% of the human genome. They facilitate genomic instability by increasing point mutations and structural variation. Numerous G4s participate in telomere maintenance and regulating transcription and replication, and evolve under purifying selection. Despite these important functions, G4s have remained under-studied in human and ape genomes due to incomplete assemblies. RESULTS:Here, we conduct a comprehensive analysis of predicted G4s (pG4s) in the recently released, telomere-to-telomere (T2T) genomes of human, bonobo, chimpanzee, gorilla, Bornean orangutan, and Sumatran orangutan. We annotate 41,232-174,442 new pG4s in these T2T compared to previous ape genome assemblies (5%-21% increase). Analyzing inter-species whole-genome alignments, we identify pG4s shared across apes (approximately one-third of all pG4s) and thousands of species-specific pG4s. pG4s accumulate and diverge at rates consistent with divergence times between species, following molecular clock. pG4s shared across apes are enriched and hypomethylated at regulatory regions-enhancers, promoters, UTRs, and origins of replication-suggesting their conserved formation and functions. Species-specific pG4s (constituting 11-27% of all pG4s) are located in regulatory regions, potentially contributing to adaptations, and in repeats, likely driving genome expansions. CONCLUSIONS:Our findings illuminate the evolutionary dynamics of G4s, conservation of their role in gene regulation, and their contributions to ape genome evolution. Our study highlights the utility of high-resolution T2T genomes in revealing elusive yet likely functionally relevant genomic features previously hidden by incomplete assemblies.
Mitochondria, cellular powerhouses, harbor DNA [mitochondrial DNA (mtDNA)] inherited from the mothers. mtDNA mutations can cause diseases, yet whether they increase with age in human oocytes remains understudied. Here, using highly accurate duplex sequencing, we detected de novo mutations in single oocytes, blood, and saliva in women 20 to 42 years of age. We found that, with age, mutations increased in blood and saliva but not in oocytes. In oocytes, mutations with high allele frequencies were less prevalent in coding than noncoding regions, whereas mutations with low allele frequencies were more uniformly distributed along the mtDNA, suggesting frequency-dependent purifying selection. Thus, mtDNA in human oocytes is protected against accumulation of mutations with aging and having functional consequences. These findings are particularly timely as humans tend to reproduce later in life.
G-quadruplexes (G4s) are functional elements of the human genome, some of which inhibit DNA replication. We investigated replication of G4s within highly abundant microsatellite (GGGA, GGGT) and transposable element (L1 and SVA) sequences. We found that genome-wide, numerous motifs are located preferentially on the replication leading strand and the transcribed strand templates. We directly tested replicative polymerase ϵ and δ holoenzyme inhibition at these G4s, compared to low abundant motifs. For all G4s, DNA synthesis inhibition was higher on the G-rich than C-rich strand or control sequence. No single G4 was an absolute block for either holoenzyme; however, the inhibitory potential varied over an order of magnitude. Biophysical analyses showed the motifs form varying topologies, but replicative polymerase inhibition did not correlate with a specific G4 structure. Addition of the G4 stabilizer pyridostatin severely inhibited forward polymerase synthesis specifically on the G-rich strand, enhancing G/C strand asynchrony. Our results reveal that replicative polymerase inhibition at every G4 examined is distinct, causing complementary strand synthesis to become asynchronous, which could contribute to slowed fork elongation. Altogether, we provide critical information regarding how replicative eukaryotic holoenzymes navigate synthesis through G4s naturally occurring thousands of times in functional regions of the human genome.
DNA secondary structures are essential elements of the genomic landscape, playing a critical role in regulating various cellular processes. These structures refer to G-quadruplexes, cruciforms, Z-DNA or H-DNA structures, amongst others (collectively called 'non-B DNA'), which DNA molecules can adopt beyond the B conformation. DNA secondary structures have significant biological roles, and their landscape is dynamic and can rearrange due to various factors, including changes in cellular conditions, temperature, and DNA-binding proteins. Understanding this dynamic nature is crucial for unraveling their functions in cellular processes. Detecting DNA secondary structures remains a challenge. Conventional methods, such as gel electrophoresis and chemical probing, have limitations in terms of sensitivity and specificity. Emerging techniques, including next-generation sequencing and single-molecule approaches, offer promise but face challenges since these techniques are mostly limited to only one type of secondary structure. Here we describe an updated version of a technique permanganate/S1 nuclease footprinting, which uses potassium permanganate to trap single-stranded DNA regions as found in many non-B structures, in combination with S1 nuclease digest and adapter ligation to detect genome-wide non-B formation. To overcome technical hurdles, we combined this method with direct adapter ligation and sequencing (PDAL-Seq). Furthermore, we established a user-friendly pipeline available on Galaxy to standardize PDAL-Seq data analysis. This optimized method allows the analysis of many types of DNA secondary structures that form in a living cell and will advance our knowledge of their roles in health and disease.
Y chromosomes of great apes harbor A mpliconic G enes (YAGs)—multi-copy gene families ( BPY2 , CDY , DAZ , HSFY , PRY , RBMY , TSPY , VCY , and XKRY ) that encode proteins important for spermatogenesis. Previous work assembled YAG transcripts based on their targeted sequencing but not using reference genome assemblies, potentially resulting in an incomplete transcript repertoire. Here we used the recently produced gapless telomere-to-telomere (T2T) Y chromosome assemblies of great ape species (bonobo, chimpanzee, human, gorilla, Bornean orangutan, and Sumatran orangutan) and analyzed RNA data from whole-testis samples for the same species. We generated hybrid transcriptome assemblies by combining targeted long reads (Pacific Biosciences), untargeted long reads (Pacific Biosciences) and untargeted short reads (Illumina)and mapping them to the T2T reference genomes. Compared to the results from the reference-free approach, average transcript length was more than two times higher, and the total number of transcripts decreased three times, improving the quality of the assembled transcriptome. The reference-based transcriptome assemblies allowed us to differentiate transcripts originating from different Y chromosome gene copies and from their non-Y chromosome homologs. We identified two sources of transcriptome diversity—alternative splicing and gene duplication with subsequent diversification of gene copies. For each gene family, we detected transcribed pseudogenes along with protein-coding gene copies. We revealed previously unannotated gene copies of YAGs as compared to currently available NCBI annotations, as well as novel isoforms for annotated gene copies. This analysis paves the way for better understanding Y chromosome gene functions, which is important given their role in spermatogenesis.
Childhood obesity represents a significant global health concern and identifying its risk factors is crucial for developing intervention programs. Many "omics" factors associated with the risk of developing obesity have been identified, including genomic, microbiomic, and epigenomic factors. Here, using a sample of 48 infants, we investigated how the methylation profiles in cord blood and placenta at birth were associated with weight outcomes (specifically, conditional weight gain, body mass index, and weight-for-length ratio) at age six months. We characterized genome-wide DNA methylation profiles using the Illumina Infinium MethylationEpic chip, and incorporated information on child and maternal health, and various environmental factors into the analysis. We used regression analysis to identify genes with methylation profiles most predictive of infant weight outcomes, finding a total of 23 relevant genes in cord blood and 10 in placenta. Notably, in cord blood, the methylation profiles of three genes (PLIN4, UBE2F, and PPP1R16B) were associated with all three weight outcomes, which are also associated with weight outcomes in an independent cohort suggesting a strong relationship with weight trajectories in the first six months after birth. Additionally, we developed a Methylation Risk Score (MRS) that could be used to identify children most at risk for developing childhood obesity. While many of the genes identified by our analysis have been associated with weight-related traits (e.g., glucose metabolism, BMI, or hip-to-waist ratio) in previous genome-wide association and variant studies, our analysis implicated several others, whose involvement in the obesity phenotype should be evaluated in future functional investigations.
Mitochondria, cellular powerhouses, harbor DNA (mtDNA) inherited from the mothers. MtDNA mutations can cause diseases, yet whether they increase with age in human germline cells-oocytes-remains understudied. Here, using highly accurate duplex sequencing of full-length mtDNA, we detected de novo mutations in single oocytes, blood, and saliva in women between 20 and 42 years of age. We found that, with age, mutations increased in blood and saliva but not in oocytes. In oocytes, mutations with high allele frequencies (≥1%) were less prevalent in coding than non-coding regions, whereas mutations with low allele frequencies (<1%) were more uniformly distributed along mtDNA, suggesting frequency-dependent purifying selection. In somatic tissues, mutations caused elevated amino acid changes in protein-coding regions, suggesting positive or destructive selection. Thus, mtDNA in human oocytes is protected against accumulation of mutations having functional consequences and with aging. These findings are particularly timely as humans tend to reproduce later in life.
We present haplotype-resolved reference genomes and comparative analyses of six ape species, namely: chimpanzee, bonobo, gorilla, Bornean orangutan, Sumatran orangutan, and siamang. We achieve chromosome-level contiguity with unparalleled sequence accuracy (<1 error in 500,000 base pairs), completely sequencing 215 gapless chromosomes telomere-to-telomere. We resolve challenging regions, such as the major histocompatibility complex and immunoglobulin loci, providing more in-depth evolutionary insights. Comparative analyses, including human, allow us to investigate the evolution and diversity of regions previously uncharacterized or incompletely studied without bias from mapping to the human reference. This includes newly minted gene families within lineage-specific segmental duplications, centromeric DNA, acrocentric chromosomes, and subterminal heterochromatin. This resource should serve as a definitive baseline for all future evolutionary studies of humans and our closest living ape relatives.
Apes possess two sex chromosomes-the male-specific Y and the X shared by males and females. The Y chromosome is crucial for male reproduction, with deletions linked to infertility. The X chromosome carries genes vital for reproduction and cognition. Variation in mating patterns and brain function among great apes suggests corresponding differences in their sex chromosome structure and evolution. However, due to their highly repetitive nature and incomplete reference assemblies, ape sex chromosomes have been challenging to study. Here, using the state-of-the-art experimental and computational methods developed for the telomere-to-telomere (T2T) human genome, we produced gapless, complete assemblies of the X and Y chromosomes for five great apes (chimpanzee, bonobo, gorilla, Bornean and Sumatran orangutans) and a lesser ape, the siamang gibbon. These assemblies completely resolved ampliconic, palindromic, and satellite sequences, including the entire centromeres, allowing us to untangle the intricacies of ape sex chromosome evolution. We found that, compared to the X, ape Y chromosomes vary greatly in size and have low alignability and high levels of structural rearrangements. This divergence on the Y arises from the accumulation of lineage-specific ampliconic regions and palindromes (which are shared more broadly among species on the X) and from the abundance of transposable elements and satellites (which have a lower representation on the X). Our analysis of Y chromosome genes revealed lineage-specific expansions of multi-copy gene families and signatures of purifying selection. In summary, the Y exhibits dynamic evolution, while the X is more stable. Finally, mapping short-read sequencing data from >100 great ape individuals revealed the patterns of diversity and selection on their sex chromosomes, demonstrating the utility of these reference assemblies for studies of great ape evolution. These complete sex chromosome assemblies are expected to further inform conservation genetics of nonhuman apes, all of which are endangered species.
G-quadruplexes (G4s) are non-canonical DNA structures that can form at approximately 1% of the human genome. G4s contribute to point mutations and structural variation and thus facilitate genomic instability. They play important roles in regulating replication, transcription, and telomere maintenance, and some of them evolve under purifying selection. Nevertheless, the evolutionary dynamics of G4s has remained underexplored. Here we conducted a comprehensive analysis of predicted G4s (pG4s) in the recently released, telomere-to-telomere (T2T) genomes of human and other great apes-bonobo, chimpanzee, gorilla, Bornean orangutan, and Sumatran orangutan. We annotated tens of thousands of new pG4s in T2T compared to previous ape genome assemblies, including 41,236 in the human genome. Analyzing species alignments, we found approximately one-third of pG4s shared by all apes studied and identified thousands of species- and genus-specific pG4s. pG4s accumulated and diverged at rates consistent with divergence times between the studied species. We observed a significant enrichment and hypomethylation of pG4 shared across species at regulatory regions, including promoters, 5' and 3'UTRs, and origins of replication, strongly suggesting their formation and functional role in these regions. pG4s shared among great apes displayed lower methylation levels compared to species-specific pG4s, suggesting evolutionary conservation of functional roles of the former. Many species-specific pG4s were located in the repetitive and satellite regions deciphered in the T2T genomes. Our findings illuminate the evolutionary dynamics of G4s, their role in gene regulation, and their potential contribution to species-specific adaptations in great apes, emphasizing the utility of high-resolution T2T genomes in uncovering previously elusive genomic features.
The ape sex chromosomes have now been fully sequenced. Rapid evolution has led to extreme differences in the Y chromosome between species, whereas the X chromosome experienced much less dynamic changes. The palindromic structure of the Y chromosome has preserved genes relevant to fertility, hidden in the repetitive DNA.
The Javan gibbon, Hylobates moloch, is an endangered gibbon species restricted to the forest remnants of western and central Java, Indonesia, and one of the rarest of the Hylobatidae family. Hylobatids consist of 4 genera (Holoock, Hylobates, Symphalangus, and Nomascus) that are characterized by different numbers of chromosomes, ranging from 38 to 52. The underlying cause of this karyotype plasticity is not entirely understood, at least in part, due to the limited availability of genomic data. Here we present the first scaffold-level assembly for H. moloch using a combination of whole-genome Illumina short reads, 10X Chromium linked reads, PacBio, and Oxford Nanopore long reads and proximity-ligation data. This Hylobates genome represents a valuable new resource for comparative genomics studies in primates.
Y chromosomal ampliconic genes (YAGs) are important for male fertility, as they encode proteins functioning in spermatogenesis. The variation in copy number and expression levels of these multicopy gene families has been studied in great apes; however, the diversity of splicing variants remains unexplored. Here, we deciphered the sequences of polyadenylated transcripts of all nine YAG families (BPY2, CDY, DAZ, HSFY, PRY, RBMY, TSPY, VCY, and XKRY) from testis samples of six great ape species (human, chimpanzee, bonobo, gorilla, Bornean orangutan, and Sumatran orangutan). To achieve this, we enriched YAG transcripts with capture probe hybridization and sequenced them with long (Pacific Biosciences) reads. Our analysis of this data set resulted in several findings. First, we observed evolutionarily conserved alternative splicing patterns for most YAG families except for BPY2 and PRY. Second, our results suggest that BPY2 transcripts and proteins originate from separate genomic regions in bonobo versus human, which is possibly facilitated by acquiring new promoters. Third, our analysis indicates that the PRY gene family, having the highest representation of noncoding transcripts, has been undergoing pseudogenization. Fourth, we have not detected signatures of selection in the five YAG families shared among great apes, even though we identified many species-specific protein-coding transcripts. Fifth, we predicted consensus disorder regions across most gene families and species, which could be used for future investigations of male infertility. Overall, our work illuminates the YAG isoform landscape and provides a genomic resource for future functional studies focusing on infertility phenotypes in humans and critically endangered great apes.
In addition to the canonical right-handed double helix, other DNA structures, termed 'non-B DNA', can form in the genomes across the tree of life. Non-B DNA regulates multiple cellular processes, including replication and transcription, yet its presence is associated with elevated mutagenicity and genome instability. These discordant cellular roles fuel the enormous potential of non-B DNA to drive genomic and phenotypic evolution. Here we discuss recent studies establishing non-B DNA structures as novel functional elements subject to natural selection, affecting evolution of transposable elements (TEs), and specifying centromeres. By highlighting the contributions of non-B DNA to repeated evolution and adaptation to changing environments, we conclude that evolutionary analyses should include a perspective of not only DNA sequence, but also its structure.
Obesity is a highly heritable condition that affects increasing numbers of adults and, concerningly, of children. However, only a small fraction of its heritability has been attributed to specific genetic variants. These variants are traditionally ascertained from genome-wide association studies (GWAS), which utilize samples with tens or hundreds of thousands of individuals for whom a single summary measurement (e.g., BMI) is collected. An alternative approach is to focus on a smaller, more deeply characterized sample in conjunction with advanced statistical models that leverage longitudinal phenotypes. Novel functional data analysis (FDA) techniques are used to capitalize on longitudinal growth information from a cohort of children between birth and three years of age. In an ultra-high dimensional setting, hundreds of thousands of single nucleotide polymorphisms (SNPs) are screened, and selected SNPs are used to construct two polygenic risk scores (PRS) for childhood obesity using a weighting approach that incorporates the dynamic and joint nature of SNP effects. These scores are significantly higher in children with (vs. without) rapid infant weight gain—a predictor of obesity later in life. Using two independent cohorts, it is shown that the genetic variants identified in very young children are also informative in older children and in adults, consistent with early childhood obesity being predictive of obesity later in life. In contrast, PRSs based on SNPs identified by adult obesity GWAS are not predictive of weight gain in the cohort of young children. This provides an example of a successful application of FDA to GWAS. This application is complemented with simulations establishing that a deeply characterized sample can be just as, if not more, effective than a comparable study with a cross-sectional response. Overall, it is demonstrated that a deep, statistically sophisticated characterization of a longitudinal phenotype can provide increased statistical power to studies with relatively small sample sizes; and shows how FDA approaches can be used as an alternative to the traditional GWAS.