Summary:Recent developments in genome sequencing and assembly technologies have enabled the automated assembly of vertebrate chromosomes from telomere to telomere. However, for some long, highly similar repeats, genome assemblers may lack sufficient information to unambiguously resolve the sequence, leaving tangles in the assembly graph and gaps in the final assembly. In recently published genomes, such gaps are often closed by manual graph curation, a process that is labor-intensive, error-prone, and sometimes infeasible. This can leave important genomic repeats, such as recently duplicated genes, misassembled or excluded from the final assembly. Here we present the Trivial Tangle Traverser (TTT) algorithm that finds optimized resolutions of assembly graph tangles. TTT uses depth of coverage and read-to-graph alignment information in a two-stage process to identify evidence-based traversals that are consistent with the underlying data. First, sequence multiplicities are estimated through mixed-integer linear programming, after which an Eulerian path is found in the derived multigraph and optimized through a gradient-descent-like approach. We evaluate TTT traversals on the HG002 human reference genome and demonstrate its use to characterize a previously unassembled amplified gene array in the zebra finch genome. Availability:TTT is available at https://github.com/marbl/TTT.
Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.
Recent developments have enabled the automated assembly of vertebrate chromosomes from telomere to telomere. However, for long, highly similar repeats, genome assemblers may leave tangles in the assembly graph and gaps in the assembly. In recently published genomes, such gaps are closed by manual graph curation, a process that is labor intensive, error prone, and sometimes infeasible. Consequently, important genomic regions may be misassembled or omitted. Here, we present the trivial tangle traverser (TTT) algorithm that finds optimized resolutions of assembly graph tangles. TTT uses depth of coverage and read-to-graph alignment information in a two-stage process to estimate sequence multiplicities and identify traversals that are consistent with the underlying data. We evaluate TTT traversals on the HG002 human reference genome, compare TTT with a state-of-the-art assembler on the giraffe T2T assembly, and demonstrate its use to characterize a previously unassembled amplified p21-activated serine/threonine kinase 3-like (PAK3L) gene array in the zebra finch genome.
Balanced Robertsonian translocation (ROB) is the most common chromosomal rearrangement in humans, with an estimated occurrence of 1 in 800 in newborn studies. Carriers are at increased risk of cancer and often diagnosed at fertility clinics after facing recurrent miscarriages, infertility, or aneuploid offspring. Genotyping carriers with DNA sequencing has been challenging because of gaps and misrepresentation of the translocation fusion site in the human reference genome. Only recently, telomere-to-telomere (T2T) human genomes successfully revealed sequences of the acrocentric short arms, including the most common ROB fusion site. A ROB results in loss of two ribosomal DNA (rDNA) arrays and its adjacent distal sequences, including the highly conserved distal junction (DJ). Here, we present a novel method to type ROB carriers directly from short sequencing reads by estimating DJ copy number. We demonstrate that our method successfully genotypes ROBs using a reference-free approach or alignments to either T2T-CHM13v2 or GRCh38. Applying the method to a cohort of healthy newborns and family members (n=4,172) as well as the UK Biobank (n=490,416), we find candidate ROBs at a frequency consistent with the previously reported 1 in 800 incidence (0.11-0.12%). In addition to ROB carriers, we report the frequency of one DJ loss (9, 2.8-3.4%) or gain (11+, 8.4-9.3%) from the two cohorts and the 1000 Genomes Project (n=3,202), and characterize the underlying structural variation in near-T2T genome assemblies from the Human Pangenome Reference Consortium. Importantly, our method provides the first sequencing-based diagnostic for Robertsonian chromosomes and can be applied to low-coverage sequencing data, enhancing its clinical applicability and enabling new studies of structural variation on the acrocentric chromosomes.
Cancer genome sequencing is essential for understanding tumor evolution and advancing precision medicine.1 However, reference gaps and germline variants obscure detection of small and large somatic variants and methylation in repetitive regions.1-3 It is common for tumor cells to gain or lose chromosome arms due to somatic structural changes that occur inside highly repetitive satellite DNA sequences in the centromeres.4 To identify the full spectrum of somatic variants, including complex rearrangements, we construct and curate near-complete, haplotype-resolved assemblies of the most recent common ancestor of an early-passage broadly-consented hypodiploid pancreatic cancer cell line and matched normal tissues. The tumor assembly completely recapitulates all 35 tumor chromosomes observed with karyotyping, with multiple translocation-induced hybrid chromosomes. The hybrid chromosomes contain putative functional dicentric and fused centromeres, nested foldback inversions causing 14 breakpoints with a haplotype switch in a single event, and centromeric satellite tandem duplications up to 136 kbp. Direct comparison of tumor and normal assembly haplotypes uncovers >7,000 variants altering >1 Mbp of sequence in repetitive regions that have been hidden by reference gaps and germline variants. 44 % of somatic small variants change representation because they alter germline variants on GRCh38, impacting mutational signatures and kataegis/omikli clusters. Most somatic LINE insertions originate from two hypomethylated non-reference germline LINE insertions, highlighting their impact on insertion mutation burden. These assemblies demonstrate that centromeric, acrocentric, and telomeric regions conventionally excluded from analysis harbor extensive somatic and epigenetic changes. Resolving complete tumor genomes enables a deeper understanding of cancer structural plasticity and the endpoints of breakage-fusion-bridge cycles. These assembled, curated paired normal-tumor benchmarks will serve as a critical foundation for developing future algorithms to characterize the most intractable regions of cancer genomes.
A pangenome reference overcomes the inherent limitation of any individual reference genome by integrating the variation present in a population. We present the Human Pangenome Reference Consortium's (HPRC) Release 2 (HPRC2), an openly available, second phase pangenome that is an approximately fivefold expansion in genome number over HPRC Release 1 (HPRC1) and measurable improvement in genome completeness, contiguity, and accuracy. Selecting samples with a principled algorithm prioritising common variant coverage, HPRC2 contributes 460 haplotypes that together capture over 99% of common variation observed in the All of Us Research Program v8 cohort. Combining high-coverage long and ultra-long reads with modern assemblers and polishers, we produce thousands of telomere-to-telomere (T2T) chromosomes, and relative to HPRC1 halve the number of structurally unreliable regions as well as individual base errors per haplotype. We complement the assemblies with whole genome multiple alignments and gene annotations, and derive formal pangenome coordinate systems for addressing off-reference variation, demonstrating that individual human genomes contain more than one hundred thousand variants not succinctly described with respect to existing reference genomes. We also present the first matched long-read backed pantranscriptome and panepigenome at this scale, provide continuous local-ancestry estimates spanning every genome, and outline a host of new tools and applications that leverage the pangenome resource for improved genomics analysis.
Bird genomes are the smallest among amniotes; however, they remain challenging to assemble due to their structural complexity. This study presents a fully phased diploid telomere-to-telomere reference genome for the zebra finch (Taeniopygia guttata), a model organism for neuroscience and evolutionary genomics. Combining sequencing strategies allowed the closing of nearly all gaps, adding ∼90 Mbp of sequence (7.8%). The assembly contains gapless sequences for all microchromosomes, including the elusive dot chromosomes with their distinctive architecture. The genome was comprehensively annotated for genes, repeats, and structural variants. Complete centromeres were identified, along with candidate DNA-binding centromere protein B (CENP-B) homologs, suggesting that birds possess a kinetochore-associated CENP system similar to that of mammals. Relative to the previous reference generated by the Vertebrate Genomes Project, 2,710 (8.13%) previously unassembled or unannotated genes were identified. This complete genome of a songbird serves as a public reference and illuminates avian genome architecture and function.
While induced pluripotent stem cells (iPSCs) have gained popularity in studying neurodegenerative diseases, the heterogeneity of stem cells used across studies impacts cross-study comparison. The iPSC Neurodegenerative Disease Initiative (iNDI) selected the KOLF2.1J cell line and prioritized its use as a reference standard for studying the effects of pathogenic variants on cell biology due to its stability and neutral neurodegenerative disease genetic risk. This cell line, and its derivatives expressing over 100 variants related to Alzheimer's disease, Parkinson's disease, and other neurological diseases, are available for academic and industry access. Current genomic data analyses are limited by the use of a human reference genome that does not capture the complete genetic background of a given iPSC line. While in the future this issue may be partially mitigated by the creation of a comprehensive human pangenome, previous work has shown that generating custom genomes is of value both to characterize the variation present and to serve as a more appropriate genomic reference. Here, we generated and characterized a custom complete genome assembly from KOLF2.1J. Mapping of sequencing reads to a personalized diploid assembly results in more comprehensive mapping compared to traditional linear references (i.e GRCh38). In addition, we provide a comprehensive custom gene annotation along with isoform expression and differential methylation analyses across multiple cell types. The assembly and all additional data is browsable and publicly available. This resource will enable more accurate investigation of the KOLF2.1J cell line and any genomics data generated compared to using traditional generalized references, while also serving as a foundational approach for establishing custom reference assemblies for other high-value iPSC lines.
Ribosomal RNA (rRNA) genes are organized in tandem arrays known as ribosomal DNA (rDNA) on multiple chromosomes in Hominidae genomes. We measured copy number and transcriptional activity status of rRNA gene arrays across multiple individual genomes, revealing an identifiable fingerprint of rDNA copy number and activity. In some cases, entire arrays were transcriptionally silent, characterized by high DNA methylation across the rRNA gene, inaccessible chromatin, and the absence of transcription factors and transcripts. Silent arrays showed reduced association with the nucleolus and decreased interchromosomal interactions, consistent with the model that nucleolar organizer function depends on transcriptional activity. Removing rDNA methylation activated silent arrays. Array activity status remained stable through induced pluripotent stem cell reprogramming and differentiation into cerebral and intestinal organoids. Haplotype tracing in two unrelated family trios showed paternal transmission of silent arrays. We propose that the epigenetic state buffers rRNA gene dosage, specifies nucleolar organizer function, and can propagate transgenerationally.
The Telomere-to-Telomere Consortium recently finished the first truly complete sequence of a human genome. To resolve the most complex repeats, this project relied on the semimanual combination of long, accurate Pacific Biosciences (PacBio) HiFi and ultralong Oxford Nanopore Technologies sequencing reads. The Verkko assembler later automated this process, achieving complete assemblies for approximately half of the chromosomes in a diploid human genome. However, the first version of Verkko was computationally expensive and could not resolve all regions of a typical human genome. Here we present Verkko2, which implements a more efficient read correction algorithm, improves repeat resolution and gap closing, introduces proximity-ligation-based haplotype phasing and scaffolding, and adds support for multiple long-read data types. These enhancements allow Verkko2 to assemble all regions of a diploid human genome, including the short arms of the acrocentric chromosomes and both sex chromosomes. Together, these changes increase the number of telomere-to-telomere scaffolds by twofold, reduce runtime by fourfold, and improve assembly correctness. On a panel of 19 human genomes, Verkko2 assembles an average of 39 of 46 complete chromosomes as scaffolds, with 21 of these assembled as gapless contigs. Together, these improvements enable telomere-to-telomere comparative genomics and pangenomics, at scale.
The most dynamic and repetitive regions of great ape genomes have traditionally been excluded from comparative studies 1–3 . Consequently, our understanding of the evolution of our species is incomplete. Here we present haplotype-resolved reference genomes and comparative analyses of six ape species: chimpanzee, bonobo, gorilla, Bornean orangutan, Sumatran orangutan and siamang. We achieve chromosome-level contiguity with substantial sequence accuracy (<1 error in 2.7 megabases) and completely sequence 215 gapless chromosomes telomere-to-telomere. We resolve challenging regions, such as the major histocompatibility complex and immunoglobulin loci, to provide in-depth evolutionary insights. Comparative analyses enabled investigations of the evolution and diversity of regions previously uncharacterized or incompletely studied without bias from mapping to the human reference genome. Such regions include newly minted gene families in lineage-specific segmental duplications, centromeric DNA, acrocentric chromosomes and subterminal heterochromatin. This resource serves as a comprehensive baseline for future evolutionary studies of humans and our closest living ape relatives.
Pedigree analysis remains the gold standard for rare disease diagnostics, yet whole genome sequencing studies typically omit critical regions like centromeres, telomeres, and acrocentric chromosome p-arms. Here, we present telomere-to-telomere (T2T) reference genomes for four self-identified African American individuals of admixed ancestry spanning three generations. Our parent-of-origin assigned, chromosome-level assemblies revealed precise meiotic recombination breakpoints in previously inaccessible regions, including recombination events across acrocentric and subtelomeric sequences. Centromeric regions were highly stable, with multi-megabase arrays inherited intact across three generations, while the position of kinetochore assembly sites remained consistent and predominantly associated with the p-arm proximal region. The relative lengths of telomeres on individual chromosomes were maintained across generations. Using a targeted rDNA assembly approach, we reconstructed a complete megabase-scale ribosomal DNA (rDNA) array corresponding to the paternal chromosome 14. This openly available pedigree provides a benchmark dataset for studying recombination and genetic and epigenetic variation across the complete genome.
We present haplotype-resolved reference genomes and comparative analyses of six ape species, namely: chimpanzee, bonobo, gorilla, Bornean orangutan, Sumatran orangutan, and siamang. We achieve chromosome-level contiguity with unparalleled sequence accuracy (<1 error in 500,000 base pairs), completely sequencing 215 gapless chromosomes telomere-to-telomere. We resolve challenging regions, such as the major histocompatibility complex and immunoglobulin loci, providing more in-depth evolutionary insights. Comparative analyses, including human, allow us to investigate the evolution and diversity of regions previously uncharacterized or incompletely studied without bias from mapping to the human reference. This includes newly minted gene families within lineage-specific segmental duplications, centromeric DNA, acrocentric chromosomes, and subterminal heterochromatin. This resource should serve as a definitive baseline for all future evolutionary studies of humans and our closest living ape relatives.
Evaluating metagenomic software is key for optimizing metagenome interpretation and focus of the Initiative for the Critical Assessment of Metagenome Interpretation (CAMI). The CAMI II challenge engaged the community to assess methods on realistic and complex datasets with long- and short-read sequences, created computationally from around 1,700 new and known genomes, as well as 600 new plasmids and viruses. Here we analyze 5,002 results by 76 program versions. Substantial improvements were seen in assembly, some due to long-read data. Related strains still were challenging for assembly and genome recovery through binning, as was assembly quality for the latter. Profilers markedly matured, with taxon profilers and binners excelling at higher bacterial ranks, but underperforming for viruses and Archaea. Clinical pathogen detection results revealed a need to improve reproducibility. Runtime and memory usage analyses identified efficient programs, including top performers with other metrics. The results identify challenges and guide researchers in selecting methods for analyses.
Bacteriophages play key roles in the dynamics of the human microbiome. By far the most abundant components of the human gut virome are tailed bacteriophages of the realm Duplodnaviria, in particular, crAss-like phages. However, apart from duplodnaviruses, the gut virome has not been dissected in detail. Here we report a comprehensive census of a minor component of the gut virome, the tailless bacteriophages of the realm Varidnaviria. Tailless phages are primarily represented in the gut by prophages, that are mostly integrated in genomes of Alphaproteobacteria and Verrucomicrobia and belong to the order Vinavirales, which currently consists of the families Corticoviridae and Autolykiviridae. Phylogenetic analysis of the major capsid proteins (MCP) suggests that at least three new families should be established within Vinavirales to accommodate the diversity of prophages from the human gut virome. Previously, only the MCP and packaging ATPase genes were reported as conserved core genes of Vinavirales. Here we report an extended core set of 12 proteins, including MCP, packaging ATPase, and previously undetected lysis enzymes, that are shared by most of these viruses. We further demonstrate that replication system components are frequently replaced in the genomes of Vinavirales, suggestive of selective pressure for escape from yet unknown host defenses or avoidance of incompatibility with coinfecting related viruses. The results of this analysis show that, in a sharp contrast to marine viromes, varidnaviruses are a minor component of the human gut virome. Moreover, they are primarily represented by prophages, as indicated by the analysis of the flanking genes, suggesting that there are few, if any, lytic varidnavirus infections in the gut at any given time. These findings complement the existing knowledge of the human gut virome by exploring a group of viruses that has been virtually overlooked in previous work.
Porphyromonas gingivalis is associated with human periodontal disease. We cloned and sequenced the gene for heat shock protein 60 (GroEL, HSP60) from P. gingivalis FDC381. The identified clone carried a 2.6 kb DNA fragment which contained two open reading frames (ORFs) encoding a 9.6- and a 58.4-kDa protein. The translated amino acid sequence of these ORFs showed a high degree of homology with known sequences for GroES and GroEL from several bacterial species and humans. Escherichia coli carrying this clone expressed a 65-kDa protein which was recognized by anti-Mycobacterium leprae HSP60 monoclonal antibody. We purified the 65-kDa protein by DEAE-sepharose chromatography and hydroxyapatite chromatography. This protein was immunogenic and was recognized by sera from a number of patients with periodontal disease. This immunological reactivity and the existence of molecular mimicry between the P. gingivalis GroEL and other HSP homologs may indicate an important role for this molecule in periodontal lesion.
Although the use of long-read sequencing improves the contiguity of assembled viral genomes compared to short-read methods, assembling complex viral communities remains an open problem. We describe the viralFlye tool for identification and analysis of metagenome-assembled viruses in long-read assemblies. We show it significantly improves viral assemblies and demonstrate that long-reads result in a much larger array of predicted virus-host associations as compared to short-read assemblies. We demonstrate that the identification of novel CRISPR arrays in bacterial genomes from a newly assembled metagenomic sample provides information for predicting novel hosts for novel viruses.
CrAssphage is the most abundant human-associated virus and the founding member of a large group of bacteriophages, discovered in animal-associated and environmental metagenomes, that infect bacteria of the phylum Bacteroidetes. We analyze 4907 Circular Metagenome Assembled Genomes (cMAGs) of putative viruses from human gut microbiomes and identify nearly 600 genomes of crAss-like phages that account for nearly 87% of the DNA reads mapped to these cMAGs. Phylogenetic analysis of conserved genes demonstrates the monophyly of crAss-like phages, a putative virus order, and of 5 branches, potential families within that order, two of which have not been identified previously. The phage genomes in one of these families are almost twofold larger than the crAssphage genome (145-192 kilobases), with high density of self-splicing introns and inteins. Many crAss-like phages encode suppressor tRNAs that enable read-through of UGA or UAG stop-codons, mostly, in late phage genes. A distinct feature of the crAss-like phages is the recurrent switch of the phage DNA polymerase type between A and B families. Thus, comparative genomic analysis of the expanded assemblage of crAss-like phages reveals aspects of genome architecture and expression as well as phage biology that were not apparent from the previous work on phage genomics.
The datasets supporting the conclusions of this article "SPAligner: alignment of long diverged molecular sequences to assembly graphs"
Background: Double-stranded DNA bacteriophages (dsDNA phages) play pivotal roles in structuring human gut microbiomes; yet, the gut virome is far from being fully characterized, and additional groups of phages, including highly abundant ones, continue to be discovered by metagenome mining. A multilevel framework for taxonomic classification of viruses was recently adopted, facilitating the classification of phages into evolutionary informative taxonomic units based on hallmark genes. Together with advanced approaches for sequence assembly and powerful methods of sequence analysis, this revised framework offers the opportunity to discover and classify unknown phage taxa in the human gut. Results: A search of human gut metagenomes for circular contigs encoding phage hallmark genes resulted in the identification of 3738 apparently complete phage genomes that represent 451 putative genera. Several of these phage genera are only distantly related to previously identified phages and are likely to found new families. Two of the candidate families, "Flandersviridae" and "Quimbyviridae", include some of the most common and abundant members of the human gut virome that infect Bacteroides, Parabacteroides, and Prevotella. The third proposed family, "Gratiaviridae," consists of less abundant phages that are distantly related to the families Autographiviridae, Drexlerviridae, and Chaseviridae. Analysis of CRISPR spacers indicates that phages of all three putative families infect bacteria of the phylum Bacteroidetes. Comparative genomic analysis of the three candidate phage families revealed features without precedent in phage genomes. Some "Quimbyviridae" phages possess Diversity-Generating Retroelements (DGRs) that generate hypervariable target genes nested within defense-related genes, whereas the previously known targets of phage-encoded DGRs are structural genes. Several "Flandersviridae" phages encode enzymes of the isoprenoid pathway, a lipid biosynthesis pathway that so far has not been known to be manipulated by phages. The "Gratiaviridae" phages encode a HipA-family protein kinase and glycosyltransferase, suggesting these phages modify the host cell wall, preventing superinfection by other phages. Hundreds of phages in these three and other families are shown to encode catalases and iron-sequestering enzymes that can be predicted to enhance cellular tolerance to reactive oxygen species. Conclusions: Analysis of phage genomes identified in whole-community human gut metagenomes resulted in the delineation of at least three new candidate families of Caudovirales and revealed diverse putative mechanisms underlying phage-host interactions in the human gut. Addition of these phylogenetically classified, diverse, and distinct phages to public databases will facilitate taxonomic decomposition and functional characterization of human gut viromes.