Humans possess human-specific traits 1,2 such as spoken language and lineage-specific traits such as ape-specific taillessness 3,4 . Previous efforts to identify the DNA sequences responsible for such human traits were limited by necessary accommodations for poor genome assembly quality and lack of population genomic sampling 5-16 . Here, we implement new algorithms that combine the near-complete human reference pangenome alignment with a new near-complete simian cross-species alignment to define human- and lineage-specific DNA sequences fixed across human haplotypes. Previously reported FOXP2/NOVA1 17,18 amino acid substitutions linked to human spoken language and TBXT transposable element insertion contributing to ape taillessness 3,4 were unique to their respective clades and fixed in sampled humans. In contrast, widely used sets of candidate human-mutated loci showed limited enrichment for either human specificity or fixation. Integration with candidate cis -regulatory elements 19-22 identified putative regulatory sequences specific to humans and linked to human-specific traits like hair reduction 23,24 and brain transcriptome patterning 25 . Although brain-associated fixed regulatory changes were present in all lineages, enrichment for spoken language was human-specific and enrichment for receptive language was ape-specific. This study provides a new pangenome-aware comparative framework and catalogs of candidate genomic loci to trace the evolutionary origins of common human traits and disease risks.
Human speech likely arose from regulatory changes for speech-related brain regions, yet causal variants and mechanisms remain unclear. RBFOX1 is a prime candidate, showing specialized expression in vocal learning circuits of human and zebra finch brains and carrying a promoter deletion linked to autism spectrum disorder (ASD) with language dysfunction. Here, we perform integrative analyses with cross-species brain single-cell multi-omic data and the more complete genomes of the Vertebrate Genomes Project. We identify a human-specific CCG insertion in the RBFOX1 promoter, creating a human-unique CCG-repeated motif. This motif is fixed in both archaic and modern humans but is disrupted by rare clinical variants that exhibit language-related phenotypes and autism. Binding motif models predicted, and reporter assays reveal that this human allele drives stronger EGR1 -dependent transcription than its chimpanzee allele. Genome-wide, 107 other genes have core promoters with the identical motif; enriched for postsynapse and implicated in ASD, including PTCHD1 . At the PTCHD1 promoter, an ASD-causative CCG-repeated variant enhances EGR1 -dependent promoter activity, and its activating effects are predicted in human brain regions using AlphaGenome. Our findings suggest that small variations in the number of CCG repeats in promoters can exert a large regulatory effect on complex traits and their associated disorders.
Bird genomes are the smallest among amniotes; however, they remain challenging to assemble due to their structural complexity. This study presents a fully phased diploid telomere-to-telomere reference genome for the zebra finch (Taeniopygia guttata), a model organism for neuroscience and evolutionary genomics. Combining sequencing strategies allowed the closing of nearly all gaps, adding ∼90 Mbp of sequence (7.8%). The assembly contains gapless sequences for all microchromosomes, including the elusive dot chromosomes with their distinctive architecture. The genome was comprehensively annotated for genes, repeats, and structural variants. Complete centromeres were identified, along with candidate DNA-binding centromere protein B (CENP-B) homologs, suggesting that birds possess a kinetochore-associated CENP system similar to that of mammals. Relative to the previous reference generated by the Vertebrate Genomes Project, 2,710 (8.13%) previously unassembled or unannotated genes were identified. This complete genome of a songbird serves as a public reference and illuminates avian genome architecture and function.
Genomic promoters are crucial gene regulatory elements1,2. Yet, comparative analyses of promoter architecture have been constrained by the limited resolution of GC-rich regions in short-read-based genome resources3-6. The Vertebrate Genomes Project (VGP) provides more complete long-read-based assemblies7, which further detect 5-methylcytosine signals directly from PacBio HiFi circular consensus reads8,9. Here, we developed a scalable computational framework to characterize DNA methylomes from HiFi data on high-quality Phase I VGP assemblies with RefSeq gene annotations for 82 vertebrate species spanning seven major taxonomic classes: mammals, birds, reptiles, amphibians, lobe-finned fishes, ray-finned fishes, and cartilaginous fishes. We observed a conserved, transcription start site-centered hypomethylation signature in promoters across all vertebrates, and an unexpected hypermethylation signature near gene boundaries that is discordant with transcripts. In addition to this conserved pattern, there were lineage-specific differences in promoter methylation profiles, with birds showing the most diverse patterns. These epigenetic landscapes track phylogenetic relationships more closely than tissue-type methylation differences and infer lineage-dependent widths of core promoters and broader promoters across major vertebrate classes. Our findings establish a comparative epigenomic framework for profiling promoter methylomes from long-read sequencing data.
The Vertebrate Genomes Project (VGP) aims to produce complete and near-error-free reference genomes for all ~70,000 extant vertebrate species1. Organized in four phases, it progressively targets all vertebrate orders, families, genera, and eventually all species. Here we present the completion of VGP Phase I, delivering reference genomes for ~95% of vertebrate orders, along with additional lineages within those orders, totaling 816 species and 1.6 trillion base pairs of main haplotype sequence. These genomes were assembled and annotated over an 8-year period (2018-2026) of rapid advances in genome sequencing, assembly, and annotation methods2-4, alongside the growth of associated consortium initiatives and international collaborations5-9. They represent some of the highest-quality vertebrate genomes currently available, and most have become the primary reference for their respective species in public databases. Comparative analyses across a subset of 579 species when we reached a threshold of 85% of orders allowed us to reconstruct the genome of the last common ancestor of all vertebrates 500 million years ago, identify diverse modes of sex chromosome evolution, reveal clade-specific three-dimensional genome architecture, discover methylated epigenetic landscapes across vertebrates, and provide a framework for studying gene and pseudogene evolution, immune loci, cancer-associated genes, and other trait-associated loci. Approximately a quarter of this subset are listed as Vulnerable to Critically Endangered by the IUCN Red List of Threatened Species, and have enabled more advanced genomic investigations of extinction risk. VGP Phase I delivers a reference backbone for vertebrate genomics, enabling discoveries that would otherwise remain out of reach across evolution, conservation, and medicine.
Vocal learning, the ability to imitate sounds, is a complex convergent trait crucial for spoken language and observed in a few independent lineages of mammals and birds. While convergences in gene expression have been found in vocal learning brain regions, amino acid convergences remain unclear. Here, we investigated whether avian vocal learning clades have amino acid convergences linked to their specialized trait. We developed a tool, Convergent Variant Finder, and applied it to an alignment of 48 species representing nearly all bird orders to identify convergent single amino acid variants among vocal learners and over 8,000 other polyphyletic species combinations. We discovered that the number of convergent variants was associated with the product of branch lengths of the most recent common ancestors of each species combination. The number of convergent variants in vocal learning clades did not exceed that of control species combinations. However, a subset of genes with vocal learner-specific convergent amino acid variants was enriched in the "learning" process, under positive selection, and significantly overlapped with gene sets for FOXP2 targets, singing-induced regulation in vocal learning nuclei, and differentially expressed in vocal learning nuclei. Moreover, we confirmed that the majority of convergent patterns in vocal learners were in the genomes of 363 species densely sampled across the avian tree. We propose that amino acid and nucleotide convergence accumulates at a steady state, with the rate proportional to divergence time. Selection associated with convergent traits, such as vocal learning, then likely acts on a subset of these changes.
Human immunodeficiency virus-1 (HIV-1) exploits the viral gp120 protein and host CD4/CCR5 receptors for the pandemic infection to humans. The host co-receptors of not only humans but also several primates and HIV-model mice can interact with the HIV receptor. However, the molecular mechanisms of these interactions remain unclear. Using Shaik et al. (2019)'s gp120/CD4/CCR5 structure of HIV-1B and human, here, we investigate the molecular dynamics between HIV sub-lineages (B, C, N, and O) and potential hosts in Euarchontoglires (primates and rodents). Although both host genes show similar protein structures conserved in all animals, CD4 gene demonstrates significantly stronger binding affinities in Catarrhini (apes and Old-World monkeys). Its known candidate residues interacted with gp120 fail to explain these affinity variations. Therefore, we identified novel candidate sites under positive selection on the Catarrhini lineage. Among four positively selected sites, residue R58 in humans is located within an antigen-antibody binding domain, exhibiting apomorphic amino acid substitutions as Arginine (R) in Catarrhini, which are mutually exclusive to the other animals where Lysine (K) is prevalent. Applying for artificial mutation test, we validated that K to R substitutions can lead stronger binding affinities of Catarrhini. Ecologically, these dynamics may relate to shared equatorial habitats in Africa and Asia. Our findings suggest a new candidate site R58 driven by the lineage-specific evolution as a molecular foundation on HIV infection.
Non-canonical (non-B) DNA motifs are genomic sequences capable of folding into three-dimensional structures distinct from the canonical right-handed helix. These structures regulate gene expression but also serve as mutation hotspots and are linked to cancer. Because non-B DNA is difficult to sequence, its annotations have been incomplete in most genome assemblies. Telomere-to-telomere (T2T) assemblies now overcome this limitation. Here, we provide a comprehensive analysis of eight types of non-B DNA motifs (e.g., G-quadruplexes and Z-DNA) in the zebra finch T2T genome. Motif content varied strongly by chromosome categories; gene-rich dot chromosomes showed the highest motif levels (22.8-40.5%), microchromosomes intermediate levels (9.8-24.8%), and macrochromosomes the lowest (9.1-10.1%). Within chromosomes, Z-DNA was enriched at centromeres, and G-quadruplexes were enriched at promoters and 5'UTRs. Low methylation at G-quadruplexes suggests they can form and contribute to gene regulation in these regions. Comparable patterns of non-B DNA distribution were observed in the near T2T chicken genome, except that A-phased repeats and not Z-DNA were enriched at chicken centromeres. Overall, our findings indicate that the non-B DNA distribution reflects the distinctive architecture of avian genomes, implicating non-canonical DNA in gene expression and centromere organization. The unusually high density on dot chromosomes is negatively correlated with PacBio sequencing depth, and thus helps explain why these chromosomes have posed exceptional challenges for sequencing. Highlights:We present the first analysis of sequences with the potential to adopt non-canonical (non-B) DNA conformations within a telomere-to-telomere (T2T) assembly of a bird genome, the zebra finch, and compare it to that in the near T2T assembly of chicken.Non-B DNA, particularly G-quadruplexes, is markedly enriched at regulatory regions such as promoters and 5'UTRs, suggesting its role in regulating gene expression in bird genomes.Z-DNA shows strong enrichment at centromeric regions, implying a contribution to centromere architecture and function in the zebra finch.The short, gene-rich, and highly recombining dot chromosomes have a strong overrepresentation of non-B DNA, which may act as a tunable regulator of euchromatin activity, but is also correlated with low sequencing depth.
The most dynamic and repetitive regions of great ape genomes have traditionally been excluded from comparative studies 1–3 . Consequently, our understanding of the evolution of our species is incomplete. Here we present haplotype-resolved reference genomes and comparative analyses of six ape species: chimpanzee, bonobo, gorilla, Bornean orangutan, Sumatran orangutan and siamang. We achieve chromosome-level contiguity with substantial sequence accuracy (<1 error in 2.7 megabases) and completely sequence 215 gapless chromosomes telomere-to-telomere. We resolve challenging regions, such as the major histocompatibility complex and immunoglobulin loci, to provide in-depth evolutionary insights. Comparative analyses enabled investigations of the evolution and diversity of regions previously uncharacterized or incompletely studied without bias from mapping to the human reference genome. Such regions include newly minted gene families in lineage-specific segmental duplications, centromeric DNA, acrocentric chromosomes and subterminal heterochromatin. This resource serves as a comprehensive baseline for future evolutionary studies of humans and our closest living ape relatives.
We present haplotype-resolved reference genomes and comparative analyses of six ape species, namely: chimpanzee, bonobo, gorilla, Bornean orangutan, Sumatran orangutan, and siamang. We achieve chromosome-level contiguity with unparalleled sequence accuracy (<1 error in 500,000 base pairs), completely sequencing 215 gapless chromosomes telomere-to-telomere. We resolve challenging regions, such as the major histocompatibility complex and immunoglobulin loci, providing more in-depth evolutionary insights. Comparative analyses, including human, allow us to investigate the evolution and diversity of regions previously uncharacterized or incompletely studied without bias from mapping to the human reference. This includes newly minted gene families within lineage-specific segmental duplications, centromeric DNA, acrocentric chromosomes, and subterminal heterochromatin. This resource should serve as a definitive baseline for all future evolutionary studies of humans and our closest living ape relatives.
Recent improvements in genome sequencing and assembly promise to generate high-quality reference genomes for many species. Yet the genome assembly process is still laborious and costly, requires substantial expertise, and is generally not scalable to the goals of very large multispecies scientific efforts. To democratise the training and assembly process at scale, we continue to expand and improve the Vertebrate Genomes Project assembly pipeline in Galaxy, published earlier this year (Larivière et al. 2024) . The automated pipeline performs de novo assembly based on PacBio HiFi reads, with optional extended graph phasing using Hi-C or parental data. It supports scaffolding using Bionano optical maps data and Hi-C via modular workflows. The workflows include quality control throughout the assembly process using GenomeScope, gfastats, Merqury, BUSCO, and Pretext. Within the Vertebrate Genome Project, these workflows have already been applied to de novo assemble genomes of over a hundred species, doubling the number of species reported in Larivière et al. 2024. We will present an overview of the new species assembled in Galaxy and what we learned from them. We will also present the updates to the published workflows and new workflows added to the VGP suite to assist with manual curation, including Galaxy workflow implementation producing most of the visualisation tracks from Treeval (Pointon, Eagles, and Sims, n.d.) . The VGP's long-term goal is to use these workflows to generate high-quality, complete reference genomes for all of the roughly 70,000 extant vertebrate species, facilitating a new era of discovery across the life sciences. We will highlight how we use new developments in Planemo (Bray et al. 2023) to accelerate the production of genome assemblies through the command line. Bray, Simon, John Chilton, Matthias Bernt, Nicola Soranzo, Marius van den Beek, Bérénice Batut, Helena Rasche, et al. 2023. “The Planemo Toolkit for Developing, Deploying, and Executing Scientific Data Analyses in Galaxy and beyond.” Genome Research 33 (2): 261–68. Larivière, Delphine, Linelle Abueg, Nadolina Brajuka, Cristóbal Gallardo-Alba, Bjorn Grüning, Byung June Ko, Alex Ostrovsky, et al. 2024. “Scalable, Accessible and Reproducible Reference Genome Assembly and Evaluation in Galaxy.” Nature Biotechnology , January. https://doi.org/ 10.1038/s41587-023-02100-3 . Pointon, D. L., W. Eagles, and Y. Sims. n.d. “Sanger-Tol/treeval v1. 0.0–Ancient Atlantis. 2023.” Publisher Full Text .
Chub mackerels (Scomber japonicus) are a migratory marine fish widely distributed in the Indo-Pacific Ocean. They are globally consumed for their high Omega-3 content, but their population is declining due to global warming. Here, we generated the first chromosome-level genome assembly of chub mackerel (fScoJap1) using the Vertebrate Genomes Project assembly pipeline with PacBio HiFi genomic sequencing and Arima Hi-C chromosome contact data. The final assembly is 828.68 Mb with 24 chromosomes, nearly all containing telomeric repeats at their ends. We annotated 31,656 genes and discovered that approximately 2.19% of the genome contained DNA transposon elements repressed within duplicated genes. Analyzing 5-methylcytosine (5mC) modifications using HiFi reads, we observed open/close chromatin patterns at gene promoters, including the FADS2 gene involved in Omega-3 production. This chromosome-level reference genome provides unprecedented opportunities for advancing our knowledge of chub mackerels in biology, industry, and conservation.
The little skate gene annotation file of the article "Little skate genome provides insights into genetic programs essential for limb-based locomotion".
BackgroundEnterococcus faecium (E. faecium) is a member of symbiotic lactic acid bacteria in gastrointestinal tract and it was successfully used to treat diarrhea cases in humans. For a lactobacteria to survive during the pasteurization process, resistance of proteins to denaturation at high temperatures is crucial. Pyruvate kinase (PYK) is one of the proteins possessing such property. It plays a major role during glycolysis by producing pyruvate and adenosine triphosphate (ATP).ObjectiveTo assess the acquired thermostability of PYK of ALE strain using in silico methods.MethodsFirst, we predicted and assessed tertiary structures of our proteins using SWISS-MODEL homology modelling server. Second, we then applied molecular dynamics (MD) simulation to simulate and assess multiple properties of molecules. Therefore, we implemented comparative MD to evaluate thermostability of PYK of recently developed high temperature resistant strain of E. faecium using Adaptive Laboratory Evolution (ALE) method. After 20ns of simulation at different temperatures, we observed that ALE enhanced strain demonstrated slightly better stability at 300, 340 and 350 K compared to that of the wild type (WT) strain.ResultsWe collected the results of MD simulation at four temperature points: 300, 340, 350 and 400 K. Our results showed that the protein demonstrated increased stability at 340 and 350 K.ConclusionResults of these study suggest that PYK of ALE enhanced strain of E. faecium demonstrates overall better stability at elevated temperatures compared to that of WT strain.
Background Monkeypox is endemic to African region and has become of Global concern recently due to its outbreaks in non-endemic countries. Although, the disease was first recorded in 1970, no monkeypox specific drug or vaccine exists as of now. Methods We applied drug repositioning method, testing effectiveness of currently approved drugs against emerging disease, as one of the most affordable approaches for discovering novel treatment measures. Techniques such as virtual ligand-based and structure-based screening were applied to identify potential drug candidates against monkeypox. Results We narrowed down our results to 6 antiviral and 20 anti-tumor drugs that exhibit theoretically higher potency than tecovirimat, the currently approved drug for monkeypox disease. Conclusions Our results indicated that selected drug compounds displayed strong binding affinity for p37 receptor of monkeypox virus and therefore can potentially be used in future studies to confirm their effectiveness against the disease.
Improvements in genome sequencing and assembly are enabling high-quality reference genomes for all species. However, the assembly process is still laborious, computationally and technically demanding, lacks standards for reproducibility, and is not readily scalable. Here we present the latest Vertebrate Genomes Project assembly pipeline and demonstrate that it delivers high-quality reference genomes at scale across a set of vertebrate species arising over the last ~500 million years. The pipeline is versatile and combines PacBio HiFi long-reads and Hi-C-based haplotype phasing in a new graph-based paradigm. Standardized quality control is performed automatically to troubleshoot assembly issues and assess biological complexities. We make the pipeline freely accessible through Galaxy, accommodating researchers even without local computational resources and enhanced reproducibility by democratizing the training and assembly process. We demonstrate the flexibility and reliability of the pipeline by assembling reference genomes for 51 vertebrate species from major taxonomic groups (fish, amphibians, reptiles, birds, and mammals).
The little skate Leucoraja erinacea , a cartilaginous fish, displays pelvic fin driven walking-like behaviors using genetic programs and neuronal subtypes similar to those of land vertebrates. However, mechanistic studies on little skate motor circuit development have been limited, due to a lack of high-quality reference genome. Here, we generated an assembly of the little skate genome, containing precise gene annotation and structures, which allowed post-genome analysis of spinal motor neurons (MNs) essential for locomotion. Through interspecies comparison of mouse, skate and chicken MN transcriptomes, shared and divergent MN expression profiles were identified. Conserved MN genes were enriched for early-stage nervous system development. Comparison of accessible chromatin regions between mouse and skate MNs revealed conservation of the potential regulators with divergent transcription factor (TF) networks through which expression of MN genes is differentially regulated. TF networks in little skate MNs are much simpler than those in mouse MNs, suggesting a more fine-grained control of gene expression operates in mouse MNs. These findings suggest conserved and divergent mechanisms controlling MN development system of vertebrates during evolution and the contribution of intricate gene regulatory networks in the emergence of sophisticated motor system in tetrapods.
Background Many short-read genome assemblies have been found to be incomplete and contain mis-assemblies. The Vertebrate Genomes Project has been producing new reference genome assemblies with an emphasis on being as complete and error-free as possible, which requires utilizing long reads, long-range scaffolding data, new assembly algorithms, and manual curation. A more thorough evaluation of the recent references relative to prior assemblies can provide a detailed overview of the types and magnitude of improvements. Results Here we evaluate new vertebrate genome references relative to the previous assemblies for the same species and, in two cases, the same individuals, including a mammal (platypus), two birds (zebra finch, Anna’s hummingbird), and a fish (climbing perch). We find that up to 11% of genomic sequence is entirely missing in the previous assemblies. In the Vertebrate Genomes Project zebra finch assembly, we identify eight new GC- and repeat-rich micro-chromosomes with high gene density. The impact of missing sequences is biased towards GC-rich 5′-proximal promoters and 5′ exon regions of protein-coding genes and long non-coding RNAs. Between 26 and 60% of genes include structural or sequence errors that could lead to misunderstanding of their function when using the previous genome assemblies. Conclusions Our findings reveal novel regulatory landscapes and protein coding sequences that have been greatly underestimated in previous assemblies and are now present in the Vertebrate Genomes Project reference genomes.
The little skate Leucoraja erinacea, a cartilaginous fish, displays pelvic fin driven walking-like behavior using genetic programs and neuronal subtypes similar to those of land vertebrates. However, mechanistic studies on little skate motor circuit development have been limited, due to a lack of high-quality reference genome. Here, we generated an assembly of the little skate genome, with precise gene annotation and structures, which allowed post-genome analysis of spinal motor neurons (MNs) essential for locomotion. Through interspecies comparison of mouse, skate and chicken MN transcriptomes, shared and divergent gene expression profiles were identified. Comparison of accessible chromatin regions between mouse and skate MNs predicted shared transcription factor (TF) motifs with divergent ones, which could be used for achieving differential regulation of MN-expressed genes. A greater number of TF motif predictions were observed in MN-expressed genes in mouse than in little skate. These findings suggest conserved and divergent molecular mechanisms controlling MN development of vertebrates during evolution, which might contribute to intricate gene regulatory networks in the emergence of a more sophisticated motor system in tetrapods.
BACKGROUND:The severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) pandemic began in 2019 but it remains as a serious threat today. To reduce and prevent spread of the virus, multiple vaccines have been developed. Despite the efforts in developing vaccines, Omicron strain of the virus has recently been designated as a variant of concern (VOC) by the World Health Organization (WHO). OBJECTIVE:To develop a vaccine candidate against Omicron strain (B.1.1.529, BA.1) of the SARS-CoV-19. METHODS:We applied reverse vaccinology methods for BA.1 and BA.2 as the vaccine target and a control, respectively. First, we predicted MHC I, MHC II and B cell epitopes based on their viral genome sequences. Second, after estimation of antigenicity, allergenicity and toxicity, a vaccine construct was assembled and tested for physicochemical properties and solubility. Third, AlphaFold2, RaptorX and RoseTTAfold servers were used to predict secondary structures and 3D structures of the vaccine construct. Fourth, molecular docking analysis was performed to test binding of our construct with angiotensin converting enzyme 2 (ACE2). Lastly, we compared mutation profiles on the epitopes between BA.1, BA.2, and wild type to estimate the efficacy of the vaccine. RESULTS:We collected a total of 10 MHC I, 9 MHC II and 5 B cell epitopes for the final vaccine construct for Omicron strain. All epitopes were predicted to be antigenic, non-allergenic and non-toxic. The construct was estimated to have proper stability and solubility. The best modelled tertiary structures were selected for molecular docking analysis with ACE2 receptor. CONCLUSIONS:These results suggest the potential efficacy of our newly developed vaccine construct as a novel vaccine candidate against Omicron strain of the coronavirus.