Bird genomes are the smallest among amniotes; however, they remain challenging to assemble due to their structural complexity. This study presents a fully phased diploid telomere-to-telomere reference genome for the zebra finch (Taeniopygia guttata), a model organism for neuroscience and evolutionary genomics. Combining sequencing strategies allowed the closing of nearly all gaps, adding ∼90 Mbp of sequence (7.8%). The assembly contains gapless sequences for all microchromosomes, including the elusive dot chromosomes with their distinctive architecture. The genome was comprehensively annotated for genes, repeats, and structural variants. Complete centromeres were identified, along with candidate DNA-binding centromere protein B (CENP-B) homologs, suggesting that birds possess a kinetochore-associated CENP system similar to that of mammals. Relative to the previous reference generated by the Vertebrate Genomes Project, 2,710 (8.13%) previously unassembled or unannotated genes were identified. This complete genome of a songbird serves as a public reference and illuminates avian genome architecture and function.
Reference genome assemblies are essential infrastructure for investigating phylogeny and population/conservation genetics of wild organisms. Birds serve as model vertebrates in ecology and evolutionary biology due to their well-documented natural histories and extensive community science data. We release a set of 350 newly assembled avian genomes, which, when combined with 97 previously published genomes, represent 447 of the bird species recorded in Denmark, the Faroe Islands, and Greenland-the largest regional dataset of a vertebrate group to date. These genomes are published for various research activities. This data release advances the global effort to build comprehensive and accessible biodiversity genomic resources for the research community.
The Vertebrate Genomes Project (VGP) aims to produce complete and near-error-free reference genomes for all ~70,000 extant vertebrate species1. Organized in four phases, it progressively targets all vertebrate orders, families, genera, and eventually all species. Here we present the completion of VGP Phase I, delivering reference genomes for ~95% of vertebrate orders, along with additional lineages within those orders, totaling 816 species and 1.6 trillion base pairs of main haplotype sequence. These genomes were assembled and annotated over an 8-year period (2018-2026) of rapid advances in genome sequencing, assembly, and annotation methods2-4, alongside the growth of associated consortium initiatives and international collaborations5-9. They represent some of the highest-quality vertebrate genomes currently available, and most have become the primary reference for their respective species in public databases. Comparative analyses across a subset of 579 species when we reached a threshold of 85% of orders allowed us to reconstruct the genome of the last common ancestor of all vertebrates 500 million years ago, identify diverse modes of sex chromosome evolution, reveal clade-specific three-dimensional genome architecture, discover methylated epigenetic landscapes across vertebrates, and provide a framework for studying gene and pseudogene evolution, immune loci, cancer-associated genes, and other trait-associated loci. Approximately a quarter of this subset are listed as Vulnerable to Critically Endangered by the IUCN Red List of Threatened Species, and have enabled more advanced genomic investigations of extinction risk. VGP Phase I delivers a reference backbone for vertebrate genomics, enabling discoveries that would otherwise remain out of reach across evolution, conservation, and medicine.
Non-canonical (non-B) DNA motifs are genomic sequences capable of folding into three-dimensional structures distinct from the canonical right-handed helix. These structures regulate gene expression but also serve as mutation hotspots and are linked to cancer. Because non-B DNA is difficult to sequence, its annotations have been incomplete in most genome assemblies. Telomere-to-telomere (T2T) assemblies now overcome this limitation. Here, we provide a comprehensive analysis of eight types of non-B DNA motifs (e.g., G-quadruplexes and Z-DNA) in the zebra finch T2T genome. Motif content varied strongly by chromosome categories; gene-rich dot chromosomes showed the highest motif levels (22.8-40.5%), microchromosomes intermediate levels (9.8-24.8%), and macrochromosomes the lowest (9.1-10.1%). Within chromosomes, Z-DNA was enriched at centromeres, and G-quadruplexes were enriched at promoters and 5'UTRs. Low methylation at G-quadruplexes suggests they can form and contribute to gene regulation in these regions. Comparable patterns of non-B DNA distribution were observed in the near T2T chicken genome, except that A-phased repeats and not Z-DNA were enriched at chicken centromeres. Overall, our findings indicate that the non-B DNA distribution reflects the distinctive architecture of avian genomes, implicating non-canonical DNA in gene expression and centromere organization. The unusually high density on dot chromosomes is negatively correlated with PacBio sequencing depth, and thus helps explain why these chromosomes have posed exceptional challenges for sequencing. Highlights:We present the first analysis of sequences with the potential to adopt non-canonical (non-B) DNA conformations within a telomere-to-telomere (T2T) assembly of a bird genome, the zebra finch, and compare it to that in the near T2T assembly of chicken.Non-B DNA, particularly G-quadruplexes, is markedly enriched at regulatory regions such as promoters and 5'UTRs, suggesting its role in regulating gene expression in bird genomes.Z-DNA shows strong enrichment at centromeric regions, implying a contribution to centromere architecture and function in the zebra finch.The short, gene-rich, and highly recombining dot chromosomes have a strong overrepresentation of non-B DNA, which may act as a tunable regulator of euchromatin activity, but is also correlated with low sequencing depth.
Accurately predicting species' responses to anthropogenic climate change is hampered by limited knowledge of their spatiotemporal ecological and evolutionary dynamics. We combine landscape genomics, demographic reconstructions, and species distribution models to assess the eco-evolutionary responses to past climate fluctuations and to future climate of an Afro-Palaearctic migratory raptor, the lesser kestrel (Falco naumanni). We uncover two evolutionarily and ecologically distinct lineages (European and Asian), whose demographic history, evolutionary divergence, and historical distribution range were profoundly shaped by past climatic fluctuations. Using future climate projections, we find that the Asian lineage is at higher risk of range contraction, increased migration distance, climate maladaptation, and consequently greater extinction risk than the European lineage. Our results emphasise the importance of providing historical context as a baseline for understanding species' responses to contemporary climate change, and illustrate how incorporating intraspecific genetic variation improves the ecological realism of climate change vulnerability assessments.
Complete datasets of genetic variants are key to biodiversity genomic studies. Long-read sequencing technologies allow the routine assembly of highly contiguous, haplotype-resolved reference genomes. However, even when complete, reference genomes from a single individual may bias downstream analyses and fail to adequately represent genetic diversity within a population or species. Pangenome graphs assembled from aligned collections of high-quality genomes can overcome representation bias by integrating sequence information from multiple genomes from the same population, species or genus into a single reference. Here, we review the available tools and data structures to build, visualize and manipulate pangenome graphs while providing practical examples and discussing their applications in biodiversity and conservation genomics across the tree of life. Pangenomes integrate multiple genomes to mitigate reference bias. This Review presents tools to build, visualize and manipulate pangenome graphs and also highlights pangenome applications in biodiversity and conservation genomics.
Vocal rhythm plays a fundamental role in sexual selection and species recognition in birds, but little is known of its genetic basis due to the confounding effect of vocal learning in model systems. Uncovering its genetic basis could facilitate identifying genes potentially important in speciation. Here we investigate the genomic underpinnings of rhythm in vocal non-learning Pogoniulus tinkerbirds using 135 individual whole genomes distributed across a southern African hybrid zone. We find rhythm speed is associated with two genes that are also known to affect human speech, Neurexin-1 and Coenzyme Q8A. Models leveraging ancestry reveal these candidate loci also impact rhythmic stability, a trait linked with motor performance which is an indicator of quality. Character displacement in rhythmic stability suggests possible reinforcement against hybridization, supported by evidence of asymmetric assortative mating in the species producing faster, more stable rhythms. Because rhythm is omnipresent in animal communication, candidate genes identified here may shape vocal rhythm across birds and other vertebrates.
Recent improvements in genome sequencing and assembly promise to generate high-quality reference genomes for many species. Yet the genome assembly process is still laborious and costly, requires substantial expertise, and is generally not scalable to the goals of very large multispecies scientific efforts. To democratise the training and assembly process at scale, we continue to expand and improve the Vertebrate Genomes Project assembly pipeline in Galaxy, published earlier this year (Larivière et al. 2024) . The automated pipeline performs de novo assembly based on PacBio HiFi reads, with optional extended graph phasing using Hi-C or parental data. It supports scaffolding using Bionano optical maps data and Hi-C via modular workflows. The workflows include quality control throughout the assembly process using GenomeScope, gfastats, Merqury, BUSCO, and Pretext. Within the Vertebrate Genome Project, these workflows have already been applied to de novo assemble genomes of over a hundred species, doubling the number of species reported in Larivière et al. 2024. We will present an overview of the new species assembled in Galaxy and what we learned from them. We will also present the updates to the published workflows and new workflows added to the VGP suite to assist with manual curation, including Galaxy workflow implementation producing most of the visualisation tracks from Treeval (Pointon, Eagles, and Sims, n.d.) . The VGP's long-term goal is to use these workflows to generate high-quality, complete reference genomes for all of the roughly 70,000 extant vertebrate species, facilitating a new era of discovery across the life sciences. We will highlight how we use new developments in Planemo (Bray et al. 2023) to accelerate the production of genome assemblies through the command line. Bray, Simon, John Chilton, Matthias Bernt, Nicola Soranzo, Marius van den Beek, Bérénice Batut, Helena Rasche, et al. 2023. “The Planemo Toolkit for Developing, Deploying, and Executing Scientific Data Analyses in Galaxy and beyond.” Genome Research 33 (2): 261–68. Larivière, Delphine, Linelle Abueg, Nadolina Brajuka, Cristóbal Gallardo-Alba, Bjorn Grüning, Byung June Ko, Alex Ostrovsky, et al. 2024. “Scalable, Accessible and Reproducible Reference Genome Assembly and Evaluation in Galaxy.” Nature Biotechnology , January. https://doi.org/ 10.1038/s41587-023-02100-3 . Pointon, D. L., W. Eagles, and Y. Sims. n.d. “Sanger-Tol/treeval v1. 0.0–Ancient Atlantis. 2023.” Publisher Full Text .
Insights into the evolution of non-model organisms are limited by the lack of reference genomes of high accuracy, completeness, and contiguity. Here, we present a chromosome-level, karyotype-validated reference genome and pangenome for the barn swallow (Hirundo rustica). We complement these resources with a reference-free multialignment of the reference genome with other bird genomes and with the most comprehensive catalog of genetic markers for the barn swallow. We identify potentially conserved and accelerated genes using the multialignment and estimate genome-wide linkage disequilibrium using the catalog. We use the pangenome to infer core and accessory genes and to detect variants using it as a reference. Overall, these resources will foster population genomics studies in the barn swallow, enable detection of candidate genes in comparative genomics studies, and help reduce bias toward a single reference genome.
The availability of public genomic resources can greatly assist biodiversity assessment, conservation, and restoration efforts by providing evidence for scientifically informed management decisions. Here we survey the main approaches and applications in biodiversity and conservation genomics, considering practical factors, such as cost, time, prerequisite skills, and current shortcomings of applications. Most approaches perform best in combination with reference genomes from the target species or closely related species. We review case studies to illustrate how reference genomes can facilitate biodiversity research and conservation across the tree of life. We conclude that the time is ripe to view reference genomes as fundamental resources and to integrate their use as a best practice in conservation genomics.
Bird song mediates speciation but little is known about its genetic basis because of the confounding effect of vocal learning in model systems. Rhythm, in particular, transcends acoustic communication across the animal kingdom and plays a fundamental role in sexual selection and species recognition in birds. Here we investigated the genomic underpinnings of rhythm in vocal non-learning Pogoniulus tinkerbirds using a reference we assembled and 134 further individual whole genomes distributed across a Southern African hybrid zone. We show that rhythm speed is associated with two genes that affect speech in humans, Neurexin-1 and Coenzyme Q8A. Leveraging ancestry, we find that rhythmic stability is also associated with these candidate loci. Furthermore, a pattern of character displacement in rhythmic stability in the contact zone suggests there is reinforcement against hybridization, supported by evidence of assortative mating. Assortative mating is asymmetric, however, occurring only in the species that produces faster, more stable rhythms. Rhythmic stability reflects motor performance, a trait long regarded as an indicator of quality. Because rhythm is an omnipresent trait in animal communication, candidate genes shaping vocal rhythm identified here may play a pivotal role in speciation across birds and other vertebrates.
Improvements in genome sequencing and assembly are enabling high-quality reference genomes for all species. However, the assembly process is still laborious, computationally and technically demanding, lacks standards for reproducibility, and is not readily scalable. Here we present the latest Vertebrate Genomes Project assembly pipeline and demonstrate that it delivers high-quality reference genomes at scale across a set of vertebrate species arising over the last ~500 million years. The pipeline is versatile and combines PacBio HiFi long-reads and Hi-C-based haplotype phasing in a new graph-based paradigm. Standardized quality control is performed automatically to troubleshoot assembly issues and assess biological complexities. We make the pipeline freely accessible through Galaxy, accommodating researchers even without local computational resources and enhanced reproducibility by democratizing the training and assembly process. We demonstrate the flexibility and reliability of the pipeline by assembling reference genomes for 51 vertebrate species from major taxonomic groups (fish, amphibians, reptiles, birds, and mammals).
The barn swallow (Hirundo rustica) poses a number of fascinating scientific questions, including the taxonomic status of postulated subspecies. Here, we obtained and assessed the sequence variation of 411 complete mitogenomes, mainly from the European H. r. rustica, but other subspecies as well. In almost every case, we observed subspecies-specific haplogroups, which we employed together with estimated radiation times to postulate a model for the geographical and temporal worldwide spread of the species. The female barn swallow carrying the Hirundo rustica ancestral mitogenome left Africa (or its vicinity) around 280 thousand years ago (kya), and her descendants expanded first into Eurasia and then, at least 51 kya, into the Americas, from where a relatively recent (<20 kya) back migration to Asia took place. The exception to the haplogroup subspecies specificity is represented by the sedentary Levantine H. r. transitiva that extensively shares haplogroup A with the migratory European H. r. rustica and, to a lesser extent, haplogroup B with the Egyptian H. r. savignii. Our data indicate that rustica and transitiva most likely derive from a sedentary Levantine population source that split at the end of the Younger Dryas (YD) (11.7 kya). Since then, however, transitiva received genetic inputs from and admixed with both the closely related rustica and the adjacent savignii. Demographic analyses confirm this species’ strong link with climate fluctuations and human activities making it an excellent indicator for monitoring and assessing the impact of current global changes on wildlife.
When vertebrates face stressful events, the hypothalamic-pituitary-adrenal (HPA) axis is activated, generating a rapid increase in circulating glucocorticoid (GC) stress hormones followed by a return to baseline levels. However, repeated activation of HPA axis may lead to increase in oxidative stress. One target of oxidative stress is telomeres, nucleoprotein complexes at the end of chromosomes that shorten at each cell division. The susceptibility of telomeres to oxidizing molecules has led to the hypothesis that increased GC levels boost telomere shortening, but studies on this link are scanty. We studied if, in barn swallows Hirundo rustica, changes in adult erythrocyte telomere length between 2 consecutive breeding seasons are related to corticosterone (CORT) (the main avian GC) stress response induced by a standard capture-restraint protocol. Within-individual telomere length did not significantly change between consecutive breeding seasons. Second-year individuals showed the highest increase in circulating CORT concentrations following restraint. Moreover, we found a decline in female stress response along the breeding season. In addition, telomere shortening covaried with the stress response: a delayed activation of the negative feedback loop terminating the stress response was associated with greater telomere attrition. Hence, among-individual variation in stress response may affect telomere dynamics.
Insights into the evolution of non-model organisms are often limited by the lack of reference genomes. As part of the Vertebrate Genomes Project, we present a new reference genome and a pangenome produced with High-Fidelity long reads for the barn swallow Hirundo rustica . We then generated a reference-free multialignment with other bird genomes to identify genes under selection. Conservation analyses pointed at genes enriched for transcriptional regulation and neurodevelopment. The most conserved gene is CAMK2N2 , with a potential role in fear memory formation. In addition, using all publicly available data, we generated a comprehensive catalogue of genetic markers. Genome-wide linkage disequilibrium scans identified potential selection signatures at multiple loci. The top candidate region comprises several genes and includes BDNF , a gene involved in stress response, fear memory formation, and tameness. We propose that the strict association with humans in this species is linked with the evolution of pathways typically under selection in domesticated taxa.
We present a genome assembly from an individual female Caprimulgus europaeus (the European nightjar; Chordata; Aves; Caprimulgiformes; Caprimulgidae). The genome sequence is 1,178 megabases in span. The majority of the assembly (99.33%) is scaffolded into 37 chromosomal pseudomolecules, including the W and Z sex chromosomes.
We present a genome assembly from an individual female Caprimulgus europaeus (the European nightjar; Chordata; Aves; Caprimulgiformes; Caprimulgidae). The genome sequence is 1,178 megabases in span. The majority of the assembly (99.33%) is scaffolded into 37 chromosomal pseudomolecules, including the W and Z sex chromosomes.
High-quality and complete reference genome assemblies are fundamental for the application of genomics to biology, disease, and biodiversity conservation. However, such assemblies are available for only a few non-microbial species 1–4 . To address this issue, the international Genome 10K (G10K) consortium 5,6 has worked over a five-year period to evaluate and develop cost-effective methods for assembling highly accurate and nearly complete reference genomes. Here we present lessons learned from generating assemblies for 16 species that represent six major vertebrate lineages. We confirm that long-read sequencing technologies are essential for maximizing genome quality, and that unresolved complex repeats and haplotype heterozygosity are major sources of assembly error when not handled correctly. Our assemblies correct substantial errors, add missing sequence in some of the best historical reference genomes, and reveal biological discoveries. These include the identification of many false gene duplications, increases in gene sizes, chromosome rearrangements that are specific to lineages, a repeated independent chromosome breakpoint in bat genomes, and a canonical GC-rich pattern in protein-coding genes and their regulatory regions. Adopting these lessons, we have embarked on the Vertebrate Genomes Project (VGP), an international effort to generate high-quality, complete reference genomes for all of the roughly 70,000 extant vertebrate species and to help to enable a new era of discovery across the life sciences.