Early vertebrate autotetraploidization events may have enabled major innovations by expanding the genetic material for functional diversification, yet their ancient timing obscures how genome doubling reshaped gene regulatory evolution. Salmonids provide a unique window to these mechanisms, because they experienced a comparatively recent autotetraploidization and are earlier in the rediploidization process - which creates new genes and regulatory elements during evolution. Here, using large-scale multiomics spanning embryonic and adult tissues in two salmonids, we investigate gene regulatory evolution following genome doubling and rediploidization, which we show is governed by developmental and tissue-specific context, with a period of maximal constraint at advanced stages of embryogenesis. This work advances understanding of vertebrate genome evolution, while providing an open resource supporting salmonid aquaculture and conservation.
Novel traits enable many rodents to thrive in extreme environmental niches. Predatory grasshopper mice ( Onychomys sp.) have co-evolved resistance to painful and lethal neurotoxins produced by their scorpion prey. Previous work reported that grasshopper mice have structural and functional modifications in sodium channel Nav1.8 that block the effect of painful toxins. However, key questions remain about the molecular adaptations underlying toxin resistance. We produced the first high-quality reference genomes and annotations for Onychomys species and Peromyscus eremicus . We implemented a comprehensive pipeline to detect positive selection across genome-scale datasets and identified Onychomys-specific mutations in Nav1.3 ( Scn3a), a sodium channel gene expressed in the central nervous system and the peripheral sensory system after nerve damage. We detected an Onychomys -specific tandem gene duplication of the Cblif gene, which encodes a glycoprotein crucial for vitamin B12 absorption. This adaptation likely supports the species’ dietary specialisation and modified stomach morphology, where parietal cells expressing Cblif are especially numerous. Our study provides a key step for establishing the Onychomys species as a model system for studying toxin resistance, alternative pain phenotypes, and behavioural traits related to predator-prey interactions.
The details of how macrophages control different healing trajectories (regeneration vs. scar formation) remain poorly defined. Spiny mice (Acomys spp.) can regenerate external ear pinnae tissue, whereas lab mice (Mus musculus) form scar tissue in response to an identical injury. Here, we used this dual species system to dissect macrophage phenotypes between healing modes. We identified secreted factors from activated Acomys macrophages that induce a pro -regenerative phenotype in fibroblasts from both species. Transcriptional profiling of Acomys macrophages and subsequent in vitro tests identified VEGFC, PDGFA, and Lactotransferrin (LTF) as potential pro -regenerative modulators. Examining macrophages in vivo, we found that Acomys-resident macrophages secreted VEGFC and LTF, whereas Mus macrophages do not. Lastly, we demonstrate the requirement for VEGFC during regeneration and find that interrupting lymphangiogenesis delays blastema and new tissue formation. Together, our results demonstrate that cell -autonomous mechanisms govern how macrophages react to the same stimuli to differentially produce factors that facilitate regeneration.
Ensembl (https://www.ensembl.org) has produced high-quality genomic resources for vertebrates and model organisms for more than twenty years. During that time, our resources, services and tools have continually evolved in line with both the publicly available genome data and the downstream research and applications that utilise the Ensembl platform. In recent years we have witnessed a dramatic shift in the genomic landscape. There has been a large increase in the number of high-quality reference genomes through global biodiversity initiatives. In parallel, there have been major advances towards pangenome representations of higher species, where many alternative genome assemblies representing different breeds, cultivars, strains and haplotypes are now available. In order to support these efforts and accelerate downstream research, it is our goal at Ensembl to create high-quality annotations, tools and services for species across the tree of life. Here, we report our resources for popular reference genomes, the dramatic growth of our annotations (including haplotypes from the first human pangenome graphs), updates to the Ensembl Variant Effect Predictor (VEP), interactive protein structure predictions from AlphaFold DB, and the beta release of our new website.
Summary Butterflies and moths (Lepidoptera) are one of the most ecologically diverse and speciose insect orders, with more than 157,000 described species. However, the abundance and diversity of Lepidoptera are declining worldwide at an alarming rate. As few Lepidoptera are explicitly recognised as at risk globally, the need for conservation is neither mandated nor well-evidenced. Large-scale biodiversity genomics projects that take advantage of the latest developments in long-read sequencing technologies offer a valuable source of information. We here present a comprehensive, reference-free, whole-genome, multiple sequence alignment of 88 species of Lepidoptera. We show that the accuracy and quality of the alignment is influenced by the contiguity of the reference genomes analysed. We explored genomic signatures that might indicate conservation concern in these species. In our dataset, which is largely from Britain, many species, in particular moths, display low heterozygosity and a high level of inbreeding, reflected in medium (0.1 - 1 Mb) and long (> 1 Mb) runs of homozygosity. Many species with low inbreeding display a higher masked load, estimated from the sum of rejected substitution scores at heterozygous sites. Our study shows that the analysis of a single diploid genome in a comparative phylogenetic context can provide relevant genetic information to prioritise species for future conservation investigation, particularly for those with an unknown conservation status.
Ensembl (https://www.ensembl.org) is unique in its flexible infrastructure for access to genomic data and annotation. It has been designed to efficiently deliver annotation at scale for all eukaryotic life, and it also provides deep comprehensive annotation for key species. Genomes representing a greater diversity of species are increasingly being sequenced. In response, we have focussed our recent efforts on expediting the annotation of new assemblies. Here, we report the release of the greatest annual number of newly annotated genomes in the history of Ensembl via our dedicated Ensembl Rapid Release platform (http://rapid.ensembl.org). We have also developed a new method to generate comparative analyses at scale for these assemblies and, for the first time, we have annotated non-vertebrate eukaryotes. Meanwhile, we continually improve, extend and update the annotation for our high-value reference vertebrate genomes and report the details here. We have a range of specific software tools for specific tasks, such as the Ensembl Variant Effect Predictor (VEP) and the newly developed interface for the Variant Recoder. All Ensembl data, software and tools are freely available for download and are accessible programmatically.
SummaryAlthough macrophages play an essential role in tissue regeneration and fibrotic repair, the ability to pinpoint how macrophages contribute to one healing type over the other has been hampered by the rarity of model systems. Providing such a system, Acomys species (aka spiny mice) can regenerate skin and the complex tissue architecture of the external ear pinna whereas lab mice ( M. musculus ) form a scar in response to an identical injury. Here we used in vitro and in vivo methods in this dual species system to dissect macrophage phenotype between healing modes. We used ear pinna fibroblasts and conditioned media from macrophage subtypes and identified secreted factors from classically activated spiny mouse macrophages that uniquely induced a pro-regenerative phenotype in fibroblasts from Acomys and Mus . By transcriptionally profiling macrophages in vitro, we found that Pdgfa , Lactotransferrin (Ltf) , Vegfc and Il1a were uniquely expressed in spiny mouse macrophages and that PDGF-A and LTF could recapitulate the pro-regenerative phenotype. Examining macrophages in vivo using temporal scRNA-seq, we found that circulating and resident macrophages participate in tissue healing regardless of the healing outcome. Importantly, we demonstrate that Pdgfa , and Vegfc uniquely mark spiny mouse resident macrophages while their receptors are expressed in fibroblasts and endothelial cells respectively. Together, our results demonstrate that cell autonomous mechanisms govern how macrophages react to the same stimuli and differentially produce factors that facilitate regeneration.
Abstract The Orthology Benchmark Service (https://orthology.benchmarkservice.org) is the gold standard for orthology inference evaluation, supported and maintained by the Quest for Orthologs consortium. It is an essential resource to compare existing and new methods of orthology inference (the bedrock for many comparative genomics and phylogenetic analysis) over a standard dataset and through common procedures. The Quest for Orthologs Consortium is dedicated to maintaining the resource up to date, through regular updates of the Reference Proteomes and increasingly accessible data through the OpenEBench platform. For this update, we have added a new benchmark based on curated orthology assertion from the Vertebrate Gene Nomenclature Committee, and provided an example meta-analysis of the public predictions present on the platform.
The COVID-19 pandemic has seen unprecedented use of SARS-CoV-2 genome sequencing for epidemiological tracking and identification of emerging variants. Understanding the potential impact of these variants on the infectivity of the virus and the efficacy of emerging therapeutics and vaccines has become a cornerstone of the fight against the disease. To support the maximal use of genomic information for SARS-CoV-2 research, we launched the Ensembl COVID-19 browser; the first virus to be encompassed within the Ensembl platform. This resource incorporates a new Ensembl gene set, multiple variant sets, and annotation from several relevant resources aligned to the reference SARS-CoV-2 assembly. Since the first release in May 2020, the content has been regularly updated using our new rapid release workflow, and tools such as the Ensembl Variant Effect Predictor have been integrated. The Ensembl COVID-19 browser is freely available at https://covid-19.ensembl.org.
The Ensembl project (https://www ensembl org) annotates genomes and disseminates genomic data for vertebrate species We create detailed and comprehensive annotation of gene structures, regulatory elements and variants, and enable comparative genomics by inferring the evolutionary history of genes and genomes Our integrated genomic data are made available in a variety of ways, including genome browsers, search interfaces, specialist tools such as the Ensembl Variant Effect Predictor, download files and programmatic interfaces Here, we present recent Ensembl developments including two new website portals Ensembl Rapid Release (http://rapid ensembl org) is designed to provide core tools and services for genomes as soon as possible and has been deployed to support large biodiversity sequencing projects Our SARS-CoV-2 genome browser (https://covid-19 ensembl org) integrates our own annotation with publicly available genomic data from numerous sources to facilitate the use of genomics in the international scientific response to the COVID-19 pandemic We also report on other updates to our annotation resources, tools and services All Ensembl data and software are freely available without restriction
Ensembl (www.ensembl.org) produces integrated comparative genomics resources including homology annotation, whole-genome alignments (WGA), and synteny reports for over 1,700 species across the eukaryotic tree of life. The number of genomes processed by Ensembl will grow rapidly over the next few years with the advent of ambitious large-scale biodiversity sequencing projects such as the Darwin Tree of Life project (DToL). DToL aims to sequence and annotate the genomes of ~66,000 species found in the British Isles. To accommodate this rapid growth, we continue to improve our methods for processing, storing and displaying comparative genomics data. Numerous algorithmic and infrastructural improvements have been made to cope with these demands. K-mer-based methods estimate species divergence quickly and reliably. HTSlib enables random access to sequence files with virtually no penalty. Hidden Markov Model profiles grant faster and sensitive classifications of sequences for tree reconstruction. To further prepare for this avalanche of data, we have recently leveraged an extremely fast distributed file-system to achieve a 5-fold speed increase of our homology processing, enabling the annotation of over a billion homologous relationships across >400 plants and vertebrates in just a few weeks. Fresh development to improve homology upscaling to hundreds of thousands of species includes a new gene tree inference pipeline implementing linear and iterative reconstruction approaches. We are also assessing the role that highly-scalable, cloud-based solutions can play in improving the performance of our pipelines. We plan to develop and deploy the multiple WGA software Cactus, to handle thousands of genomes.
Almost one third of Earth’s land surface is arid, with deserts alone covering more than 46 million square kilometres. Nearly 2.1 billion people inhabit deserts or drylands and these regions are also home to a great diversity of plant and animal species including many that are unique to them. Aridity is a multifaceted environmental stress combining a lack of water with limited food availability and typically extremes of temperature, impacting animal species across the planet from polar cold valleys, to Andean deserts and the Sahara. These harsh environments are also home to diverse microbial communities, demonstrating the ability of bacteria, fungi and archaea to settle and live in some of the toughest locations known. We now understand that these microbial ecosystems i.e. microbiotas, the sum total of microbial life across and within an environment, interact across both the environment, and the macroscopic organisms residing in these arid environments. Although multiple studies have explored these microbial communities in different arid environments, few studies have examined the microbiota of animals which are themselves arid-adapted. Here we aim to review the interactions between arid environments and the microbial communities which inhabit them, covering hot and cold deserts, the challenges these environments pose and some issues arising from limitations in the field. We also consider the work carried out on arid-adapted animal microbiotas, to investigate if any shared patterns or trends exist, whether between organisms or between the animals and the wider arid environment microbial communities. We determine if there are any patterns across studies potentially demonstrating a general impact of aridity on animal-associated microbiomes or benefits from aridity-adapted microbiomes for animals. In the context of increasing desertification and climate change it is important to understand the connections between the three pillars of microbiome, host genome and environment.
Pseudogenes are ideal markers of genome remodeling. In turn, the mouse is an ideal platform for studying them, particularly with the availability of developmental transcriptional data and the sequencing of 18 strains. Here, we present a comprehensive genome-wide annotation of the pseudogenes in the mouse reference genome and associated strains. We compiled this by combining manual curation of over 10,000 pseudogenes with results from automatic annotation pipelines. Also, by comparing the human and mouse, we annotated 165 unitary pseudogenes in mouse, and 303 unitaries in human. We make all our annotation available through mouse.pseudogene.org . The overall mouse pseudogene repertoire (in the reference and strains) is similar to human in terms of overall size, biotype distribution (~80% processed/~20% duplicated) and top family composition (with many GAPDH and ribosomal pseudogenes). However, notable differences arise in the pseudogene age distribution, with multiple retro-transpositional bursts in mouse evolutionary history and only one in human. Furthermore, in each strain about a fifth of the pseudogenes are unique, reflecting strain-specific functions and evolution. Additionally, we find that ~15% of the pseudogenes are transcribed, a fraction similar to that for human, and that pseudogene transcription exhibits greater tissue and strain specificity compared to protein-coding genes. Finally, we show that highly transcribed parent genes tend to give rise to processed pseudogenes.
Numerous factors have been shown to influence microbiome composition, including host genetics, diet and environmental variables. Recent studies have explored how host genetics influence human gut microbial communities, and those of pre-clinical models e.g. mice. However to date no studies have determined the influence of host genetics and extreme environments on wild animal microbiomes. Furthermore, when ‘wild’ species have been utilised this has typically involved working on captive individuals, restricted to a single species or limited to single sampling without replicates. Here we present work which takes advantage of a unique opportunity to directly investigate host organism genetic influences and environment on the resident microbiome. Our dataset comprised samples from individuals of two closely related sympatric, independently adapted arid mouse species, subject to the same stresses and diet; Acomys cahirinus and Acomys russatus. These desert dwellers are very highly adapted to the extreme environment they live in, and therefore represent an exciting opportunity to probe key microbiome-environment-host genetic questions. Samples are replicates from wild individuals, separated by a period of four months. Wild individuals were captured on two occasions and faecal samples collected. DNA extracted and subjected to shotgun metagenomic sequencing, generating NGS data. Preliminary analysis allowed us to probe community composition and relative abundance within individuals, and between members of the same species. Utilising Kraken, Centrifuge and Metaphlan we have found that abundant taxa include Lactobacillus and Roseburia. Ongoing analysis will enable us to establish the relative influence of host genetics on the metagenome composition of the two species.
Understanding the mechanisms driving lineage-specific evolution in both primates and rodents has been hindered by the lack of sister clades with a similar phylogenetic structure having high-quality genome assemblies. Here, we have created chromosome-level assemblies of the Mus caroli and Mus pahari genomes. Together with the Mus musculus and Rattus norvegicus genomes, this set of rodent genomes is similar in divergence times to the Hominidae (human-chimpanzee-gorilla-orangutan). By comparing the evolutionary dynamics between the Muridae and Hominidae, we identified punctate events of chromosome reshuffling that shaped the ancestral karyotype of Mus musculus and Mus caroli between 3 and 6 million yr ago, but that are absent in the Hominidae. Hominidae show between four- and sevenfold lower rates of nucleotide change and feature turnover in both neutral and functional sequences, suggesting an underlying coherence to the Muridae acceleration. Our system of matched, high-quality genome assemblies revealed how specific classes of repeats can play lineage-specific roles in related species. Recent LINE activity has remodeled protein-coding loci to a greater extent across the Muridae than the Hominidae, with functional consequences at the species level such as reproductive isolation. Furthermore, we charted a Muridae-specific retrotransposon expansion at unprecedented resolution, revealing how a single nucleotide mutation transformed a specific SINE element into an active CTCF binding site carrier specifically in Mus caroli , which resulted in thousands of novel, species-specific CTCF binding sites. Our results show that the comparison of matched phylogenetic sets of genomes will be an increasingly powerful strategy for understanding mammalian biology.
Despite the rapid development of sequencing technologies, the assembly of mammalian-scale genomes into complete chromosomes remains one of the most challenging problems in bioinformatics. To help address this difficulty, we developed Ragout 2, a reference-assisted assembly tool that works for large and complex genomes. By taking one or more target assemblies (generated from an NGS assembler) and one or multiple related reference genomes, Ragout 2 infers the evolutionary relationships between the genomes and builds the final assemblies using a genome rearrangement approach. By using Ragout 2, we transformed NGS assemblies of 16 laboratory mouse strains into sets of complete chromosomes, leaving <5% of sequence unlocalized per set. Various benchmarks, including PCR testing and realigning of long Pacific Biosciences (PacBio) reads, suggest only a small number of structural errors in the final assemblies, comparable with direct assembly approaches. We applied Ragout 2 to the Mus caroli and Mus pahari genomes, which exhibit karyotype-scale variations compared with other genomes from the Muridae family. Chromosome painting maps confirmed most large-scale rearrangements that Ragout 2 detected. We applied Ragout 2 to improve draft sequences of three ape genomes that have recently been published. Ragout 2 transformed three sets of contigs (generated using PacBio reads only) into chromosome-scale assemblies with accuracy comparable to chromosome assemblies generated in the original study using BioNano maps, Hi-C, BAC clones, and FISH.
We report full-length draft de novo genome assemblies for 16 widely used inbred mouse strains and find extensive strain-specific haplotype variation. We identify and characterize 2,567 regions on the current mouse reference genome exhibiting the greatest sequence diversity. These regions are enriched for genes involved in pathogen defence and immunity and exhibit enrichment of transposable elements and signatures of recent retrotransposition events. Combinations of alleles and genes unique to an individual strain are commonly observed at these loci, reflecting distinct strain phenotypes. We used these genomes to improve the mouse reference genome, resulting in the completion of 10 new gene structures. Also, 62 new coding loci were added to the reference genome annotation. These genomes identified a large, previously unannotated, gene (Efcab3-like) encoding 5,874 amino acids. Mutant Efcab3-like mice display anomalies in multiple brain regions, suggesting a possible role for this gene in the regulation of brain development.
Noncoding regulatory variants play a central role in the genetics of human diseases and in evolution. Here we measure allele-specific transcription factor binding occupancy of three liver-specific transcription factors between crosses of two inbred mouse strains to elucidate the regulatory mechanisms underlying transcription factor binding variations in mammals. Our results highlight the pre-eminence of cis-acting variants on transcription factor occupancy divergence. Transcription factor binding differences linked to cis-acting variants generally exhibit additive inheritance, while those linked to trans-acting variants are most often dominantly inherited. Cis-acting variants lead to local coordination of transcription factor occupancies that decay with distance; distal coordination is also observed and may be modulated by long-range chromatin contacts. Our results reveal the regulatory mechanisms that interplay to drive transcription factor occupancy, chromatin state, and gene expression in complex mammalian cell states.
Background The genomes of laboratory rat strains are characterised by a mosaic haplotype structure caused by their unique breeding history. These mosaic haplotypes have been recently mapped by extensive sequencing of key strains. Comparison of genomic variation between two closely related rat strains with different phenotypes has been proposed as an effective strategy for the discovery of candidate strain-specific regions involved in phenotypic differences. We developed a method to prioritise strain-specific haplotypes by integrating genomic variation and genomic regulatory data predicted to be involved in specific phenotypes. Specifically, we aimed to identify genomic regions associated with Metabolic Syndrome (MetS), a disorder of energy utilization and storage affecting several organ systems. Results We compared two Lyon rat strains, Lyon Hypertensive (LH) which is susceptible to MetS, and Lyon Low pressure (LL), which is susceptible to obesity as an intermediate MetS phenotype, with a third strain (Lyon Normotensive, LN) that is resistant to both MetS and obesity. Applying a novel metric, we ranked the identified strain-specific haplotypes using evolutionary conservation of the occupancy three liver-specific transcription factors (HNF4A, CEBPA, and FOXA1) in five rodents including rat. Consideration of regulatory information effectively identified regions with liver-associated genes and rat orthologues of human GWAS variants related to obesity and metabolic traits. We attempted to find possible causative variants and compared them with the candidate genes proposed by previous studies. In strain-specific regions with conserved regulation, we found a significant enrichment for published evidence to obesity—one of the metabolic symptoms shown by the Lyon strains—amongst the genes assigned to promoters with strain-specific variation. Conclusions Our results show that the use of functional regulatory conservation is a potentially effective approach to select strain-specific genomic regions associated with phenotypic differences among Lyon rats and could be extended to other systems.
Gene duplication and loss are major sources of genetic polymorphism in populations, and are important forces shaping the evolution of genome content and organization. We have reconstructed the origin and history of a 127-kbp segmental duplication, R2d, in the house mouse (Mus musculus). R2d contains a single protein-coding gene, Cwc22. De novo assembly of both the ancestral (R2d1) and the derived (R2d2) copies reveals that they have been subject to nonallelic gene conversion events spanning tens of kilobases. R2d2 is also a hotspot for structural variation: its diploid copy number ranges from zero in the mouse reference genome to >80 in wild mice sampled from around the globe. Hemizygosity for high copy-number alleles of R2d2 is associated in cis with meiotic drive; suppression of meiotic crossovers; and copy-number instability, with a mutation rate in excess of 1 per 100 transmissions in some laboratory populations. Our results provide a striking example of allelic diversity generated by duplication and demonstrate the value of de novo assembly in a phylogenetic context for understanding the mutational processes affecting duplicate genes.