Mustelinae are among the most diverse and taxonomically complex subfamilies within the Mustelidae, yet their evolutionary history and genetic diversity remain largely unexplored at the whole-genome level. Here, we present the first comprehensive comparative and phylogenomic study of this lineage, integrating nuclear and mitochondrial genomes from 10 species across the Holarctic and Indomalayan realms. Our dataset includes two novel genome assemblies (Mustela strigidorsa, M. sibirica) and an improved genome for M. nivalis, enabling robust cross-species analyses of genome size, chromosomal evolution, genetic diversity, and demographic history. We uncover striking inter- and intraspecific variation in genome-wide heterozygosity and genome size, with evidence of marked homozygosity in some Asian lineages (M. eversmanii, M. sibirica, M. strigidorsa) and remarkable genetic diversity in widespread species such as M. nivalis and M. erminea. Phylogenomic results support the previously suggested split of M. richardsonii from M. erminea, but we found no evidence for speciation within M. nivalis. Ancestral reconstruction of chromosomal rearrangements revealed key chromosomal fissions that shaped the Mustelinae radiation, including early events predating the divergence of modern Mustela species. The results confirmed the suggested ancestral karyotypes of Mustela (2n = 44) and Mustelinae (2n = 42). Finally, demographic reconstructions exposed species-specific responses to Quaternary climatic cycles, ranging from long-term resilience in M. nivalis to repeated population bottlenecks in M. putorius and M. sibirica. Collectively, our findings establish a genomic foundation for future evolutionary and conservation genomic research on this emblematic Mustelidae lineage.
High-quality reference genome assemblies have become essential for deepening our understanding of biodiversity, yet obtaining them for many species remains surprisingly challenging. Drawing on experiences from the European Reference Genome Atlas (ERGA) community, we focus on permit and sample-handling procedures leading up to nucleic acid sequencing, covering tasks such as ensuring ethical and legal compliance, verifying accurate species identification, maintaining sample integrity during transport, and isolating high-quality DNA or nuclei. While many of the challenges and solutions we discuss are broadly relevant, our regulatory and logistical examples are primarily from Europe. By synthesising practical guidance, we highlight the crucial importance of taxonomic expertise, proper vouchering and biobanking, rigorous cold-chain management or alternative preservation methods, and emphasise adherence to packaging and shipping requirements for biological materials. We showcase examples spanning diverse regions, taxa and source materials, which underscore the importance of context-specific strategies and internationally harmonised protocols, particularly for metadata reporting. Our recommendations aim to support both small-scale projects and large initiatives, directing collective efforts to facilitate efficient sampling, vouchering and sample processing for future genomic studies.
The Biodiversity Genomics Europe (BGE) Project has the overarching aim of accelerating the use of genomic science to enhance understanding of biodiversity, monitor biodiversity change, and guide interventions to address its decline. The BGE Project comprises activities focused on DNA Barcoding (Barcoding Stream) and Reference Genome Generation (Genomes Stream) for eukaryotic species across Europe, bringing together two European networks: the International Barcode of Life in Europe (iBOL Europe) and the European Reference Genome Atlas (ERGA). This publication is an abridged version of the successful grant proposal developed jointly by iBOL Europe and ERGA in response to the Horizon Europe call HORIZON-CL6-2021-BIODIV-01-01. Two key strands of genomic science form the basis of this proposal: DNA barcoding - sequencing short, standardised genomic regions to tell the world’s species apart, transforming the speed of completion of the inventory of life on Earth and providing the foundations of a global bio-surveillance system for biodiversity; and genome sequencing - generating high-quality complete reference genomes for all species on Earth, transforming understanding of biodiversity at the genetic level, and delivering fundamental knowledge of how biological systems function and how species respond and adapt to environmental change. The BGE Project objectives are focused on (i) Capacity: To establish functioning biodiversity genomics networks at the European level to connect and grow community capacity to use genomic tools to tackle the biodiversity crisis; (ii) Production: To establish and implement large-scale biodiversity genomic data generation pipelines for Europe to accelerate the production and accessibility of genomic data for biodiversity characterisation, conservation, and biomonitoring; and (iii) Application: To apply genomic tools to enhance understanding of pan-European biodiversity and biodiversity declines to improve the efficacy of management interventions and biomonitoring programmes.
Drosophila suzukii is a globally invasive fruit fly that causes significant economic impacts on soft-skinned fruit crops. While its invasion history has been studied on a global scale, regional-scale demographic and selective dynamics remain less understood. Portugal is at the westernmost limit of the European expansion, a key region to investigate genetic diversity and adaptation at the most recently colonized areas. Previous whole-genome sequencing of Portuguese populations suggested a Mediterranean invasion route and identified candidate variants potentially involved in local adaptation. Here, we re-sample the Portuguese populations two to four years later, to examine temporal changes in the invasion dynamics. Our results confirm the Mediterranean invasion route and show that the extremely low genetic diversity in one early-sampled population (PT-VM19) reflects a founder event. We show that this population gained genetic diversity over time, likely through gene flow from neighboring populations. We detect signatures of selection along the genome and find that most overlap with previously reported candidate genes is driven by PT-VM19, suggesting that many earlier signals were false positives associated with its low diversity. Yet, nine genes show consistent signatures of selection over time in Portuguese populations, identifying them as potential candidates for local adaptation. These genes are involved in neural development, embryogenesis, metabolism, and chromatin organization. Based on these, we developed and validated a set of SNP markers for monitoring D. suzukii. Our study underscores how temporal genomic data from recently invaded regions can uncover demographic recovery and ongoing adaptation and contributes genomic tools for applied population monitoring.
The common hamster is currently one of the fastest-declining mammals in Europe, with its protection facing multiple challenges due to the complexity of the threats posed by factors such as population fluctuations, agricultural intensification, and reduced reproductive output. Overcoming these challenges requires population genomic data to inform targeted conservation measures. Therefore, the Cricetus cricetus reference genome will serve as an important genomic resource for conservation genomic works aiming to halt the decline of this critically endangered species of the western part of the Eurasian steppe. Consistently with the reported karyotype, a total of 12 contiguous chromosomal pseudomolecules (10 autosomes and 2 sex chromosomes) and one mitochondrial genome were assembled from the genome sequence. This chromosome-level assembly encompasses 2.5 Gb, composed of 179 contigs and 49 scaffolds, with contig and scaffold N50 values of 47.7 Mb and 305.5 Mb, respectively.
Human-driven environmental change is reshaping ecosystems and challenging species’ ability to adapt. Understanding how genetic variation enables adaptation is crucial for conservation and requires exemplary systems to test hypotheses and make predictions. One particularly suitable model for studying climate-driven adaptation is seasonal color change (SCC), a phenological trait in which individuals transition between summer-dark and winter-white pelage/plumage to maintain camouflage. This review evaluates SCC as a model for predicting adaptive responses to climate change. First, we address the vulnerability of phenological traits to climate change, due to their dependence on photoperiodic cues and complex molecular regulation. Second, we review SCC literature across all 21 SCC species, summarizing knowledge on its regulation, the fitness costs of mismatch induced by snow loss, and the limited role of plasticity in buffering these effects. Third, we review recent findings on the genetic basis of SCC polymorphism that have linked adaptation to selection on pigmentation alleles with multiple evolutionary origins (including introgression and de novo mutations). Finally, we discuss the implications of the genetic architecture of SCC polymorphism for evolutionary rescue and conservation strategies, as well as methods for testing adaptation conditions using modeling approaches. While past research on SCC already showcased how predictive evolution can be incorporated into conservation action, we identify research gaps, including limited fitness data, taxonomic biases, and the need for real-time ecological and genomic monitoring. Addressing these gaps will improve the accuracy of predictive models and the success of management strategies aiming at protecting species’ resilience to rapid environmental change.
In large-scale biodiversity genomics projects, the number of species that could be sequenced exceeds the resources available. Species selection is therefore a crucial component, requiring clear criteria and procedures. In a bottom-up approach, the Biodiversity Genomics Europe (BGE) project implemented an Automated Decision-Making (ADM) process for species selection based on objective criteria and tested it on simulated and empirical data. Here, we present this species ranking ADM process, which includes three stages: exclusion, ranking, and feasibility check. The composition of selected species retained the diversity of the community-nominated species pool for key taxonomic, geographic, and demographic assessment criteria while reducing bias. Feasibility and funding limits influenced the final selection more than other factors, indicating that investments in these areas would improve available reference-genome diversity. The ADM achieved species selection for genome sequencing in a large-scale biodiversity project in a relatively objective manner consistent with the broader European biodiversity genomic community’s priorities.
Unraveling how adaptive traits originate and evolve is key to understanding the mechanisms shaping species' diversity and their adaptive potential. Seasonal color molts, from summer-brown to winter-white, evolved in at least 21 mammals and birds to maintain camouflage in environments with seasonal snow, but the occurrence of winter-brown morphs reflects seemingly convergent local adaptation to distinct snow conditions. In the least weasel (Mustela nivalis), alternative winter morphs map to the pigmentation gene MC1R, but the evolutionary history and functional basis of this variation remained unknown. Using in vitro cellular assays, we show that winter-brown coats are caused by a derived protein-coding amino acid substitution that reduces MC1R affinity to its ligands, ASIP and α-MSH. Using targeted enrichment and sequencing, we find that this mutation arose de novo within the species, around one million years ago, and was maintained across the geographically structured populations formed during its evolution in Europe. Using simulations, we show that genetic drift is unlikely to explain the long-term maintenance of this variant at intermediate frequencies, which can be driven by spatially varying selection, anchoring local adaptive responses. Our results underscore how long-standing adaptive variation can fuel recurrent adaptation to heterogeneous environments through time.
Understanding the architecture of biological adaptations is a major endeavor of evolutionary biology. Using Natural History collections, we study the genetic basis and evolution of white/brown winter coat color variation in the long-tailed weasel (Neogale frenata), a crucial phenological adaptation for camouflage in habitats with seasonal snow. We produced whole-genome sequencing data for museum specimens, along two winter color morph transition areas in North America, at the West and East coasts. Genome-wide association scans identified a single genomic region linked to color variation polymorphism with approximately 300 kb and 200 kb in the West and East regions, respectively, which included the pigmentation gene MC1R. We identified three MC1R alleles, two of which with deletions of nine or eight amino acids, alternatively associated with the winter brown morphs in the West and East, respectively. These deletions affect the second transmembrane domain, and in one case also the first extracellular loop, which in silico analyses predicted to impact the protein's function. Our findings show alternative intraspecific evolutionary solutions for environmental adaptation in long-tailed weasels, building on the evidence that major genes of the melanin production pathway are hotspots for recurrent and independent evolution of winter camouflage adaptation. This adaptive variation may be crucial to anchor adaptive responses facing future environmental change.
The present study aimed to provide insight into the genetic diversity, phylogenetic history, and maternal origin of the native sheep populations from the Arabian Peninsula and test hypotheses regarding the possible introgression of mtDNA from Arabian Peninsula sheep into African and Asian populations. The mtDNA sequencing of 17 populations from Arabia and Africa was conducted, and mtDNA data on six populations from Asia and Africa were compared with our data. Measurements of genetic diversity indices (h and pi) were higher for populations from the Arabian Peninsula than for those from Africa. Based on the estimated population structure, comparing pairwise FST and AMOVA values between Arabian and African populations indicated low genetic differentiation. According to phylogenetic relationship analysis (neighbour-joining (NJ) tree), the Arabian Peninsula sheep population sequences were grouped into four maternal haplogroups (HPGs): A, B, C, and E. Among these groups, HPG B was predominant, HPG A was the second most common HPG, and the remaining HPGs (C and E) exhibited lower frequencies in the studied Arabian sheep breeds. In addition, median-joining network analysis provided strong evidence of previous introgression between Arabian, African and Asian sheep, which might have arisen through seafaring trade or the migratory movements of ancient humans. Finally, Approximate Bayesian Computation support the colonization of Africa from the Arabian Peninsula via two colonization routes, an older one in the north and a more recent one in the south.
Pleistocene climatic fluctuations have often driven range shifts and hybridization among related species, leaving present-day genomic footprints. In the Iberian Peninsula, repeated and transient post-glacial contacts among hare species have left extensive mitochondrial DNA traces, but the genomic correlates and underlying biogeographic scenarios are still incompletely understood. Here, we study genome admixture in the broom hare, Lepus castroviejoi, endemic to the Cantabrian region, using its non-Iberian sister species, L. corsicanus, for contrast. Coalescent analyses of 10 genomes estimate that these species remained isolated since their divergence around 50,000 years ago, consistent with their current allopatry. Further analyses with 25 additional genomes indicate that small fractions of the L. castroviejoi genome originate from L. granatensis, L. timidus, and L. europaeus (0.72%, 0.08%, and 0.04%, respectively). Introgression dating based on tract lengths suggests L. granatensis was already admixed with L. timidus when it hybridized with L. castroviejoi, which could explain granatensis-timidus ancestry tract junctions detected in L. castroviejoi. Genomic segments with such junctions contain genes enriched for cell signaling and olfactory receptor activity, suggesting functional drivers facilitating genetic exchange. This research demonstrates that genomic ancestry inferences can reveal complex multiway admixture histories, which can be used to illuminate complex past biogeographic events.
Demographic declines have important consequences for population viability, since they can lead to losses in genome diversity, as well as increased inbreeding and expression of deleterious mutations. Scandinavia was colonized by the Arctic fox (Vulpes lagopus) at the Pleistocene/Holocene transition, and the population has since been on the periphery of the global distribution. The Scandinavian population became even more fragmented in the early 1900s due to human persecution, and experienced an additional decline in the 1980s. We generated high-coverage genomes from pre-bottleneck, as well as modern Scandinavian and Russian specimens, and found that genome-wide diversity was lower and inbreeding higher in Scandinavia compared to the Siberian population, even prior to the historical bottleneck, most likely reflecting the long-term partial isolation and recent postglacial origin of the Scandinavian population. The southern subpopulation has the highest inbreeding levels, likely due to having been recently founded and highly isolated. Our results also show that although inbreeding increased substantially over the past century, the amount of total genetic load did not change. Overall, these findings illustrate the utility of a temporal approach to disentangle the genomic consequences of recent declines from ancient biogeographic processes.
Biodiversity resilience relies on genetic diversity, which sustains the evolutionary potential of organisms in dynamic ecosystems. Genomics is a powerful tool for accurately estimating genetic diversity across genomes of species and populations. However, integration of genomic data into conservation efforts faces challenges due to the heterogeneity of approaches employed. Establishing common sets of standards for genomic data production and analysis is essential to consistently interpret results and clearly communicate outcomes to stakeholders. While the European Reference Genome Atlas (ERGA) community has contributed significantly to the standardisation of reference genome methodologies in synergy with other initiatives, there is now an urgent need to extend these principles to downstream analyses. ERGA aims to build on its experience to help establish harmonised approaches in applied biodiversity genomics research, aligned with ongoing efforts to define standardised metrics for measuring and reporting genetic diversity. Establishing consensus on best practices for genome-wide data generation methods and applications will substantially increase accuracy, interpretability, and comparability, together with enhanced stakeholder capacities. By identifying key opportunities and challenges, as well as conducting preliminary stakeholder mapping and examining case studies, the goal is to build an inclusive framework that ensures the relevance and widespread adoption of these best practices: fostering trust and confidence in genomics research practices to meet stakeholder needs in biodiversity conservation. We call upon the broader research community to join efforts in establishing these approaches, recognising the importance of participation of end-users, to foster the integration of genomic data into the toolkit for measuring and reporting genetic diversity.
Pleistocene climatic fluctuations have often driven range shifts and hybridization among related species, leaving present-day genomic footprints. In the Iberian Peninsula, Lepus timidus , after its post-deglaciation retreat, has left extensive mitochondrial DNA traces in three other hare species, but the genomic correlates and underlying biogeographic scenarios are still incompletely understood. This study focuses on Lepus castroviejoi , endemic to the Cantabrian region, using its non-Iberian sister species, L. corsicanus , for comparison. By analyzing coalescent patterns from 10 genomes, we estimate that these species remained isolated since their divergence, around 50,000 years ago, consistent with their current allopatry. Further analyses with 25 additional genomes indicate that small fractions of the L. castroviejoi genome originate from L. granatensis , L. timidus , and L. europaeus (0.72%, 0.08%, and 0.04%, respectively). Introgression dating based on tract lengths suggests L. granatensis was already admixed with L. timidus when it hybridized with L. castroviejoi , which could explain the granatensis - timidus ancestry tract junctions detected in L. castroviejoi . Genomic segments with such junctions contain genes enriched for cell signaling and olfactory receptor activity, possibly facilitating genetic exchange. This research demonstrates how genomic ancestry inferences can reveal complex multiway admixture histories and illuminate past biogeographic events. ### Competing Interest Statement The authors have declared no competing interest.
The broom hare (Lepus castroviejoi) is a threatened Iberian endemic, for which there is limited knowledge. We use genetic non-invasive sampling (gNIS; N = 185 faeces samples) and specimens from hunting and roadkills (N = 22) in conjunction with a 15-microsatellite panel and a 541-bp fragment of cytochrome-b to assess the genetic diversity, population structure and evolutionary history of this species. Populations from the other four European hare species were also analysed to accurately compare the genetic diversity patterns and infer admixture. Species identification from gNIS was inferred using small fragments of cytochrome-b and transferrin genes and individual identification was obtained using microsatellites. The broom hare population showed the lowest level of nuclear DNA diversity of all analysed hare species (N = 76; Na = 2.53, H-e = 0.186 and F-is = 0.341) and very low mitochondrial DNA diversity (N = 64; H-d = 0.743 and & pi; = 0.01543). Only the Italian hare (L. corsicanus) showed a similar pattern of low genetic diversity. No hybridization with the neighbouring hare species was detected. However, two mitochondrial DNA lineages, corresponding to two ancient events of introgression of mountain hare (L. timidus) origin, were characterized. There was evidence for shallow spatial population differentiation of the broom hare. The described reduced genetic diversity, associated with a narrow distribution range and recent population declines, represents a risk of population extinction, and highlights the need for conservation measures of this endemic threatened hare species.
Guanylate binding proteins (GBPs) are an evolutionarily ancient family of proteins that are widely distributed among eukaryotes. They belong to the dynamin superfamily of GTPases, and their expression can be partially induced by interferons (IFNs). GBPs are involved in the cell-autonomous innate immune response against bacterial, parasitic and viral infections. Evolutionary studies have shown that GBPs exhibit a pattern of gene gain and loss events, indicative for the birth-and-death model of evolution. Most species harbor large GBP gene clusters that encode multiple paralogs. Previous functional and in-depth evolutionary studies have mainly focused on murine and human GBPs. Since rabbits are another important model system for studying human diseases, we focus here on lagomorphs to broaden our understanding of the multifunctional GBP protein family by conducting evolutionary analyses and performing a molecular and functional characterization of rabbit GBPs. We observed that lagomorphs lack GBP3, 6 and 7. Furthermore, Leporidae experienced a loss of GBP2, a unique duplication of GBP5 and a massive expansion of GBP4. Gene expression analysis by reverse transcriptase quantitative polymerase chain reaction (RT-qPCR) and transcriptome data revealed that leporid GBP expression varied across tissues. Overexpressed rabbit GBPs localized either uniformly and/or discretely to the cytoplasm and/or to the nucleus. Oryctolagus cuniculus (oc)GBP5L1 and rarely ocGBP5L2 were an exception, colocalizing with the trans-Golgi network (TGN). In addition, four ocGBPs were IFN-inducible and only ocGBP5L2 inhibited furin activity. In conclusion, from an evolutionary perspective, lagomorph GBPs experienced multiple gain and loss events, and the molecular and functional characteristics of ocGBP suggest a role in innate immunity.
A genomic database of all Earth's eukaryotic species could contribute to many scientific discoveries; however, only a tiny fraction of species have genomic information available. In 2018, scientists across the world united under the Earth BioGenome Project (EBP), aiming to produce a database of high-quality reference genomes containing all ~1.5 million recognized eukaryotic species. As the European node of the EBP, the European Reference Genome Atlas (ERGA) sought to implement a new decentralised, equitable and inclusive model for producing reference genomes. For this, ERGA launched a Pilot Project establishing the first distributed reference genome production infrastructure and testing it on 98 eukaryotic species from 33 European countries. Here we outline the infrastructure and explore its effectiveness for scaling high-quality reference genome production, whilst considering equity and inclusion. The outcomes and lessons learned provide a solid foundation for ERGA while offering key learnings to other transnational, national genomic resource projects and the EBP.
The European Reference Genome Atlas (ERGA) consortium aims to generate a reference genome catalogue for all of Europe's eukaryotic biodiversity. The biological material underlying this mission, the specimens and their derived samples, are provided through ERGA's pan-European network. To demonstrate the community's capability and capacity to realise ERGA's ambitious mission, the ERGA Pilot project was initiated. In support of the ERGA Pilot effort to generate reference genomes for European biodiversity, the ERGA Sampling and Sample Processing committee (SSP) was formed by volunteer experts from ERGA's member base. SSP aims to aid participating researchers through (i) establishing standards for and collecting of sample/specimen metadata; (ii) prioritisation of species for genome sequencing; and (iii) development of taxon-specific collection guidelines including logistics support. SSP serves as the entry point for sample providers to the ERGA genomic resource production infrastructure and guarantees that ERGA's high-quality standards are upheld throughout sample collection and processing. With the volume of researchers, projects, consortia, and organisations with interests in genomics resources expanding, this manuscript shares important experiences and lessons learned during the development of standardised operational procedures and sample provider support. The manuscript details our experiences in incorporating the FAIR and CARE principles, species prioritisation, and workflow development, which could be useful to individuals as well as other initiatives.