DNA barcode reference libraries provide useful tools for specimen identification, highlighting potential new species and detecting introduced ones. Here, we present a comprehensive DNA barcode library for European ants and, in order to tackle the Linnean, Wallacean and Darwinian shortfalls of this group, we provide an updated checklist, distribution data, mitochondrial genetic diversity maps and mitochondrial gene trees. The European ant fauna is here established to include 55 genera and 650 species (587 of which are native), including one species newly recorded for Europe and novel citations for 26 species from 11 countries. Our genetic dataset includes 6530 georeferenced COI sequences (62.1% d e novo) for 506 species (77.8%) across all genera. On average, 12.9 sequences were obtained per species, and 209 species were sequenced for the first time. We generated intra- and interspecific genetic distance estimates, 52 genus-level trees, mitochondrial genetic diversity and specimen maps for 384 species, as well as haplotype networks for 289 species, available in the Atlas V1.0 'The Mitochondrial Genetic Diversity Maps of European Ants'. We estimate that 56.3% of European ants are monophyletic with respect to the COI gene and can be unambiguously identified by DNA barcoding, though performance varies widely among genera. We observed moderate levels of barcode sharing (19.3%) and of barcode gap presence (47.6%), as well as high levels of intraspecific divergences (up to 17.9%). These findings likely reflect both biological and operational factors and highlight the existence of potential cryptic taxa and the need for taxonomic revisions. The framework presented here aims to facilitate future research, species discovery and conservation of European ants.
Lebanon's diverse landscapes and rich ecological zones support a high diversity of insect species, many of which remain taxonomically unresolved. To address this knowledge gap, we conducted the first comprehensive DNA barcoding survey of Lebanese insect diversity. Specimens collected using Malaise traps deployed at four sites between 2019 and 2021 were analyzed for sequence variation in the 658 bp barcode region of the mitochondrial cytochrome c oxidase subunit I (COI) gene. The 58 123 insect specimens included representatives of 5747 Barcode Index Numbers (BINs), and 56% of them are unique to Lebanon. By integrating DNA barcoding with morphological taxonomy, we assigned each BIN to one of 20 insect orders and identified 1224 species. Most specimens (95%) belonged to five dominant orders: Diptera, Hymenoptera, Hemiptera, Coleoptera, and Lepidoptera. The high proportion of unique BINs suggests high endemism within Lebanon and the broader Levant region. By generating a foundational inventory for Lebanon's insect fauna, this study has provided baseline data critical for future biodiversity monitoring and conservation planning in the Eastern Mediterranean while also making a substantial contribution to the global DNA barcode reference library.
A bstract Wild silkmoths (Saturniidae) are one of the most emblematic and most studied families of moths. Yet, the absence of a robust phylogenetic framework based on a comprehensive taxonomic sampling impedes our understanding of their evolutionary history. We analyzed 1,024 ultraconserved elements (UCEs) and their flanking regions to infer the relationships among 338 species of Saturniidae representing all described subfamilies, tribes, and genera. We investigated systematic biases in genomic data and performed dating and historical biogeographic analyses to reconstruct the evolutionary history of wild silkmoths in space and time. Using Gene Genealogy Interrogation, we showed that saturation of nucleotide sequence data blurred our understanding of early divergences and first biogeographic events. Our analyses support a Neotropical origin of saturniids, shortly after the Cretaceous-Paleogene extinction event ( ca 64.0 [stem] - 52.0 [crown] Ma), and two independent colonization events of the Old World during the Eocene, presumably through the Bering Land Bridge. Early divergences strongly shaped the distribution of extant subfamilies as they showed very limited mobility across biogeographical regions, except for Saturniinae, a subfamily now present on all continents but Antarctica. Overall, our results provide a framework for in-depth investigations into the spatial and temporal dynamics of all saturniid lineages and for the integration of their evolutionary history into further global studies of biodiversity and conservation. Rather unexpectedly for a taxonomically well-known family such as Saturniidae, the proper alignment of taxonomic divisions and ranks with our phylogenetic results leads us to propose substantial rearrangements of the family classification, including the description of one new subfamily and two new tribes.
Although insects are fundamental to understanding and conserving global biodiversity, they are vastly understudied. Here, we present a national inventory of Costa Rican insects based upon 3.78 million DNA barcodes representing 152 891 Barcode Index Numbers (BINs, proxies for species) from 28 localities sampled from 2017 to 2023 through the national BioAlfa program of Costa Rica. Although only 3.6% of BINs are linked to Linnean species, barcode-based community analyses revealed strong, consistent ecogeographic structure. Clustering of BIN data using bootstrapped Jaccard distances revealed seven distinct mainland assemblages and a distinct island cluster, shaped primarily by Costa Rica’s mountain ranges, elevation, and slope orientation. Separate analyses for Coleoptera, Diptera, Hemiptera, Hymenoptera, and Lepidoptera coupled with analyses focused on some of their largest families (e.g., Braconidae, Cecidomyiidae, Cicadellidae, Erebidae, and Staphylinidae) confirmed these patterns and further revealed extremely high species turnover with most BINs being exclusive to a single region or locality. Our results reveal limited overlap of insect communities across ecosystems, implying that each life zone harbors unique taxonomic assemblages. Large-scale DNA barcoding has detected fine-grained spatial structure, providing a genomic framework for biodiversity monitoring and conservation in diverse tropical regions undergoing rapid environmental change.
How best to rapidly assess the diversity of terrestrial arthropod communities has attracted considerable debate. In particular, what sampling methods should be combined to capture maximum species richness and sampling efficiency in surveys of arthropods. We tested the efficiency and complementarity of a standardised combination of six commonly used sampling methods (Malaise trap, flight intercept trap, pan trap, pitfall trap, Berlese funnel, and sweep net) at 13 locations across Canada. About 55% of species were collected by only one method; Malaise traps, sweep nets, and flight intercept traps contributed the highest diversity. We suggest the combined use of these collecting methods, possibly extended by other methods, for studies that seek to survey the canopy and soil fauna.
Abstract Understanding global biodiversity patterns and their drivers is a prerequisite for countering the biodiversity crisis. In this paper, we introduce a novel generalized linear model, Hubbell regression, to estimate a key biodiversity descriptor, the fundamental biodiversity number. This can be converted into a set of biodiversity descriptors, including Shannon and Simpson indices, and more. Hence, quantifying the impact of environmental conditions on the fundamental biodiversity number allows us to predict the general properties of local biodiversity in any setting. In addition to having a strong mathematical foundation, Hubbell regression consistently outperformed current state-of-the-art models in predicting global biodiversity. We apply the method to arthropods, which account for the majority of terrestrial biodiversity. By parameterizing the models using samples of 1.78 million arthropods from 2415 samples collected at 135 sites spanning all continents, we pinpoint the drivers of arthropod biodiversity and its features at the global scale. We find that actual evapotranspiration is the single largest predictor of arthropod diversity and explains nearly 30% of the variation in richness. Moreover, we infer that high human activity has led to a 21.3 % and 29.2% decrease in potential insect richness in tropical and dry zones, respectively, but increased insect richness in polar regions. These insights bring a new foundation for biodiversity research and action.
ABSTRACT DNA barcoding involves the recovery of a DNA sequence for a target gene region from its source specimen. This process gains complexity when multiple sequences are recovered from a specimen, as is often the case when data are generated by high-throughput sequencers. This diversity can reflect both methodological artifacts (e.g., chimeras, PCR errors, sequencing errors, tag jumps) and real template diversity in the DNA extract (e.g., contamination, endosymbionts, NUMTs, parasites). To support analysis of the sequence data from three million specimens annually, the Centre for Biodiversity Genomics (CBG) has developed BIP, the Barcode Inference Pipeline. Compatible with all sequencing platforms, BIP processes .fastq files and returns both target DNA barcodes and non-target sequences. To generate results, BIP implements quality and size filtration, demultiplexing, primer trimming, chimera scanning, sequence error correction, OTU delineation, and sequence identification. When analysis targets the cytochrome c oxidase 1 (COI) barcode region, BIP also assigns each OTU to a known BIN or identifies its nearest neighbour BIN. As final output, BIP returns summary files ready for upload to BOLD or for other downstream analyses. They include a taxonomic assignment for each OTU, generated by comparison with a DNA barcode reference library. We describe BIP’s flexibility and structure, then demonstrate its functionality by analyzing COI sequence data from 100K specimens. Because of its capacity to disentangle target and non-target sequences, BIP outperforms an alternative software package, ONTbarcoder, in several important ways. To ease access, installation, and functionality, BIP is provided as a Docker container (github.com/cbg-innov/BIP).
ABSTRACT Aim Insects provide essential ecosystem services, but there is a concern that insect abundance is in rapid decline. Identifying trends in insect biodiversity is challenging because we lack baseline data for most biogeographic regions, and most insect species remain undescribed. Here, we aim to report the first nationwide inventory of insect biodiversity in South Africa. Location Eleven National Botanical Gardens representing four major South African biomes: Savanna, Grassland, Fynbos and Albany Thicket. Time Period Sampling from 1 February 2022 to 31 January 2023. Major Taxa Studied Insects (Arthropods) and vascular plants. Methods We deployed two Malaise traps per botanical garden and used DNA barcoding to identify insect specimens. We reconstructed phylogenies for insect molecular clusters (BINs) and 12,149 vascular plant species. Insect and plant community structure were compared using dissimilarity metrics and Mantel tests. Results We generated DNA barcodes for 337,809 specimens, representing 29,919 unique molecular clusters, as species proxies. We estimated that the true richness of above ground insect biodiversity in South Africa may be closer to 100,000 species, yet 97% of barcoded taxa could not be assigned to species level. At current rates of taxonomic description, formally cataloguing this dark diversity would take more than 200 years. Estimated insect species richness varied strongly among biomes, with higher insect diversity in the Savanna biome than the floristically rich Fynbos. Turnover in plant lineages was associated with insect community turnover, particularly insect herbivores and predators. Main Conclusions We show that plant species richness alone is a poor predictor of insect species richness, with biome‐specific factors playing an important role in shaping insect richness gradients. However, turnover in plant lineages was a strong predictor of insect compositional differences between gardens. Our study provides both critical baseline data on insect diversity in South Africa and new insights into the ecological and evolutionary drivers of insect biodiversity gradients.
We present the first dataset collection of arthropod diversity from ten terrestrial sites in the Neotropical Eastern Pacific Bioregion, in the territories of Costa Rica, Ecuador, Panama, and the Galapagos and Cocos Islands. The twenty-five datasets comprise 1-7 years of sampling. Together, we document a total of 952,087 arthropod specimens and 45,812 BINs. The datasets are distributed in 1 phylum, 7 classes, 39 orders, 556 families, and 2,318 genera. In the first year of sampling, the datasets revealed unprecedented arthropod diversity; notably, at Mashpi in Ecuador (n = 7,848 BINs) and Guanacaste in Costa Rica (n = 7,790 BINs), both exhibited the highest number of unique BINs. Arthropod abundance was greatest at Quetzales (n = 130,193 records) and Baru (n = 111,044 records) in Costa Rica. The five most abundant orders were Diptera (n = 20,901 BINs), Hymenoptera (n = 8,964 BINs), Coleoptera (n = 4,272 BINs), Lepidoptera (n = 3,634 BINs), and Hemiptera (n = 2,166 BINs). The dataset collection provides a robust baseline for future arthropod biodiversity research.
Species descriptions remain the foundation of biodiversity science, but alpha taxonomy faces a twofold challenge: accelerating the pace of species discovery while ensuring that descriptions produce data that are usable, comparable, and reproducible. This is particularly acute for “dark taxa”—small, hyperdiverse, and morphologically challenging groups—for which morphology-based workflows alone often fail to deliver scalable and reliable identifications. As a result, most species remain undescribed, and many described species are difficult or impossible to identify. We argue that species descriptions must be integrative in a way that meets the practical requirements of modern taxonomy, including scalability, reliable identification, reproducibility, and applicability across life stages. At present, standardized molecular data in the form of DNA barcoding is the only approach that consistently satisfies these criteria. Emerging approaches based on robotics, imaging and artificial intelligence may eventually provide complementary frameworks, but currently lack the standardization and interoperability required for routine application at scale. In contrast, barcode data already enable high-throughput species delimitation, straightforward identification, and direct comparability across studies, supported by global reference libraries and standardized protocols. We therefore propose that DNA barcodes be treated as a baseline component of species descriptions for invertebrates. Standardizing molecular data in taxonomy will accelerate species discovery, improve reproducibility and stability, and ensure that newly described taxa remain accessible in an increasingly DNA-based research landscape.
Viral domestication, the co-option of integrated viral genes for host functions, has been repeatedly documented in endoparasitoid Hymenoptera. Whether this phenomenon extends to other parasitoid insects remains to be assessed. Here, we tested this hypothesis in tachinid flies, the largest group of non-hymenopteran parasitoids. To this end, we investigated patterns of viral endogenization in 52 genomes, including 37 newly sequenced species. Similarly to hymenopteran endoparasitoids, tachinid genomes were found to display numerous endogenous viral elements (EVEs), primarily derived from insect-infecting viruses with either DNA or RNA genomes. However, the majority of integrated sequences were species-specific, with no evidence of major ancient gene domestication events. In addition, the EVE content did not differ between dipteran parasitoids and their free-living Diptera counterparts. These findings contrast with the recurrent viral domestication observed in endoparasitoid Hymenoptera. This pattern is discussed in light of important attributes of tachinid flies including their lack of venom and perforating ovipositor, and their avoidance of the host immune system.
This study presents the first comprehensive molecular assessment of the Lepidoptera fauna of Austria based on DNA barcodes (cytochrome c oxidase I, COI; 658 bp Folmer region). The barcode reference library comprises approximately 23,500 sequences, representing 3591 Linnaean species or about 85% of the known national species (ca. 4200 species). Congruence between morphological species identifications under the Linnaean system and barcode data was evaluated using the Barcode Index Number (BIN) system. A total of 244 species could not be unambiguously assigned, showing two to seven BINs that exhibit elevated genetic divergence and may partially represent cryptic diversity. These taxa, together with 40 currently unnamed lineages, require further integrative taxonomic assessment. The distinctiveness of the Austrian Lepidoptera fauna is discussed in the context of endemic genetic diversity. Finally, 17 new faunistic records for Austria are reported.
DNA barcode data are essential infrastructure for biodiversity science as they enable scalable species identification and integrative analyses that link sequences to specimens, taxonomy, and geography. The Barcode of Life Data System (BOLD) is a centralized bioinformatics workbench that supports the full barcode data lifecycle: including acquisition, storage, validation, analysis, and dataset publication within a secure collaboration model. While BOLD's web interface is optimized for interactive dataset assembly and publication, many researchers require programmatic access to both public data and permissioned private records to build reproducible pipelines for curation and analysis prior to release. We introduce BOLDconnectR, an R package that provides authenticated access to BOLD, returns records in the Barcode Core Data Model (BCDM), performs automated transformation into commonly used R data structures, and enables customizable analytic workflows that generate publication-ready outputs.
ABSTRACT Aim Our research aimed to assess the beta diversity of Chironomidae (Diptera) and faunal connections across springs, streams, and rivers within the Skadar Lake basin, a Mediterranean biodiversity hotspot, using DNA barcoding. Location Skadar Lake basin—a Mediterranean biodiversity hotspot. Methods We analysed mitochondrial cytochrome c oxidase subunit I sequences from larvae collected at 56 sites, employing Barcode Index Numbers (BINs) as proxies for species. Results Our analysis revealed high overall diversity (224 BINs) along high site heterogeneity and a large proportion of rare BINs. Nearly two‐thirds of BINs were represented by fewer than four specimens, and 84% occurred at three or fewer sites. We observed extremely low faunal overlap between ecosystems, with only nine BINs shared among springs, streams, and rivers. Springs had the highest richness (169 BINs) and exceptional uniqueness, hosting 126 exclusive BINs (ca. 56%). No significant large‐scale geographical or altitudinal diversity gradients were detected, suggesting local factors override broader patterns. Main Conclusions The major conclusion is that hydrogeological isolation within the karst landscape profoundly shapes chironomid diversity, creating unique and isolated spring communities. These distinct, often rare assemblages indicate that chironomids are a sensitive indicator of habitat fragmentation and the vulnerability of these unique Mediterranean freshwater ecosystems to environmental change and localised impacts.
The International Barcode of Life (iBOL) initiative is building a globally accessible DNA-based system for species identification and discovery. This paper outlines the mission and strategic priorities for the iBOL community in Europe (iBOL Europe), set in a global context. The mission of iBOL Europe is to produce, curate, and provide access to a complete DNA barcode reference library of European eukaryotic biodiversity, catalyzing species discovery and enabling comprehensive, harmonized species identification and biomonitoring, and supporting the global iBOL program. Immediate objectives include completing reference libraries for priority taxa, democratizing access to sequencing technologies, and strengthening a distributed community of practice. Key actions identified span five thematic areas: community building, sample collection and taxonomic verification, sequencing infrastructure, data management, and mainstreaming DNA-based approaches to meet societal needs. The strategy emphasizes integration with European research infrastructures to ensure long-term sustainability and resilience for biodiversity genomics in Europe.
Estimating the number of insect species on Earth is a daunting challenge. The current consensus estimate—about six million species—is likely far too low, as we will show. Our estimate of the global number of insect species rests on a sample of more than 1,600,000 DNA-barcoded insect specimens representing 53,945 species from 15 “core” Malaise traps deployed in dry forest, cloud forest, and rainforest ecosystems of the Área de Conservación Guanacaste (ACG) in Costa Rica. Even this massive sample fails to reveal the full extent of ACG insect species richness. To estimate total ACG insect richness, we adjust the observed count of insect species by an “undersampling ratio,” computed for a hyperdiverse subfamily of parasitoid wasps (Braconidae: Microgastrinae). The ratio compares microgastrine richness from the core Malaise traps to a lower-bound estimate of true microgastrine richness—including undetected species—based on 21,669 specimens from three sources: the 15 core Malaise traps, 15 “peripheral” Malaise traps spanning all three ecosystems, and 11,373 DNA-barcoded specimens reared from some 1,500 species of microgastrine-parasitized caterpillars (Lepidoptera). To estimate global insect richness, we apply Earth/ACG ratios for tree species and several animal taxa to upscale our estimate of ACG insect richness (nearly 333,000 species). Adopting conservative assumptions, we reach an estimate of 14 to 20 million insect species on Earth, depending on the upscaling group—two to three times the current consensus estimates. Upscaling instead from a point estimate of ACG richness with a wide CI, global estimates reach nearly 30 million species.
ABSTRACT Current metabarcoding pipelines are inflexible with respect to study design and are poorly suited to long-read sequence data. To address these limitations, we developed MAP, the Metabarcoding Analysis Pipeline, which is a sequence-to-answer workflow supporting the analysis of amplicons from highly multiplexed and replicated study designs. Although MAP can analyze amplicons of any length from any genetic marker, it includes several features tailored to long-read COI metabarcoding. MAP installs from a Docker container and requires only sequence data, a parameters file, and a reference library. It produces intuitive reports, enabling users to evaluate their data immediately after analysis. We validate MAP by showing that it generates biodiversity estimates that correspond closely to a ground-truth dataset of single-specimen DNA barcode data and by demonstrating that it outperforms alternative platforms for COI metabarcoding. MAP is free, open-source, and available from: https://github.com/cbg-innov/MAP .
PCR amplification of mixed-template samples can generate chimeric sequences, molecules that combine partial sequences from two or more source organisms. These artefacts pose a significant challenge for high-throughput sequencing studies, generating both type I errors (true species mistakenly rejected) and type II errors (chimeras incorrectly viewed as valid species). Despite their impact, the incidence, structural properties, and factors driving chimera formation remain poorly understood. Using nanopore sequencing of 531 nematode-infected arthropod specimens, chimera formation was characterized in two widely used gene regions: the 658 bp of cytochrome c oxidase subunit I (COI) barcode and a 900 bp segment of 18S ribosomal DNA (18S rDNA). Chimeras were detected for both gene regions despite the deep sequence divergences between members of these two phyla. However, the incidence of chimeric molecules was higher for 18S rDNA than COI with chimera formation correlating with local variation in DNA stability and conserved sequence tracts between parent species. By clarifying the associations between DNA stability, sequence conservation, and chimera formation, this study offers insights that can inform the development of more robust approaches to chimera detection.
Human lice can be classified into several mitochondrial clades with distinct geographic distributions, reflecting long-term coevolution with ancient hominids and possibly with human populations. Analysis of ancient lice DNA from archaeological human remains can provide insights into the geographic origins of these clades and the timing of their dispersal. This study analyzes the COI barcode region of lice recovered from the ~1500-year-old salt-preserved mummy (Salt Man 2) at the Chehrabad salt mine in Zanjan Province, Iran. The sequences were compared with those of recent lice from Iran and other regions, available in the GenBank database, to determine clade affiliation and possible historical relationships. Our phylogenetic analysis recovered the five recognized mitochondrial clades (A–E). Ancient lice clustered within clade A but displayed novel haplotypes. Recent nits from Gilan and Kurdistan provinces were grouped within clades A and B, the only clades reported in Iran. The most complete ancient sequence (MLICE005-18) showed close similarity to sequences from Iraq, Pakistan, Egypt, and Iran. These findings, along with future research, could help determine when ancient lice lineages were established or introduced into specific regions. They could also clarify whether they represent haplotypes that have since disappeared or a lineage that still exists in understudied human populations.
The Thai species of the genus Pristaulacus Kieffer, 1900 are reviewed, and 17 species are recognized. We describe and illustrate a new species, P. chiangmaiensis Turrisi & Ghafouri Moghaddam, sp. nov., using an integrative taxonomic approach. One species, P. vilhelmseni Turrisi & Smith, 2011, is newly recorded from Thailand. We also provide diagnoses, illustrations, and an illustrated and updated identification key for Thai Pristaulacus species. Additionally, DNA barcoding data are presented for several Thai and other Pristaulacus species viz., P. chiangmaiensis, P. erythrocephalus Cameron, 1905, P. fasciatus (Say, 1829), P. flavicrurus (Bradley, 1901), P. gloriator (Fabricius, 1804), P. intermedius Uchida, 1932, P. cf. rufipes Enderlein, 1912, P. rufitarsis (Cresson, 1864), P. stigmaterus (Cresson, 1864), P. vilhelmseni Turrisi & Smith, 2011, which are included in a phylogenetic analysis for the first time. A comprehensive distribution map of all known Thai Pristaulacus species is also provided. An updated checklist of Pristaulacus species from the Oriental region is tabulated and briefly discussed.