Insects comprise millions of species, many experiencing severe population declines under environmental and habitat changes. High-throughput approaches are crucial for accelerating our understanding of insect diversity, with DNA barcoding and high-resolution imaging showing strong potential for automatic taxonomic classification. However, most image-based approaches rely on individual specimen data, unlike the unsorted bulk samples collected in large-scale ecological surveys. We present the Mixed Arthropod Sample Segmentation and Identification (MassID45) dataset for training automatic classifiers of bulk insect samples. It uniquely combines molecular and imaging data at both the unsorted sample level and the full set of individual specimens. Human annotators, supported by an AI-assisted tool, performed two tasks on bulk images: creating segmentation masks around each individual arthropod and assigning taxonomic labels to over 17 000 specimens. Combining the taxonomic resolution of DNA barcodes with precise abundance estimates of bulk images holds great potential for rapid, large-scale characterization of insect communities. This dataset pushes the boundaries of tiny object detection and instance segmentation, fostering innovation in both ecological and machine learning research.
Caterpillar–food plant records collected over approximately 38 years in the Area de Conservación Guanacaste (ACG) in northwestern Costa Rica are described and summarized. The data comprise 431,212 individual rearing records, 197,366 of which represent unique plant–herbivore associations, i.e., same species pair found on separate dates and at different plants of the same species. These represent 29,187 different caterpillar–food plant associations between 2,489 plant and 7,160 Lepidoptera species. We evaluate changes in the taxonomic composition of the food plant flora and Lepidoptera fauna between 1990 and 2020 and across habitat/community types. Food plant and caterpillar community species richness in the rain forest changed considerably over the first 10 years but remained more stable since. Dry forest communities were more consistent than in rain forest. The cloud forest biota was the most consistent between 1995 and 2010, but as in dry forest, the caterpillar fauna changed considerably during 2015–2020. Plant species composition was more constant than caterpillar composition. The taxonomic distributions of diet specialists and generalists are explored. Most of the species-rich Lepidoptera families contain many specialists, variously concentrated throughout each family, though highly polyphagous collectively. The exceptions include Sphingidae, which show preference for Rubiaceae, Hesperiinae for monocotyledons, and non-Hesperiinae skippers for Fabaceae. Among plant families for which there are over 1,000 independent rearings, Acanthaceae, Apocynaceae, Arecaceae, Costaceae, Melastomataceae, Moraceae, Piperaceae, Poaceae, Rubiaceae, Rutaceae, and Solanaceae hosted the greatest proportion of specialists. However, the level at which dietary specialization corresponds to taxonomic rank varies with both caterpillar and plant taxon. Most fern-feeders are polyphagous with respect to fern families but still specialists on Polypodiopsida. A selection of plant families with conspicuous allelochemical and/or structural defenses and a selection of caterpillars and caterpillar families with equally conspicuous counter-defenses were examined. We determined that (1) unpalatable, aposematic herbivores tend to be specialists and (2) families of plants predominantly consumed by highly defended caterpillars host fewer polyphagous herbivores than families with less conspicuously defended plants. Highly toxic plant families with the fewest rearings, such as Aristolochiaceae and Zamiaceae, hosted many monophagous caterpillars. Biochemical and structural plant defenses appear to mediate herbivore diet breadth for many plant families.
DNA-based biodiversity surveys result in massive-scale data, including up to millions of species-of which, most are rare. Making the most of such data for inference and prediction requires modeling approaches that can relate species occurrences to environmental and spatial predictors, while incorporating information about their taxonomic or phylogenetic placement. Even if the scalability of joint species distribution models to large communities has greatly advanced, incorporating hundreds of thousands of species has not been feasible to date, leading to compromised analyses. Here we present a 'common to rare transfer learning' (CORAL) approach, based on borrowing information from the common species to enable statistically and computationally efficient modeling of both common and rare species. We illustrate that CORAL leads to much improved prediction and inference in the context of DNA metabarcoding data from Madagascar, comprising 255,188 arthropod species detected in 2,874 samples.
Global biodiversity gradients are generally expected to reflect greater species replacement closer to the equator. However, empirical validation of global biodiversity gradients largely relies on vertebrates, plants, and other less diverse taxa. Here we assess the temporal and spatial dynamics of global arthropod biodiversity dynamics using a beta-diversity framework. Sampling includes 129 sampling sites whereby malaise traps are deployed to monitor temporal changes in arthropod communities. Overall, we encountered more than 150,000 unique barcode index numbers (BINs) (i.e. species proxies). We assess between site differences in community diversity using beta-diversity and the partitioned components of species replacement and richness difference. Global total beta-diversity (dissimilarity) increases with decreasing latitude, greater spatial distance and greater temporal distance. Species replacement and richness difference patterns vary across biogeographic regions. Our findings support long-standing, general expectations of global biodiversity patterns. However, we also show that the underlying processes driving patterns may be regionally linked.
Olixon testaceum is a widely distributed species of brachypterous parasitoid wasp (Vespoidea: Rhopalosomatidae) occurring in Meso- and South America, but little is known of its biology. Here, the first known host of O. ? testaceum is identified as the cricket Anaxipha sp. (Grylloidea: Trigonidiidae) through DNA barcoding of six Olixon larvae and their hosts. Barcoding results also indicated substantial genetic diversity within nominal O. testaceum specimens. The number of species and statistical significance of these groups were tested using Maximum Likelihood phylogenies, distance-based cluster analyses, and coalescence models. All analyses revealed at least six distinct lineages, which suggests six or more cryptic species within O. ? testaceum. Combined with what is currently known about Rhopalosoma host use, these results indicate that rhopalosomatids may be generalist rather than specialist parasitoids, and further confirm the benefits of open global collaboration and DNA barcoding in advancing taxonomic knowledge.
In Project LIFEPLAN, bulk samples of arthropods are collected with Malaise traps and the species are identified with a non-destructive COI metabarcoding process. This protocol describes the steps for lysis, DNA extraction, PCR amplification, library construction, sequencing and sequence analysis, going from bulk arthropod samples to COI sequences, aiming at 1 million reads per bulk sample.
Natural history collections are the physical repositories of our knowledge on species, the entities of biodiversity. Making this knowledge accessible to society – through, for example, digitisation or the construction of a validated, global DNA barcode library – is of crucial importance. To this end, we developed and streamlined a workflow for ‘museum harvesting’ of authoritatively identified Diptera specimens from the Smithsonian Institution’s National Museum of Natural History. Our detailed workflow includes both on-site and off-site processing through specimen selection, labelling, imaging, tissue sampling, databasing and DNA barcoding. This approach was tested by harvesting and DNA barcoding 941 voucher specimens, representing 32 families, 819 genera and 695 identified species collected from 100 countries. We recovered 867 sequences (> 0 base pairs) with a sequencing success of 88.8% (727 of 819 sequenced genera gained a barcode > 300 base pairs). While Sanger-based methods were more effective for recently-collected specimens, the methods employing next-generation sequencing recovered barcodes for specimens over a century old. The utility of the newly-generated reference barcodes is demonstrated by the subsequent taxonomic assignment of nearly 5000 specimen records in the Barcode of Life Data Systems.
The parasitoid wasp genusAlphomelonMason, 1981 is revised, based on a combination of basic morphology (dichotomous key and brief diagnostic descriptions), DNA barcoding, biology (host data and wasp cocoons), and distribution data. A total of 49 species is considered; the genus is almost entirely Neotropical (48 species recorded from that region), but three species reach the Nearctic, with one of them extending as far north as 45° N in Canada.Alphomelonparasitizes exclusively Hesperiinae caterpillars (Lepidoptera: Hesperiidae), mostly feeding on monocots in the families Arecaceae, Bromeliaceae, Cannaceae, Commelinaceae, Heliconiaceae, and Poaceae. Most wasp species parasitize either on one or very few (2–4) host species, usually within one or two hesperiine genera; but some species can parasitize several hosts from up to nine different hesperiine genera. Among species with available data for their cocoons, roughly half weave solitary cocoons (16) and half are gregarious (17); cocoons tend to be surrounded by a rather distinctive, coarse silk (especially in solitary species, but also distinguishable in some gregarious species). Neither morphology nor DNA barcoding alone was sufficient on its own to delimit all species properly; by integrating all available evidence (even if incomplete, as available data for every species is different) a foundation is provided for future studies incorporating more specimens, especially from South America. The following 30new speciesare described:cruzi,itatiaiensis, andpalomae, authored by Shimbori & Fernandez-Triana; andadrianguadamuzi,amazonas,andydeansi,calixtomoragai,carolinacanoae,christerhanssoni,diniamartinezae,duvalierbricenoi,eldaarayae,eliethcantillanoae,gloriasihezarae,guillermopereirai,hazelcambroneroae,josecortesi,keineraragoni,luciarosae,manuelriosi,mikesharkeyi,osvaldoespinozai,paramelanoscelis,paranigriceps,petronariosae,ricardocaleroi,rigoi,rostermoragai,sergioriosi, andyanayacu, authored by Fernandez-Triana & Shimbori.
Lifeplan is a global biodiversity monitoring project with the aim of assessing the current state of biodiversity worldwide, and using this knowledge to generate predictions of how biodiversity might look in the future. In this protocol we describe the materials and method used to sample flying insects with a Malaise trap for the Global Malaise Trap Project and the LIFEPLAN project on a global scale and in a wide variety of environmental conditions and habitats. The aim is to identify species in further analysis (e.g. image recognition, DNA sequencing) and create species lists for different locations across the globe. This protocol contains a detailed description from setup of the Malaise traps in the field to weekly data collection, as well as steps to reduce ethanol in the samples and collect necessary information to help in species identification. We identify the equipment used in Lifeplan, but also give technical specifications of that equipment so that other users of this protocol can find equivalent alternative equipment. We also specify what metadata should be collected with the Malaise trap data. The technical solution we use to collect metadata in the Lifeplan project is described in detail in the full Lifeplan protocol. It is critical that we employ standardized operating procedures for the Malaise trapping. Our coordinated efforts will ensure specimen preservation for sequence analysis and high data quality, permitting the comparison of sites at a global scale. For global standardization with the BIOSCANinitiative, of which LIFEPLAN is a part. LIFEPLAN is based on bulk processing (metabarcoding) of samples and automatic image recognition which are outside of the scope of this protocol and will be described elsewhere.
Introduction: Species of Mesochorus are found worldwide and members of this genus are primarily hyperparasitoids of Ichneumonoidea and Tachinidae. Objectives: To describe species of Costa Rican Mesochorus reared from caterpillars and to a lesser extent Malaise-trapped. Methods: The species are diagnosed by COI mtDNA barcodes, morphological inspection, and host data. A suite of images and host data (plant, caterpillar, and primary parasitoid) are provided for each species. Results: A total of 158 new species of Mesochorus. Sharkey is the taxonomic authority for all. Conclusions: This demonstrates a practical application of DNA barcoding that can be applied to the masses of undescribed neotropical insect species in hyperdiverse groups.
The use of DNA barcoding has revolutionised biodiversity science, but its application depends on the existence of comprehensive and reliable reference libraries. For many poorly known taxa, such reference sequences are missing even at higher-level taxonomic scales. We harvested the collections of the Smithsonian's National Museum of Natural History (USNM) to generate DNA barcoding sequences for genera of terrestrial arthropods previously not recorded in one or more major public sequence databases. Our workflow used a mix of Sanger and Next-Generation Sequencing (NGS) approaches to maximise sequence recovery while ensuring affordable cost. In total, COI sequences were obtained for 5,686 specimens belonging to 3,737 determined species in 3,886 genera and 205 families distributed in 137 countries. Success rates varied widely according to collection data and focal taxon. NGS helped recover sequences of specimens that failed a previous run of Sanger sequencing. Success rates and the optimal balance between Sanger and NGS are the most important drivers to maximise output and minimise cost in future projects. The corresponding sequence and taxonomic data can be accessed through the Barcode of Life Data System, GenBank, the Global Biodiversity Information Facility, the Global Genome Biodiversity Network Data Portal and the NMNH data portal.