Widespread degradation of nature has increased pressure on corporations and financial institutions to assess and mitigate their biodiversity impact, however, collecting relevant local data can be costly. The increasing availability of biodiversity data and Earth observation (EO) data provides a route to cost effective impact assessment via extrapolation using existing data and statistical modeling. Through a review of the datasets and tools currently used by corporations and financial institutions we show that extrapolation to local sites from global datasets, or using only proxies, is the dominant approach. We test the reliability of such assessments by combining high resolution earth observation time series data with extensive biodiversity data from recent environmental DNA (eDNA) surveys of two countries with widely varying conditions, Sweden and Madagascar. We use machine learning in combination with high-quality biodiversity data to predict five essential biodiversity variables (EBVs) for local sites, using cross-validation to test prediction accuracy. The results show that reasonably accurate EBV predictions can be obtained for sites with some local data, but performance declines considerably when modelling summary measures at new sites. Moreover, the quality of predictions, both within sites and at new sites, is dependent on the EBV and local context. To address the concerns over the reliability of model-based EBV assessments, we propose a biodiversity data hierarchy framework, which can be used by organisations to track stepwise improvements in the data sources underpinning their biodiversity impact assessments.
Abstract DNA metabarcoding—high‐throughput sequencing of barcode regions from bulk samples—has become a key tool for insect biodiversity assessment. Yet, how methodological choices affect the accuracy of metabarcoding data remains insufficiently explored. In this paper, we ask: (1) How does the lysis method (non‐destructive lysis vs. destructive homogenization) affect community recovery? (2) How comprehensively does metabarcoding capture species richness? (3) To what extent can spike‐ins improve abundance estimates? (4) How accurately can species abundances be estimated? We evaluated the accuracy of insect metabarcoding using 4749 bulk samples from a large‐scale biodiversity survey subjected to mild lysis. Of these samples, 856 were also homogenized, allowing a systematic comparison of the effect of alternative treatments. To potentially improve abundance estimates, we added six biological spike‐ins (i.e. foreign insects) to all samples, and two synthetic spike‐ins (artificial DNA fragments) to the homogenization treatment. In addition, we established the contents of 15 samples by individually barcoding all specimens, enabling direct assessment of occurrence and abundance estimates. Our results revealed consistent differences between destructive and non‐destructive treatments. While both methods reliably detected the majority of species, small and soft‐bodied taxa were more often recovered after mild lysis than after homogenization, while the reverse was true for heavily sclerotized, hairy and large taxa. Using biological spike‐ins for calibration reduced the variance in read numbers per specimen considerably, especially in homogenized samples, while synthetic spike‐ins were less effective. In a Bayesian analysis, where species data were matched to the best‐fitting spike‐in calibration curve, accurate abundance estimates (±1 individual) were obtained for 72.9% of species occurrences. Our results show that it is possible to obtain reasonably accurate abundance estimates from metabarcoding data and that mild lysis and homogenization result in different taxon‐specific biases in terms of occurrence data, with neither method outperforming the other. Abundance accuracy is improved by homogenization rather than mild lysis of samples, and by the use of biological rather than synthetic spike‐ins. Together, these findings provide a major step towards robust, quantitative biodiversity monitoring using DNA‐metabarcoding.
This is an erratum for the article “Annotated Automatic Pruning of Universal Probabilistic Programming Languages” published in ACM Trans. Probab. Mach. Learn. 1, 3, Article 15 (August 2025), 33 pages.
An important aspect of making inference based on a probabilistic program practical is efficiency; faster evaluation enables more work per unit of time, which can be translated into more precision. Inference via Markov chain Monte Carlo has a property that can be favorably exploited for efficiency: most proposed samples are computed as minor variations of previous samples, i.e., a clever implementation can skip computations pertaining to what is unchanged. This paper provides an approach for automatically translating a probabilistic program to a dynamic graph, reminiscent of functional reactive programming, that explicitly represents data dependencies, enabling proposals to only recompute the parts of the graph that depend on redrawn random variables. The graph-building interface follows familiar functional programming interfaces, which also connect to their expressiveness in terms of probabilistic programming: models using the applicative functor portion express Bayesian networks, while those using monads represent universal probabilistic programming languages.
How communities are structured into functional groups and trophic layers is key to understanding ecosystem functioning. Nonetheless, we lack insights about spatiotemporal variation in guild composition of communities and its causes. To investigate spatial and temporal patterns and drivers of variation in insect feeding guilds, we combined data from a nationwide survey of Swedish insects using Malaise traps and DNA metabarcoding with a comprehensive trait database. We assigned species into one of three feeding guilds (phytophages, saprophages, predators) or into one of three associated parasitoid guilds. We then analysed patterns in species richness for each guild. Species richness declined with latitude in all guilds. Beyond this gradient, local variation in species richness matched between hosts and their parasitoids. Yet, hosts and their parasitoids responded differently to habitat. The phenological peak of parasitoid species richness appeared later than the peak of their hosts, but the length of time lags varied among guilds. Spatiotemporal patterns were driven by guild-specific responses to temperature, though much variation remained between seasons and locations even when controlling for temperature. Overall, these patterns suggest that shifts in both climate and land use may alter the synchrony of insect trophic layers, with unknown consequences.
Biodiversity impact assessments aim to enable market actors, regulators, and political agents to effectively steer human activities in a more sustainable direction. However, current biodiversity impact assessments often rely on biased, incomplete, or indirect data. We review the potential of addressing these shortcomings using emerging methods based on environmental DNA (eDNA). The eDNA technologies are developing rapidly, and DNA metabarcoding is now sufficiently mature to allow cost-effective, standardized recording of detailed local biodiversity data. The eDNA data allow computation of a wide range of the essential biodiversity variables, and open data reporting mechanisms are already in place. If companies were required to collect and openly report eDNA data documenting their impact, they could optimize biodiversity outcomes in relation to productivity and other factors in a fast development cycle. Simultaneously, publicly funded research could focus on analyzing the data and successively refining actionable metrics based on fundamental ecological principles.
Deep metabarcoding offers an efficient and reproducible approach to biodiversity monitoring, but noisy data and incomplete reference databases challenge accurate diversity estimation and taxonomic annotation. Here, we introduce a novel algorithm, NEEAT, for removing spurious operational taxonomic units (OTUs) originating from nuclear-embedded mitochondrial DNA sequences (NUMTs) or sequencing errors. It integrates 'echo' signals across samples with the identification of unusual evolutionary patterns among similar DNA sequences. We also extensively benchmark current tools for chimera removal, taxonomic annotation and OTU clustering of deep metabarcoding data. The best performing tools/parameter settings are integrated into HAPP, a high-accuracy pipeline for processing deep metabarcoding data. Tests using CO1 data from BOLD and large-scale metabarcoding data on insects demonstrate that HAPP significantly outperforms existing methods, while enabling efficient analysis of extensive datasets by parallelizing computations across taxonomic groups.
1. The private sector is increasingly aware of its dependence on biodiversity and the financial risks and opportunities involved. This has generated a lot of demand for investing in nature-positive solutions. There is an obvious and non-negotiable basis for such initiatives: biodiversity data. Without this data and the tools built from it, no actor can assess the effects on the ecosystems they rely on. We identify two key barriers to corporate biodiversity action: (1) lack of biodiversity data and (2) challenges with biodiversity data literacy, i.e. the domain knowledge necessary to apply data products for decision making in appropriate contexts. Building on this, we present an end-to-end framework mapping biodiversity data to data products and business use cases, to establish a shared language between business and biodiversity research. 2. First, we provide examples of new technologies for generating biodiversity data at unprecedented scales, such as environmental DNA, computer vision and audio monitoring. We discuss the large amount of biodiversity data available in open databases, with a focus on the Global Biodiversity Information Facility (GBIF), including their origins, limitations, and biases. We highlight the one billion untapped primary biodiversity data points in natural history collections, and the opportunity to mobilise them into open databases using technology at relatively low cost. 3. Second, we discuss biodiversity data products, focusing on the ability to interpret, communicate, and effectively apply biodiversity models, metrics, and tools in relevant contexts. We address the challenges posed by the complexity of biodiversity, the importance of its definitions, and the use of aggregated metrics for biodiversity and ecosystem services in reporting, including the role of nature tech. 4. Third, we present the business case for investing in more and open biodiversity data, with examples of actions by companies and the finance sector. We also propose a mechanism to incentivise and reward direct investments in biodiversity data mobilisation. In conclusion, we call on businesses to prioritise financial investment in biodiversity data collection and mobilisation, to create better data products that can accelerate deployment of solutions to the biodiversity crisis.
DNA metabarcoding of species-rich taxa is becoming a popular high-throughput method for biodiversity inventories. Unfortunately, its accuracy and efficiency remain unclear, as results mostly pertain to poorly known taxa in underexplored regions. This study evaluates what an extensive sampling effort combined with metabarcoding can tell us about the lepidopteran fauna of Sweden-one of the best-understood insect taxa in one of the most-surveyed countries of the world. We deployed 197 Malaise traps across Sweden for a year, generating 4749 bulk samples for metabarcoding, and compared the results to existing data sources. We detected more than half (1535) of the 2990 known Swedish lepidopteran species and 323 species not reported during the sampling period by other data providers. Full-length barcoding confirmed three new species for the country, substantial range extensions for two species and eight genetically distinct barcode variants potentially representing new species, one of which has since been described. Most new records represented small, inconspicuous species from poorly surveyed regions, highlighting components of the fauna overlooked by traditional surveying. These findings demonstrate that DNA metabarcoding is a highly efficient and accurate biodiversity sampling method, capable of yielding significant new discoveries even for the most well known of insect faunas.
Global change threatens a vast number of species with severe population declines or even extinction. The threat status of an organism is often designated based on geographic range, population size, or declines in either. However, invertebrates, which comprise the bulk of animal diversity, are conspicuously absent from global frameworks that assess extinction risk. Many invertebrates are hard to study, and it has been questioned whether current risk assessments are appropriate for the majority of these organisms. As the majority of invertebrates are rare, we contend that the lack of data for these organisms makes current criteria hard to apply. Using empirical evidence from one of the largest terrestrial arthropod surveys to date, consisting of over 33 000 species collected from over a million hours of survey effort, we demonstrate that estimates of trends based on low sample sizes are associated with major uncertainty and a risk of misclassification under criteria defined by the IUCN. We argue that even the most ambitious monitoring efforts are unlikely to produce enough observations to reliably estimate population sizes and ranges for more than a fraction of species, and there is likely to be substantial uncertainty in assessing risk for the majority of global biodiversity using species-level trends. In response, we discuss the need to focus on metrics we can currently measure when conducting risk assessments for these organisms. We highlight modern statistical methods that allow quantification of metrics that could incorporate observations of rare invertebrates into global conservation frameworks, and suggest how current criteria might be adapted to meet the needs of the majority of global biodiversity.
We present the data from the Insect Biome Atlas project (IBA), characterizing the terrestrial arthropod faunas of Sweden and Madagascar. Over 12 months, Malaise trap samples were collected weekly (biweekly or monthly in the winter, when feasible) at 203 locations within 100 sites in Sweden and weekly at 50 locations within 33 sites in Madagascar; this was complemented by soil and litter samples from each site. The field samples comprise 4,749 Malaise trap, 192 soil and 192 litter samples from Sweden and 2,566 Malaise trap and 190 litter samples from Madagascar. Samples were processed using mild lysis or homogenization, followed by DNA metabarcoding of CO1 (418 bp). The data comprise 698,378 non-chimeric sequence variants from Sweden and 687,866 from Madagascar, representing 33,989 (33,046 Arthropoda) and 77,599 (77,380 Arthropoda) operational taxonomic units, respectively. These are the most comprehensive data presented on these faunas so far, allowing unique analyses of the size, composition, spatial turnover and seasonal dynamics of the sampled communities. They also provide an invaluable baseline against which to gauge future changes.
Obtaining genome-wide data from complex samples, such as environmental material or bulk species collections, is increasingly feasible, yet inferring species presence and population genomic insights remains challenging. We applied metagenomic sequencing to 40 arthropod bulk samples collected with Malaise traps across Sweden and compared results with metabarcoding of the same material. Using a custom genome database, we achieved genus-level classification largely consistent with metabarcoding. While metagenomics detected all genera identified by metabarcoding, conservative filtering thresholds designed to minimise false positives also excluded some true signals, particularly for low-abundance taxa. Taxonomic overlap between methods was further constrained by limited reference database representation. Beyond taxonomic assignment, metagenomic sequencing yielded genome-level information: we inferred haplotype diversity, heterozygosity and geographic population structure for several abundant species, including variable degrees of hybrid origin in red wood ants and the genetic distinctiveness of Gotland bumblebees. Finally, by-catch plant DNA present in the bulk samples revealed plausible arthropod-plant interactions, several of which align with known ecological associations. Together, these results demonstrate the potential of metagenomics for biodiversity monitoring and population genomics, while underscoring the importance of filtering criteria and comprehensive reference databases.
The more insects there are, the more food there is for insectivores and the higher the likelihood for insect-associated ecosystem services. Yet, we lack insights into the drivers of insect biomass over space and seasons, for both tropical and temperate zones. We used 245 Malaise traps, managed by 191 volunteers and park guards, to characterize year-round flying insect biomass in a temperate (Sweden) and a tropical (Madagascar) country. Surprisingly, we found that local insect biomass was similar across zones. In Sweden, local insect biomass increased with accumulated heat and varied across habitats, while biomass in Madagascar was unrelated to the environmental predictors measured. Drivers behind seasonality partly converged: In both countries, the seasonality of insect biomass differed between warmer and colder sites, and wetter and drier sites. In Sweden, short-term deviations from expected season-specific biomass were explained by week-to-week fluctuations in accumulated heat, rainfall and soil moisture, whereas in Madagascar, weeks with higher soil moisture had higher insect biomass. Overall, our study identifies key drivers of the seasonal distribution of flying insect biomass in a temperate and a tropical climate. This knowledge is key to understanding the spatial and seasonal availability of insects—as well as predicting future scenarios of insect biomass change.
Any single ecosystem will provide many ecosystem functions. Whether these functions tend to increase in concert or trade off against each other is a question of much current interest. Equally topical are the drivers behind ecosystem function rates. Yet, we lack large-scale systematic studies that investigate how abiotic factors can directly or indirectly — via effects on biodiversity — drive ecosystem functioning. In this study, we assessed the impact of climate, landscape and biotic community on ecosystem functioning and multifunctioning in the temperate and tropical zone, and investigated potential trade-offs among ecosystem functions in both zones. To achieve this, we measured a diverse set of insect-related ecosystem functions — including herbivory, seed dispersal, predation, decomposition and pollination — at 50 sites across Madagascar and 171 sites across Sweden, and characterized the insect community at each site using Malaise traps. We used structural equations models to infer causality of the effects of climate, landscape, and biodiversity on ecosystem functioning. For the temperate zone, we found that abiotic factors were more important than biotic factors in driving ecosystem functioning, while in the tropical zone, effects of biotic drivers were most pronounced. In terms of trade-offs among functions, in the temperate zone, only seed dispersal and predation were positively correlated, while all other functions were uncorrelated. By contrast, in the tropical zone, most ecosystem functions increased in concert, highlighting that tropical ecosystems can simultaneously provide a diverse set of functions. These correlated functions in Madagascar could for the most part be explained by similar responses to local climate, landscape, and biota. Our study suggests that the functioning of temperate and tropical ecosystems differs fundamentally in patterns and drivers. Without a better understanding of these differences, it will be impossible to correctly predict shifts in ecosystem functioning in response to environmental disturbances. To identify global patterns and drivers of ecosystem functioning, we will next need replicate sampling across biomes – as here achieved for two regions, thus paving the road and setting the baseline expectations. ### Competing Interest Statement The authors have declared no competing interest.
Among the most widely used information underpinning international conservation efforts is the IUCN Red List of endangered species. The Red List designates species extinction risk based on geographic range, population size, or declines in either. However, the Red-List has poor representation of invertebrates which comprise the majority of animal diversity, and it has frequently been questioned whether Red List criteria are appropriate for these organisms. Due to their small size, difficulty in identification, and general rarity, many invertebrates are hard to study, making Red List criteria difficult to apply. Here we discuss these criticisms in the context of empirical evidence from one of the largest terrestrial arthropod surveys to date, documenting the abundance and distribution of over 13,000 species in Sweden. Using simple empirical examples from these data, we argue that even the most ambitious monitoring efforts are unlikely to produce enough observations to reliably estimate population sizes and ranges for more than a fraction of species. Thus, there is likely to be substantial uncertainty in classifying most species according to current criteria. In response, we discuss the introduction of potential new IUCN criteria to more accurately capture the conservation needs of invertebrates, and to increase the representation of invertebrates on the IUCN Red List.
Gall wasps (Hymenoptera: Cynipidae) comprise 13 distinct tribes whose interrelationships remain incompletely understood. Recent analyses of ultra‐conserved elements (UCEs) represent the first attempt at resolving these relationships using phylogenomics. Here, we present the first analysis based on protein‐coding sequences from genome and transcriptome assemblies. Unlike UCEs, these data allow more sophisticated substitution models, which can potentially resolve issues with long‐branch attraction. We include data for 37 cynipoid species, including two tribes missing in the UCE analysis: Aylacini (s. str.) and Qwaqwaiini. Our results confirm the UCE result that Cynipidae are not monophyletic. Specifically, the Paraulacini and Diplolepidini + Pediaspidini fall outside a core clade (Cynipidae s. str.), which is more closely related to the insect‐parasitic Figitidae, and this result is robust to the exclusion of long‐branch taxa that could mislead the analysis. Given this, we here divide the Cynipidae into three families: the Paraulacidae stat. prom., Diplolepididae stat. prom. and Cynipidae (s. str.). Our results suggest that the Eschatocerini are the sister group of the remaining Cynipidae (s. str.). Within the Cynipidae (s. str.), the Aylacini (s. str.) are more closely related to oak gall wasps (Cynipini) and some of their inquilines (Ceroptresini) than to other herb gallers (Aulacideini and Phanacidini), and the Qwaqwaiini likely form a clade together with Synergini (s. str.) and Rhoophilini. Several alternative scenarios for the evolution of cynipid life histories are compatible with the relationships suggested by our analysis, but all are complex and require multiple shifts among parasitoids, inquilines and gall inducers.
Sampling of species-rich taxa followed by DNA metabarcoding is quickly becoming a popular high-throughput method for biodiversity inventories. Unfortunately, we know little about its accuracy and efficiency, as the results mostly pertain to poorly-known organism groups in underexplored environments or regions of the world. Here we ask what an extensive sampling effort based on Malaise trapping and metabarcoding can tell us about the lepidopteran fauna of Sweden - one of the best-understood insect taxa in one of the most-surveyed countries of the world. Specifically, we deployed 197 Malaise traps for a single year across Sweden in a systematic sampling design, then metabarcoded the resulting 4,749 bulk samples, and compared the results to existing data sources. We detected more than half (1,535) of the 2,990 lepidopteran species ever recorded as occurring in Sweden, and 323 species not reported during the sampling period by other data providers. Full-length barcoding of individual specimens confirmed three new species for the country and extensive range extensions for two species. It also corroborated eight genetically distinct COI variants that may represent new species to science, one of which has since been described. Most of the new records are for small and inconspicuous species and poorly surveyed regions, suggesting that they represent previously overlooked components of the fauna. Our findings, corroborated by independent metagenomic analyses, show that DNA metabarcoding can be a highly efficient and accurate method of biodiversity sampling, to the extent that it can generate significant new discoveries even for the most well-known of insect faunas. ### Competing Interest Statement The authors have declared no competing interest.
Insects are diverse and sustain essential ecosystem functions, yet remain understudied. Recent reports about declines in insect abundance and diversity have highlighted a pressing need for comprehensive large-scale monitoring. Metabarcoding (high-throughput bulk sequencing of marker gene amplicons) offers a cost-effective and relatively fast method for characterizing insect community samples. However, the methodology applied varies greatly among studies, thus complicating the design of large-scale and repeatable monitoring schemes. Here we describe a non-destructive metabarcoding protocol that is optimized for high-throughput processing of Malaise trap samples and other bulk insect samples. The protocol details the process from obtaining bulk samples up to submitting libraries for sequencing. It is divided into four sections: 1) Laboratory workspace preparation; 2) Sample processing-decanting ethanol, measuring the wet-weight biomass and the concentration of the preservative ethanol, performing non-destructive lysis and preserving the insect material for future work; 3) DNA extraction and purification; and 4) Library preparation and sequencing. The protocol relies on readily available reagents and materials. For steps that require expensive infrastructure, such as the DNA purification robots, we suggest alternative low-cost solutions. The use of this protocol yields a comprehensive assessment of the number of species present in a given sample, their relative read abundances and the overall insect biomass. To date, we have successfully applied the protocol to more than 7000 Malaise trap samples obtained from Sweden and Madagascar. We demonstrate the data yield from the protocol using a small subset of these samples.