
In species distribution modeling (SDM), the environmental representativeness effect hinders the comparison and generalization of discrimination statistics, as their values are context-dependent and may vary without reflecting actual differences in model accuracy. To address this issue, a harmonization approach based on the uniform distribution of suitability values has been proposed, demonstrating effectiveness when true absence data are available. However, in most cases, models solely rely on presence records and use background points instead of true absences, posing additional validation challenges. This study simulates habitat suitability and presence-absence data, and evaluates the robustness of several validation indices: background-based AUC (AUCb), its harmonized version (uAUCb), and two variations of the Boyce index. The results show that the Boyce index remains unaffected by the representativeness effect, whereas AUCb varies across scenarios, confirming the influence of the representativeness effect on SDM results with background data. Harmonization through uAUCb successfully makes values comparable, but its reliability depends on sample size, requiring at least 100 presences and 10000 background points.
The global biodiversity crisis, driven by habitat loss, climate change, and overexploitation, demands efficient data management and accessibility to inform conservation decisions. While open-access biodiversity data have increased, a substantial portion, termed "dark data," remains unpublished and inaccessible. We explored the mobilization of dark data existing in Environmental Assessment (EA) reports in Spain to improve the Digital Accessible Knowledge for 43 field-recorded threatened species. Based on the IUCN Red List criteria, we evaluated the impact of EA-related dark data on two required metrics: the extent of occurrence (EOO) and area of occupancy (AOO). Our results show that integrating dark data increased EOO and AOO for 23% and 93% of the field-recorded threatened species, respectively, including endangered species like Aquila adalberti. We highlight the importance of mobilizing data collected during the EA processes and ensuring that data are made findable, accessible, interoperable, and reusable (FAIR) to inform conservation decisions. Mobilizing EA-related dark data offers a cost-effective approach by which to strengthen biodiversity management, facilitate decision-making, and justify biodiversity conservation.
The niche centrality hypothesis predicts that individuals near the niche centroid have higher fitness due to more favorable environmental conditions. In this study, we evaluated the effect of the niche structure of the Antioquia Wren on 20 different traits grouped into four datasets: morphology, genetic diversity, coloration and acoustics. We tested the relationship with distinct niche configurations using minimum volume ellipsoids, and evaluated the relationship between the distance to the niche centroid and trait values through ordinary linear and generalized linear squared regression models. We found positive relationships for variables associated with beak morphology and crown feathers hue. Conversely, we found a negative relationship with birdsong frequency and the distance between longest primary and first secondary flight feathers. All effects of the niche structure on traits were weak (<0.2). We found no consistency in the relationship between niche structure and the remaining 13 traits. We identify potential mechanisms underlying both positive relationships and the absence of trait-niche relationships. Our findings emphasize that factors such as biotic interactions, climatic heterogeneity, range size, niche breadth, centroid position, and intrinsic trait variability are likely to shape how species conform to the niche centrality hypothesis.
A discussion is presented of the use of citizen science data to study trends in population abundance and on numbers of species. The discussion is framed around two widely accepted models of indices of biological signal per unit of effort. When the trend is number of species per unit of effort, the model is the MichaelisMenten equation, that has an asymptote. When the trend is in abundance per unit effort, the model is the simple linear "catch/effort" model of fisheries. It is found that number of species per unit of effort has an artifactual tendency to negative trends, while abundance per unit effort is more likely to be less influenced by artifacts of the metrics. A calculation shows that, under certain assumptions, the number of records in iNaturalist is a good approximation to number of different individuals sighted. These results are checked by simulation and exemplified using data for three families of butterflies of Mexico.
Downloading images of preserved specimens in bulk is becoming increasingly important for many research projects, especially those connected with machine learning , image analysis. A useful source of images is the standard biodiversity aggregator, the Global Biodiversity Information Facility (GBIF). Here we identify four major issues connected to GBIF image downloads, distinct from those associated with text downloads. These are specialIntscript license considerations, specialIntscript citation issues, specialIntscript restricting to specific providers for project reasons or cybersecurity concerns, and, finally, specialIntscript attempting to use links that are no longer functioning (often referred to as "link rot" or "data rot"). We suggest an incremental approach to downloading and suggest techniques for improved image download. We provide an implementation of our suggestions in Python (gbif- image-downloader).
In recent years, the expansion of public biodiversity platforms and associated datasets has greatly improved access to ecological and biological information. These resources now cover vast geographic areas, extended temporal scales, and diverse taxonomic groups, becoming essential for ecological studies by enabling more comprehensive analyses and novel hypotheses testing. Concurrently, the development of programming packages has facilitated data access and interaction, optimizing their retrieval processes. However, most existing tools are limited to specific databases, posing challenges for studies requiring smooth data integration from multiple sources. The growing availability of biodiversity records highlight the urgent need for robust tools to efficiently download, process, and analyse, ecological and biological information. To address this limitation, we introduce biodumpy, a new Python package developed for the retrieval, management, and integration of biological data from several public databases. biodumpy provides access to up-to-date and comprehensive datasets spanning genetic, geographic, taxonomic, and bibliographic sources. It includes specialized modules for efficient data retrieval across taxonomic lists, with the possibility to process multiple modules simultaneously. By integrating diverse data sources, biodumpy enhances data acquisition, providing researchers with a powerful framework for comprehensive analyses and supporting ecological research to tackle complex environmental challenges.
Review of the book The Game of Species. An Introduction to Biodiversity, by Sara Gamboa
The prediction of grasslands plant diversity using satellite images has been intensively studied. However, the accuracy of functional diversity (FD) is still unknown. Therefore, high spatial resolution Worldview-3 (WV-3) multiple spectral data were used to predict species and FD at the pixel scale (1.2 × 1.2 m) over central Hunshandak Sandland, Inner Mongolia, north China. Data acquired from 120 field plots (6 × 6 m) were used to train and validate several statistical learning methods with a primary objective of linking the satellite spectral and texture indices to the plant diversity indices. Among the several diversity indices tested, functional trait diversity, in particularly Functional Attribute Diversity (FAD1), Modified Functional Attribute Diversity (MFAD) were best predicted (coefficient of determination approximately 0.29 and 0.14, respectively, n=48) using texture indices. However, species diversity (richness, H, E, or D) and other FDs haven’t not been well predicted by WV-3 data. WV data did not significantly improve the prediction accuracy for plant diversity in sandy grassland. Further, high plot-level vegetation coverage can improve the performance of spectral indices for predicting H, E, D and FD. These results highlighted the assessing variability across field conditions and demonstrated the capacity of high spatial-spectral satellite images to monitor plant functional diversity in sandy grasslands.
Taxonomy is a highly dynamic science upon which most biodiversity studies rely. Constant revisions of species delimitation hypotheses, using ever-growing amounts of data and tools cause species numbers and identities to continuously and rapidly change. Reptiles are the most species rich terrestrial vertebrate group and are amongst the most threatened and least known vertebrate taxa, representing nearly half of all datadeficient terrestrial vertebrate species. Every year hundreds of new species are described and dozens are revised, resulting in synonymizations, splittings, generic reassignments, or elevation from synonymy or from subspecies into species status. The nomenclature of this group is therefore highly dynamic and consequently, to integrate available reptile datasets generally requires extensive nomenclature review, especially for broad scale analyses. letsRept is a new R package that integrates the Reptile Database - the best curated and reliable global taxonomic reference for reptiles - into the R programming environment. Its main functions allow users to retrieve the most up-to-date taxonomic information in real time, to compare lists of species names to current nomenclature, and to detect names that have been changed by either lumping or splitting, all through web scraping techniques. Additional functions allow to produce quick taxonomic summaries, access species accounts, retrieve full reference lists and more. By permitting to embed the Reptile Database directly into R workflows, the letsRept package improves the integration of datasets from different sources, with authoritative taxonomy, reducing data loss due to nomenclature mismatch and improving the consistency in biodiversity analyses.
The prediction of grassland plant diversity using satellite imagery has been the subject of intensive research. However, the accuracy of functional diversity (FD) predictions remains unclear. To address this, high-spatial-resolution WorldView-3 (WV-3) multispectral data were used to predict species diversity and FD at the pixel scale (1.2 x 1.2 m) in the central Hunshandak Sandland, Inner Mongolia, northern China. Data collected from 120 field plots (6 x 6 m) were employed to train and validate several statistical learning methods, with the primary objective of establishing links between 156 satellite-derived spectral and texture indices and 6 plant diversity indices. Among the various diversity indices tested, functional trait diversity-specifically Functional Attribute Diversity (FAD1) and Modified Functional Attribute Diversity (MFAD)-were predicted most effectively (with coefficients of determination of approximately 0.29 and 0.14, respectively; n=48) using texture indices. In contrast, species diversity (richness, H, E, or D) and other FD metrics were not well predicted by WV-3 data. Overall, WV data did not significantly improve the accuracy of plant diversity predictions in sandy grasslands. Additionally, high plot-level vegetation coverage was found to enhance the performance of spectral indices in predicting H, E, D, and FD. These results underscore the importance of accounting for variability across field conditions and demonstrate the potential of high-spatial-and-spectral-resolution satellite imagery for monitoring plant functional diversity in sandy grasslands.
Mexico hosts a great diversity of reptile species; yet, many reptiles are either threatened or endangered. Complete and updated information is required to implement appropriate management and conservation actions; however, species inventories can include taxonomic, geographic, and temporal gaps. Therefore, this study aimed to evaluate the magnitude of these gaps in digitally accessible information on reptiles from the state of Nayarit, located in northwestern Mexico. A database was generated using information from the National Biodiversity Information System (SNIB) of the National Commission for the Knowledge and Use of Biodiversity (CONABIO). The growth rate of new species descriptions was calculated, and the completeness of the inventory was evaluated in 10-km grid cells across various time periods, considering biogeographic and physiographic regions. The species description growth rate was low. In addition, approximately 40% of the surface of Nayarit exhibited information gaps among reptile records, particularly in mountainous and hard-toreach areas. Notably, the least amount of information was recorded between 1981 and 2000. Our results lay the groundwork for future research and the development of effective strategies to conserve and manage the natural resources of Nayarit.
Wildlife disease surveillance has received considerable attention following recent emergence of high-consequence zoonotic pathogens in humans. Increased portability and affordability of sequencing technologies over the last decade have made real-time sequencing of wild animals and their pathogens a reality. Wildlife samples screened for pathogens, however, are rarely permanently archived in museum biorepositories, which limits potential for scientific validation and prevents extension by related disciplines (e.g., ecology, evolution, conservation). To better connect biodiversity and biomedical sciences, the Museums and Emerging Pathogens in the Americas (MEPA) network developed the Field+Genomics Workshop to build capacity for surveillance of wildlife and their pathogens in biodiverse countries. Here, we share workshop resources, in English and Spanish, to facilitate reproducibility and expansion of the workshop into the future. The workshop lasted 10 days, 6 days of fieldwork and 4 days of molecular lab and bioinformatic techniques. The field component emphasized the importance of holistic collecting-that is, permanently preserving many parts and symbionts from each sampled organism-as a critical step in wildlife and pathogen surveillance and to build foundational scientific infrastructure. The molecular component of the workshop used samples collected during the field portion to identify hosts and pathogens in real-time. For this component, we trained participants in methods of DNA extraction, library preparation, and Nanopore Adaptive Sampling (a software feature for real-time selective enrichment or depletion of target sequences). Bioinformatic training consisted of a basic introduction to computational genomics, a worked example to analyze a small sequence dataset, and an exercise using data generated from samples collected during the workshop. In total, the workshop cost similar to$37K (similar to$3K per participant), however, similar to 25% of those funds are invested in basic equipment and infrastructure that is reusable in future workshops (e.g., sequencer, computer, etc.). This workshop highlights the effort and expertise required to conduct voucher-backed surveillance of wildlife and their pathogens and the many benefits of uniting biodiversity and biomedical sciences to build local capacity.
Ecological niche modeling (ENM) is a widely used analytical approach for predicting species distributions and has been applied to study spatial epidemiology of infectious diseases. Nevertheless, research evaluating the key components and assumptions of ENM in disease systems remains limited, raising concerns about its robustness, reproducibility, and transparency. To address this limitation, we conducted a systematic review and evaluated articles on ENM applications to infectious diseases between 2020 and 2022. We reviewed 78 articles to extract information following a standard protocol for reporting ENM analysis and summarized the information for each component (e.g., study subject, location, duration). The spatial extent of study areas varied from village to global scales, temporal duration ranged from 1 to 101 years, and the organismal levels ranged from individuals (57.7%) to populations (33.3%). Less frequently reported components included temporal autocorrelation tests (2.7%), algorithmic uncertainty (28.2%), temporal resolution (35.9%), background data selection (44.9%), coordinate reference system (41.0%), model performance from validation data (46.2%), and model averaging (20.5%). Our findings highlight a lack of consistency and transparency in disease ecology and disease biogeography studies, which may lead to misleading ENM applications in spatial epidemiology. Researchers and reviewers applying ENM to disease systems should clearly report key modeling components to ensure biologically sound outputs. This article identified trends and gaps in reporting ENM protocols for mapping disease transmission risk.
Mexico hosts a great diversity of reptile species; however, many reptiles are either threatened or endangered. Complete and updated information is required to implement appropriate management and conservation actions; however, species inventories can include taxonomic, geographic, and temporal gaps. Therefore, this study aimed to evaluate the magnitude of these gaps in digitally accessible information on reptiles from the state of Nayarit, located in northwestern Mexico. A database was generated using information from the National Biodiversity Information System (SNIB) of the National Commission for the Knowledge and Use of Biodiversity (CONABIO). The growth rate of new species descriptions was calculated, and the completeness of the inventory was evaluated in 10-km grid cells across various time periods, considering biogeographic and physiographic regions. The species description growth rate was low. In addition, approximately 40% of the surface of Nayarit exhibited information gaps among reptile records, particularly in mountainous and hard-to-reach areas. Notably, the least amount of information was recorded between 1981 and 2000. Our results lay the groundwork for future research and the development of effective strategies to conserve and manage the natural resources of Nayarit.
Machine-learning emerged as an excellent alternative to understanding ecological patterns and processes at different spatiotemporal scales. Our study aimed to offer a pictorial overview of the status quo on the use of machine-learning in ecology and conservation globally. Using keywords in the Scopus engine, we indexed all publications in ecology and conservation using machine-learning. We employed descriptive statistics and regressions models to provide an overview and predict geopolitical patterns. The majority of manuscripts were condensed in economically affluent countries, such as the United States (USA) and China (CHN) which together amount to 91 (36.8%) studies. There is a spatial aggregation in the authors’ affiliations, once 182 (73.7%) studies derived from both Nearctic and Palearctic teams, whereas Tropical teams published 65 (26.3%) manuscripts and the most-cited papers also are concentrated in northern regions. In ecology and conservation, machine-learning first appear in the literature in 2003. Yet, increased exponentially since the 2010s. In 2010, this overview indicated nine manuscripts, whereas 10-yrs later reached 120 publications. Most studies (N = 173; 70.1%) are focused on landscape and vertebrate ecology. The primary aims of the publications were widely variable but strongly adherent to providing the best-information on both landscape-scale classifications and species distribution modeling. The manuscripts encompass different methods, from maximum entropy to boosted regression trees and random forest, sometimes using a gamma of deep-learning architectures. Finally, the predictive variables (i.e., mammal diversity and per capita GDP) do not exert significant influences on the number of studies published.
Despite the large body of literature on avian migratory behavior, there is little information about stopover sites during bird movement, including the population-level drivers of breeding grounds and wintering grounds. Stopovers play an essential role in bird migratory site chains for energy supply and rest. There is an urgent need to detect and protect stopover sites to secure the long-term sustainability of migratory network connectivity and robustness. To address this challenge, we reconstructed a migration network and identified geographic hotspots denoted as stopover sites by analyzing the high-density population movements of 52 focal migratory bird species with observation data from eBird through PageRank algorithm. Furthermore, potential alternative stopover sites were explored using a word embedding technique based on geo-functional similarity. Our study was conducted in North and Central America during a three-year period and revealed three key areas, including Florida peninsula and its inland, the region of Central America, and the region near Puget Sound. Results from this study can be used for conservation prioritization guidance.
The National Ecological Observatory Network (NEON) is a long-term monitoring program at the continental scale designed to understand and forecast ecological responses to environmental change at local to broad scales. However, despite robust and nearly continuous collections, NEON mosquito data have been underused in downstream analyses. Here, we provide species-level estimated abundances for nighttime collected female mosquitoes derived from the mosquitoes sampled from CO2 traps (DP1.10043.001) (RELEASE-2024; NEON, 2024). By including zero counts, our derived data complement existing data sets and provide an analysis-ready time series useful for investigating mosquito phenology, abundances, and diversity at the species or community level. We also outline a set of considerations specific to filtering NEON mosquito data by sex and for day or nighttime collections, highlighting factors that could introduce uncertainty to abundance estimates. Along with the data set, we provide an R Markdown file that includes annotated code and documents our data filtering and QC/QA steps, as well as data files used to filter the mosquito data based on QC/QA criteria. All files are freely available for download through the Environmental Data Initiative data portal. Our reproducible and fully documented workflow can be easily adapted for specific needs or other NEON surveillance data. Our work aims to enhance the accessibility and use of NEON’s rich, long-term monitoring data.
Understanding how deforestation and changes in habitat boundaries affect biodiversity is essential for developing conservation solutions. These topics are central to biology and ecology programs, where students learn to apply their knowledge in real-world conservation efforts. Higher education plays a crucial role in strengthening this understanding, particularly in life sciences programs. Given the complexity of ecological processes in altered landscapes, agent-based modeling provides an interactive and engaging way to simplify and visualize the effects of land use changes. In this study, we integrate Amazonian anurans, highly sensitive to temperature and humidity fluctuations, with an Agent-based model to simulate the impacts of deforestation, habitat restoration, and land abandonment on species survival and movement. Their ectothermic nature and dependence on pulmocutaneous respiration make them especially vulnerable to the drier and more variable conditions caused by deforestation. Integrating this model into conservation biology courses has enhanced learning by encouraging independent exploration, both in and out of the classroom. This tool, an agent-based model, is particularly suited for university-level ecology and conservation courses, and can also serve as an effective awareness tool in environmental education and decision-making workshops, highlighting the negative effects of human-made habitat changes on biodiversity.
Genetic diversity is fundamental to biological diversity, vital for species’ health and adaptation to environmental change. Under the recently adopted Kunming-Montreal Global Biodiversity Framework (GBF), 196 Parties committed to report the status of genetic diversity for both wild and domesticated species. For this, three genetic diversity indicators were developed, two of which focus on processes contributing to genetic diversity conservation: ensuring that populations are large enough to maintain genetic diversity (effective population size Ne 500 indicator) and maintaining genetically distinct populations (populations maintained, PM indicator). A third indicator focuses on the number of species being monitored using DNA-based methods. Adopted by 196 CBD Parties in December 2022, GBF integrated Ne 500 and PM as headline and complementary indicators, respectively. To aid nations in quantifying these indicators, a detailed set of guideline materials was developed, encompassing species selection, data compilation, and indicator computation. These guidelines draw from the collaborative efforts of the first multinational assessment of genetic diversity indicators that was recently completed and that will be refined continually through a versioning system, as more experience is gained and shared. The materials aim to support the global monitoring framework established by the CBD and are accessible online for utilization and updates. The guidelines are available at this link.
The availability of biodiversity databases is expanding at unprecedented rates. Nevertheless, species occurrence data can be intrinsically biased and contain uncertainties that impact the accuracy and reliability of biodiversity estimates. In this study, we developed a reproducible framework to assess three dimensions of bias—taxonomic, spatial, and temporal—as well as temporal uncertainty associated with data collections. We utilized the vegetation plot data located in Europe, from sPlotOpen, an open-access database, as a case study. The metrics proposed for estimating bias include completeness of the species richness for taxonomic bias, Nearest Neighbor Index for spatial bias, and Pielou’s index for temporal bias. Additionally, we introduced a new method based on a negative exponential curve to model the temporal decay in biodiversity data, aiming to quantify temporal uncertainty. Finally, we assessed the sampling bias considering the influence of various spatial variables (i.e, road density, human population count, Natura 2000 network and topographic roughness). We discovered that the facets of bias and the temporal uncertainty varied throughout Europe, as did the different roles played by spatial variables in determining biases. sPlotOpen showed a clustered distribution of the vegetation plots, and an uneven distribution in sampling completeness, year of sampling and temporal uncertainty. The facets of bias were significantly explained mainly by the presence of Natura 2000 network and marginally by the human population count. These results suggest that employing an efficient procedure to examine biases and uncertainties in data collections can enhance data quality and provide more reliable biodiversity estimates.