Premise Plant phenology dictates many aspects of community function and ecosystem dynamics. Yet, global phenology data are still limited, especially in areas lacking monitoring programs. Here we present a new data resource, PhenoVision-Leaf, which extends a computer vision pipeline utilizing iNaturalist digital image vouchers to produce global-scale leaf phenophase data for deciduous woody genera. Methods We first discuss our implementation of a new human annotation framework for leaf phenology on iNaturalist, aligning with phenophase definitions used by the larger phenology community. We then showcase the use of 165,988 crowdsourced annotated records to train a Vision Transformer model with a two-stage regime to maximize precision across single- and multi-image records. This approach extends PhenoVision from scoring individual images to aggregating at the iNaturalist record level, better aligning with human annotation processes. Results Post-hoc validation showed high performance for detecting present green and colored leaves (>98% precision) and reasonable precision for breaking leaf buds (>87% precision). Applying PhenoVision-Leaf to over six million iNaturalist records yielded 5.6 million record-level phenology observations across 6500 species and 57 families, filling geographic and taxonomic gaps. Discussion These data, now accessible through the Phenobase web application, establish a foundation for near real-time monitoring of leaf phenology, supporting global-scale synthesis analyses.
Herbarium specimens represent critical historical records of plant phenology, yet automating annotation of reproductive structures remains challenging given the diversity of floral morphologies, specimen age and quality, and image quality. Here, we present a machine learning pipeline that uses an ensemble modeling approach to detect flowers on herbarium specimens and deliver these data to the phenology research community. After testing multiple strategies for generating training data, we found in-house expert-curated annotations were essential for producing reliable results. Expert validation found relatively strong accuracy for detecting present floral structures, but still had moderately high false negative rates. Applying the ensemble to our filtered final image dataset of 22 million records resulted in 11.1 million records labeled with flowers present. However, only 2.9 million of these contained complete metadata necessary for downstream phenology research, highlighting the need for full label digitization efforts. Still, this dataset represents a large compilation of historical herbarium-derived phenology records available as a resource for the phenology community. We end by demonstrating how integrating these machine-labeled records into Phenobase, a publicly-available phenology database, expands taxonomic and temporal coverage for large-scale phenological analyses, and discuss remaining challenges and next steps. ### Competing Interest Statement The authors have declared no competing interest. National Science Foundation, USA, DBI2223512, DBI2223508
Genetic diversity within species represents a fundamental yet underappreciated level of biodiversity. Because genetic diversity can indicate species resilience to changing climate, its measurement is relevant to many national and global conservation policy targets. Many studies produce large amounts of genome-scale genetic diversity data for wild populations, but most (87%) do not include the associated spatial and temporal metadata necessary for them to be reused in monitoring programs or for acknowledging the sovereignty of nations or Indigenous peoples. We undertook a distributed datathon to quantify the availability of these missing metadata and to test the hypothesis that their availability decays with time. We also worked to remediate missing metadata by extracting them from associated published papers, online repositories, and direct communication with authors. Starting with 848 candidate genomic data sets (reduced representation and whole genome) from the International Nucleotide Sequence Database Collaboration, we determined that 561 contained mostly samples from wild populations. We successfully restored spatiotemporal metadata for 78% of these 561 data sets (n = 440 data sets with data on 45,105 individuals from 762 species in 17 phyla). Examining papers and online repositories was much more fruitful than contacting 351 authors, who replied to our email requests 45% of the time. Overall, 23% of our email queries to authors unearthed useful metadata. The probability of retrieving spatiotemporal metadata declined significantly as age of the data set increased. There was a 13.5% yearly decrease in metadata associated with published papers or online repositories and up to a 22% yearly decrease in metadata that were only available from authors. This rapid decay in metadata availability, mirrored in studies of other types of biological data, should motivate swift updates to data-sharing policies and researcher practices to ensure that the valuable context provided by metadata is not lost to conservation science forever.
Genetic diversity within species represents a fundamental yet underappreciated level of biodiversity. Because genetic diversity can indicate species and population resilience to changing climate, its measurement is relevant to many national and global conservation policy targets. Many studies of evolutionary biology, molecular ecology and conservation genetics produce large amounts of genome-scale genetic diversity data for wild populations. While open data policies have ensured an abundance of freely available genomic data stored in the databases of the International Nucleotide Sequence Database Collaboration (INSDC), only about 13% of current accessions have the associated spatial and temporal metadata in INSDC necessary to be reused in monitoring programs, macrogenetic studies, or for acknowledging the sovereignty of nations or Indigenous Peoples. We undertook a “distributed datathon” to quantify the availability of these missing metadata in sources external to the INSDC and to test the hypothesis that these metadata decay with time. We also worked to remediate these missing metadata by extracting them, when present, from associated published papers, online repositories, and/or from direct communication with authors. Starting with 848 programmatically identified candidate datasets (INSDC BioProjects), we manually determined that 561 contained samples from wild populations. We successfully restored spatiotemporal metadata (locality name and/or geospatial coordinates and collection year) for 78% of these 561 datasets (N = 440 BioProjects comprising 45,105 individuals or BioSamples from 762 species in 17 phyla). We also quantified the availability of 33 additional categories of metadata in sources external to the INSDC. Information about associated publications and the type of habitat from which the samples were taken was the most easily found; information about sampling permits was the most challenging to locate. Looking at papers and online repositories was much more fruitful than contacting authors, who only replied to our email requests 45% of the time. Overall, 23% of our email queries to authors discovered useful metadata. Importantly, we found that the probability of retrieving spatiotemporal metadata declines significantly with the age of the dataset, with a 13.5% yearly decrease for metadata located in published papers or online repositories and up to a 22% yearly decrease for metadata that were only available from authors. This observable metadata decay, mirrored in studies of other types of biological data, should motivate swift updates to data sharing policies and researcher practices to ensure that the valuable context provided by metadata is not lost forever.
Understanding variation of traits within and among species through time and across space is central to many questions in biology. Many resources assemble species-level trait data, but the data and metadata underlying those trait measurements are often not reported. Here, we introduce FuTRES (Functional Trait Resource for Environmental Studies; pronounced few-tress), an online datastore and community resource for individual-level trait reporting that utilizes a semantic framework. FuTRES already stores millions of trait measurements for paleobiological, zooarchaeological, and modern specimens, with a current focus on mammals. We compare dynamically derived extant mammal species' body size measurements in FuTRES with summary values from other compilations, highlighting potential issues with simply reporting a single mean estimate. We then show that individual-level data improve estimates of body mass-including uncertainty-for zooarchaeological specimens. FuTRES facilitates trait data integration and discoverability, accelerating new research agendas, especially scaling from intra- to interspecific trait variability.
Emerging infectious diseases have been especially devastating to amphibians, the most endangered class of vertebrates. For amphibians, the greatest disease threat is chytridiomycosis, caused by one of two chytridiomycete fungal pathogens Batrachochytrium dendrobatidis (Bd) and Batrachochytrium salamandrivorans (Bsal). Research over the last two decades has shown that susceptibility to this disease varies greatly with respect to a suite of host and pathogen factors such as phylogeny, geography (including abiotic factors), host community composition, and historical exposure to pathogens; yet, despite a growing body of research, a comprehensive understanding of global chytridiomycosis incidence remains elusive. In a large collaborative effort, Bd-Maps was launched in 2007 to increase multidisciplinary investigations and understanding using compiled global Bd occurrence data (Bsal was not discovered until 2013). As its database functions aged and became unsustainable, we sought to address critical needs utilizing new technologies to meet the challenges of aggregating data to facilitate research on both Bd and Bsal. Here, we introduce an advanced central online repository to archive, aggregate, and share Bd and Bsal data collected from around the world. The Amphibian Disease Portal (https://amphibiandisease.org) addresses several critical community needs while also helping to build basic biological knowledge of chytridiomycosis. This portal could be useful for other amphibian diseases and could also be replicated for uses with other wildlife diseases. We show how the Amphibian Disease Portal provides: (1) a new repository for the legacy Bd-Maps data; (2) a repository for sample-level data to archive datasets and host published data with permanent DOIs; (3) a flexible framework to adapt to advances in field, laboratory, and informatics technologies; and (4) a global aggregation of Bd and Bsal infection data to enable and accelerate research and conservation. The new framework for this project is built using biodiversity informatics best practices and metadata standards to ensure scientific reproducibility and linkages across other biological and biodiversity repositories.
Material samples are vital across multiple scientific disciplines with samples collected for one project often proving valuable for additional studies.The Internet of Samples (iSamples) project aims to integrate large, diverse, cross-discipline sample repositories and enable access and discovery of material samples as FAIR data (Findable, Accessible, Interoperable, and Reusable).Here we report our recent progress in controlled vocabulary development and mapping.In addition to a core metadata schema to integrate SESAR, GEOME, Open Context, and Smithsonian natural history collections, three small but important controlled vocabularies (CVs) describing specimen type, material type, and sampled feature were created.
Understanding variation of traits within and among species through time and across space is central to many questions in biology. Many resources have been developed to assemble trait data at the species level, but the underlying data and metadata about those trait measurements are often not reported, limiting broadest utility. Here we introduce FuTRES (Functional Trait Resource for Environmental Studies), a datastore and community resource for individual-level trait reporting that utilizes a strong semantic framework and best practices approach to overcome previous limitations. FuTRES already stores millions of trait measurements that span across multiple time scales, including zooarchaeological and paleobiological specimens, with a current focus on mammals. Two case studies showcase the promise of FuTRES. The first compares dynamically derived extant mammal species' body size estimates with summary values from other compilations, highlighting potential issues with simply reporting a single mean estimate. The second shows that FuTRES data improves estimates of body mass – including uncertainty – for zooarchaeological specimens. FuTRES facilitates trait data integration and discoverability, accelerating new research agendas, especially scaling from intra- to interspecific trait variability.
Material samples form an important portion of the data infrastructure for many disciplines. Here, a material sample is a physical object, representative of some physical thing, on which observations can be made. Material samples may be collected for one project initially, but can also be valuable resources for other studies in other disciplines. Collecting and curating material samples can be a costly process. Integrating institutionally managed sample collections, along with those sitting in individual offices or labs, is necessary to faciliate large-scale evidence-based scientific research. Many have recognized the problems and are working to make data related to material samples FAIR: findable, accessible, interoperable, and reusable. The Internet of Samples (i.e., iSamples) is one of these projects. iSamples was funded by the United States National Science Foundation in 2020 with the following aims: enable previously impossible connections between diverse and disparate sample-based observations; support existing research programs and facilities that collect and manage diverse sample types; facilitate new interdisciplinary collaborations; and provide an efficient solution for FAIR samples, avoiding duplicate efforts in different domains (Davies et al. 2021) enable previously impossible connections between diverse and disparate sample-based observations; support existing research programs and facilities that collect and manage diverse sample types; facilitate new interdisciplinary collaborations; and provide an efficient solution for FAIR samples, avoiding duplicate efforts in different domains (Davies et al. 2021) The initial sample collections that will make up the internet of samples include those from the System for Earth Sample Registration (SESAR), Open Context, the Genomic Observatories Meta-Database (GEOME), and Smithsonian Institution Museum of Natural History (NMNH), representing the disciplines of geoscience, archaeology/anthropology, and biology. To achieve these aims, the proposed iSamples infrastructure (Fig. 1) has two key components: iSamples in a Box (iSB) and iSamples Central (iSC). The iSC component will be a permanent Internet service that preserves, indexes, and provides access to sample metadata aggregated from iSBs. It will also ensure that persistent identifiers and sample descriptions assigned and used by individual iSBs are synchronized with the records in iSC and with identifier authorities like International Geo Sample Number (IGSN) or Archival Resource Key (ARK). The iSBs create and maintain identifiers and metadata for their respective collection of samples. While providing access to the samples held locally, an iSB also allows iSC to harvest its metadata records. The metadata modeling strategy adopted by the iSamples project is a metadata profile-based approach, where core metadata fields that are applicable to all samples, form the core metadata schema for iSamples. Each individual participating collectionis free to include additional metadata in their records, which will also be harvested by iSC and are discoverable through the iSC user interface or APIs (Application Programming Interfaces), just like the core. In-depth analysis of metadata profiles used by participating collections, including Darwin Core, has resulted in an iSamples core schema currently being tested and refined through use. See the current version of the iSamples core schema. A number of properties require a controlled vocabulary. Controlled vocabularies used by existing records are kept, while new vocabularies are also being developed to support high-level grouping with consistent semantics across collection types. Examples include vocabularies for Context Category, Material Category, and Specimen Type (Table 1). These vocabularies were also developed in a bottom-up manner, based on the terms used in the existing collections. For each vocabulary, a decision tree graph was created to illustrate relations among the terms, and a card sorting exercise was conducted within the project team to collect feedback. Domain experts are invited to take part in this exercise here, here, and here. These terms will be used as upper-level terms to the existing category terms used in the participating collections and hence create connections among individual participating collections. iSample project members are also active in the TDWG Material Sample Task Group and the global consultation on Digital Extended Specimens. Many members of the iSamples project also lead or participate in a sister research coordination network (RCN), Sampling Nature. The goal of this RCN is to develop and refine metadata standards and controlled vocabularies for the iSamples and other projects focusing on material samples. We cordially invite you to participate in the Sampling Nature RCN and help shape the future standards for material samples. Contact Sarah Ramdeen (sramdeen@ideo.columbia.edu) to engage with the RCN.
Sampling the natural world and built environment underpins much of science, yet systems for managing material samples and associated (meta)data are fragmented across institutional catalogs, practices for identification, and discipline-specific (meta)data standards. The Internet of Samples (iSamples) is a standards-based collaboration to uniquely, consistently, and conveniently identify material samples, record core metadata about them, and link them to other samples, data, and research products. iSamples extends existing resources and best practices in data stewardship to render a cross-domain cyberinfrastructure that enables transdisciplinary research, discovery, and reuse of material samples in 21st century natural science.
Genetic data represent a relatively new frontier for our understanding of global biodiversity. Ideally, such data should include both organismal DNA-based genotypes and the ecological context where the organisms were sampled. Yet most tools and standards for data deposition focus exclusively either on genetic or ecological attributes. The Genomic Observatories Metadatabase (GEOME: geome-db.org) provides an intuitive solution for maintaining links between genetic data sets stored by the International Nucleotide Sequence Database Collaboration (INSDC) and their associated ecological metadata. GEOME facilitates the deposition of raw genetic data to INSDCs sequence read archive (SRA) while maintaining persistent links to standards-compliant ecological metadata held in the GEOME database. This approach facilitates findable, accessible, interoperable and reusable data archival practices. Moreover, GEOME enables data management solutions for large collaborative groups and expedites batch retrieval of genetic data from the SRA. The article that follows describes how GEOME can enable genuinely open data workflows for researchers in the field of molecular ecology.
Ideally, an information system that automates the integration of disparate datasets should be able to minimize the loss of information from any one dataset, achieve computational complexity suitable for working with large datasets, be flexible enough to easily incorporate new data sources, and produce output that is easily analyzed and understood by data users. Achieving all of these goals within highly heterogeneous and highly complex data domains is a major challenge. In this talk, we present the results of our recent efforts to develop such a system for data about plant phenology. Our data integration system, which is built around the Plant Phenology Ontology, currently supports semantically fine-grained integration of phenological data from both field observations and herbarium specimens. We show that even with a heavily axiomatized ontology and sophisticated, machine-reasoning-based data analysis, it is possible to implement a high-throughput data integration pipeline capable of processing millions of individual records in a matter of minutes while running on modest, server-class hardware. Success requires careful ontology design and judicious application of machine reasoning techniques. We also discuss some of the many challenges that remain for designing efficient, general-purpose data integration systems.
Functional traits are the features of organisms that directly interact with the environment. Studying change and variation in these traits across space, time, and taxonomy can inform how species have responded to environmental and climatic change, how communities are assembled, and other eco-evolutionary questions. Trait data are collected at the individual level; however, animal trait databases often report these data at the species level, undermining their value for researchers who want to look at variation within species and rendering trait data ambiguous when taxonomy is updated. Additionally, these data are often recorded in auxiliary fields, such as “field notes” or hidden in supplementary materials or published tables, making them difficult to recover by researchers. Furthermore, animal trait data from paleontological, zooarchaeological (from archaeological sites), and neontological specimens are typically curated in separate forums and formats that are not easily integrated to provide perspective across the entire range of time. We are developing a toolkit to overcome these challenges called FuTRES: Functional Trait Resource for Environmental Studies. We seek to make these data accessible, standardize trait descriptions across Vertebrata, and teach (future) scientists how to create FAIR (findable, accessible, interoperable, and reusable) trait datasets. To make data more FAIR, FuTRES employs ontologies, a logical framework for relating terms to search datasets and standardizing traits across datasets. FuTRES builds off existing ontologies and standards, such as UBERON for anatomical terms and PATO and OBA for trait terms, as well as create new terms that are general enough to be used for all vertebrates and multiple disciplines. This talk will showcase the semantic framework underpinning FuTRES, describe how we are linking diverse trait datasets to ontologies and, therefore, each other, and report the results of a preliminary analysis of integrated datasets.
PREMISE OF THE STUDY:The Plant Phenology Ontology (PPO) was originally developed to integrate phenology observations of whole plants across different global observation networks. Here we describe a new release of the PPO and associated data pipelines that supports integration of phenology observations from herbarium specimens, which provide historical and modern phenology data.METHODS AND RESULTS:Critical changes to the PPO include key terms that describe how measurements from parts of plants, which are captured in most imaged herbarium specimens, relate to whole plants. We provide proof of concept for ingesting annotations from imaged herbarium sheets of Prunus serotina, the common black cherry. We then provide an example analysis of changes in flowering timing over the past 125 years, demonstrating the value of integrating herbarium and observational phenology data sets.CONCLUSIONS:These conceptual and technical advances will support the addition of phenology data from herbaria, but also could be expanded upon to facilitate the inclusion of data from photograph-based citizen science platforms. With the incorporation of herbarium phenology data, new historical baseline data will strengthen the capability to monitor, model, and forecast plant phenology changes.
Four years after the Genomic Observatories Network was formally established as a collaboration between the Group on Earth Observations Biodiversity Observation Network and the Genomic Standards Consortium, we review the development of the network. Considering institutional infrastructure, we note the growing role of omic observation in active and increasingly interlinked marine networks, with examples such as EMBRC/ASSEMBLE, International Long Term Ecological Research Network, AtlantOS, National Association of Marine Labs, Smithsonian MarineGEO, and Partnership on Observation of the Global Oceans. We also note some key human elements essential to meeting the networks' goals, address how the community is evolving, and why performing seemingly simple tasks within a broadly distributed community presents significant challenges even among those who have agreed to use standards. From the perspectives above, we review lessons learned from use cases that leverage Genomic Observatories Network, such as the Autonomous Reef Monitoring Structures (ARMS), Ocean Sampling Day (OSD) and myOSD, which included experiences with citizen science. Looking forward, we survey 1) promising new technologies for in situ biological observation (e.g., cheap 3D printed omics samplers), 2) progress towards adoption of omics methods in marine policy and conservation programs, and 3) opportunities that a Genomic Observatory brings, alone or embedded in a network, to address novel scientific questions and support Essential Biodiversity Variables, Essential Ocean Variables, and indices such as the Ocean Health Index. Given the intensive nature of omics investigation, we note emerging cyberinfrastructure solutions, such as the Genomic Observatories Metadatabase (GeOMe), an open-access repository for geographic and ecological metadata associated with biosamples, and predictive modeling efforts, such as those of the Island Digital Ecosystem Avatar (IDEA) Consortium. Finally, we explore the potential of Genomic Observatories as components of high-resolution calibration sites. Such observatories would provide super-contextualized data trusts for machine learning and artificial intelligence applications that draw on multi-omic observation.
The BioCollections Ontology (BCO) is a semantic model for applied biodiversity science, geared toward data integration. It is based on common semantics shared by ontologies in the Open Biological and Biomedical (OBO) Foundry, thus enhancing interoperability with data linked to other life science ontologies. BCO is also compatible with more general semantic web ontologies. By clarifying the semantics of biodiversity science activities such as specimen collection, trait observation, and taxonomic identification - at a level that is useful for practical applications - the BCO supports the types of data management systems needed in the age of large, distributed, and diverse data. This chapter describes the key elements of the BCO and how it complements biodiversity standards efforts such as Darwin Core and MIxS, as well as other relevant ontologies. Concrete examples demonstrate how BCO, in conjunction with other tools and semantic technologies, is addressing scientific and informatic challenges through the enhancement of biodiversity data.
The Genomic Observatories Metadatabase (GeOMe, http://www.geome-db.org/) is an open access repository for geographic and ecological metadata associated with biosamples and genetic data. It contributes to the informatics stack – Biocode Commons – of the Genomic Observatories Network (https://gigascience.biomedcentral.com/articles/10.1186/2047-217X-3-2). The GeOMe project interface enables administrators to plan and execute field based sample collection efforts. GeOMe projects specify a core set of sample metadata fields based on community standard vocabularies and also includes plugins for associating samples with photos, subsamples, NextGen sequence metadata, and permits. Users can upload their own expedition-specific metadata, which contributes to the overall project dataset while providing the user a convenient method for updating and refining their contributed data. GeOMe provides connection points to the Global Biodiversity Information Facility and archived genetic data stored in the National Center for Biotechnology Information's (NCBI's) Sequence Read Archive (SRA), linking specimens and seqeuences via unique persistent identifiers.