Abstract Camera trapping has revolutionized wildlife ecology and conservation by providing automated data acquisition, leading to the accumulation of massive amounts of camera trap data worldwide. Although management and processing of camera trap‐derived Big Data are becoming increasingly solvable with the help of scalable cyber‐infrastructures, harmonization and exchange of the data remain limited, hindering its full potential. There is currently no widely accepted standard for exchanging camera trap data. The only existing proposal, “Camera Trap Metadata Standard” (CTMS), has several technical shortcomings and limited adoption. We present a new data exchange format, the Camera Trap Data Package (Camtrap DP), designed to allow users to easily exchange, harmonize and archive camera trap data at local to global scales. Camtrap DP structures camera trap data in a simple yet flexible data model consisting of three tables (Deployments, Media and Observations) that supports a wide range of camera deployment designs, classification techniques (e.g., human and AI, media‐based and event‐based) and analytical use cases, from compiling species occurrence data through distribution, occupancy and activity modeling to density estimation. The format further achieves interoperability by building upon existing standards, Frictionless Data Package in particular, which is supported by a suite of open software tools to read and validate data. Camtrap DP is the consensus of a long, in‐depth, consultation and outreach process with standard and software developers, the main existing camera trap data management platforms, major players in the field of camera trapping and the Global Biodiversity Information Facility (GBIF). Under the umbrella of the Biodiversity Information Standards (TDWG), Camtrap DP has been developed openly, collaboratively and with version control from the start. We encourage camera trapping users and developers to join the discussion and contribute to the further development and adoption of this standard.
Access to high-quality ecological data is critical to assessing and modeling biodiversity and its changes through space and time. The Darwin Core standard has proven to be immensely helpful in sharing species occurrence data (see Wieczorek et al. 2012, Global Biodiversity Information Facility, GBIF) and promoting biodiversity research following the FAIR principles of findability, accessibility, interoperability and reusability (Wilkinson et al. 2016). However, it is limited in its ability to fully accommodate inventory data (i.e., linked records of multiple taxa at a specific place and time). Information about the inventory processes is often either unreported or described in an unstructured manner, limiting its potential re-use for larger-scale analyses. Two key aspects that are not captured in a structured manner yet are: i) information about the species that were not detected during an inventory, and ii) ancillary information about sampling effort and completeness. Non-detections (i.e., reported counts of zero) potentially enable more accurate and precise estimates of distribution, abundance, and changes in abundance. This becomes possible when variation in effort is used to estimate the likelihood that a non-detection represents a true absence of that taxon during the inventory. Currently, ecological inventory data, when shared at all, are typically discoverable through dataset catalogs (e.g., governmental data repositories) and supplementary materials to publications. With few exceptions, indexing of such data with the detail and structure needed has not been attempted at broad temporal and spatial scales, despite the potentially high value resulting from making inventory data more readily accessible. To address these limitations in documenting inventory data using the Darwin Core, Guralnick et al. (2018) proposed the Humboldt Core. Subsequent discussions within the biodiversity standards community made it clear that greater integration could be achieved by creating an extension of the Darwin Core, rather than developing a new standard in isolation. Extension design work began in 2021 and progress has been reported by Brenton (2021) and Sica et al. (2022). Over the last year the Humboldt Extension Task Group has sought advice from data providers and aggregators and updated its vocabulary terms. A challenging aspect has been creating terminology for the parent-child relationships (see Properties of Hierarchical Events) needed to describe surveys that may be as simple as a collection of checklists (one level of hierarchy) or as complex as species records from traps within plots along transects across habitats over multiple years (at least four levels of hierarchy). The Task Group has committed to completing a User Guide for the Humboldt Extension. Group members who contributed to the Darwin Core (Darwin Core Task Group 2009) and the Vocabulary Maintenance Specification (Vocabulary Maintenance Specification Task Group 2017) have provided valuable expertise on term refinement and process. Through ratification of the Humboldt Extension as a Darwin Core Event extension, we expect to provide the community with a usable solution, tied to well-established data publication mechanisms, for sharing and using inventory data. This effort promises to overcome a key bottleneck in the sharing of critically important ecological data, enhancing data discoverability, interoperability and re-use while lowering reporting burden and data and metadata heterogeneity. Global data aggregation initiatives, such as GBIF, will benefit from this development as they develop their data models and the range of standards and extensions they support. We anticipate that the Humboldt Extension will be attractive both to data publishers and data users, by facilitating the representation and indexing of data in richer, more meaningful ways. Despite the data-intensive nature of fundamental ecological research and applied monitoring for management and policy, ecological data have remained as one of the FAIR data frontiers. We anticipate that the Humboldt Extension will address most data exchange needs of all professional communities involved.
The Audubon Core vocabulary terms subjectPart and subjectOrientation are used to describe the depicted part of an organism and its orientation in an image. We describe the criteria and process for developing controlled vocabularies for these two terms. The vocabularies take the form of Simple Knowledge Organization System (SKOS) concept schemes and their terms are categorized using SKOS collections to allow users to select from particular sets of values appropriate for particular organism groups and their parts. We also report the results of implementation testing used to determine the usability of the proposed terms with actual images of living organisms and preserved specimens.
The W3C Generating RDF from Tabular Data on the Web Recommendation provides a mechanism for mapping CSV-formatted data to any RDF graph model. Since the Wikibase data model used by Wikidata can be expressed as RDF, this Recommendation can be used to document tabular snapshots of parts of the Wikidata knowledge graph in a simple form that is easy for humans and applications to read. Those snapshots can be used to document how subgraphs of Wikidata have changed over time and can be compared with the current state of Wikidata using its Query Service to detect vandalism and value added through community contributions.
Access to high-quality ecological data is pivotal to assessing and modeling biodiversity and its change through space and time. Inventory data (i.e., recording multiple species at specific places and times) are particularly relevant to monitoring species distributions and abundance, but their reliability for use in downstream models depends on reporting the methodology implemented and associated sampling effort and completeness. This information about the inventory processes is often either not reported or described in an unstructured manner, greatly limiting potential re-use for larger-scale analyses. In order to support the reuse of inventories and to assure better standardization of newly collected data, we developed a framework to standardize inventory data reporting that is general enough for broad use. Guralnick et al. (2018) introduced the Humboldt Core as a proof of concept. In 2021, the TDWG Humboldt Core Task Group was established to review how to best integrate the terms proposed in the original publication with existing standards and implementation schemas. In the context of sharing data using the Darwin Core standard (DwC), different types of inventories can be represented as Events with different nesting levels. Therefore, it was deemed appropriate to develop an extension to DwC that allows capturing the details of the inventory process. The Task Group members revised all original terms, reformulated definitions, and discarded or added new terms where needed. We are developing a user guide and reaching out to the larger biodiversity community to test the Humboldt Extension with real-world case study datasets using a test instance of the GBIF Integrated Publishing Toolkit (IPT). In this presentation, we will review the development process, give an overview of how the Humboldt Extension can be used to report key information on the inventory process, and provide example cases. After testing with real world cases, our next step will be to seek ratification of Humboldt as a Darwin Core Event extension following the Vocabulary Maintenance Standard. We expect that this will help to overcome a key bottleneck in the sharing of critically important ecological data, enhancing data discoverability, interoperability and re-use while lowering reporting burden.
The Biodiversity Information Standards (TDWG) Material Sample Task Group*1 kicked off in the third quarter of 2021. The group’s initial focus was to 1) achieve a clear conceptual delineation between the terms MaterialSample, PreservedSpecimen, LivingSpecimen, and FossilSpecimen (the terms used in basisOfRecord in the current DwC-A provided to the Integrated Publishing Toolkit (IPT) for describing physical material) 2) define the conceptual relationship between these terms and the term Organism 3) consider the possible implications of the activities towards the diversification of the Global Biodiversity Information Facility (GBIF) data model*2 and what standards already exist that should inform our work. Based on this conceptual work, the group is now developing a concrete proposal for a clarification of a MaterialSample class with its own properties. Our presentation will provide a brief review of the task group's progress and our thoughts about what comes next.
WikiProject Clinical Trials is a Wikidata community project to integrate clinical trials metadata with the Wikipedia ecosystem. Using Wikidata methods for data modeling, import, querying, curating, and profiling, the project brought ClinicalTrials.gov records into Wikidata and enriched them. The motivation for the project was gaining the benefits of hosting in Wikidata, which include distribution to new audiences and staging the content for the Wikimedia editor community to develop it further. Project pages present options for engaging with the content in the Wikidata environment. Example applications include generation of web-based profiles of clinical trials by medical condition, research intervention, research site, principal investigator, and funder. The project's curation workflows including entity disambiguation and language translation could be expanded when there is a need to make subsets of clinical trial information more accessible to a given community. This project's methods could be adapted for other clinical trial registries, or as a model for using Wikidata to enrich other metadata collections.
The Art in the Christian Tradition image collection and database was developed to support the Vanderbilt Divinity School’s Revised Common Lectionary website, but it has become an important source of images in its own right. To improve discoverability, we started a project to create Wikidata items for all works in the collection, with the goal of linking those items to as many images in Wikimedia Commons as possible. We describe several challenges we faced in the early stages of the project. We provide an introduction to creating an item and refer to some useful tools for learning how to edit Wikidata and to perform bulk uploads. We end by describing how features of Wikidata can make artworks more discoverable and how we hope to improve the quality of our image metadata in the future.
When the Audubon Core Multimedia Resources Metadata Schema*1 was ratified, it included two terms for describing what was being viewed in an image of an organism: ac:subjectPart, to indicate the morphological component of the organism included in the view, and ac:subjectOrientation, to describe the direction or viewing angle of the subject part relative to the image aquisition device. Although it was recommended that values for those terms come from controlled vocabularies, no such vocabularies had been created by TDWG. In 2019, the Views Controlled Vocabularies Task Group*2 was chartered to develop controlled vocabularies for these two terms. The result was two Simple Knowledge Organization System*3 (SKOS) concept schemes*4, 5, and a mechanism for determining which subjectOrientation values are appropriate for a given subjectPart and which subjectParts are appropriate for various organism groups. In this presentation, we briefly review the vocabulary development process, key features of the vocabularies, and give an overview of how the vocabularies can be used in several example cases.
One impediment to the uptake of linked data technology is developers’ unfamiliarity with typical Resource Description Framework (RDF) serializations like Turtle and RDF/XML. JSON for Linking Data (JSON-LD) is designed to bypass this problem by expressing linked data in the well-known Javascript Object Notation (JSON) format that is popular with developers. JSON-LD is now Google’s preferred format for exposing Schema.org structured data in web pages for search optimization, leading to its widespread use by web developers. Another successful use of JSON-LD is by the International Image Interoperability Framework (IIIF), which limits its use to a narrow design pattern, which is readily consumed by a variety of applications. This presentation will show how a similar design pattern has been used in Audubon Core and with Biodiversity Information Standards (TDWG) controlled vocabularies to serialize data in a manner that is both easily consumed by conventional applications, but which also can be seamlessly loaded as RDF into triplestores or other linked data applications. The presentation will also suggest how JSON-LD might be used in other contexts within TDWG vocabularies, including with the Darwin Core Resource Relationship terms.
Users may be more likely to understand and utilize standards if they are able to read labels and definitions of terms in their own languages. Increasing standards usage in non-English speaking parts of the world will be important for making biodiversity data from across the globe more uniformly available. For these reasons, it is important for Biodiversity Information Standards (TDWG) to make its standards widely available in as many languages as possible. Currently, TDWG has six ratified controlled vocabularies*1, 2, 3, 4, 5, 6 that were originally available only in English. As an outcome of this workshop, we have made term labels and definitions in those vocabularies available in the languages of translators who participated in its sessions. In the introduction, we reviewed the concept of vocabularies, explained the distinction between term labels and controlled value strings, and described how multilingual labels and definitions fit into the standards development process. The introduction was followed by working sessions in which individual translators or small groups working in a single language filled out Google Sheets with their translations. The resulting translations were compiled along with attribution information for the translators and made freely available in JavaScript Object Notation (JSON) and comma separated values (CSV) formats.*7
Because TDWG vocabularies change and grow as they are developed by the community, it is nearly impossible to document their version history and generate both machine and human readable documentation by manual editing of multiple documents in several formats. In this talk, I will provide an overview of the workflow that has been established to maintain vocabularies in accordance with the TDWG Standards Documentation and Vocabulary Maintenance specifications. I will show how vocabulary creators and maintainers can use simple CSV spreadsheets to create new vocabularies or to update existing ones. I will also provide an overview of the Python scripts that TDWG infrastructure maintainers use to process those simple spreadsheets to turn them into the authoritative files in TDWG's rs.tdwg.org GitHub repository, which serves as the data source for both machine readable serializations of the vocabularies and human readable standards documents.
Digitisation and publication of museum specimen data is happening worldwide, but far from complete. Museums can start by sharing what they know about their holdings at a higher level, long before each object has its own record. Information about what is held in collections worldwide is needed by many stakeholders including collections managers, funders, researchers, policy-makers, industry, and educators. To aggregate this information from collections, the data need to be standardised (Johnston and Robinson 2002). So, the Biodiversity Information Standards (TDWG) Collection Descriptions (CD) Task Group is developing a data standard for describing collections, which gives the ability to provide: automated metrics, using standardised collection descriptions and/or data derived from specimen datasets (e.g., counts of specimens) and a global registry of physical collections (i.e., digitised or non-digitised). automated metrics, using standardised collection descriptions and/or data derived from specimen datasets (e.g., counts of specimens) and a global registry of physical collections (i.e., digitised or non-digitised). Outputs will include a data model to underpin the new standard, and guidance and reference implementations for the practical use of the standard in institutional and collaborative data infrastructures. The Task Group employs a community-driven approach to standard development. With international participation, workshops at the Natural History Museum (London 2019) and the MOBILISE workshop (Warsaw 2020) allowed over 50 people to contribute this work. Our group organized online "barbecues" (BBQs) so that many more could contribute to standard definitions and address data model design challenges. Cloud-based tools (e.g., GitHub, Google Sheets) are used to organise and publish the group's work and make it easy to participate. A Wikibase instance is also used to test and demonstrate the model using real data. There are a range of global, regional, and national initiatives interested in the standard (see Task Group charter). Some, like GRSciColl (now at the Global Biodiversity Information Facility (GBIF)), Index Herbariorum (IH), and the iDigBio US Collections List are existing catalogues. Others, including the Consortium of European Taxonomic Facilities (CETAF) and the Distributed System of Scientific Collections (DiSSCo), include collection descriptions as a key part of their near-term development plans. As part of the EU-funded SYNTHESYS+ project, GBIF organized a virtual workshop: Advancing the Catalogue of the World's Natural History Collections to get international input for such a resource that would use this CD standard. Some major complexities present themselves in designing a standardised approach to represent collection descriptions data. It is not the first time that the natural science collections community has tried to address them (see the TDWG Natural Collections Description standard). Beyond natural sciences, the library community in particular gave thought to this (Heaney 2001, Johnston and Robinson 2002), noting significant difficulties. One hurdle is that collections may be broken down into different degrees of granularity according to different criteria, and may also overlap so that a single object can be represented in more than one collection description. Managing statistics such as numbers of objects is complex due to data gaps and variable degrees of certainty about collection contents. It also takes considerable effort from collections staff to generate structured data about their undigitised holdings. We need to support simple, high-level collection summaries as well as detailed quantitative data, and to be able to update as needed. We need a simple approach, but one that can also handle the complexities of data, scope, and social needs, for digitised and undigitised collections. The data standard itself is a defined set of classes and properties that can be used to represent groups of collection objects and their associated information. These incorporate common characteristics ('dimensions') by which we want to describe, group and break down our collections, metrics for quantifying those collections, and properties such as persistent identifiers for tracking collections and managing their digital counterparts. Existing terms from other standards (e.g. Darwin Core, ABCD) are re-used if possible. The data model (Fig. 1) underpinning the standard defines the relationships between those different classes, and ensures that the structure as well as the content are comparable across different datasets. It centres around the core concept of an 'object group', representing a set of physical objects that is defined by one or more dimensions (e.g., taxonomy and geographic origin), and linked to other entities such as the holding institution. To the object group, quantitative data about its contents are attached (e.g. counts of objects or taxa), along with more qualitative information describing the contents of the group as a whole. In this presentation, we will describe the draft standard and data model with examples of early adoption for real-world and example data. We will also discuss the vision of how the new standard may be adopted and its potential impact on collection discoverability across the collections community.
To improve the suitability of the Darwin Core standard for the research and management of alien species, the standard needs to express the native status of organisms, how well established they are and how they came to occupy a location. To facilitate this, we propose: 1. To adopt a controlled vocabulary for the existing Darwin Core term dwc:establishmentMeans 2. To elevate the pathway term from the Invasive Species Pathways extension to become a new Darwin Core term dwc:pathway maintained as part of the Darwin Core standard 3. To adopt a new Darwin Core term dwc:degreeOfEstablishment with an associated controlled vocabulary These changes to the standard will allow users to clearly state whether an occurrence of a species is native to a location or not, how it got there (pathway), and to what extent the species has become a permanent feature of the location. By improving Darwin Core for capturing and sharing these data, we aim to improve the quality of occurrence and checklist data in general and to increase the number of potential uses of these data.
For the last 15 years, Biodiversity Information Standards (TDWG) has recognized two competing standards for organism occurrence data, ABCD (Access to Biological Collections Data; Holetschek et al. 2012) and DarwinCore (Wieczorek et al. 2012). These two representations emerged from contrasting strategies for mobilizing information about organism occurrences (also commonly called species occurrence data). ABCD was capable of representing details of more kinds of information, but was necessarily more complicated. DarwinCore, on the other hand, was simpler but more limited in its ability to represent data of different kinds and formats. TDWG endorsed both standards because the different projects and communities that generated them remained dedicated to their different strategies and tool sets, and the Global Biodiversity Information Facility (GBIF) developed the ability to integrate data published in either standard. Since their inceptions, DarwinCore and ABCD have become more similar. DarwinCore has gotten more complicated through the addition of terms and has begun to assign terms to classes. ABCD is now expressed in RDF (Resource Description Framework), potentially enabling re-use of terms with alternative structures among classes. At the same time, methodologies for conceptual modeling and representing complex scientific data have continued to evolve. In particular, a suite of modeling and data representation methods related to linked data and the semantic web, i.e., RDF, SKOS (Simple Knowledge Organization System), and OWL (web Ontology Language), promise to make it easier for us to reconcile shared concepts among different representations or schemas. A mapping between ABCD 2.1 and DarwinCore has existed since before 2005.*1 ABCD 3.0 and DarwinCore are both now represented in RDF. In addition, the BioCollections Ontology (BCO) covers many of the shared concepts and is derived from the Basic Formal Ontology (BFO), an upper level ontology that has oriented many other biomedical ontologies. Reconciling ABCD and DarwinCore through alignment with BCO (in the OBO Foundry; Smith et al. 2007) would better connect TDWG standards to other domains in biology. We appreciate that many working scientists and data managers perceive ontologies as overly complicated. To mitigate the steep learning curve associated with ontologies, we expect to create simpler application profiles or schemas to guide and serve narrower communities of practice within the wider biodiversity domain. We also plan to integrate the current work of the Taxonomic Names and Concepts Interest Group and thereby eliminate the redundancy between DarwinCore and Taxonomic Concepts Transfer Schema (TCS; Kennedy et al. 2006). At the time of this writing, we have only agreements from the authors (i.e., conveners of relevant TDWG Interest Groups and other key stakeholders) to collaborate in pursuit of these common goals. In this presentation we will give a more detailed description of our objectives and products, the methods we are using to achieve them, and our progress to date.
Knowledge graphs have the potential to unite disconnected digitized biodiversity data, and there are a number of efforts underway to build biodiversity knowledge graphs. More generally, the recent popularity of knowledge graphs, driven in part by the advent and success of the Google Knowledge Graph, has breathed life into the ongoing development of semantic web infrastructure and prototypes in the biodiversity informatics community. We describe a one week training event and hackathon that focused on applying three specific knowledge graph technologies – the Neptune graph database; Metaphactory; and Wikidata - to a diverse set of biodiversity use cases. We give an overview of the training, the projects that were advanced throughout the week, and the critical discussions that emerged. We believe that the main barriers towards adoption of biodiversity knowledge graphs are the lack of understanding of knowledge graphs and the lack of adoption of shared unique identifiers. Furthermore, we believe an important advancement in the outlook of knowledge graph development is the emergence of Wikidata as an identifier broker and as a scoping tool. To remedy the current barriers towards biodiversity knowledge graph development, we recommend continued discussions at workshops and at conferences, which we expect to increase awareness and adoption of knowledge graph technologies.