Taxonomic lists are important tools for efficient communication about biodiversity. The processes by which they are created and maintained need to be robust, scientifically sound, and transparent. Articulating and scoring a set of governance quality indicators provides a way to assess the relative strengths of list management, gives list users a means to assess the quality of this process, augments the information available to list aggregators, and allows patterns to be measured over time and among forms of life. Based on published principles, we created 12 governance quality indicators which we tested on 16 lists spanning a range of taxonomic groups. Independence of taxonomy from nomenclature scored most strongly, but scores for local and regional involvement were lower. The governance quality indicators may eventually provide a rubric for assessing best practice species list governance but now need a period of further testing, review, and refinement before they are institutionalized.
Against a background of the climate and biodiversity crises, there is an urgent need for robust and citable biodiversity information for policy and management decisions. Species are fundamental units of biodiversity and underpin communication in biology. Delineating, describing, and naming species provide the foundation for tracking biodiversity. Taxonomists recognise over 2 million described species, the scientific names of which follow provisions of codes of nomenclature, providing stability for communication about biodiversity. However, described species represent only a fraction of global biodiversity. Current advances in the fields of molecular biology and the growing use of image-based identifications have resulted in an explosion of informal species names globally, herein referred to as temporary names, increasing the rate of discovery of undescribed species and cryptic species complexes. We define two categories of temporary names: Type 1 names that are delineated in a local context but not further assessed; and Type 2 names that have been taxonomically assessed and recognised as either new or part of an unresolved species complex. We explore the different types and uses of temporary names, indicate how they can be managed in a robust and standardised manner and demonstrate how biodiversity databases, such as WoRMS, can be expanded to allow the tracking of both formal and informal scientific names. We propose a solution for the expanding problem of temporary names by defining and recommending the addition of Type 2 temporary names to nomenclatural databases such as WoRMS. We provide practical recommendations on how such names should be selected for entry and then entered to databases in a standardised way. These recommendations are a small step forward, but their broad adoption would support the robust integration of informal and formal taxonomies.
Based on a thorough literature review and expert consultation, this study provides an inventory of all introduced non-indigenous species (iNIS) reported for Belgian marine and brackish waters. The data indicate a strong increase in iNIS in the study area from the 1990s onward, averaging 2.2 newly detected species per year, with a cumulative total of 108 iNIS between 1800 and 2024. The majority of these iNIS have the Northwestern Pacific or Northwestern Atlantic as their native region and are primarily introduced in Western Europe via shipping or aquaculture. In addition to compiling the inventory, the context in which the iNIS are detected is examined, distinguishing between official monitoring programs, project-based data collection efforts, and citizen science initiatives. Our findings indicate that while the EU aims to promote coordination between its Marine Strategy Framework Directive (MSFD) and Water Framework Directive (WFD), a misalignment occurs in the practical implementation of iNIS monitoring at the Belgian level. For example, a coherent and integrated monitoring framework across marine, brackish, and freshwater systems is still lacking. Furthermore, despite the EU’s ambition to ensure comprehensive iNIS monitoring, no legal framework currently mandates targeted monitoring in coastal ports, despite their well-documented role as hotspots for new marine introductions. After all, iNIS monitoring is only mandatory under the MSFD, which in essence applies only seaward from the coastal baseline and therefore does not cover waters within these ports. In addition, while the EU’s IAS Regulation has recently incorporated a few marine species on the Union list, its monitoring requirements remain primarily focused on terrestrial and freshwater species. As a result, observations published by citizens with significant expertise in the field represent the primary source of marine iNIS data in coastal port areas in recent decades in Belgium. The fragmentary nature of iNIS data complicates the efficient flow of information to international or European iNIS reference databases that support policy and decision-making. Yet, even species officially reported by Member States under MSFD Descriptor 2 are not always included in these reference databases. Nonetheless, accurate data on iNIS presence and distribution are essential for effectively targeting and managing iNIS.
The European Ocean Biodiversity Information System (EurOBIS) was established in 2004, as part of the Marine Biodiversity and Ecosystem Functioning European Union Network of Excellence (MarBEF) project. One of the key project tasks was to integrate different resources on marine biodiversity. This gave birth to EurOBIS, a data system to capture, integrate and present distribution of marine species from individual datasets. Integration and consolidation aimed to provide a better understanding of long-term, large-scale patterns in European marine waters. The first marine biogeographic data went live in August 2004, and its data content has been growing steadily. The general EurOBIS goal is to help fill gaps in our scientific knowledge by making diverse biogeographic data on marine species freely available and accessible online. It is part of the OBIS network, focusing on data collected within European marine waters, or collected by European researchers and institutes outside Europe. EurOBIS closely collaborates with other regional OBIS nodes in Europe. Over time, EurOBIS formed alliances with European initiatives as a supporting infrastructure and network, including serving as the backbone of the European Marine Observations and Data Network Biology (EMODnet Biology) since 2009 (Perez Perez et al. 2023) and being part of the central LifeWatch Species Information Backbone since 2014. Both projects ensure a constant flow of marine species occurrence data to EurOBIS. Nowadays, European Horizon project calls include a clause requiring that marine data generated within these projects need to become part of European data flows, specifically becoming part of EMODnet. This involves an increase in needed support of the EurOBIS Data Management Team (DMT), and additional training of data providers on how to deal with marine biodiversity data to make them suitable for a (semi-)automated flow to EMODnet Biology. These needs are addressed in several ways, ranging from in-person data training to online training courses being offered through the Ocean Teacher Global Academy (OTGA). The EurOBIS infrastructure is dynamic, keeping track of recent developments in the field of data formats and standards, compatible with DarwinCore (DwC) (Wieczorek et al. 2012). While EurOBIS started out as capturing presence and abundance data of species, it follows trends and needs in biodiversity informatics, allowing it to deal with all biodiversity-related measurements and DwC changes, including e.g., the development and implementation of the Extended Measurement or Fact extension (De Pooter et al. 2017). There is a parallel evolution in the way data are being mobilized and the tools offered to make data submission as easy as possible. Early on, data were submitted largely through email or a File Transfer Protocol (FTP) server. The odd data transfer used Distributed Generic Information Retrieval (DiGIR). Since the ratification of DwC as a TDWG standard, a gradual transfer was observed from FTP and DiGIR to the usage of the Integrated Publishing Toolkit (Robertson et al. 2014) to share data. The DMT still receives datasets through email, although this is becoming an exception. The DMT also invests in online tools that help data suppliers to control their data, in terms of format, quality and adhering to accepted standards, increasing both the data quality and ease of sharing. Over time, the DMT observed several changes within the data-providing community. The rather reluctant attitude towards sharing data of the late 1990s/early 2000s has decreased, as data management and FAIR (Findable, Accessible, Interoperable, Reusable) data (Wilkinson et al. 2016) became more embedded in the daily work of scientists. The DMT distinguishes several types of data providers: the perfectionists, the enthusiasts, the pragmatic minimalists, the legacy heirs. the perfectionists, the enthusiasts, the pragmatic minimalists, the legacy heirs. Each provider type poses challenges, balancing the need to capture the minimum information for the data to become FAIR and be useable by others, and a desire to be perfect, where perfection can be hard to reach and out of the control of the provider. The biggest struggle is not the barrier of sharing data, but dealing with the illustrious monsters “quality control” and “standard vocabularies.” There is still a need to train researchers on how to process and manage their data, which is a field where the regional and global OBIS networks can play a continuing important role.
GuardIAS is a three-year Horizon Europe project starting in January 2025, uniting diverse expertise to address aquatic invasive alien species (IAS) management. This multidisciplinary initiative comprises seven interconnected work packages targeting all invasion stages (pre-border, border, post-border) to develop tools for disrupting invasions. GuardIAS will employ Artificial Intelligence and data workflows to enhance biodiversity databases with species distributions, environmental tolerances, traits, and genetic information, thereby improving the European Alien Species Information Network (EASIN) and harmonizing key datasets. The citizen science platform iNaturalist will be enriched with expert-verified images of aquatic IAS for early detection and monitoring of geographic spread. An Early Warning System focused on IAS of EU concern will be developed and integrated into EASIN. To prevent hull biofouling-a major IAS introduction pathway-GuardIAS will explore nanotechnology-based antifouling coatings. The project will also investigate recreational boat movements along European coastlines, an understudied factor in IAS secondary dispersal. An eDNA reference library and assay panel will be developed for effective IAS detection. Advanced models, such as the Nobel Prize winning Multi-Region Input-Output analysis, will assess IAS risks, including impacts on threatened species and critical habitats under current and future scenarios. Systematic conservation planning tools will prioritize IAS monitoring and management actions based on their impacts. GuardIAS will enhance data collection, monitoring, early detection, and public awareness through innovative citizen science initiatives like BioArtBlitz events-where arts serve as a communication vehicle- eDNA sampling, sound analysis projects on Zooniverse, and marina events for boaters. Stakeholder engagement will be fostered through applied games. Collaborating with environmental authorities, industry, and aquatic managers, the project will co-design and implement eradication and control efforts in marine and freshwater environments. By integrating Social Sciences and Humanities, GuardIAS will promote collaborative knowledge creation, understand public perceptions on IAS management, and facilitate exploitation of the project's outcomes.
The World Register of Marine Species (WoRMS) started in 2007 with the question “how many species live in our oceans?”. Now, a little over 15 years later, WoRMS is able to answer several questions related to marine species discovery rates and provides a dynamic number of existing marine species, based on the information provided by hundreds of taxonomic experts worldwide, who have proven to be diverse and dynamic. We present basic statistics on marine species discovery rates based on the currently available content of WoRMS, as well as insights in the day-to-day activities and dynamics of our editorial board and the progress made so far on the content priorities as defined by the WoRMS Steering Committee. As for all dynamic systems, WoRMS is not complete and faces challenges. As an endorsed project of the UN Ocean Decade, WoRMS aims to tackle a number of these challenges and knowledge-gaps by 2030, including detailed documentation of authorships and original descriptions, and will provide continuous support to all marine initiatives, programs and projects that rely on WoRMS as an authoritative classification and catalogue of marine names.
Based on the World Register of Marine Species (WoRMS), there are currently c. 242,000 known valid marine species living in the world's oceans and marine biota continue to be discovered and named steadily at a current average of 2,332 new species per year. The “average” newly described marine species is a benthic crustacean, annelid, or mollusc between 2 and 10 mm in size, living in the tropics at depths of 0–60 m, and represented in the description by 7–19 specimens. It is described after a shelf life of 13.5 years in an article with two to three authors in a journal with an IF <1, published by an academic institution or society or a small commercial publisher. It is highly likely that the description is not accompanied by molecular data and that its authors do not work in an institution in a region of the world where the new species comes from. At the current pace of discovery and characterization, it will take several hundred years to describe the remaining 1–2 million unknown marine species. With increased facilitation of access to literature, marine taxonomy will increasingly rely on retired professionals and citizen scientists. The barriers to new marine species descriptions are in part technological (access to habitats that are difficult to sample) and educational (training to generate and use molecular barcodes), but mostly institutional (funding of taxonomic work) and regulatory (restrictions imposed by access and benefit sharing legislation).
EMODnet Biology (hosted and coordinated by the Flanders Marine Institute (VLIZ)) is one of the seven themes within the European Marine Observation and Data network (EMODnet). The EMODnet Biology consortium aims to facilitate the accessibility and usage of marine biodiversity data. With the principle of "collect once, use many times" at its core, EMODnet Biology fosters collaboration across various sectors, including research, policy-making, industry, and individual citizens, to enhance knowledge sharing and inform decision-making. EMODnet Biology focuses on providing free and open access to comprehensive historical and recent data on the occurrence of marine species and their traits in all European regional seas. It achieves this through partnerships and collaboration with diverse international initiatives, such as the World Register of Marine Species (WoRMS), Marine Regions and the European node of the Ocean Biodiversity Information System (EurOBIS) among others. By promoting the usage of the Darwin Core Standard (Wieczorek et al. 2012), EMODnet Biology fosters data interoperability and ensures seamless integration with wider networks such as the Global Biodiversity Information Facility (GBIF) and the Ocean Biodiversity Information System (OBIS), serving as a significant data provider of the latter, as it is responsible for most of its data generated in Europe. Since its inception, EMODnet Biology has undertaken actions covering various areas, including providing access to marine biological data with spatio-temporal, taxonomic, environmental- and sampling-related information among others; developing an exhaustive data quality control tool based on the Darwin Core standard, the British Oceanographic Data Centre and Natural Environment Research Council Vocabulary Server (BODC NVS2) parameters and other controlled vocabularies used; creating and providing training courses to guide data providers; performing gap analyses to identify data quality and coverage shortcomings; creating and publishing marine biological distribution maps for various species or species groups; and interacting with international and European initiatives, projects and organizations. providing access to marine biological data with spatio-temporal, taxonomic, environmental- and sampling-related information among others; developing an exhaustive data quality control tool based on the Darwin Core standard, the British Oceanographic Data Centre and Natural Environment Research Council Vocabulary Server (BODC NVS2) parameters and other controlled vocabularies used; creating and providing training courses to guide data providers; performing gap analyses to identify data quality and coverage shortcomings; creating and publishing marine biological distribution maps for various species or species groups; and interacting with international and European initiatives, projects and organizations. Furthermore, EMODnet Biology contributes to the overall EMODnet initiative, which covers multidisciplinary data and products. Thanks to the use of standard protocols and tools across disciplines, EMODnet Biology products can contribute to multidisciplinary analysis of pressures and impacts on key marine species and habitats, and, lastly, support a better management and planning of the maritime space. In conclusion, EMODnet Biology plays a pivotal role in biodiversity informatics by providing users with a wealth of accessible and reusable marine biodiversity data and products. Its collaborative approach, extensive partnerships, and adherence to the FAIR (Findable, Accessible, Interoperable, Reusable) data principles (Wilkinson et al. 2016) as well as to the Infrastructure for Spatial Information in Europe (INSPIRE) metadata technical guidelines (European Commission Joint Research Centre 2013) and the Open Geospatial Consortium (OGC) standards make it a valuable resource for advancing knowledge, informing policies, and supporting sustainable management of marine ecosystems.
The analysis of biological and ecological traits has a long history in evolutionary and ecological research. However, trait data are often scattered and standardised terminology that transcends taxonomic and biogeographical context are generally missing. As part of the development of a global trait database of marine species, we collated trait information for European seaweeds and structured the data within the standardised framework of the World Register of Marine Species (WoRMS). We collected 45 175 trait records for 21 biologically and ecologically relevant traits of seaweeds. This resulted in a trait database for 1745 European seaweed species of which more than half (56 %) of the records were documented at the species level, while the remaining 44 % were documented at a higher taxonomic level and subsequently inherited at lower levels. The trait database for European seaweeds will serve as a foundation for future research on diversity and evolution of seaweeds and their responses to global changes. The data will contribute to developing detailed trait-based ecosystem models and will be an important tool to inform marine conservation policies. The data are publicly accessible through the AlgaeTraits portal, https://doi.org/10.14284/574 (AlgaeTraits, 2022).
We provide an overview of the World Amphipoda Database (WAD), a global species database that is part of the World Register of Marine Species (WoRMS). Launched in 2013, the database contains entries for over 10,500 accepted species names. Edited currently by 31 amphipod taxonomists, following WoRMS priorities, the WAD has at least one editor per major group. All accepted species are checked by the editors, as is the authorship available for all of the names. The higher classification is documented for every species and a type species is recorded for every genus name. This constitutes five of the 13 priorities for completion, set by WoRMS. In 2015, five LifeWatch grants were allocated for WAD activities. These included a general training workshop in 2016, together with data input for the superfamily Lysianassoidea and for a number of non-marine groups. Philanthropy grants in 2019 and 2021 covered more important gaps across the whole group. Further work remains to complete the linking of unaccepted names, original descriptions, and environmental information. Once these tasks are completed, the database will be considered complete for 8 of the 13 priorities, and efforts will continue to input new taxa annually and focus on the remaining priorities, particularly the input of type localities. We give an overview of the current status of the order Amphipoda, providing counts of the number of genera and species within each family belonging to the six suborders currently recognized.
Historical biodiversity documents comprise an important link to the long-term data life cycle and provide useful insights on several aspects of biodiversity research and management. However, because of their historical context, they present specific challenges, primarily time- and effort-consuming in data curation. The data rescue process requires a multidisciplinary effort involving four tasks: (a) Document digitisation (b) Transcription, which involves text recognition and correction, and (c) Information Extraction, which is performed using text mining tools and involves the entity identification, their normalisation and their co-mentions in text. Finally, the extracted data go through (d) Publication to a data repository in a standardised format. Each of these tasks requires a dedicated multistep methodology with standards and procedures. During the past 8 years, Information Extraction (IE) tools have undergone remarkable advances, which created a landscape of various tools with distinct capabilities specific to biodiversity data. These tools recognise entities in text such as taxon names, localities, phenotypic traits and thus automate, accelerate and facilitate the curation process. Furthermore, they assist the normalisation and mapping of entities to specific identifiers. This work focuses on the IE step (c) from the marine historical biodiversity data perspective. It orchestrates IE tools and provides the curators with a unified view of the methodology; as a result the documentation of the strengths, limitations and dependencies of several tools was drafted. Additionally, the classification of tools into Graphical User Interface (web and standalone) applications and Command Line Interface ones enables the data curators to select the most suitable tool for their needs, according to their specific features. In addition, the high volume of already digitised marine documents that await curation is amassed and a demonstration of the methodology, with a new scalable, extendable and containerised tool, “DECO” (bioDivErsity data Curation programming wOrkflow) is presented. DECO’s usage will provide a solid basis for future curation initiatives and an augmented degree of reliability towards high value data products that allow for the connection between the past and the present, in marine biodiversity research.
The rise in demand for more FAIR (Findable, Accessible, Interoperable, and Reusable) data is being answered by increasingly automated ways to capture, process, publish and register biodiversity datasets. Coupled with the increasing possibilities for detecting hundreds of species in a single sample/event (i.e., eDNA), this results in taxonomic information that is multiple levels of magnitude higher than it was a couple of years ago. This spike in content has an adverse effect on the ability of researchers to find relevant datasets within catalogues, due to the limitations in storing and displaying the taxonomic metadata (e.g., in the real-estate of a webpage, in the timeframe required to access the quantity of information, in displaying information to users in a comprehensive and comprehensible way). WoRMS (World Register of Marine Species) is a taxonomic backbone that provides species information. One user of WoRMS is the Integrated Marine Information System (IMIS), a metadata catalogue for marine data, which is also the metadata catalogue of the European node of the Ocean Biodiversity Information System (EurOBIS). Taxonomic information added to the metadata records in IMIS are linked to WoRMS Persitent Identifiers or PIDs (AphiaIDs). Tension between providing all the taxonomic metadata while not overloading the catalogue is being addressed for the use-case of WoRMS+IMIS+EurOBIS. Our approach is to apply a filter-and-replace algorithm during the automated registration of the taxonomic metadata to describe available datasets. This technique reduces the detailed taxonomic information of actual occurrences in the dataset content into practical (good enough) metadata. It takes as input all the species in the dataset along with their hierarchical structure, as well as a configuration parameter allowing for an upper bound to the acceptable number of taxa to be output. The core principle of the algorithm is to start off with the minimal result set containing only the hierarchical root (always "biota") of the complete taxonomy in the dataset, and then to gradually consider replacing each element with its children one level deeper, as long as that replacement keeps fitting the upper bound for the total set. This approach ensures that no coverage is lost, meaning every taxon in the actual dataset is represented in the result, although possibly through one of its parents, X layers up. Note that as long as only one child is underlying, the switch will always happen. So by nature, it will go down to the lowest relevant detail without challenging the upper bound limit. It also allows for variation in the actual processing by allowing for different ordering strategies on the current result-set. Ordering strategies under test are: Order (descending) by weight of underlying available children → favouring more detail in those parts of the tree that have the most members in the setWith weight defined as the count of all underlying available species, ORWeight defined as the sum-product of those species with their actual occurrences in the dataset (thus further favouring detail to those parts of the tree that are more prevalent in the samples) Order (descending) by number of direct children → favouring fanning-out over available high-level siblings Order (ascending) by ratio of present children over available children in the taxa → favouring replacing too vague parents with the more specific sublevels that are actually in the dataset Order (descending) by weight of underlying available children → favouring more detail in those parts of the tree that have the most members in the setWith weight defined as the count of all underlying available species, ORWeight defined as the sum-product of those species with their actual occurrences in the dataset (thus further favouring detail to those parts of the tree that are more prevalent in the samples) With weight defined as the count of all underlying available species, OR Weight defined as the sum-product of those species with their actual occurrences in the dataset (thus further favouring detail to those parts of the tree that are more prevalent in the samples) Order (descending) by number of direct children → favouring fanning-out over available high-level siblings Order (ascending) by ratio of present children over available children in the taxa → favouring replacing too vague parents with the more specific sublevels that are actually in the dataset Alongside this basic approach, additional pre-processing of the taxa in the dataset can apply some form of "pruning". In this approach all nodes in the available taxon tree of the dataset are ordered (descending) by the weight of underlying children, and those at the end—below some defined cutoff ratio (extra parameter to the algorithm)—are simply discarded. Again, the interpretation of this weight (only species count, or multiplied by occurrence count) yields to variants of the algorithm to be tested. Applying this pruning means a deliberate departure is taken from the full-coverage guarantee mentioned earlier. Caution should be applied, of course, but removing more irrelevant (low occurrence) parts of the tree will allow for making space in the bounded result-set for more detail in the parts that are relevant. In order to create an objective basis for comparing the resulting variants, a number of "qualification" parameters are considered to quantify the effect of the suggested reduction. Based on the datasets in EurOBIS, the variants of this algorithm are being applied and results will be presented on how they affect the various qualification parameters. It is worth observing that both datasets and their metadata records are distinct resources that are being linked to species (taxa references) and that the different purposes they serve require different levels of detail to be presented. As other types of entities (publications, habitats, experts, geography, traits, etc.) are considered for linking to species, we believe similar reduction algorithms will be necessary.
DiSSCo Flanders aims at developing a standardised natural science collections management infrastructure, ensuring proper long-term conservation, and future re-usage of the collections. Meise Botanic Garden coordinates the Flemish consortium. This four-year project, funded by the FWO (Research Foundation – Flanders), started in January 2021. The consortium brings together both the more classical ‘museum’ collections (Meise Botanic Garden, Ghent University Museum), with research collections (Research Institute for Nature and Forest, Flanders Research Institute for Agriculture, Fisheries and Food, Flanders Marine Institute, universities), and living collections (Belgian Association of Botanic Gardens and Arboreta, Zoo of Antwerp (Poo 2022)). Many of the research collections are smaller orphan collections, lacking a (full-time) curator, for which collection management is not the core business of the hosting institute. The Belgian federal DiSSCo members are associated with the project, allowing close collaboration between European, national and regional (digital) collection management inititatives. Building on the expertise of the European DiSSCo-related projects, an initial high-level inventory and assessment of the collections has been made and will provide a better understanding of the collections landscape in Flanders (Van Baelen 2022). This step in the digitization will increase the findability of the diverse regional collections and highlight the available knowledge to the research community. Close collaboration with the TDWG Collections Descriptions interest group and the development of the Latimer Core standard (Woodburn 2021) will ensure maximal interoperability of the collection data. In order to be relevant for research, the Flemish collections need to be more interconnected and linked to other data sources. To ensure that collection data is adhering to the FAIR principles (Findable, Accessible, Interoperable and Resuable) and ready to connect to the DiSSCo research infrastructure, DiSSCo Flanders will have a large focus on the collection(s) management system (CMS). Depending on the needs and specificities of each collection, a strategy will be chosen to implement and/or collection data will be migrated to (an) optimised system(s). Parallel to the digital inventory and CMS choice, the consortium addresses specific themes in the format of working groups. Most of the institutes have identified that their molecular collection (DNA, tissues) has not been properly acknowledged as a separate long-term collection. Challenges such as the storage and curation of e-DNA and environmental samples (e.g., soil) have been identified. Very often it will only be feasible to add basic data in the CMS and essential data is not available to the researchers. DiSSCo Flanders invests in the enrichment of specimen data to ensure that a maximal amount of information becomes available to the community. This can be accomplished either through the use of citizen science/crowdsourcing using the DoeDat platform (Groom 2018), or by using machine learning techniques performed by the IDLab at Ghent University (Thirukokaranam Chandrasekar 2021). Besides the technical aspects of the DiSSCo research infrastructure, the consortium is covering topics such as data publication, legal aspects of collections, standard operating procedures, etc. through knowledge sharing and active dialog.
The World Register of Marine Species (WoRMS) is an authoritative classification and catalogue of marine names. The WoRMS portal and available web-services are a gateway to access a treasure-chest of information, not only on taxon names themselves, but also on their mutual relations (e.g., original names, accepted versus unaccepted names, taxonomic classification), and related information such as ecological traits, distributions and linked literature. Over its fifteen years of existence, WoRMS has not only been growing in content and quality, thanks to the voluntary efforts of more than 300 experts worldwide, it has also kept a technical trajectory that involves adapting to new standards and technologies. Although WoRMS has always been relatively easily accessible through its portal and web services, and applied the basic data-sharing principles of FAIR (Findable, Accessible, Interoperable, and Reusable), there is still room for improvement. The recently growing call for globally uniform identifiers coming from the application of the FAIR data sharing principles, and the growing investment into globally open and interlinked "digital twin" representations of our oceans and the organisms found in them, have introduced the fundamentals for growing a marine knowledge graph, and as a consequence, has directed some technical attention towards applying semantic web technologies. WoRMS plays a key role in the field of (marine) biodiversity, as this research field strongly relies on the correct usage of species names, and understanding the taxonomic relationships between taxon names. As WoRMS is regarded as the authoritative resource for marine names, it is also heavily used as a quality-control tool for the correct usage of taxon names within various European and global initiatives. WoRMS provides support to global databases and infrastructures that use (or are in need of) a marine taxonomic backbone, such as the LifeWatch Species Information Backbone, the Ocean Biodiversity Information System (OBIS) and the Global Ocean Observing System (GOOS), as well as improves the content and strengthens relationships with environment-independent initiatives and infrastructures such as the Catalogue of Life (COL), the Barcode of Life Data System (BoLD) & GenBank. In addition to the taxonomic value of WoRMS, it is also highly valued for its available information on species traits, which form a critical component in ecological marine research. In its role as a marine taxonomic backbone, along with being linked to numerous other environment-independent initiatives and infrastructures, WoRMS has always required an adaptability towards the challenging new ways specific applications and concrete research have been choosing to apply the identifiers affixed by WoRMS. It is in this tradition we now announce and describe our approach to publish the content of the register as fully linked open data, using semantic web technologies. We describe in some detail the choices made to select and apply specific vocabularies that already exist for the description and interconnected linking of taxon names, to address the tension between hanging on to an historic (URN) persistent identifier and providing a dereferenceable URI that supports the appreciated "follow your nose" property, to link to other relevant registries, to design meaningful predicates for inbound links to the register and, more technically to divide the available content into meaningful sub-sections for retrieval of optional detail, and to provide a roadmap for meaningful fragmentations of the full register to allow for an effective consumption of the relevant (e.g., newly updated) parts into specific data-consumption scenarios. to select and apply specific vocabularies that already exist for the description and interconnected linking of taxon names, to address the tension between hanging on to an historic (URN) persistent identifier and providing a dereferenceable URI that supports the appreciated "follow your nose" property, to link to other relevant registries, to design meaningful predicates for inbound links to the register and, more technically to divide the available content into meaningful sub-sections for retrieval of optional detail, and to provide a roadmap for meaningful fragmentations of the full register to allow for an effective consumption of the relevant (e.g., newly updated) parts into specific data-consumption scenarios. We believe this work to be an important step towards achieving some future goals. It should further facilitate the production of automated, managed or hybrid crosswalks between various taxonomic registers and classifications. To be especially considered here, is helping to make omics taxonomic references more meaningfully comparable with WoRMS. In the process of others linking their digital objects (e.g., services, datasets, publications, experts) to taxon IDs in WoRMS, they are effectively also linking to each other, which opens the doors between apparently unconnected bodies. Its continuing use as a global standard and trustworthy reference of community-accepted names for biological taxa becomes the essential glue connecting all sorts of services, initiatives, communities.
Biological ocean science has a long history; it goes back millennia, whereas the related data services have emerged in the recent digital era of the past decades. To understand where we come from—and why data services are so important—we will start by taking you back to the rise in the study of marine biology—marine biodiversity—and its key players, before immersing ourselves in the data life cycle, past and present joint global initiatives, and systems that allow(ed) scientists to more easily access biological data, online services through some simple keyboard strokes, and the many challenges we still encounter on a daily basis when dealing with these types of data.
Integrative ZoologyEarly View LETTER TO THE EDITOR WoRMS needs YOU! A Reply to Collareta et al. 2020 Tammy HORTON, Corresponding Author tammy.horton@noc.ac.uk National Oceanography Centre, Southampton, UK Correspondence: Tammy Horton, National Oceanography Centre, European Way Southampton SO14 3ZH, United Kingdom of Great Britain and Northern Ireland. Email: tammy.horton@noc.ac.ukSearch for more papers by this authorAndreas KROH, Natural History Museum Vienna, AustriaSearch for more papers by this authorLeen VANDEPITTE, Flanders Marine Institute (VLIZ), Wandelaarkaai 7, Oostende, BelgiumSearch for more papers by this author Tammy HORTON, Corresponding Author tammy.horton@noc.ac.uk National Oceanography Centre, Southampton, UK Correspondence: Tammy Horton, National Oceanography Centre, European Way Southampton SO14 3ZH, United Kingdom of Great Britain and Northern Ireland. Email: tammy.horton@noc.ac.ukSearch for more papers by this authorAndreas KROH, Natural History Museum Vienna, AustriaSearch for more papers by this authorLeen VANDEPITTE, Flanders Marine Institute (VLIZ), Wandelaarkaai 7, Oostende, BelgiumSearch for more papers by this author First published: 07 January 2021 https://doi.org/10.1111/1749-4877.12519 Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onEmailFacebookTwitterLinked InRedditWechat No abstract is available for this article. Early ViewOnline Version of Record before inclusion in an issue RelatedInformation
This paper recommends best practice for the use of open nomenclature (ON) signs applicable to image-based faunal analyses. It is one of numerous initiatives to improve biodiversity data input to improve the reliability of biological datasets and their utility in informing policy and management. Image-based faunal analyses are increasingly common but have limitations in the level of taxonomic precision that can be achieved, which varies among groups and imaging methods. This is particularly critical for deep-sea studies owing to the difficulties in reaching confident species-level identifications of unknown taxa. ON signs indicate a standard level of identification and improve clarity, precision and comparability of biodiversity data. Here we provide examples of recommended usage of these terms for input to online databases and preparation of morphospecies catalogues. Because the processes of identification differ when working with physical specimens and with images of the taxa, we build upon previously provided recommendations for specific use with image-based identifications.
A major historical challenge for the management of anthropogenic introductions of species has been the absence of a globally standardised system for species nomenclature. For over a decade, the World Register of Marine Species (WoRMS) has provided a taxonomically authoritative classification and designation of the currently accepted names for all known marine species. However, WoRMS mainly focuses on taxonomy and does not specifically address species introductions. Here, we introduce the World Register of Introduced Marine Species (WRiMS), a database directly linked to WoRMS that includes all introduced marine species, distinguishing native and introduced geographic ranges. Both the WoRMS and WRiMS contents are continually updated by specialists who add citations of original species descriptions, key taxonomic literature, images and notes on native and introduced geographic distributions. WRiMS editors take responsibility for assessing the validity of species records by critically evaluating if a species has been introduced to a region, erroneously identified and/or potentially naturally present in a region but previously unnoticed. WRiMS currently contains 2,714 introduced species. The amount and quality of the information entered depend on the availability of experts to update its contents. Because WRiMS is global and it combines species taxonomic and geographic information with links to other resources and expertise, it is currently the most comprehensive standardised database of marine introduced species. In addition, WRiMS forms the basis for a future global early warning system of marine species introductions.