In October 2024, Parties to the United Nations Convention on Biological Diversity agreed to a new multilateral mechanism to fund biodiversity conservation through the sharing of benefits from open biodiversity data. Biological databases hosting genetic and other biological data, known as digital sequence information (DSI), are central to the implementation of the mechanism. This paper assesses the new international agreement and its implications for DSI databases. We walk through the database provisions in COP16 Decision 16/2, which include notifying users and submitters about the mechanism, improving metadata on geographical location of sample collection, and consistency with open access, as well as consideration of the FAIR, CARE, and TRUST principles. Drawing on surveys, interviews, and a workshop with biological database managers, we identify practical and scalable measures including updating terms of use, revising submission procedures, and strengthening user communication. We also propose approaches to capture and report non-monetary benefits such as capacity building, publications, interoperability, and training. These actions illustrate how DSI databases can remain open, sustainable, and globally connected while supporting benefit-sharing from the use of DSI on genetic resources.
The network of the national COVID-19 Data Portals was developed and linked to the COVID-19 Data Portal (https://www.covid19dataportal.org/) in response to the need for rapid data sharing and analysis during the 2020-2022 SARS-CoV-2 pandemic. Built on open-source code developed by the Swedish COVID-19 Data Portal (now the Swedish Pathogens Portal, www.pathogens.se) the network included 12 national portals addressing demand for local open data sharing and access, across data types and resources. It provides a robust case study of national initiatives for FAIR (Findable, Accessible, Interoperable and Reusable) resources and a foundation for future pandemic preparedness across pathogens globally. In this paper we outline the structure of the origins of the network of National COVID-19 Datal Portals, the technical aspects and code originating from the Swedish Portal and provide an overview of the services and tools offered by each Portal. The paper showcases the process and operation of four Portals: Sweden, Poland, Spain, Norway and The Netherlands. It considers useful lessons and approaches for future pandemic preparedness which enable researchers to easily identify and obtain the key data from resources on a national and international level.
The network of the national COVID-19 Data Portals was developed and linked to the COVID-19 Data Portal (https://www.covid19dataportal.org/)inresponsetothe need for rapid data sharing and analysis during the 2020–2022 SARS-CoV-2 pandemic. Built on open-source code developed by the Swedish COVID-19 Data Portal (now the Swedish Pathogens Portal, www.pathogens.se) the network included 12 national portals addressing demand for local open data sharing and access, across data types and resources. It provides a robust case study of national initiatives for FAIR (Findable, Accessible, Interoperable and Reusable) resources and a foundation for future pandemic preparedness across pathogens globally. In this paper we outline the structure of the origins of the network of National COVID-19 Datal Portals, the technical aspects and code originating from the Swedish Portal and provide an overview of the services and tools offered by each Portal. The paper showcases the process and operation of four Portals: Sweden, Poland, Spain, Norway and The Netherlands. In this study, we observe that pandemic response greatly benefits from an established infrastructure that can be quickly mobilised, developed and extended. Collaborations and preparation built on solid foundations over several years, supported by investment in the form of national and international research grants, is key for sustainability, continuation and readiness to deploy such efforts.
The Biodiversity Genomics Europe (BGE) Project has the overarching aim of accelerating the use of genomic science to enhance understanding of biodiversity, monitor biodiversity change, and guide interventions to address its decline. The BGE Project comprises activities focused on DNA Barcoding (Barcoding Stream) and Reference Genome Generation (Genomes Stream) for eukaryotic species across Europe, bringing together two European networks: the International Barcode of Life in Europe (iBOL Europe) and the European Reference Genome Atlas (ERGA). This publication is an abridged version of the successful grant proposal developed jointly by iBOL Europe and ERGA in response to the Horizon Europe call HORIZON-CL6-2021-BIODIV-01-01. Two key strands of genomic science form the basis of this proposal: DNA barcoding - sequencing short, standardised genomic regions to tell the world’s species apart, transforming the speed of completion of the inventory of life on Earth and providing the foundations of a global bio-surveillance system for biodiversity; and genome sequencing - generating high-quality complete reference genomes for all species on Earth, transforming understanding of biodiversity at the genetic level, and delivering fundamental knowledge of how biological systems function and how species respond and adapt to environmental change. The BGE Project objectives are focused on (i) Capacity: To establish functioning biodiversity genomics networks at the European level to connect and grow community capacity to use genomic tools to tackle the biodiversity crisis; (ii) Production: To establish and implement large-scale biodiversity genomic data generation pipelines for Europe to accelerate the production and accessibility of genomic data for biodiversity characterisation, conservation, and biomonitoring; and (iii) Application: To apply genomic tools to enhance understanding of pan-European biodiversity and biodiversity declines to improve the efficacy of management interventions and biomonitoring programmes.
Motivation:Biodata resources constitute a critical, large-scale, and globally distributed infrastructure underpinning life science research, yet their organic growth has hindered efforts to quantify key indicators needed to justify sustainable support, including usage, impact, and interdependencies. Here, we present an updated Global Biodata Coalition inventory alongside a Total Resource Usage (TRU) dataset that integrates this inventory with two complementary literature-derived sources: data citations and informal resource name mentions extracted from full-text articles using a fine-tuned machine learning model. A unified database schema enables cross-resource comparisons, dependency network analyses, and evaluation of resource name distinctiveness. Results:The combined dataset captures 11.5 million formal and informal references, revealing that most resources are acknowledged informally within article text. Network analysis indicates a densely interconnected ecosystem in which Global Core Biodata Resources function as key providers and integrators, underscoring their foundational role. While full resource names are generally distinctive, widespread use of acronyms limits detectability through text mining. Together, these findings provide robust empirical evidence of a highly utilized and interconnected biodata infrastructure, highlight limitations of single-metric assessments, and underscore the need for multi-dimensional evaluation frameworks and more consistent data citation practices to support informed decision-making and long-term sustainability. Availability and implementation:The database and analytical code described here are available on https://github.com/globalbiodata.
Abstract Members of the International Nucleotide Sequence Database Collaboration (INSDC; https://www.insdc.org/) collect, exchange, and preserve comprehensive open nucleotide sequence information and provide tools for its access. The INSDC has stated its commitment to welcoming new members into the collaboration to be more representative of the global community of data and users. To reach this goal, a comprehensive definition of the INSDC data model and minimum requirements for data acceptance have been established. Here we describe the processes used to arrive upon these INSDC Specifications and lay out strategies for their continued upkeep to remain current and relevant. Database URL: https://www.insdc.org/
The European Nucleotide Archive (ENA; https://www.ebi.ac.uk/ena), hosted at the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI), remains a global, open-access platform for the submission, archiving, dissemination, and reuse of nucleotide sequence data. In 2025, ENA continues to advance its mission of fostering FAIR (findable, accessible, interoperable, reusable) data principles through innovations in interoperability, scalability, and global engagement, providing infrastructure for a rapidly growing volume of data across diverse domains. This article highlights the key developments in 2025, including the progress of the technical transformation, enhanced support for large-scale biodiversity projects, and the implementation of the International Nucleotide Sequence Database Collaboration Global Participation Initiative. We also discuss infrastructure enhancements to handle exponential data growth and improve user experiences and data discovery.
The success of environmental DNA (eDNA) approaches for species detection has revolutionized biodiversity monitoring and distribution mapping. Targeted eDNA amplification approaches, such as quantitative PCR, have improved our understanding of species distribution, and metabarcoding-based approaches have enabled biodiversity assessment at unprecedented scales and taxonomic resolution. eDNA datasets, however, are often scattered across repositories with inconsistent formats, varying access restrictions, and inadequate metadata; this limits their interoperation, reuse, and overall impact. Adopting FAIR (Findable, Accessible, Interoperable, and Reusable) data practices with eDNA data can transform the monitoring of biodiversity and individual species and support data-driven biodiversity management across broad scales. FAIR practices remain underdeveloped in the eDNA community, partly due to gaps in adapting existing vocabularies, such as Darwin Core (DwC) and Minimum Information about any (x) Sequence (MIxS), to eDNA-specific needs and workflows. To address these challenges, we propose a comprehensive FAIR eDNA (FAIRe) Metadata Checklist, which integrates existing data standards and introduces new terms tailored to eDNA workflows. Metadata are systematically linked to both raw data (e.g., metabarcoding sequences, Ct/Cq values of targeted qPCR assays) and derived biological observations (e.g., Amplicon Sequence Variant (ASV)/Operational Taxonomic Unit (OTU) tables, species presence/absence). Along with formatting guidelines, tools, templates, and example datasets, we introduce a standardized, ready-to-use approach for FAIR eDNA practices. Through broad collaboration, we seek to integrate these guidelines into established biodiversity and molecular data standards, promote journal data policies, and foster user-driven improvements and uptake of FAIR practices among eDNA data producers. In proposing this standardized approach and developing a long-term plan with key databases and data standard organizations, the goal is to enhance accessibility, maximize reuse, and elevate the scientific impact of these valuable biodiversity data resources.
RNAcentral was founded in 2014 to serve as a comprehensive database of non-coding RNA sequences. It began by providing a single unified interface to more specialised resources, and now contains 45 million sequences. It has grown beyond providing a single interface to many specialised resources and now provides several services and analyses. These include secondary structure prediction with R2DT, sequence search, and analysis with Rfam. Since its last publication in 2021, RNAcentral has developed two major features. First, literature integration with the development of LitScan and LitSumm. LitScan automatically identifies and links relevant publications to RNA entries, while LitSumm uses natural language processing to generate functional summaries from the literature. Together, these tools address the critical challenge of connecting sequence data with scattered functional knowledge across thousands of publications. Secondly, RNAcentral has created gene level entries. Gene level entries represent a large structural change to RNAcentral. While RNAcentral previously organized data exclusively at the sequence level, we now group related transcripts into gene-centric views. This allows researchers to explore all isoforms, splice variants, and related sequences for a gene in a unified interface, better reflecting biological organization and facilitating comparative analyses. RNAcentral is freely available at: https://rnacentral.org .
The European Nucleotide Archive (ENA, https://www.ebi.ac.uk/ena), maintained at the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI) provides freely accessible services, both for deposition of, and access to, open nucleotide sequencing data. Open scientific data are of paramount importance to the scientific community and contribute daily to the acceleration of scientific advance. Outlined here are changes to and updates on the ENA service in 2024, aligning with the broad goals of enhancing interoperability, globalisation of the service and scaling the platform to meet current and future needs.
Understanding how environmental and ecological factors shape variability in soil-associated microbial communities is a complex problem, particularly on islands, which contain a wide range of diverse and unique geology, fauna, and flora. The island of Crete features sharp altitudinal gradients, diverse landscapes, and distinct ecological zones shaped by its complex geological history making it an ideal natural laboratory for studying how environmental variation influences soil microbial communities. In this study, we characterized the soil microbial communities across Crete’s ecozones and identify environmental factors associated with their diversity and composition. We performed a single-day, island-wide soil microbiota investigation, the first of its kind, to address this challenge by eliminating sources of variability including seasonality, weather conditions, anthropogenic or land use changes over time, and ecological succession of microbial communities. This island collection event (Island Sampling Day, ISD) was conducted in conjunction with the annual meeting of the Genomic Standards Consortium, on the island of Crete, and utilized standard data and metadata collection protocols. We generated amplicon sequences (V3-V4 regions of the 16 S ribosomal RNA gene) and a metadata-enriched dataset from 435 soil samples across 72 sites and four distinct ecozones for future whole-island microbiome studies. Here we report on the study design and sample collection process along with our initial examination of the ecological drivers of soil microbial community variability (e.g., elevation, soil types, soil pH, soil moisture, vegetation type, land use) across the Crete ecozones (defined by elevation and distinct habitats).
Mission Microbiomes AtlantECO (MMA) is an international study of the most fundamental fabric of the ocean — the ocean microbiome — aiming to understand its structure, functioning and connectivity in the South Atlantic Ocean. It is an initiative of the European Union’s research and innovation project AtlantECO, and an initiative of the All Atlantic Ocean Research and Innovation Alliance (AAORIA), which aims at sharing capacity along and across the Atlantic, and providing knowledge-based resources to help design policies for the management and protection of Atlantic Ecosystem Services. The success and legacy of MMA on science and society relies on the fair, open and inclusive collaboration among its partners from South America, South Africa and Europe. We will present MMA’s Data Sharing and Publication Best Practices (https://doi.org/10.5281/zenodo.7791092). With respect to the Convention on Biological Diversity (both the Nagoya Protocol & BBNJ Agreement), equitable access and shared capacity can be evidenced by statistics about where digital marine genetic resources originate within exclusive economic zones and the high seas, and who is accessing and exploiting these resources. We will present a few services and statistics that can be generated by the European Bioinformatics Institute (EMBL-EBI) in support to fair, open and inclusive collaboration.
The COVID-19 pandemic proved how sharing of genomic sequences in a timely manner, as well as early detection and surveillance of variants and characterization of their clinical impacts, helped to inform public health responses. However, the area of (re)emerging infectious diseases and our global connectivity require interdisciplinary collaborations to happen at local, national and international levels and connecting data to understand the linkages between all factors involved. Here, we describe experiences and lessons learned from a COVID-19 pilot study aimed at developing a model for storage and sharing linked laboratory data and clinical-epidemiological data using European open science infrastructure. We provide insights into the barriers and complexities of internationally sharing linked, complex cohort datasets from opportunistic studies for connected data analyses. An analytical timeline of events, describing key actions and delays in the execution of the pilot, and a critical path, defining steps in the process of internationally sharing a linked cohort dataset are included. The pilot showed how building on existing infrastructure that had previously been developed within the European Nucleotide Archive at the European Molecular Biology Laboratory-European Bioinformatics Institute for pathogen genomics data sharing, allowed the rapid development of connected “data hubs.” These data hubs were required to link human clinical-epidemiological data under controlled access with open high dimensional laboratory data, under FAIR (Findable, Accessible, Interoperable, Reusable) principles. Based on our own experiences, we call for action and make recommendations to support and to improve data sharing for outbreak preparedness and response.
We present the Aquatic Symbiosis Genomics Project, a global collaboration to generate high quality genome sequences for a wide range of eukaryotes and their microbial symbionts. Launched under the Symbiosis in Aquatic Systems Initiative of the Gordon and Betty Moore Foundation, the ASG Project brings together researchers from across the globe who hope to use these reference genomes to augment and extend their analyses of the dynamics, mechanisms and environmental importance of symbioses. Applying large-scale, high-throughput sequencing and assembly technologies, the ASG collaboration will assemble and annotate the genomes of 500 symbiotic organisms – both the “hosts” and the microbial symbionts with which they associate. These data will be released openly to benefit all who work on symbioses, from conservation geneticists to those interested in the origin of the eukaryotic cell.
The European Nucleotide Archive (ENA; https://www.ebi.ac.uk/ena) is maintained by the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI). The ENA is one of the three members of the International Nucleotide Sequence Database Collaboration (INSDC). It serves the bioinformatics community worldwide via the submission, processing, archiving and dissemination of sequence data. The ENA supports data types ranging from raw reads, through alignments and assemblies to functional annotation. The data is enriched with contextual information relating to samples and experimental configurations. In this article, we describe recent progress and improvements to ENA services. In particular, we focus upon three areas of work in 2023: FAIRness of ENA data, pandemic preparedness and foundational technology. For FAIRness, we have introduced minimal requirements for spatiotemporal annotation, created a metadata-based classification system, incorporated third party metadata curations with archived records, and developed a new rapid visualisation platform, the ENA Notebooks. For foundational enhancements, we have improved the INSDC data exchange and synchronisation pipelines, and invested in site reliability engineering for ENA infrastructure. In order to support genomic surveillance efforts, we have continued to provide ENA services in support of SARS-CoV-2 data mobilisation and have adapted these for broader pathogen surveillance efforts.
The COVID-19 pandemic has seen large-scale pathogen genomic sequencing efforts, becoming part of the toolbox for surveillance and epidemic research. This resulted in an unprecedented level of data sharing to open repositories, which has actively supported the identification of SARS-CoV-2 structure, molecular interactions, mutations and variants, and facilitated vaccine development and drug reuse studies and design. The European COVID-19 Data Platform was launched to support this data sharing, and has resulted in the deposition of several million SARS-CoV-2 raw reads. In this paper we describe (1) open data sharing, (2) tools for submission, analysis, visualisation and data claiming (e.g. ORCiD), (3) the systematic analysis of these datasets, at scale via the SARS-CoV-2 Data Hubs as well as (4) lessons learnt. This paper describes a component of the Platform, the SARS-CoV-2 Data Hubs, which enable the extension and set up of infrastructure that we intend to use more widely in the future for pathogen surveillance and pandemic preparedness.
Molecular sequencing data generation is being driven by global and regional efforts to discover, understand and monitor biodiversity. To fully explore this data in biodiversity research we need a network of connected data resources, linking sequence data with natural history collections, taxonomy and literature. The BiCIKL project (Biodiversity Community Integrated Knowledge Library, Penev et al. 2022) has set the groundwork towards creating this network of linked data and fostering FAIR (Findable, Accessible, Interoperable and Reusable) practices in the biodiversity domain. Connecting biodiversity and molecular data along the biodiversity research cycle requires a foundation of well-structured and rich metadata in the molecular sequence databases. Referencing the physical specimens is important as this provides context about the source of the material that was used for generating the molecular sequence data, including information about origin and species identification. To connect biodiversity and molecular data, we developed tools and workflows for improving and standardising metadata, federated searches and validations for specimen reference in sequence data, such as the SpASe tool, which enables the discovery of links between natural history collections and sequences, and the European Nucleotide Archive Source Attribute Helper API, which facilitates the construction of specimen attributes in a structured format. This work was done in close collaboration with DiSSCo (Distributed System of Scientific Collections) and some biodiversity genomics projects (e.g. Biodiversity Genomics Europe, BGE). Furthermore, we enabled community curation of biological source annotations such as specimen references in sequence data through the PlutoF platform and the ELIXIR Contextual Data Clearinghouse (Abarenkov et al. 2021, Balavenkataraman Kadhirvelu et al. 2022) and increased bidirectional linking from sequences in the European Nucleotide Archive (ENA) to collections, taxonomy and literature services (e.g., Plazi TreatmentBank, OpenBioDiv). We also worked closely with the community to enable the structured publication of environmental DNA data, promoting and engaging in the definition of standards and developing tools to facilitate data deposition and retrieval. Overall, the project has contributed significantly to strengthen the connections between the biodiversity and genomics communities towards higher data integration and interoperability. Structured, enriched, accessible and linked sequence data will provide a strong foundation for the application of biodiversity knowledge in the response to global challenges, such as biodiversity loss, ecosystem change and food security. Beyond BiCIKL, we will continue our work as a community to promote a culture of FAIR linked molecular data, towards a fully integrated biodiversity knowledge ecosystem.
The eDNAqua-Plan project stands as a beacon of innovation in the biomonitoring of marine and freshwater ecosystems, propelled by the urgent need to integrate DNA-based approaches in aquatic bioassessment and monitoring frameworks. The broad utilisation of cutting-edge environmental DNA (eDNA) and DNA barcoding methodologies is dependent on complete, reliable, and accessible reference DNA sequence data (Rimet et al. 2021). Complete and interoperable metadata is crucial to allow a broad reuse of (e)DNA data and analysis outputs, and for a broader uptake of results by end users. The eDNAqua-Plan project aims to address key limitations to the routine implementation of eDNA-based monitoring methods in Europe by developing plans for federated DNA barcode reference libraries and eDNA data repositories to support DNA-based environmental monitoring. This will ensure a sustainable and reliable infrastructure to underpin its broad use, thereby paving the way for more effective conservation and management strategies. The project is working towards creating a comprehensive overview of standardisation efforts and data workflows, through collaborations with other projects, initiatives and infrastructures for aquatic monitoring across the European Union (EU) and associated countries. We are analysing existing archives (e.g., International Nucleotide Sequence Database Collaboration (INSDC), Barcode of Life Data System (BOLD), Global Biodiversity Information Facility (GBIF), Ocean Biodiversity Information System (OBIS)), portals, and papers to determine current and best practices through the use of questionnaires, manual evaluation of repositories, and machine learning methods (LLMs). This includes an overview of the usage of existing metadata and data standards (e.g., Minimum Information about any (X) Sequence Specifications from the Genomics Standards Consortium (GSC), Darwin Core standard). The results are being integrated by a team of experts in marine and freshwater biomonitoring. With a diverse consortium comprising 18 partner institutions from 11 countries and one international institute, eDNAqua-Plan brings together experts in marine and freshwater monitoring, eDNA analysis, and data science. The collective effort by this consortium will lay the groundwork for the creation of a digital ecosystem of eDNA repositories and an integrated reference library of marine and freshwater species, adhering to FAIR (Findable, Accessible, Interoperable, and Reusable) principles.
We present the Aquatic Symbiosis Genomics Project, a global collaboration to generate high quality genome sequences for a wide range of eukaryotes and their microbial symbionts. Launched under the Symbiosis in Aquatic Systems Initiative of the Gordon and Betty Moore Foundation, the ASG Project brings together researchers from across the globe who hope to use these reference genomes to augment and extend their analyses of the dynamics, mechanisms and environmental importance of symbiosis. Applying large-scale, high-throughput sequencing and assembly technologies, the ASG collaboration will assemble and annotate the genomes of 500 symbiotic organisms – both the “hosts” and the microbial symbionts with which they associate. These data will be released openly to benefit all who work on symbiosis, from conservation geneticists to those interested in the origin of the eukaryotic cell.
The members of the International Nucleotide Sequence Database Collaboration (INSDC; https://insdc.org) have built systems to collect, archive and disseminate sequence data for more than four decades. The three collaborating organizations, the National Library of Medicine, National Center for Biotechnology Information (NLM-NCBI) in the United States, Research Organization of Information and Systems, National Institute of Genetics (ROIS-NIG) in Japan; and the European Molecular Biology Laboratory-European Bioinformatics Institute (EMBL-EBI) formalized their relationship through the adoption of an arrangement which documents their commitment to free and open access to genomic sequences. The INSDC is committed to expand the collaboration to be more representative of the global community of sequences and users. Diversifying participation through new membership will advance open science and data sharing and, in turn, drive innovation. This expansion will additionally benefit the INSDC and its broad user base by providing additional diverse perspectives as it explores emerging areas of data management, including federation, attribution and management.