Environmental DNA (eDNA)-derived data, particularly DNA sequences and species occurrence records, are powerful resources for various applications, including biodiversity monitoring, biosecurity, and agricultural bio-surveillance. An important part of eDNA data’s value lies in its reusability, which can be enhanced through the use of metadata that adheres to FAIR (Findable, Accessible, Interoperable, Reusable) data principles (Wilkinson et al. 2016, Abarenkov et al. 2023). While FAIR practices are gaining recognition within the eDNA community, adoption remains limited. A key barrier is that existing data standards, including Darwin Core (DwC, Darwin Core Task Group 2009) and the Minimum Information about any (x) Sequence (MIxS), lack sufficient terms to fully describe eDNA workflows, from sampling and processing to bioinformatic outputs (Takahashi et al. 2025). To address this challenge, we have initiated a collaborative eDNA Task Group supported by the Biodiversity Information Standards (TDWG) and the Genomic Standards Consortium (GSC) organisations. We build upon current standardisation efforts to develop a comprehensive and interoperable eDNA metadata checklist for formal integration into the MIxS and DwC frameworks. See the task group descriptions at TDWG*2 and GSC.*1 Our goal is to ensure compatibility with, and seamless adoption by, major biodiversity and sequence data infrastructures, such as Global Biodiversity Information Facility (GBIF), Ocean Biodiversity Information System (OBIS), and International Nucleotide Sequence Database Collaboration (INSDC). This talk will present progress of the TDWG/GSC joint eDNA initiative, including our collaborative efforts with related working groups and initiatives, such as eDNAqua-plan, the Darwin Core Data Package, DNA-derived data extension to DwC, Minimum Information about any Ancient Sequence (MInAS), and the FAIR eDNA (FAIRe) initiatives (Takahashi et al. 2025). We invite contributions and collaboration to help shape a shared metadata framework that supports the growth, reuse, and long-term utility of eDNA data worldwide.
Environmental DNA (eDNA) has emerged as a transformative tool for monitoring aquatic biodiversity, offering a non-invasive and highly sensitive approach to detecting organisms across diverse ecosystems. However, its effective downstream application across Europe in environmental management is hindered by inconsistencies in data standardisation, metadata reporting, and accessibility. This perspective comprehensively evaluates current data repositories, data submission workflows, and standardisation efforts within the European aquatic eDNA landscape. By employing a multi-method approach, including an inventory of eDNA databases, a metadata assessment, a stakeholder questionnaire, and a generative Artificial Intelligence (AI)-driven analysis of scientific literature, our findings reveal substantial variability in metadata reporting practices, with several areas misaligned with Findable, Accessible, Interoperable, and Reusable (FAIR) principles. While some repositories demonstrate strong data curation and accessibility, others lack essential metadata descriptors, limiting interoperability. We identify critical gaps in metadata submission, particularly concerning sampling methods and wet lab workflows, which heavily impact data reusability. The use of generative AI in this study further enabled large-scale identification of recurring reporting weaknesses, highlighting structural challenges that extend beyond individual studies. Addressing these gaps and leveraging advanced computational approaches through international standards and harmonised guidelines represents a clear way forward, as articulated in the recent “Making eDNA FAIR” paper by Takahashi et al. (2025), which is based on the use of Darwin Core (DwC) and Genomics Standards Consortium (GSC) MIxS standards, as well as Global Biodiversity Information Facility (GBIF)’s “Publishing DNA-derived data through biodiversity data platforms” guidelines. Furthermore, additional complementary principles strengthen this framework. The Collective benefit, Authority to control, Responsibility, Ethics (CARE) principles emphasise Indigenous data governance and responsible sample stewardship, while the Transparency, Responsibility, User focus, Sustainability, Technology (TRUST) principles provide criteria for repository reliability and long-term digital preservation. Together, the combined application of FAIR, CARE, and TRUST principles provides a structured foundation for ensuring robust, interoperable, and ethically managed eDNA data that support aquatic biodiversity research, management, and conservation across Europe.
The European Nucleotide Archive (ENA; https://www.ebi.ac.uk/ena), hosted at the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI), remains a global, open-access platform for the submission, archiving, dissemination, and reuse of nucleotide sequence data. In 2025, ENA continues to advance its mission of fostering FAIR (findable, accessible, interoperable, reusable) data principles through innovations in interoperability, scalability, and global engagement, providing infrastructure for a rapidly growing volume of data across diverse domains. This article highlights the key developments in 2025, including the progress of the technical transformation, enhanced support for large-scale biodiversity projects, and the implementation of the International Nucleotide Sequence Database Collaboration Global Participation Initiative. We also discuss infrastructure enhancements to handle exponential data growth and improve user experiences and data discovery.
The success of environmental DNA (eDNA) approaches for species detection has revolutionized biodiversity monitoring and distribution mapping. Targeted eDNA amplification approaches, such as quantitative PCR, have improved our understanding of species distribution, and metabarcoding-based approaches have enabled biodiversity assessment at unprecedented scales and taxonomic resolution. eDNA datasets, however, are often scattered across repositories with inconsistent formats, varying access restrictions, and inadequate metadata; this limits their interoperation, reuse, and overall impact. Adopting FAIR (Findable, Accessible, Interoperable, and Reusable) data practices with eDNA data can transform the monitoring of biodiversity and individual species and support data-driven biodiversity management across broad scales. FAIR practices remain underdeveloped in the eDNA community, partly due to gaps in adapting existing vocabularies, such as Darwin Core (DwC) and Minimum Information about any (x) Sequence (MIxS), to eDNA-specific needs and workflows. To address these challenges, we propose a comprehensive FAIR eDNA (FAIRe) Metadata Checklist, which integrates existing data standards and introduces new terms tailored to eDNA workflows. Metadata are systematically linked to both raw data (e.g., metabarcoding sequences, Ct/Cq values of targeted qPCR assays) and derived biological observations (e.g., Amplicon Sequence Variant (ASV)/Operational Taxonomic Unit (OTU) tables, species presence/absence). Along with formatting guidelines, tools, templates, and example datasets, we introduce a standardized, ready-to-use approach for FAIR eDNA practices. Through broad collaboration, we seek to integrate these guidelines into established biodiversity and molecular data standards, promote journal data policies, and foster user-driven improvements and uptake of FAIR practices among eDNA data producers. In proposing this standardized approach and developing a long-term plan with key databases and data standard organizations, the goal is to enhance accessibility, maximize reuse, and elevate the scientific impact of these valuable biodiversity data resources.
The evolution of wastewater genomic surveillance (WWGS) has led to the development of many new methodologies, allowing for the broad application of WWGS for detection and monitoring of diverse pathogens and genetic markers. Variability in techniques and approaches creates challenges for data integration and interoperability that hinder analyses necessary for public health insights. Here, the Public Health Alliance for Genomic Epidemiology (PHA4GE) – in collaboration with scientists and stakeholders from over 20 countries, as well as global data repositories – presents a wastewater contextual data specification package relevant for a wide array of public health and research use cases. The PHA4GE wastewater contextual data specification is an ISO-compatible, ontology-based, modular data standard that is implemented by a free, open source data curation and validation tool called the DataHarmonizer. To facilitate interoperability and data sharing, interchange formats and instructions for automated transformations are included among the package’s supporting documentation. The specification package is part of a growing library of interoperable pathogen/target-specific standards designed upon a shared framework using semantic best practices. We hope that this standard will not only aid in the implementation of WWS, but also serve as an exemplar for the development of related data standards, such as for other environmental use cases or other metagenomic surveillance efforts.
The European Nucleotide Archive (ENA, https://www.ebi.ac.uk/ena), maintained at the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI) provides freely accessible services, both for deposition of, and access to, open nucleotide sequencing data. Open scientific data are of paramount importance to the scientific community and contribute daily to the acceleration of scientific advance. Outlined here are changes to and updates on the ENA service in 2024, aligning with the broad goals of enhancing interoperability, globalisation of the service and scaling the platform to meet current and future needs.
Mission Microbiomes AtlantECO (MMA) is an international study of the most fundamental fabric of the ocean — the ocean microbiome — aiming to understand its structure, functioning and connectivity in the South Atlantic Ocean. It is an initiative of the European Union’s research and innovation project AtlantECO, and an initiative of the All Atlantic Ocean Research and Innovation Alliance (AAORIA), which aims at sharing capacity along and across the Atlantic, and providing knowledge-based resources to help design policies for the management and protection of Atlantic Ecosystem Services. The success and legacy of MMA on science and society relies on the fair, open and inclusive collaboration among its partners from South America, South Africa and Europe. We will present MMA’s Data Sharing and Publication Best Practices (https://doi.org/10.5281/zenodo.7791092). With respect to the Convention on Biological Diversity (both the Nagoya Protocol & BBNJ Agreement), equitable access and shared capacity can be evidenced by statistics about where digital marine genetic resources originate within exclusive economic zones and the high seas, and who is accessing and exploiting these resources. We will present a few services and statistics that can be generated by the European Bioinformatics Institute (EMBL-EBI) in support to fair, open and inclusive collaboration.
The European Nucleotide Archive (ENA; https://www.ebi.ac.uk/ena) is maintained by the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI). The ENA is one of the three members of the International Nucleotide Sequence Database Collaboration (INSDC). It serves the bioinformatics community worldwide via the submission, processing, archiving and dissemination of sequence data. The ENA supports data types ranging from raw reads, through alignments and assemblies to functional annotation. The data is enriched with contextual information relating to samples and experimental configurations. In this article, we describe recent progress and improvements to ENA services. In particular, we focus upon three areas of work in 2023: FAIRness of ENA data, pandemic preparedness and foundational technology. For FAIRness, we have introduced minimal requirements for spatiotemporal annotation, created a metadata-based classification system, incorporated third party metadata curations with archived records, and developed a new rapid visualisation platform, the ENA Notebooks. For foundational enhancements, we have improved the INSDC data exchange and synchronisation pipelines, and invested in site reliability engineering for ENA infrastructure. In order to support genomic surveillance efforts, we have continued to provide ENA services in support of SARS-CoV-2 data mobilisation and have adapted these for broader pathogen surveillance efforts.
The eDNAqua-Plan project stands as a beacon of innovation in the biomonitoring of marine and freshwater ecosystems, propelled by the urgent need to integrate DNA-based approaches in aquatic bioassessment and monitoring frameworks. The broad utilisation of cutting-edge environmental DNA (eDNA) and DNA barcoding methodologies is dependent on complete, reliable, and accessible reference DNA sequence data (Rimet et al. 2021). Complete and interoperable metadata is crucial to allow a broad reuse of (e)DNA data and analysis outputs, and for a broader uptake of results by end users. The eDNAqua-Plan project aims to address key limitations to the routine implementation of eDNA-based monitoring methods in Europe by developing plans for federated DNA barcode reference libraries and eDNA data repositories to support DNA-based environmental monitoring. This will ensure a sustainable and reliable infrastructure to underpin its broad use, thereby paving the way for more effective conservation and management strategies. The project is working towards creating a comprehensive overview of standardisation efforts and data workflows, through collaborations with other projects, initiatives and infrastructures for aquatic monitoring across the European Union (EU) and associated countries. We are analysing existing archives (e.g., International Nucleotide Sequence Database Collaboration (INSDC), Barcode of Life Data System (BOLD), Global Biodiversity Information Facility (GBIF), Ocean Biodiversity Information System (OBIS)), portals, and papers to determine current and best practices through the use of questionnaires, manual evaluation of repositories, and machine learning methods (LLMs). This includes an overview of the usage of existing metadata and data standards (e.g., Minimum Information about any (X) Sequence Specifications from the Genomics Standards Consortium (GSC), Darwin Core standard). The results are being integrated by a team of experts in marine and freshwater biomonitoring. With a diverse consortium comprising 18 partner institutions from 11 countries and one international institute, eDNAqua-Plan brings together experts in marine and freshwater monitoring, eDNA analysis, and data science. The collective effort by this consortium will lay the groundwork for the creation of a digital ecosystem of eDNA repositories and an integrated reference library of marine and freshwater species, adhering to FAIR (Findable, Accessible, Interoperable, and Reusable) principles.
The mouse is an important model organism in the Human Genome Project and is set to play a pivotal role in the comparative analysis of diverse genomes that will be increasingly important for studies of gene function as we move post-genomics. In mouse genetics, it is recognised that the systematic generation of new mouse mutations along with identification of the underlying genes is an important challenge for gene function studies in the post-genomics era. Large numbers of mouse mutations affecting a plethora of biological pathways and with a diverse range of phenotypes need to be generated and mapped in the mouse genome. Many mouse mutations will be homologues of known human genetic disease loci and uncovering the underlying gene to any mouse mutation will shed light on gene function both in the mouse, human and other species. Sequencing of the mouse genome and its comparison with human genome sequence will not only aid gene identification but will provide a rapid route to uncovering genes underlying any mouse mutation. Furthermore, the sequence comparison of two mammalian genomes can be expected to provide profound new insights into gene regulation and genome evolution.
The European Nucleotide Archive (ENA; https://www.ebi.ac.uk/ena), maintained by the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI), offers those producing data an open and supported platform for the management, archiving, publication, and dissemination of data; and to the scientific community as a whole, it offers a globally comprehensive data set through a host of data discovery and retrieval tools. Here, we describe recent updates to the ENA's submission and retrieval services as well as focused efforts to improve connectivity, reusability, and interoperability of ENA data and metadata.
In this review, we provide a summary of recent progress in ontology mapping (OM) at a crucial time when biomedical research is under a deluge of an increasing amount and variety of data. This is particularly important for realising the full potential of semantically enabled or enriched applications and for meaningful insights, such as drug discovery, using machine-learning technologies. We discuss challenges and solutions for better ontology mappings, as well as how to select ontologies before their application. In addition, we describe tools and algorithms for ontology mapping, including evaluation of tool capability and quality of mappings. Finally, we outline the requirements for an ontology mapping service (OMS) and the progress being made towards implementation of such sustainable services.
The Pistoia Alliance was established nearly ten years ago to promote innovation by industry through pre-competitive collaboration to reduce the barriers to innovation. The Ontologies Mapping Project [1] was established in 2016 to enable better tools and services for mapping between ontologies and to establish best practices for ontology management in the Life Sciences. The project is now focussed on the development of an ontology mapping service (OMS).
Open PHACTS is a pre-competitive project to answer scientific questions developed recently by the pharmaceutical industry. Having high quality biological interaction information in the Open PHACTS Discovery Platform is needed to answer multiple pathway related questions. To address this, updated WikiPathways data has been added to the platform. This data includes information about biological interactions, such as stimulation and inhibition. The platform's Application Programming Interface (API) was extended with appropriate calls to reference these interactions. These new methods of the Open PHACTS API are available now.
The Pistoia Alliance was established nearly ten years ago to promote innovation by industry through pre-competitive collaboration to reduce the barriers to innovation. The Ontologies Mapping Project [1] was established in 2016 to enable better tools and services for mapping between ontologies and to establish best practices for ontology management in the Life Sciences. The project is now focussed on the development of an ontology mapping service (OMS).
BACKGROUND:The disease and phenotype track was designed to evaluate the relative performance of ontology matching systems that generate mappings between source ontologies. Disease and phenotype ontologies are important for applications such as data mining, data integration and knowledge management to support translational science in drug discovery and understanding the genetics of disease.RESULTS:Eleven systems (out of 21 OAEI participating systems) were able to cope with at least one of the tasks in the Disease and Phenotype track. AML, FCA-Map, LogMap(Bio) and PhenoMF systems produced the top results for ontology matching in comparison to consensus alignments. The results against manually curated mappings proved to be more difficult most likely because these mapping sets comprised mostly subsumption relationships rather than equivalence. Manual assessment of unique equivalence mappings showed that AML, LogMap(Bio) and PhenoMF systems have the highest precision results.CONCLUSIONS:Four systems gave the highest performance for matching disease and phenotype ontologies. These systems coped well with the detection of equivalence matches, but struggled to detect semantic similarity. This deserves more attention in the future development of ontology matching systems. The findings of this evaluation show that such systems could help to automate equivalence matching in the workflow of curators, who maintain ontology mapping services in numerous domains such as disease and phenotype.
Open PHACTS started as an Innovative Medicines Initiative (IMI) project in 2011. The objective was to integrate a wide range of pharmacologically related data into a central resource that is described according to standards of the semantic web community. Central to this unification was the development of an API to allow the project partners to build their own custom applications and workstreams. Also central to the Open PHACTS Discovery Platform is the development of identity resolution and identity mapping services and a chemistry resolution service. Over the last year the Open PHACTS Discovery Platform has matured with many exciting improvements and greater sustainability of the platform, improved code, data updates and new data sources.
Current research and development approaches to drug discovery have become less fruitful and more costly. One alternative paradigm is that of drug repositioning. Many marketed examples of repositioned drugs have been identified through serendipitous or rational observations, highlighting the need for more systematic methodologies to tackle the problem. Systems level approaches have the potential to enable the development of novel methods to understand the action of therapeutic compounds, but requires an integrative approach to biological data. Integrated networks can facilitate systems level analyses by combining multiple sources of evidence to provide a rich description of drugs, their targets and their interactions. Classically, such networks can be mined manually where a skilled person is able to identify portions of the graph (semantic subgraphs) that are indicative of relationships between drugs and highlight possible repositioning opportunities. However, this approach is not scalable. Automated approaches are required to systematically mine integrated networks for these subgraphs and bring them to the attention of the user. We introduce a formal framework for the definition of integrated networks and their associated semantic subgraphs for drug interaction analysis and describe DReSMin, an algorithm for mining semantically-rich networks for occurrences of a given semantic subgraph. This algorithm allows instances of complex semantic subgraphs that contain data about putative drug repositioning opportunities to be identified in a computationally tractable fashion, scaling close to linearly with network data. We demonstrate the utility of our approach by mining an integrated drug interaction network built from 11 sources. This work identified and ranked 9,643,061 putative drug-target interactions, showing a strong correlation between highly scored associations and those supported by literature. We discuss the 20 top ranked associations in more detail, of which 14 are novel and 6 are supported by the literature. We also show that our approach better prioritizes known drug-target interactions, than other state-of-the art approaches for predicting such interactions.