The BioStudies database is an EMBL-EBI resource, serving as a repository for descriptions of biological studies, providing links to associated data housed in both internal databases and external resources, and accommodating data that do not fit into specialised archives. BioStudies offers mechanisms for defining and implementing metadata guidelines tailored to specific data sources, projects, or emerging communities, and organises datasets into collections.The platform facilitates data acquisition across the entire scientific process lifecycle, supporting data sharing during collaborative projects, data deposition for manuscript submission, and creation of data packages post-publication. Furthermore, BioStudies has assisted various projects in transitioning from legacy data infrastructures by offering sustainable data archival and access services. This article discusses the data collections within BioStudies across these different data lifecycle stages and details the underlying technology developed for this generic data infrastructure. We report on the progress we have made since 2018 when the previous BioStudies article was published.This paper is targeted towards scientists seeking ways to publish heterogenous datasets associated to their manuscripts, data stewards looking to set up data sharing for a new project or data type, data scientists needing large volumes of life sciences data, as well as data management system developers.BioStudies is available at https://ebi.ac.uk/biostudies
The BioSamples database (https://www.ebi.ac.uk/biosamples/) at the European Bioinformatics Institute (EMBL-EBI) is a core infrastructure resource that provides a centralized platform for the storage, curation, and dissemination of biological sample metadata. BioSamples holds sample descriptions across diverse datasets, enhancing their reusability and integrability in accordance with the FAIR (findable, accessible, interoperable, and reusable) principles. BioSamples serves as the central place to store the sample metadata for multiple repositories and archives at EMBL-EBI, including the European Nucleotide Archive, ArrayExpress, and the European Phenome-Genome Archive. In this article, we outline technical updates made to support increased sample registration and provide a better user experience. We describe four use cases whereby different communities have leveraged BioSamples to connect multi-omics data and enable the creation of public data portals and data status trackers. BioSamples is freely available, and its content is distributed under the EMBL-EBI Terms of Use available at https://www.ebi.ac.uk/about/terms-of-use. The BioSamples code is available at https://github.com/EBIBioSamples/biosamples-v4 and https://doi.org/10.5281/zenodo.17304021 and is distributed under the Apache 2.0 license.
The European Nucleotide Archive (ENA; https://www.ebi.ac.uk/ena), hosted at the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI), remains a global, open-access platform for the submission, archiving, dissemination, and reuse of nucleotide sequence data. In 2025, ENA continues to advance its mission of fostering FAIR (findable, accessible, interoperable, reusable) data principles through innovations in interoperability, scalability, and global engagement, providing infrastructure for a rapidly growing volume of data across diverse domains. This article highlights the key developments in 2025, including the progress of the technical transformation, enhanced support for large-scale biodiversity projects, and the implementation of the International Nucleotide Sequence Database Collaboration Global Participation Initiative. We also discuss infrastructure enhancements to handle exponential data growth and improve user experiences and data discovery.
In this Comment, we direct attention to initial efforts to establish a high-quality databank of atomic force microscopy (AFM) data: bioAFM-DB. We outline the state of this endeavor, its challenges, and potential courses of action.
The European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI) is one of the world's leading sources of public biomolecular data. Based at the Wellcome Genome Campus in Hinxton, UK, EMBL-EBI is one of six sites of the European Molecular Biology Laboratory, Europe's only intergovernmental life sciences organization. This overview summarizes the latest developments in services that EMBL-EBI data resources provide to scientific communities globally. All of the data resources described are freely available to access and reuse at https://www.ebi.ac.uk/services.
We introduce bia-binder (BioImage Archive Binder), an open-source, cloud-architectured, and web-based coding environment tailored to bioimage analysis that is freely accessible to all researchers. The service generates easy-to-use Jupyter Notebook coding environments hosted on EMBL-EBI's Embassy Cloud, which provides significant computational resources. The bia-binder architecture is free, open-source and publicly available for deployment. It features fast and direct access to images in the BioImage Archive, the Image Data Resource, and the BioStudies databases. We believe that this service can play a role in mitigating the current inequalities in access to scientific resources across academia. As bia-binder produces permanent links to compiled coding environments, we foresee the service to become widely-used within the community and enable exploratory research. bia-binder is built and deployed using helmsman and helm and released under the MIT licence. It can be accessed at binder.bioimagearchive.org and runs on any standard web browser.
The European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI) is one of the world's leading sources of public biomolecular data. Based at the Wellcome Genome Campus in Hinxton, UK, EMBL-EBI is one of six sites of the European Molecular Biology Laboratory, Europe's only intergovernmental life sciences organization. This overview summarizes the latest developments in services that EMBL-EBI data resources provide to scientific communities globally (https://www.ebi.ac.uk/services).
The EU-ToxRisk project (2016–2021) was a large European project working towards shifting toxicological testing away from animal tests, towards a toxicological assessment based on comprehensive mechanistic understanding of cause-consequence relationships of chemical adverse effects. More than 40 partners from scientific institutions, industry and regulators coordinated their work towards this goal in a six-year long programme. The breadth and variety of data and knowledge generated, presented a challenging data management landscape.Here, we describe our approach to data management as developed under EU-ToxRisk. The main building blocks of the data infrastructure are: 1) An easy-to-use, extensible data and metadata format; 2) A flexible system with protocols for data capture and sharing from the entire consortium; 3) A methods database for describing and reviewing data generation and processing protocols; 4) Data archiving using a sustainable resource; 5) Data transformation from the archive to the system that provides granular access; 6) Application Programming Interface (API) for access to individual data points; 7) Data exploration and analysis modules, based on a «web notebook» approach to executable data processing documentation; and 8) Knowledge portal that ties together all of the above and provides a collaboration space for information exchange across the consortium. This knowledge infrastructure is being extended and refined for the support of follow-up projects (RISK-HUNT3R, ASPIS cluster, European Open Science Cloud (2021–2026)).
The main goals and challenges for the life science communities in the Open Science framework are to increase reuse and sustainability of data resources, software tools, and workflows, especially in large-scale data-driven research and computational analyses. Here, we present key findings, procedures, effective measures and recommendations for generating and establishing sustainable life science resources based on the collaborative, cross-disciplinary work done within the EOSC-Life (European Open Science Cloud for Life Sciences) consortium. Bringing together 13 European life science research infrastructures, it has laid the foundation for an open, digital space to support biological and medical research. Using lessons learned from 27 selected projects, we describe the organisational, technical, financial and legal/ethical challenges that represent the main barriers to sustainability in the life sciences. We show how EOSC-Life provides a model for sustainable data management according to FAIR (findability, accessibility, interoperability, and reusability) principles, including solutions for sensitive- and industry-related resources, by means of cross-disciplinary training and best practices sharing. Finally, we illustrate how data harmonisation and collaborative work facilitate interoperability of tools, data, solutions and lead to a better understanding of concepts, semantics and functionalities in the life sciences.
Despite the huge impact of data resources in genomics and structural biology, until now there has been no central archive for biological data for all imaging modalities. The BioImage Archive is a new data resource at the European Bioinformatics Institute (EMBL-EBI) designed to fill this gap. In its initial development BioImage Archive accepts bioimaging data associated with publications, in any format, from any imaging modality from the molecular to the organism scale, excluding medical imaging. The BioImage Archive will ensure reproducibility of published studies that derive results from image data and reduce duplication of effort. Most importantly, the BioImage Archive will help scientists to generate new insights through reuse of existing data to answer new biological questions, and provision of training, testing and benchmarking data for development of tools for image analysis. The archive is available at https://www.ebi.ac.uk/bioimage-archive/ . Highlights The BioImage Archive is a new archival data resource at the European Bioinformatics Institute (EMBL-EBI). The BioImage Archive aims to accept all biological imaging data associated with peer-reviewed publications using approaches that probe biological structure, mechanism and dynamics, as well as other important datasets that can serve as reference examples for particular biological or technical domains. The BioImage Archive aims to encourage the use of valuable imaging data, to improve reproducibility of published results that rely on image data, and to facilitate extraction of novel biological insights from existing data and development of new image analysis methods. The BioImage Archive forms the foundation for an ecosystem of related databases, supporting those resources with storage infrastructure and indexing across databases. Across this ecosystem, the BioImage Archive already stores and provides access to over 1.5 petabytes of image data from many different imaging modalities and biological domains. Future development of the BioImage Archive will support the fast-emerging next generation file formats (NGFFs) for bioimaging data, providing access mechanisms tailored toward modern visualisation and data exploration tools, as well as unlocking the power of modern AI-based image-analysis approaches.
The data currently described was generated within the EU/FP7 HeCaToS project (Hepatic and Cardiac Toxicity Systems modeling). The project aimed to develop an in silico prediction system to contribute to drug safety assessment for humans. For this purpose, multi-omics data of repeated dose toxicity were obtained for 10 hepatotoxic and 10 cardiotoxic compounds. Most data were gained from in vitro experiments in which 3D microtissues (either hepatic or cardiac) were exposed to a therapeutic (physiologically relevant concentrations calculated through PBPK-modeling) or a toxic dosing profile (IC20 after 7 days). Exposures lasted for 14 days and samples were obtained at 7 time points (therapeutic doses: 2-8-24-72-168-240-336 h; toxic doses 0-2-8-24-72-168-240 h). Transcriptomics (RNA sequencing & microRNA sequencing), proteomics (LC-MS), epigenomics (MeDIP sequencing) and metabolomics (LC-MS & NMR) data were obtained from these samples. Furthermore, functional endpoints (ATP content, Caspase3/7 and O2 consumption) were measured in exposed microtissues. Additionally, multi-omics data from human biopsies from patients are available. This data is now being released to the scientific community through the BioStudies data repository (https://www.ebi.ac.uk/biostudies/).
A gap in community practice on data citation that emerged during the AGU fall meeting 2020 Data FAIR Town Hall, “Why Is Citing Data Still Hard?” with the goal of addressing the use case of citing a large number of datasets such that credit for individual datasets is assigned properly. The discussion included the concept of a “Data Collection” and the infrastructure and guidance still needed to fully implement the capability so it is easier for researchers to use and receive credit when their data are cited in this manner. Such collections of data may contain thousands to millions of elements with a citation needing to include subsets of elements potentially from multiple collections. Such citations will be crucial to enable reproducible research and credit to data and digital object creators. To address this gap, the data citation community of practice formed including members from data centres, research journals, informatics research communities, and data citation infrastructure. The community has the goal of recommending an approach that is realistic for researchers to use and for each stakeholder to implement that leverages existing infrastructure. To achieve data citation of these subsets of large data collections the concept of a “reliquary” is introduced. In this context the reliquary is a container of persistent identifiers (PIDs) or references defining the objects used in a research study. This can include any number of elements. The reliquary can then be cited as a single entity in academic publications. The reliquary concept will enable data citation use cases such as the citation of elements within a data collection that are formed from numerous underlying datasets that have their own PIDs, unambiguous citation of data used in IPCC Assessment Reports, and citing the subsets of collections of research data that contain millions of elements. The discussions over the course of 2021 have developed a theoretical concept, at the time of writing formal use cases and initial applications are being defined. The recommendation developed by this effort will be available for review and comment by communities such as ESIP and RDA. All are welcome.
ArrayExpress (https://www.ebi.ac.uk/arrayexpress) is an archive of functional genomics data at EMBL-EBI, established in 2002, initially as an archive for publication-related microarray data and was later extended to accept sequencing-based data. Over the last decade an increasing share of biological experiments involve multiple technologies assaying different biological modalities, such as epigenetics, and RNA and protein expression, and thus the BioStudies database (https:// www.ebi.ac.uk/biostudies) was established to deal with such multimodal data. Its central concept is a study, which typically is associated with a publication. BioStudies stores metadata describing the study, provides links to the relevant databases, such as European Nucleotide Archive (ENA), as well as hosts the types of data for which specialized databases do not exist. With BioStudies now fully functional, we are able to further harmonize the archival data infrastructure at EMBL-EBI, and ArrayExpress is beingmigrated toBioStudies. In future, all functional genomics data will be archived at BioStudies. The process will be seamless for the users, who will continue to submit data using the online tool Annotare and will be able to query and download data largely in the same manner as before. Nevertheless, some technical aspects, particularly programmatic access, will change. This update guides the users through these changes.
This protocol illustrates the steps necessary to deposit correlated 3D cryo-imaging data from cryo-structured illumination microscopy and cryo-soft X-ray tomography with the BioStudies and EMPIAR deposition databases of the European Bioinformatics Institute. There is currently a real need for a robust method of data deposition to ensure unhindered access to and independent validation of correlative light and X-ray microscopy data to allow use in further comparative studies, educational activities, and data mining. For complete details on the use and execution of this protocol, please refer to Kounatidis et al. (2020).
Imaging technologies are used throughout the life and biomedical sciences to understand mechanisms in biology and diagnosis and therapy in animal and human medicine. We present criteria for globally applicable guidelines for open image data tools and resources for the rapidly developing fields of biological and biomedical imaging.
Bioimaging data have significant potential for reuse, but unlocking this potential requires systematic archiving of data and metadata in public databases. We propose draft metadata guidelines to begin addressing the needs of diverse communities within light and electron microscopy. We hope this publication and the proposed Recommended Metadata for Biological Images (REMBI) will stimulate discussions about their implementation and future extension.
Uncovering cellular responses from heterogeneous genomic data is crucial for molecular medicine in particular for drug safety. This can be realized by integrating the molecular activities in networks of interacting proteins. As proof-of-concept we challenge network modeling with time-resolved proteome, transcriptome and methylome measurements in iPSC-derived human 3D cardiac microtissues to elucidate adverse mechanisms of anthracycline cardiotoxicity measured with four different drugs (doxorubicin, epirubicin, idarubicin and daunorubicin). Dynamic molecular analysis at in vivo drug exposure levels reveal a network of 175 disease-associated proteins and identify common modules of anthracycline cardiotoxicity in vitro, related to mitochondrial and sarcomere function as well as remodeling of extracellular matrix. These in vitro-identified modules are transferable and are evaluated with biopsies of cardiomyopathy patients. This to our knowledge most comprehensive study on anthracycline cardiotoxicity demonstrates a reproducible workflow for molecular medicine and serves as a template for detecting adverse drug responses from complex omics data.
Hazard assessment, based on new approach methods (NAM), requires the use of batteries of assays, where individual tests may be contributed by different laboratories. A unified strategy for such collaborative testing is presented. It details all procedures required to allow test information to be usable for integrated hazard assessment, strategic project decisions and/or for regulatory purposes. The EU-ToxRisk project developed a strategy to provide regulatorily valid data, and exemplified this using a panel of > 20 assays (with > 50 individual endpoints), each exposed to 19 well-known test compounds (e.g. rotenone, colchicine, mercury, paracetamol, rifampicine, paraquat, taxol). Examples of strategy implementation are provided for all aspects required to ensure data validity: (i) documentation of test methods in a publicly accessible database; (ii) deposition of standard operating procedures (SOP) at the European Union DB-ALM repository; (iii) test readiness scoring accoding to defined criteria; (iv) disclosure of the pipeline for data processing; (v) link of uncertainty measures and metadata to the data; (vi) definition of test chemicals, their handling and their behavior in test media; (vii) specification of the test purpose and overall evaluation plans. Moreover, data generation was exemplified by providing results from 25 reporter assays. A complete evaluation of the entire test battery will be described elsewhere. A major learning from the retrospective analysis of this large testing project was the need for thorough definitions of the above strategy aspects, ideally in form of a study pre-registration, to allow adequate interpretation of the data and to ensure overall scientific/toxicological validity.