The 21st century drastically transformed the way scientific research is carried out. All stages, from planning to results interpretation, are heavily dependant on specialised research software, the quality of which becomes an important issue during project implementation. Establishing generally-accepted software quality metrics is an important step in software adoption, continuous improvement, and sustainability. EVERSE is an EU-funded project focusing on promoting research software as a first-class citizen in the scientific research community, providing a framework for the quality assessment and evaluation of such software. The project provides a set of Research Software Quality Indicators covering different dimensions, which goes beyond the FAIR principles applied to Research Software. In this context, the project consists of the assembling ‘resqui’ workflow, which facilitates the evaluation of a list of indicators, and the ‘dashVERSE’ web-interface which visualizes the results. OpenEBench, the ELIXIR platform supporting community-driven scientific benchmarking activities and the technical monitoring of research software, has implemented the FAIRsoft indicators. FAIRsoft high- and low-level indicators focus on automatically measuring how FAIR a given software is. Indicators represent a community-effort in translating the original FAIR principles from data to software and then proposing concrete approaches to measure such principles.. Here we present the integration of the FAIRsoft indicators, as implemented in the OpenEBench Software Observatory, into the EVERSE framework. Such integration effort aims to reduce duplicated efforts, leveraging current implementations, and facilitating specialized knowledge exchange across scientific communities with a common goal: making research software of high quality and, therefore, contributing towards its long-term adoption and sustainability.
The Spanish National Bioinformatics Institute (INB), founded in 2003 as a distributed network, is the ELIXIR Node in Spain and has two objectives: 1) deepen its involvement and leadership within ELIXIR and broaden the resources provided as part of ELIXIR infrastructure to the Life Sciences community; and 2) increase its impact within the Spanish National Health System . INB/ELIXIR-ES continues to strengthen its technological capabilities in federated data infrastructures, interoperability, and FAIR data management within ELIXIR. The Node is actively involved in several ELIXIR-driven projects and commissioned services (CoS) within the 2024-28 Work Programme. From the ELIXIR perspective, the Service Delivery Plan (SDP) maintains 40 resources offered by 24 groups belonging to 12 institutions. Regarding national activities, the INB/ELIXIR-ES leads the Translational Bioinformatics Network (TransBioNet) and serves as proxy between IMPaCT-Data activities, the Data Science pillar of the Spanish National Infrastructure for Precision Medicine, and European efforts. These activities align with major European data projects such as the Genomic Data Infrastructure (GDI), EUCAIM, and the Federated European Genome-phenome Archive (FEGA), while implementing Global Alliance for Genomics and Health (GA4GH) standards in its technological developments. The Node strongly engages within the 2024-28 Work Programme: TechnologyTier: co-leadership of Data, Tools and Training Platforms, with contributions across all Platforms. During this period, the co-led ELIXIR Beacon Network Infrastructure Service secured funding for this service, strengthening federated data discovery capabilities. ScienceTier : co-leadership of CMR and HDTR Science priority areas; co-leadership of Rare Diseases, FHD, Cancer Data and Biodiversity Communities; and Pathogens Data, and RNA Data Focus Group; leading and participating in several CoS in the HDTR and CMR areas. PeopleTier : co-leadership of the ELEAD2.0 leadership programme, a CoS built on the experiences of Bioinfo4Women, and active role in the PeoplePulse CoS. Active role in the NodeTier NSCS, together with 4 Platforms, 13 Communities, and 8 Focus Groups. Regarding the INB/ELIXIR-ES portfolio, EGA, an ELIXIR CDR, is co-developed and maintained by CRG and EMBL-EBI with BSC’s infrastructure support, with the current focus on its extension through Federated EGA. Canada joined the Federated EGA, marking the first major expansion beyond Europe and reinforcing its global dimension. Additionally, four resources are recognised as ELIXIR RIRs: 3DBIONOTES-API, FAIRtracks, FAIRCookbook and OpenEBench. Various ELIXIR Communities have adopted OpenEBench as their community-driven benchmarking platform. The INB/ELIXIR-ES continued its training activities and organised key meetings within ELIXIR. It gathered its national community in the XV Symposium on Bioinformatics (JBI2025) jointly organised with ELIXIR-PT and INSTRUCT-ES. Other ELIXIR events organised were the ELIXIR 3DBioinfo Community Annual General Meeting with the 3D-SIG Community, and the Biodiversity and Microbiome Community meetings. https://inb-elixir.es https://inb-elixir.es/resources
ABSTRACT IMPaCT-Data, the Data Science pillar of the Spanish National Infrastructure for Precision Medicine coordinated by the Barcelona Supercomputing Center (BSC), has assembled and deployed a national-scale infrastructure for clinical, genomics, and imaging workloads across major Spanish biomedical research institutions. The main technical challenge is enabling large-scale, compute-intensive workflows while complying with strict data-governance constraints, heterogeneous local infrastructures, and non-uniform compute and storage capabilities. The first layer of the platform is based on a distributed execution model in which workflows are centrally dispatched while computation is performed locally at participating institutions. A central Galaxy server instance hosted at BSC provides workflow management, provenance tracking, user access, and operational monitoring. Authentication and authorization are implemented using OpenID Connect (OIDC), with institutional identity providers federated through a central Keycloak service, allowing the preservation of local identity management policies. Galaxy Pulsar handles job dispatching and, based on asynchronous messaging (RabbitMQ/AMQPS), distributes workloads to remote federated nodes and delegates processing to site-local backends (e.g. container engines, HPC clusters). Analysis tools and reference datasets are distributed and synchronized through CVMFS, reducing the operational burden associated with maintaining synchronized analysis environments across a federated infrastructure. For highly sensitive datasets, a second execution network is being deployed, enabled and orchestrated using the Federated Execution Manager (FEM), and exposed to the researcher through a tailored virtual research environment powered by openVRE (future safeVRE). Under this architecture, participating institutions will retain full control not only over data access but also over the execution environment and approved analysis tools, making it particularly suitable for secure processing environments (SPEs). In this context, FEM allows institutions to enforce local governance policies, validate containers and workflows before execution, and control how intermediate results are generated and shared across sites. This design also supports regulated biomedical scenarios in which datasets remain in read-only mode and computation must be executed under strict auditing and traceability requirements. In addition, FEM natively supports more complex federation patterns, including federated analysis (calculator-aggregator schemes) and federated learning workflows (client-server schemes). The resulting platform provides a practical technical blueprint for biomedical data processing aligned with ELIXIR principles and other major efforts and initiatives, such as Genomics Alliance for Genomics and Health (GA4GH), EUCAIM, and GDI, among others: distributed execution close to data, standards-aligned interoperability, and reproducible workflows across a clinically governed federated network.
The rapid growth of genomic and biological data, combined with the complexity of biomedical workflows, has created challenges in achieving interoperability across High-Performance Computing (HPC), Cloud Computing facilities, and secure data platforms. Seamless integration, robust security, and efficient resource utilization are crucial for enabling the optimal use of health data. In this poster, we present OpenVRE (Open Virtual Research Environment), an open-source, cloud-based platform designed to address these challenges. OpenVRE bridges the gap between HPC resources, sensitive data infrastructures, and analytical tools and workflows, providing a flexible environment for researchers.
In the Big Data era, a change of paradigm in the use of molecular dynamics is required. Trajectories should be stored under FAIR (findable, accessible, interoperable and reusable) requirements to favor its reuse by the community under an open science paradigm.
Artificial intelligence (AI) is a powerful technology with the potential to disrupt cancer detection, diagnosis and treatment. However, the development of new AI algorithms requires access to large and complex real-world datasets. Although such datasets are constantly being generated, access to them is limited by data fragmentation across numerous repositories and sites, heterogeneity, lack of annotations, and potential privacy issues. The European Cancer Imaging Initiative is a flagship of Europe's Beating Cancer Plan, aiming to unlock the power of AI for cancer patients, clinicians, and researchers by establishing a federated European infrastructure for cancer images through the EU-funded EUropean Federation for CAncer IMages (EUCAIM) project. This infrastructure, called Cancer Image Europe, builds on the AI for Health Imaging network (AI4HI), established European Research Infrastructures (Euro-BioImaging, BBMRI-ERIC, EATRIS, ECRIN, and ELIXIR), and numerous related partners providing access to research tools, images, and related clinical, pathology and molecular data. The infrastructure targets clinicians, researchers, and innovators by providing the means to develop and validate data-intensive AI-based and other IT-enabled clinical decision-making systems supporting precision medicine. Common data models, including a linking hyperontology, quality standards, compliance with the FAIR (Findability, Accessibility, Interoperability and Reusability) principles, data annotation, curation and anonymization services are provided to ensure data quality and interoperability, consistency and privacy. In summer 2024, the EUCAIM project released the first prototype of an EU-wide infrastructure, with a comprehensive dashboard integrating applications for dataset discovery, federated search, data access request, metadata harvesting, annotation, secure processing environments and federated processing. CRITICAL RELEVANCE STATEMENT: EUCAIM's federated infrastructure for cancer image data advances medical research and related AI development in Europe. It addresses the current fragmentation and heterogeneity of data repositories is legally compliant, and facilitates collaboration among clinicians, researchers, and innovators. KEY POINTS: AI solutions to advance cancer care rely on large, high-quality real-world datasets. EUCAIM's federated infrastructure for cancer image data empowers cancer research in Europe. It provides access to research tools, images, and related clinical, pathology and molecular data.
The second edition of the IMPaCT-Data program ( IMPaCT-Data 2 ) focuses on assembling and deploying a federated infrastructure to standardize, integrate and analyze Clinical, Omics (primarily Genomics) and Medical Imaging Data. This infrastructure will enable the use of biomedical data from ISCIII IMPaCT-Cohort and IMPaCT-PMP projects, using them as use cases to demonstrate the platform's utility, and having the long-term aim of becoming a stable research infrastructure supporting biomedical research activities in Spain. The platform and available software components will provide an advanced data analysis environment linked to a hybrid computational infrastructure , combining cloud and HPC resources. Alignment with major European initiatives, such as the European Genomic Data Infrastructure (GDI) and the European Federation of Cancer Images (EUCAIM), is a key aspect of the project to guarantee technical alignment. In this context, Federated EGA (FEGA) will be essential to articulate the management of the federated data space. The platform will adopt transversal elements from the European Open Science Cloud (EOSC) and implement Global Alliance for Genomics and Health (GA4GH) standards for interoperability. This project represents the national implementation of projects like GDI, being relevant because IMPaCT-Cohort program leads the Spanish sequencing effort in the Genome of Europe (GoE) context. The project will adopt good practices for research software development as seen in ELIXIR experiences, ensuring long-term sustainability. Specific goals include: Establishing a website as an access point to the data analysis environment, featuring a catalogue of tools, resources, guides, and recommendations. Defining requirements for deploying analysis tools, including advanced statistical analysis and Artificial Intelligence. Setting the architecture of the federated network associated with the IMPaCT platform.
OpenEBench (https://openebench.bsc.es/) is the ELIXIR open-data collaborative platform to support community-driven scientific benchmarking. As part of the ELIXIR Tools Platform, OpenEBench is dedicated to advancing scientific benchmarking and technical monitoring practices of research software in Life Sciences. The open nature of the platform facilitates its use and adoption by different communities within ELIXIR and beyond. Currently, there are 12 active communities within the platform with the expectation to reach 20 in the near future, thanks partially to the engagement with different European projects. OpenEBench has captured an overall diversity of scientific benchmarking needs since its creation. This diversity is also present among the new communities, i.e. for long-term storage and results display (CAID), benchmarking workflow execution (LRGASP), periodic benchmarking events (QfO), or continuous benchmarking (CAMEO). Recently, there has been a noticeable increase of Artificial Intelligence (AI) software and models benchmarks and OpenEBench has swiftly responded to these demands by extending its backend and capabilities. Indeed, OpenEBench is part of future deployments across different projects, including cancer imaging benchmark for enhanced AI in oncology (EuCanImage), adoption and deployment of research software best practices (EOSC-EVERSE and ELIXIR STEERS), sex and gender biases in AI models for health applications (BAIHA), and benchmarking of multilingual natural language processing to standardise the structuring of cardiology reports across European regions (DataTools4Heart). In addition to this diversity, OpenEBench introduces a new feature called "Project Spaces" designed to facilitate collaboration by providing a web space where projects and communities can present their efforts, provide guidelines and share relevant information with anyone interested in engaging with them.
MOTIVATION:Software plays a crucial and growing role in research. Unfortunately, the computational component in Life Sciences research is often challenging to reproduce and verify. It could be undocumented, opaque, contain unknown errors that affect the outcome, or be directly unavailable and impossible to use for others. These issues are detrimental to the overall quality of scientific research. One step to address this problem is the formulation of principles that research software in the domain should meet to ensure its quality and sustainability, resembling the FAIR (findable, accessible, interoperable, and reusable) data principles. RESULTS:We present here a comprehensive series of quantitative indicators based on a pragmatic interpretation of the FAIR Principles and their implementation on OpenEBench, ELIXIR's open platform providing both support for scientific benchmarking and an active observatory of quality-related features for Life Sciences research software. The results serve to understand the current practices around research software quality-related features and provide objective indications for improving them. AVAILABILITY AND IMPLEMENTATION:Software metadata, from 11 different sources, collected, integrated, and analysed in the context of this manuscript are available at https://doi.org/10.5281/zenodo.7311067. Code used for software metadata retrieval and processing is available in the following repository: https://gitlab.bsc.es/inb/elixir/software-observatory/FAIRsoft_ETL.
Interactive Jupyter Notebooks in combination with Conda environments can be used to generate FAIR (Findable, Accessible, Interoperable and Reusable/Reproducible) biomolecular simulation workflows. The interactive programming code accompanied by documentation and the possibility to inspect intermediate results with versatile graphical charts and data visualization is very helpful, especially in iterative processes, where parameters might be adjusted to a particular system of interest. This work presents a collection of FAIR notebooks covering various areas of the biomolecular simulation field, such as molecular dynamics (MD), protein-ligand docking, molecular checking/modeling, molecular interactions, and free energy perturbations. Workflows can be launched with myBinder or easily installed in a local system. The collection of notebooks aims to provide a compilation of demonstration workflows, and it is continuously updated and expanded with examples using new methodologies and tools.
OpenEBench, the ELIXIR platform supporting scientific community benchmarking activities and the technical monitoring of research software in Life Sciences, has strived since its inception in 2017 to serve a broad audience of end-users. The 2024 update brings forth a variety of new features with the primary goal of enriching user engagement and fostering collaboration. The widgets deployed in this update represent an innovative approach to data visualization. These encapsulated packages of code empower the visualization of results effortlessly, grouping all the functionalities in a simple and visual layout. Additionally, while initially developed for OpenEBench, these widgets are designed in such a way that can be easily deployed and integrated in third-party websites. Moreover, the introduction of the OpenEBench intranet marks a significant shift towards strengthening internal collaboration and communication within users. In the 2024 release, it is possible to create and manage communities and events directly within the platform, automatizing those operations and reducing the interactions with the OpenEBench helpdesk. This initial release will capture users’ feedback, which will be used to improve this newly deployed functionality and as the basis to streamline other operations within the platform. The primary objective is empowering scientific communities to manage their own content while reducing error-prone manual operations. As these enhancements are adopted, we expect that OpenEBench will continue to serve as an essential resource for scientific communities benchmarking activities within and beyond ELIXIR and as a framework for monitoring the research software quality across the Life Sciences community.
RegulonDB is a database that contains the most comprehensive corpus of knowledge of the regulation of transcription initiation of Escherichia coli K-12, including data from both classical molecular biology and high-throughput methodologies. Here, we describe biological advances since our last NAR paper of 2019. We explain the changes to satisfy FAIR requirements. We also present a full reconstruction of the RegulonDB computational infrastructure, which has significantly improved data storage, retrieval and accessibility and thus supports a more intuitive and user-friendly experience. The integration of graphical tools provides clear visual representations of genetic regulation data, facilitating data interpretation and knowledge integration. RegulonDB version 12.0 can be accessed at https://regulondb.ccg.unam.mx.
The European Genomic Data Infrastructure (GDI) project aims to establish a federated and secure platform facilitating access and analysis of genomic, phenotypic, and clinical data across Europe. This initiative operates through a network, where each node represents an European country responsible for implementing the required software stack to enable data sharing and processing within this infrastructure. The IMPaCT-Data Biomedical Cloud is the Spanish national implementation of the European Genomic Data Infrastructure (GDI). It is being established to provide a scalable and flexible analysis environment, enabling the integration, management and analysis of clinical, genomic and medical imaging data available within the Spanish National Precision Medicine Infrastructure associated with Science and Technology (IMPaCT).
Molecular dynamics (MD) simulations are keeping computers busy around the world, generating a huge amount of data that is typically not open to the scientific community. Pioneering efforts to ensure the safety and reusability of MD data have been based on the use of simple databases providing a limited set of standard analyses on single-short trajectories. Despite their value, these databases do not offer a true solution for the current community of MD users, who want a flexible analysis pipeline and the possibility to address huge non-Markovian ensembles of large systems. Here we present a new paradigm for MD databases, resilient to large systems and long trajectories, and designed to be compatible with modern MD simulations. The data are offered to the community through a web-based graphical user interface (GUI), implemented with state-of-the-art technology, which incorporates system-specific analysis designed by the trajectory providers. A REST API and associated Jupyter Notebooks are integrated into the platform, allowing fully customized meta-analysis by final users. The new technology is illustrated using a collection of trajectories obtained by the community in the context of the effort to fight the COVID-19 pandemic. The server is accessible at https://bioexcel-cv19.bsc.es/#/. It is free and open to all users and there are no login requirements. It is also integrated into the simulations section of the BioExcel-MolSSI COVID-19 Molecular Structure and Therapeutics Hub: https://covid.molssi.org/simulations/ and is part of the MDDB effort (https://mddbr.eu).
The Spanish Precision Medicine Infrastructure associated with Science and Technology (IMPaCT), aims to lay the foundations for impulsing precision medicine within the Spanish National Health System. IMPaCT revolves around three main pillars: Predictive Medicine, Data Science and Genomic Medicine. As part of the Data Science program, the IMPaCT-Data Biomedical Cloud is being established for providing a scalable and flexible analysis environment, enabling the integration, management and analysis of clinical, genomic and medical imaging data available within IMPaCT. The IMPaCT-Data Biomedical Cloud presents a federated computing environment that provides a range of platforms, including UseGalaxy.* and VRE, with an authentication system based on OIDC (Keycloak) that incorporates Life Sciences Log-in capabilities and is supplemented by an authorization system based on GA4GH Passports. IMPaCT-Data Biomedical Cloud main objective is to provide researchers with a robust and efficient computing environment, enabling them to collaborate, analyze data, and share results in a secure and scalable manner and is advancing to finally provide a complete computing environment with included streamlined authentication, fine-grained authorization controls, and flexible platform integration. As part of its continuos work, two videos have been recently published showing the capabilities of the current prototype-based implementation illustrating the analysis of distributed data and the use of GA4GH-based Passports and Visas mechanisms for accessing data across different systems including the use of different identity providers across this federated infrastructure.
Exascale computing has been a dream for ages and is close to becoming a reality that will impact how molecular simulations are being performed, as well as the quantity and quality of the information derived for them. We review how the biomolecular simulations field is anticipating these new architectures, making emphasis on recent work from groups in the BioExcel Center of Excellence for High Performance Computing. We exemplified the power of these simulation strategies with the work done by the HPC simulation community to fight Covid-19 pandemics. This article is categorized under: Data Science > Computer Algorithms and Programming Data Science > Databases and Expert Systems Molecular and Statistical Mechanics > Molecular Dynamics and Monte-Carlo Methods
Mutations in the kinase domain of the epidermal growth factor receptor (EGFR) can be drivers of cancer and also trigger drug resistance in patients receiving chemotherapy treatment based on kinase inhibitors. A priori knowledge of the impact of EGFR variants on drug sensitivity would help to optimize chemotherapy and design new drugs that are effective against resistant variants before they emerge in clinical trials. To this end, we explored a variety of in silico methods, from sequence-based to "state-of-the-art" atomistic simulations. We did not find any sequence signal that can provide clues on when a drug-related mutation appears or the impact of such mutations on drug activity. Low-level simulation methods provide limited qualitative information on regions where mutations are likely to cause alterations in drug activity, and they can predict around 70% of the impact of mutations on drug efficiency. High-level simulations based on nonequilibrium alchemical free energy calculations show predictive power. The integration of these "state-of-the-art" methods into a workflow implementing an interface for parallel distribution of the calculations allows its automatic and high-throughput use, even for researchers with moderate experience in molecular simulations.
Artificial intelligence (AI) is transforming the field of medical imaging and has the potential to bring medicine from the era of ‘sick-care’ to the era of healthcare and prevention. The development of AI requires access to large, complete, and harmonized real-world datasets, representative of the population, and disease diversity. However, to date, efforts are fragmented, based on single–institution, size-limited, and annotation-limited datasets. Available public datasets ( e.g. , The Cancer Imaging Archive, TCIA, USA) are limited in scope, making model generalizability really difficult. In this direction, five European Union projects are currently working on the development of big data infrastructures that will enable European, ethically and General Data Protection Regulation-compliant, quality-controlled, cancer-related, medical imaging platforms, in which both large-scale data and AI algorithms will coexist. The vision is to create sustainable AI cloud-based platforms for the development, implementation, verification, and validation of trustable, usable, and reliable AI models for addressing specific unmet needs regarding cancer care provision. In this paper, we present an overview of the development efforts highlighting challenges and approaches selected providing valuable feedback to future attempts in the area. Key points • Artificial intelligence models for health imaging require access to large amounts of harmonized imaging data and metadata. • Main infrastructures adopted either collect centrally anonymized data or enable access to pseudonymized distributed data. • Developing a common data model for storing all relevant information is a challenge. • Trust of data providers in data sharing initiatives is essential. • An online European Union meta-tool-repository is a necessity minimizing effort duplication for the various projects in the area.
Biomaterials research output has experienced an exponential increase over the last three decades. The majority of research is published in the form of scientific articles and is therefore available as unstructured text, making it a challenging input for computational processing. Computational tools are becoming essential to overcome this information overload. Among them, text mining systems present an attractive option for the automated extraction of information from text documents into structured datasets. This work presents the first automated system for biomaterial related information extraction from the National Library of Medicine's premier bibliographic database (MEDLINE) research abstracts into a searchable database. The system is a text mining pipeline that periodically retrieves abstracts from PubMed and identifies research and clinical studies of biomaterials. Thereafter, the pipeline identifies sixteen concept types of interest in the abstract using the Biomaterials Annotator, a tool for biomaterials Named Entity Recognition (NER). These concepts of interest, along with the abstract and relevant metadata are then deposited in DEBBIE, the Database of Experimental Biomaterials and their Biological Effect. DEBBIE is accessible through a web application that provides keyword searches and displays results in an intuitive and meaningful manner, aiming to facilitate an efficient mapping and organization of biomaterials information.