In alignment with the European Health Data Space (EHDS), Spain's IMPaCT - Precision Medicine Infrastructure associated with Science and Technology - aims to establish a trusted research environment (TRE) for secure and FAIR data sharing and analysis. It is structured into three pillars: Predictive Medicine (IMPaCT-Cohort), Genomic Medicine (IMPaCT-Genomics), and Data Science (IMPaCT-Data). IMPaCT-Data leads the development of the IMPaCT Digital Platform (IDP), integrating clinical, genomic, and imaging data to support the national IMPaCT-Cohort and Personalised Medicine Projects (IMPaCT-PMPs). Its Reference Implementation defines the architecture across a federated model.
The growing volume of biomedical data requires robust frameworks for standardized representation and privacy-preserving discoverability. The Global Alliance for Genomics and Health (GA4GH) Beacon protocol enables secure, federated queries across distributed datasets, while the Observational Medical Outcomes Partnership (OMOP) Common Data Model (CDM) provides a standardized structure for harmonizing clinical data. Integrating these two approaches facilitates scalable and interoperable discovery across cohorts and institutions. Within the DATOS-CAT project, we applied this integration to the GCAT (Genomes for Life) cohort, deploying BeaconOMOP (v0.9) to enable discovery of clinical data stored in an OMOP-formatted PostgreSQL database.This implementation ensured alignment with the FAIR data principles —making data Findable, Accessible, Interoperable, and Reusable. In addition, it enhances the visibility of Catalan population-based cohorts and establishes procedures scalable to other initiatives, such as IMPaCT-Cohorte. Building on this foundation, we developed BeaconOMOP v1.0 to broaden database compatibility. By incorporating SQLAlchemy, the system started to become database-agnostic, supporting PostgreSQL and MySQL. This advancement improved flexibility and reduced query response times by approximately 50%. Additionally, we introduced a fully asynchronous version of the API, enabling efficient handling of simultaneous queries and enhancing overall scalability. In parallel, through our contribution to the UNICAS project, a national pilot initiative creating a pediatric network for personalized medicine in rare diseases across Spanish hospitals, BeaconOMOP, though a Beacon Network, was adopted as the core federated discovery platform. For this use case, we extended v1.0 to integrate Denodo as a supported data source, reflecting its widespread adoption in hospital environments (not asynchronously for now). This deployment supports not only data discovery but also patient resolution across institutions, a critical step toward federated genomic and phenotypic analysis. These continuous enhancements have made BeaconOMOP more versatile, efficient, and robust. The developments support ongoing efforts within IMPaCT project and align with broader initiatives such as the GA4GH Driver Project and the European Genomic Data Infrastructure (GDI), positioning the platform to meet the evolving needs of distributed biomedical research.
Artificial intelligence (AI) has recently seen transformative breakthroughs in the life sciences, expanding possibilities for researchers to interpret biological information at an unprecedented capacity, with novel applications and advances being made almost daily. In order to maximise return on the growing investments in AI-based life science research and accelerate this progress, it has become urgent to address the exacerbation of long-standing research challenges arising from the rapid adoption of AI methods. We review the increased erosion of trust in AI research outputs, driven by the issues of poor reusability and reproducibility, and highlight their consequent impact on environmental sustainability. Furthermore, we discuss the fragmented components of the AI ecosystem and lack of guiding pathways to best support Open and Sustainable AI (OSAI) model development. In response, this perspective introduces a practical set of OSAI recommendations directly mapped to over 300 components of the AI ecosystem. Our work connects researchers with relevant AI resources, facilitating the implementation of sustainable, reusable and transparent AI. Built upon life science community consensus and aligned to existing efforts, the outputs of this perspective are designed to aid the future development of policy and structured pathways for guiding AI implementation.
The 21st century drastically transformed the way scientific research is carried out. All stages, from planning to results interpretation, are heavily dependant on specialised research software, the quality of which becomes an important issue during project implementation. Establishing generally-accepted software quality metrics is an important step in software adoption, continuous improvement, and sustainability. EVERSE is an EU-funded project focusing on promoting research software as a first-class citizen in the scientific research community, providing a framework for the quality assessment and evaluation of such software. The project provides a set of Research Software Quality Indicators covering different dimensions, which goes beyond the FAIR principles applied to Research Software. In this context, the project consists of the assembling ‘resqui’ workflow, which facilitates the evaluation of a list of indicators, and the ‘dashVERSE’ web-interface which visualizes the results. OpenEBench, the ELIXIR platform supporting community-driven scientific benchmarking activities and the technical monitoring of research software, has implemented the FAIRsoft indicators. FAIRsoft high- and low-level indicators focus on automatically measuring how FAIR a given software is. Indicators represent a community-effort in translating the original FAIR principles from data to software and then proposing concrete approaches to measure such principles.. Here we present the integration of the FAIRsoft indicators, as implemented in the OpenEBench Software Observatory, into the EVERSE framework. Such integration effort aims to reduce duplicated efforts, leveraging current implementations, and facilitating specialized knowledge exchange across scientific communities with a common goal: making research software of high quality and, therefore, contributing towards its long-term adoption and sustainability.
The network of the national COVID-19 Data Portals was developed and linked to the COVID-19 Data Portal (https://www.covid19dataportal.org/) in response to the need for rapid data sharing and analysis during the 2020-2022 SARS-CoV-2 pandemic. Built on open-source code developed by the Swedish COVID-19 Data Portal (now the Swedish Pathogens Portal, www.pathogens.se) the network included 12 national portals addressing demand for local open data sharing and access, across data types and resources. It provides a robust case study of national initiatives for FAIR (Findable, Accessible, Interoperable and Reusable) resources and a foundation for future pandemic preparedness across pathogens globally. In this paper we outline the structure of the origins of the network of National COVID-19 Datal Portals, the technical aspects and code originating from the Swedish Portal and provide an overview of the services and tools offered by each Portal. The paper showcases the process and operation of four Portals: Sweden, Poland, Spain, Norway and The Netherlands. It considers useful lessons and approaches for future pandemic preparedness which enable researchers to easily identify and obtain the key data from resources on a national and international level.
The network of the national COVID-19 Data Portals was developed and linked to the COVID-19 Data Portal (https://www.covid19dataportal.org/)inresponsetothe need for rapid data sharing and analysis during the 2020–2022 SARS-CoV-2 pandemic. Built on open-source code developed by the Swedish COVID-19 Data Portal (now the Swedish Pathogens Portal, www.pathogens.se) the network included 12 national portals addressing demand for local open data sharing and access, across data types and resources. It provides a robust case study of national initiatives for FAIR (Findable, Accessible, Interoperable and Reusable) resources and a foundation for future pandemic preparedness across pathogens globally. In this paper we outline the structure of the origins of the network of National COVID-19 Datal Portals, the technical aspects and code originating from the Swedish Portal and provide an overview of the services and tools offered by each Portal. The paper showcases the process and operation of four Portals: Sweden, Poland, Spain, Norway and The Netherlands. In this study, we observe that pandemic response greatly benefits from an established infrastructure that can be quickly mobilised, developed and extended. Collaborations and preparation built on solid foundations over several years, supported by investment in the form of national and international research grants, is key for sustainability, continuation and readiness to deploy such efforts.
The Spanish National Bioinformatics Institute (INB), founded in 2003 as a distributed network, is the ELIXIR Node in Spain and has two objectives: 1) deepen its involvement and leadership within ELIXIR and broaden the resources provided as part of ELIXIR infrastructure to the Life Sciences community; and 2) increase its impact within the Spanish National Health System . INB/ELIXIR-ES continues to strengthen its technological capabilities in federated data infrastructures, interoperability, and FAIR data management within ELIXIR. The Node is actively involved in several ELIXIR-driven projects and commissioned services (CoS) within the 2024-28 Work Programme. From the ELIXIR perspective, the Service Delivery Plan (SDP) maintains 40 resources offered by 24 groups belonging to 12 institutions. Regarding national activities, the INB/ELIXIR-ES leads the Translational Bioinformatics Network (TransBioNet) and serves as proxy between IMPaCT-Data activities, the Data Science pillar of the Spanish National Infrastructure for Precision Medicine, and European efforts. These activities align with major European data projects such as the Genomic Data Infrastructure (GDI), EUCAIM, and the Federated European Genome-phenome Archive (FEGA), while implementing Global Alliance for Genomics and Health (GA4GH) standards in its technological developments. The Node strongly engages within the 2024-28 Work Programme: TechnologyTier: co-leadership of Data, Tools and Training Platforms, with contributions across all Platforms. During this period, the co-led ELIXIR Beacon Network Infrastructure Service secured funding for this service, strengthening federated data discovery capabilities. ScienceTier : co-leadership of CMR and HDTR Science priority areas; co-leadership of Rare Diseases, FHD, Cancer Data and Biodiversity Communities; and Pathogens Data, and RNA Data Focus Group; leading and participating in several CoS in the HDTR and CMR areas. PeopleTier : co-leadership of the ELEAD2.0 leadership programme, a CoS built on the experiences of Bioinfo4Women, and active role in the PeoplePulse CoS. Active role in the NodeTier NSCS, together with 4 Platforms, 13 Communities, and 8 Focus Groups. Regarding the INB/ELIXIR-ES portfolio, EGA, an ELIXIR CDR, is co-developed and maintained by CRG and EMBL-EBI with BSC’s infrastructure support, with the current focus on its extension through Federated EGA. Canada joined the Federated EGA, marking the first major expansion beyond Europe and reinforcing its global dimension. Additionally, four resources are recognised as ELIXIR RIRs: 3DBIONOTES-API, FAIRtracks, FAIRCookbook and OpenEBench. Various ELIXIR Communities have adopted OpenEBench as their community-driven benchmarking platform. The INB/ELIXIR-ES continued its training activities and organised key meetings within ELIXIR. It gathered its national community in the XV Symposium on Bioinformatics (JBI2025) jointly organised with ELIXIR-PT and INSTRUCT-ES. Other ELIXIR events organised were the ELIXIR 3DBioinfo Community Annual General Meeting with the 3D-SIG Community, and the Biodiversity and Microbiome Community meetings. https://inb-elixir.es https://inb-elixir.es/resources
ABSTRACT IMPaCT-Data, the Data Science pillar of the Spanish National Infrastructure for Precision Medicine coordinated by the Barcelona Supercomputing Center (BSC), has assembled and deployed a national-scale infrastructure for clinical, genomics, and imaging workloads across major Spanish biomedical research institutions. The main technical challenge is enabling large-scale, compute-intensive workflows while complying with strict data-governance constraints, heterogeneous local infrastructures, and non-uniform compute and storage capabilities. The first layer of the platform is based on a distributed execution model in which workflows are centrally dispatched while computation is performed locally at participating institutions. A central Galaxy server instance hosted at BSC provides workflow management, provenance tracking, user access, and operational monitoring. Authentication and authorization are implemented using OpenID Connect (OIDC), with institutional identity providers federated through a central Keycloak service, allowing the preservation of local identity management policies. Galaxy Pulsar handles job dispatching and, based on asynchronous messaging (RabbitMQ/AMQPS), distributes workloads to remote federated nodes and delegates processing to site-local backends (e.g. container engines, HPC clusters). Analysis tools and reference datasets are distributed and synchronized through CVMFS, reducing the operational burden associated with maintaining synchronized analysis environments across a federated infrastructure. For highly sensitive datasets, a second execution network is being deployed, enabled and orchestrated using the Federated Execution Manager (FEM), and exposed to the researcher through a tailored virtual research environment powered by openVRE (future safeVRE). Under this architecture, participating institutions will retain full control not only over data access but also over the execution environment and approved analysis tools, making it particularly suitable for secure processing environments (SPEs). In this context, FEM allows institutions to enforce local governance policies, validate containers and workflows before execution, and control how intermediate results are generated and shared across sites. This design also supports regulated biomedical scenarios in which datasets remain in read-only mode and computation must be executed under strict auditing and traceability requirements. In addition, FEM natively supports more complex federation patterns, including federated analysis (calculator-aggregator schemes) and federated learning workflows (client-server schemes). The resulting platform provides a practical technical blueprint for biomedical data processing aligned with ELIXIR principles and other major efforts and initiatives, such as Genomics Alliance for Genomics and Health (GA4GH), EUCAIM, and GDI, among others: distributed execution close to data, standards-aligned interoperability, and reproducible workflows across a clinically governed federated network.
This review examines the current landscape of federated learning frameworks to evaluate their long-term sustainability, flexibility, and usability in biomedical research, where strict data regulations limit data sharing across institutions. Through a systematic literature analysis, the study assesses these frameworks against findability, accessibility, interoperability, and reusability for research software principles and compares reported use cases to framework functionalities to identify gaps in usability and scalability. The findings reveal that while most frameworks perform well in findability and reusability, they exhibit limited interoperability both among themselves and with specific software libraries. Although often developed for particular use cases, the technical foundations of these frameworks suggest potential for broader applicability. However, the scarce integration of privacy-preserving techniques and a predominant reliance on horizontal architectures may constrain their scalability in more complex federated learning scenarios. Ultimately, this analysis highlights the necessity for federated learning frameworks to evolve toward greater interoperability, flexibility, and privacy-awareness.
The rapid growth of omics data and its increasing use in biomedical research have intensified the need for both secure processing environments (SPEs) and trusted research environments (TREs) capable of handling large-scale analyses of sensitive datasets that cannot be distributed, while ensuring security, reproducibility, and data reusability. WfExS-backend [ https://github.com/inab/WfExS-backend ] was originally introduced as a high-level workflow execution orchestrator designed to address these challenges through isolated, containerised execution, encrypted data handling, and provenance tracking. In this work, we present a proof-of-concept implementation that illustrates the integration of the WfExS-backend with GA4GH Task Execution Service (TES) standard implementations bound to SPEs and TREs. We demonstrate this approach using a fine-tuned, customised Funnel instance as the TES backend, where the WfExS-backend itself runs within a containerised environment and orchestrates workflow execution on TES-managed resources. The proof-of-concept involves executing a real-world genomics pipeline, Sarek [ https://nf-co.re/sarek ], implemented in Nextflow under nf-core standards and integrated into their ecosystem, starting from a previous execution. The details of the previous instantiation are imported directly from a Workflow Run RO-Crate (WRROC) definition along with its default parameter values, required containers, and reference datasets. Sarek’s WRROC was generated from a previous execution with WfExS. The running layout relies on a nested container model based on Singularity, where WfExS-backend runs in a container while invoking additional containers to run individual workflow steps. This approach preserves isolation, portability, and compatibility with HPC environments, while enabling execution across distributed infrastructures. Data confidentiality is maintained through Crypt4GH encryption, and workflow provenance is captured using a separate WRROC representation. The current setup executes workflows on a single TES node and job. Future work will extend this model to support execution of independent workflow steps across multiple TES nodes or jobs, improving scalability and more efficient use of distributed computing resources. These mechanisms aim to align with European regulatory frameworks for the processing of sensitive health data, including the Data Governance Act and the European Health Data Space (EHDS). This work represents a concrete step towards the practical, secure, and reproducible execution of sensitive workflows in distributed environments, aligning with the objectives of the EOSC-ENTRUST project and contributing to the development of interoperable, dedicated research environments for the secure, federated analysis of sensitive biomedical data.
Publication metrics remain the primary academic performance assessment in modern academia. This focus has fostered the publish or perish culture, negatively impacting research output quantity and quality. Consequently, a gap exists in the recognition and reward of impactful non-traditional research artefacts, such as curated datasets, research software, and training materials. To address this, we share perspectives to expand the assessment consideration for researchers producing these outputs. We also highlight the dedicated roles and need for wider consideration of career paths of digital research technical professionals, including academic data curators, research software engineers and domain expert trainers. Through a mapping of valuable non-traditional artefacts illustrated by Life Science examples of ELIXIR Europe, we aim to foster a productive and sustainable Open Science ecosystem where these outputs are made visible and considered. We examine the technical solutions for credit and sociological barriers to recognition and reward, while also considering the impacts of generative AI. Furthermore, we discuss policy and funder mandate reforms to ensure those responsible for non-traditional research artefacts are properly recognised and rewarded, to accelerate global scientific discovery.
The European Health Data Space (EHDS) will help researchers use health data across EU Member States (MS). Currently, cross-border research faces heterogeneous data access processes. Using a real-world use case, this paper analyses challenges and opportunities brought by the upcoming implementation of the EHDS, assessing the situation before and after the regulation comes into force. The use case focused on metastatic colorectal cancer, analysing the relations between mutational signatures and clinical trajectories while addressing data access procedures across MS. The regulatory landscape and the challenges that need to be addressed for the EHDS to enable the secondary use of health data, particularly genomic data, are complex and heterogeneous across MS. We describe the pathway from data application to access to pseudonymized data in secure processing environments, emphasizing the legal requirements, including the role of ethics committees. Finally, we analyse the success factors for achieving access to the data and the reasons for access denial to support shaping the upcoming EHDS implementation. Several challenges remain unaddressed for cross-border data use, especially in the context of genomic data, where the complexity and heterogeneity of informed consent can impact or even impede data-sharing efforts. While EHDS can simplify processes across MS, it is crucial to ensure that additional safeguards do not negatively impact or block access to health data and that EHDS infrastructure is ready for effective and affordable processing of large volumes of genomic and other data.
The rapid growth of genomic and biological data, combined with the complexity of biomedical workflows, has created challenges in achieving interoperability across High-Performance Computing (HPC), Cloud Computing facilities, and secure data platforms. Seamless integration, robust security, and efficient resource utilization are crucial for enabling the optimal use of health data. In this poster, we present OpenVRE (Open Virtual Research Environment), an open-source, cloud-based platform designed to address these challenges. OpenVRE bridges the gap between HPC resources, sensitive data infrastructures, and analytical tools and workflows, providing a flexible environment for researchers.
Software is an essential component of research. However, little attention has been paid to it compared with that paid to research data. Recently, there has been an increase in efforts to acknowledge and highlight the importance of software in research activities. Structured metadata from platforms like bio.tools, Bioconductor, and Galaxy ToolShed offers valuable insights into research software in the Life Sciences. Although originally intended to support discovery and integration, this metadata can be repurposed for large-scale analysis of software practices. However, its quality and completeness vary across platforms, reflecting diverse documentation practices. To gain a comprehensive view of software development and sustainability, consolidating this metadata is necessary, but requires robust mechanisms to address its heterogeneity and scale. This article presents an evaluation of instruction-tuned large language models for the task of software metadata identity resolution, a critical step in assembling a cohesive collection of research software. Such a collection is the reference component for the Software Observatory at OpenEBench, a platform that aggregates metadata to monitor the FAIRness of research software in the Life Sciences. We benchmarked multiple models against a human-annotated gold standard, examined their behavior on ambiguous cases, and introduced an agreement-based proxy for high-confidence automated decisions. The proxy achieved high precision and statistical robustness, while also highlighting the limitations of current models and the broader challenges of automating semantic judgment in FAIR-aligned software metadata across registries and repositories.
In this era of rapidly expanding human genomics in research and healthcare, efficient data reuse is essential to maximize benefits for society. In response, the Federated European Genome–Phenome Archive (FEGA) was launched in 2022, and as of 2024, the FEGA network was composed of seven national nodes. Here we describe the complexities, challenges and achievements of FEGA, unravelling the dynamic interplay of regulatory frameworks, technical challenges and the shared vision of advancing genomic research.
The rising popularity of computational workflows is driven by the need for repetitive and scalable data processing, sharing of processing know-how, and transparent methods. As both combined records of analysis and descriptions of processing steps, workflows should be reproducible, reusable, adaptable, and available. Workflow sharing presents opportunities to reduce unnecessary reinvention, promote reuse, increase access to best practice analyses for non-experts, and increase productivity. In reality, workflows are scattered and difficult to find, in part due to the diversity of available workflow engines and ecosystems, and because workflow sharing is not yet part of research practice. WorkflowHub provides a unified registry for all computational workflows that links to community repositories, and supports both the workflow lifecycle and making workflows findable, accessible, interoperable, and reusable (FAIR). By interoperating with diverse platforms, services, and external registries, WorkflowHub adds value by supporting workflow sharing, explicitly assigning credit, enhancing FAIRness, and promoting workflows as scholarly artefacts. The registry has a global reach, with hundreds of research organisations involved, and more than 800 workflows registered.
Artificial intelligence (AI) is a powerful technology with the potential to disrupt cancer detection, diagnosis and treatment. However, the development of new AI algorithms requires access to large and complex real-world datasets. Although such datasets are constantly being generated, access to them is limited by data fragmentation across numerous repositories and sites, heterogeneity, lack of annotations, and potential privacy issues. The European Cancer Imaging Initiative is a flagship of Europe's Beating Cancer Plan, aiming to unlock the power of AI for cancer patients, clinicians, and researchers by establishing a federated European infrastructure for cancer images through the EU-funded EUropean Federation for CAncer IMages (EUCAIM) project. This infrastructure, called Cancer Image Europe, builds on the AI for Health Imaging network (AI4HI), established European Research Infrastructures (Euro-BioImaging, BBMRI-ERIC, EATRIS, ECRIN, and ELIXIR), and numerous related partners providing access to research tools, images, and related clinical, pathology and molecular data. The infrastructure targets clinicians, researchers, and innovators by providing the means to develop and validate data-intensive AI-based and other IT-enabled clinical decision-making systems supporting precision medicine. Common data models, including a linking hyperontology, quality standards, compliance with the FAIR (Findability, Accessibility, Interoperability and Reusability) principles, data annotation, curation and anonymization services are provided to ensure data quality and interoperability, consistency and privacy. In summer 2024, the EUCAIM project released the first prototype of an EU-wide infrastructure, with a comprehensive dashboard integrating applications for dataset discovery, federated search, data access request, metadata harvesting, annotation, secure processing environments and federated processing. CRITICAL RELEVANCE STATEMENT: EUCAIM's federated infrastructure for cancer image data advances medical research and related AI development in Europe. It addresses the current fragmentation and heterogeneity of data repositories is legally compliant, and facilitates collaboration among clinicians, researchers, and innovators. KEY POINTS: AI solutions to advance cancer care rely on large, high-quality real-world datasets. EUCAIM's federated infrastructure for cancer image data empowers cancer research in Europe. It provides access to research tools, images, and related clinical, pathology and molecular data.
Technology is moving faster than ever. Keeping systems up to date is key to staying relevant and making sure they last in the long run. OpenEBench , ELIXIR's gateway to the community-driven benchmarking efforts and research , software monitoring for Life Science tools, workflows and other complex systems, is also part of this ongoing transformation. To improve maintainability , scalability , and performance , we are refactoring OpenEBench by integrating modern libraries, transitioning to TypeScript , and adopting a modular UI architecture . These changes make development more efficient, improve code reliability, and ensure the platform remains flexible for future updates. However, technology alone is not enough—keeping documentation updated is just as important. Clear and up-to-date documentation improves and facilitates the contribution of developers. It also facilitates engaging with the platform’s users. This poster will outline our approach to modernizing OpenEBench, covering the technical improvements and the role of documentation in sustaining an evolving platform. We’ll share the key challenges we faced, the benefits of these changes, and why regular updates—both in code and documentation—are essential to keeping OpenEBench reliable and accessible for the community.
Søren Brunak合作论文数Rigshospitalet;Novo Nordisk Foundation Center for Protein Research, University of Copenhagen;Department of Systems Biology, Technical University of Denmark5