In alignment with the European Health Data Space (EHDS), Spain's IMPaCT - Precision Medicine Infrastructure associated with Science and Technology - aims to establish a trusted research environment (TRE) for secure and FAIR data sharing and analysis. It is structured into three pillars: Predictive Medicine (IMPaCT-Cohort), Genomic Medicine (IMPaCT-Genomics), and Data Science (IMPaCT-Data). IMPaCT-Data leads the development of the IMPaCT Digital Platform (IDP), integrating clinical, genomic, and imaging data to support the national IMPaCT-Cohort and Personalised Medicine Projects (IMPaCT-PMPs). Its Reference Implementation defines the architecture across a federated model.
The Spanish National Bioinformatics Institute (INB), founded in 2003 as a distributed network, is the ELIXIR Node in Spain and has two objectives: 1) deepen its involvement and leadership within ELIXIR and broaden the resources provided as part of ELIXIR infrastructure to the Life Sciences community; and 2) increase its impact within the Spanish National Health System . INB/ELIXIR-ES continues to strengthen its technological capabilities in federated data infrastructures, interoperability, and FAIR data management within ELIXIR. The Node is actively involved in several ELIXIR-driven projects and commissioned services (CoS) within the 2024-28 Work Programme. From the ELIXIR perspective, the Service Delivery Plan (SDP) maintains 40 resources offered by 24 groups belonging to 12 institutions. Regarding national activities, the INB/ELIXIR-ES leads the Translational Bioinformatics Network (TransBioNet) and serves as proxy between IMPaCT-Data activities, the Data Science pillar of the Spanish National Infrastructure for Precision Medicine, and European efforts. These activities align with major European data projects such as the Genomic Data Infrastructure (GDI), EUCAIM, and the Federated European Genome-phenome Archive (FEGA), while implementing Global Alliance for Genomics and Health (GA4GH) standards in its technological developments. The Node strongly engages within the 2024-28 Work Programme: TechnologyTier: co-leadership of Data, Tools and Training Platforms, with contributions across all Platforms. During this period, the co-led ELIXIR Beacon Network Infrastructure Service secured funding for this service, strengthening federated data discovery capabilities. ScienceTier : co-leadership of CMR and HDTR Science priority areas; co-leadership of Rare Diseases, FHD, Cancer Data and Biodiversity Communities; and Pathogens Data, and RNA Data Focus Group; leading and participating in several CoS in the HDTR and CMR areas. PeopleTier : co-leadership of the ELEAD2.0 leadership programme, a CoS built on the experiences of Bioinfo4Women, and active role in the PeoplePulse CoS. Active role in the NodeTier NSCS, together with 4 Platforms, 13 Communities, and 8 Focus Groups. Regarding the INB/ELIXIR-ES portfolio, EGA, an ELIXIR CDR, is co-developed and maintained by CRG and EMBL-EBI with BSC’s infrastructure support, with the current focus on its extension through Federated EGA. Canada joined the Federated EGA, marking the first major expansion beyond Europe and reinforcing its global dimension. Additionally, four resources are recognised as ELIXIR RIRs: 3DBIONOTES-API, FAIRtracks, FAIRCookbook and OpenEBench. Various ELIXIR Communities have adopted OpenEBench as their community-driven benchmarking platform. The INB/ELIXIR-ES continued its training activities and organised key meetings within ELIXIR. It gathered its national community in the XV Symposium on Bioinformatics (JBI2025) jointly organised with ELIXIR-PT and INSTRUCT-ES. Other ELIXIR events organised were the ELIXIR 3DBioinfo Community Annual General Meeting with the 3D-SIG Community, and the Biodiversity and Microbiome Community meetings. https://inb-elixir.es https://inb-elixir.es/resources
MOTIVATION:Enabling clinicians and researchers to directly interact with global genomic data resources by removing technological barriers is vital for medical genomics. AskBeacon enables large language models (LLMs) to be applied to securely shared cohorts via the Global Alliance for Genomics and Health Beacon protocol. By simply "asking" Beacon, actionable insights can be gained, analyzed, and made publication-ready. RESULTS:In the Parkinson's Progression Markers Initiative (PPMI), we use natural language to ask whether the sex-differences observed in Parkinson's disease are due to X-linked or autosomal markers. AskBeacon returns a publication-ready visualization showing that for PPMI the autosomal marker occurred 1.4 times more often in males with Parkinson's disease than females, compared to no differences for the X-linked marker. We evaluate commercial and open-weight LLM models, as well as different architectures to identify the best strategy for translating research questions to Beacon queries. AskBeacon implements extensive safety guardrails to ensure that genomic data is not exposed to the LLM directly, and that generated code for data extraction, analysis and visualization process is sanitized and hallucination resistant, so data cannot be leaked or falsified. AVAILABILITY AND IMPLEMENTATION:AskBeacon is available at https://github.com/aehrc/AskBeacon.
In this era of rapidly expanding human genomics in research and healthcare, efficient data reuse is essential to maximize benefits for society. In response, the Federated European Genome–Phenome Archive (FEGA) was launched in 2022, and as of 2024, the FEGA network was composed of seven national nodes. Here we describe the complexities, challenges and achievements of FEGA, unravelling the dynamic interplay of regulatory frameworks, technical challenges and the shared vision of advancing genomic research.
Regulatory frameworks such as the General Data Protection Regulation (GDPR) and the European Health Data Space (EHDS) are reshaping how genomic and clinical data could be safely shared across Europe. In alignment with these efforts, and following the example of the 1+Million Genomes (1+MG) initiative, we are working to advance the development of secure infrastructures to facilitate cross-border access to sensitive biomedical data. A key component of this landscape is the European Genome-phenome Archive (EGA) and its Federated counterpart (FEGA), which serve as essential repositories for storing and managing genomic and pheno-clinical data. However, leveraging these databases for secure analysis within Virtual Research Environments (VREs) remains a significant challenge, and is actively tackled in projects like the European Genomic Data Infrastructure (GDI), EOSC-ENTRUST among others. Our current work focuses on expanding the secure data analysis capabilities of Galaxy in order to make it fit in this ecosystem. Through the ELIXIR Comission Service “Empowering Users: Orchestrating Sensitive Data Access for Interactive Federated Analysis in Virtual Research Environments,” we are working on enabling designated secure compute nodes to process encrypted data without exposing user private keys through GA4GH Crypt4GH protocols. This approach facilitates encrypted data access and transfer from EGA/FEGA to Galaxy while optimally routing such data to computing nodes where it will be processed. Our solution will support flexible security configurations, from encrypted data transfer with minimal restrictions to fully secure federated analysis. Furthermore, the underlying infrastructure is designed for adaptability, aiming to making it deployable in other VREs beyond Galaxy.
BACKGROUND:An unprecedented amount of personal health data, with the potential to revolutionize precision medicine, is generated at health care institutions worldwide. The exploitation of such data using artificial intelligence (AI) relies on the ability to combine heterogeneous, multicentric, multimodal, and multiparametric data, as well as thoughtful representation of knowledge and data availability. Despite these possibilities, significant methodological challenges and ethicolegal constraints still impede the real-world implementation of data models. TECHNICAL DETAILS:The EuCanImage is an international consortium aimed at developing AI algorithms for precision medicine in oncology and enabling secondary use of the data based on necessary ethical approvals. The use of well-defined clinical data standards to allow interoperability was a central element within the initiative. The consortium is focused on 3 different cancer types and addresses 7 unmet clinical needs. We have conceived and implemented an innovative process to capture clinical data from hospitals, transform it into the newly developed EuCanImage data models, and then store the standardized data in permanent repositories. This new workflow combines recognized software (REDCap for data capture), data standards (FHIR for data structuring), and an existing repository (EGA for permanent data storage and sharing), with newly developed custom tools for data transformation and quality control purposes (ETL pipeline, QC scripts) to complement the gaps. CONCLUSION:This article synthesizes our experience and procedures for health care data interoperability, standardization, and reproducibility.
In the age of data-driven biomedical research and clinical practice, the sharing of genomic and clinical data for health research and personalized medicine has become an important contributor to improved diagnosis and treatment.From the data owner's perspective, potential benefits include improved treatments, personalization of healthcare practice, and more effective control of disease proliferation.However, the requirement for high levels of data security to protect sensitive information presents a barrier to data discovery and sharing [1].Beacon is designed to enable the benefits of data discovery while minimizing the associated risks.It is a Global Alliance for Genomics and Health (GA4GH) API specification (see Box 1 for a definition of key Beacon terminology), allowing easy discovery of sensitive data that require controlled (authorized) access [2].It uses simple concepts and can be adapted to different use cases.The protocol is designed to respond to queries, such as the following:"Can you provide data about males, diagnosed with Type 2 diabetes, whose age of onset is below 30 years, and who carry mutations in the APOE gene?"Depending on the data controller's preferences over response granularity, the response options range from, "Yes, our data includes one or more" (boolean response), "Yes, we have 125" (count response), to "Yes, and here are some details about the 125 individuals that match your request" (detailed "record level" response).Discovery is the necessary first step in the sharing and reuse of data and other assets, and Beacon facilitates this by enabling federated discovery in any number of networks, as a complement to unwieldy central catalogs.Many beacons have already been successfully "lit" (deployed) across the globe (https://public.tableau.com/app/profile/elixir/viz/ELIXIRBeaconNetwork/Sheet1).This article is written to support data owners that might be interested in deploying a beacon to make their data discoverable while keeping them secure.Whatever your background is, this article will provide you with some tips to get you started.Specifically, we will review important steps to complete-and pitfalls to avoid-when deploying a beacon. Tip #1: Evaluate the value of your data to your research or clinical domainDeploying a Beacon instance will help you keep the data both as open as possible and as restricted as necessary (Tip #9).That said, you might want to take into account the perceived value of your dataset, which depends on the target communities.For example, a dataset on the
An unprecedented amount of personal health data, with the potential to revolutionise precision medicine, is generated at healthcare institutions worldwide. The exploitation of such data using artificial intelligence relies on the ability to combine heterogeneous, multicentric, multimodal and multiparametric data, as well as thoughtful representation of knowledge and data availability. Despite these possibilities, significant methodological challenges and ethico-legal constraints still impede the real-world implementation of data models. The EuCanImage is an international consortium aimed at developing AI algorithms for precision medicine in oncology and enabling secondary use of the data based on necessary ethical approvals. The use of well-defined clinical data standards to allow interoperability was a central element within the initiative. The consortium is focused on three different cancer types and addresses seven unmet clinical needs. This article synthesises our experience and procedures for healthcare data interoperability and standardisation. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This project has received funding from the European Union's Horizon 2020 research and innovation programme under grant agreement No 952103. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes This study describes a new process to harmonize and standardize clinical data. The data will be available upon request to the authors.
Often, research data is available through repositories and archives with varying access mechanisms depending on the nature of the data. However, it is not usual for such repositories to have the capabilities to process, integrate, and visualise a particular dataset of interest. To facilitate these operations, a framework is necessary for reproducible results and interoperability. We present in this work a technical demonstrator showcasing the mobilisation of data across platforms, illustrating a workflow from a data repository, such as the European Genome-phenome Archive (EGA) or NextCloud to cBioPortal via Galaxy. This particular demonstrator begins with extracting raw synthetic data from where it is stored, followed by its transfer to Galaxy for subsequent analysis. Within Galaxy, the dataset is processed by two distinct workflows to convert it into the sequence variant format required by cBioPortal. The first workflow aligns raw FASTA data and identifies somatic variants, generating a Variant Call Format (VCF) file. The VCF file is annotated using the second workflow and transformed into Mutation Annotation Format (MAF). This MAF file and its accompanying clinical data are then uploaded to cBioPortal for further analysis and visualisation. Furthermore, here, we present a methodology for creating genomic and clinical synthetic data, providing insights about generating representative datasets for testing and validation purposes. Through this demonstrator, we showcase an approach to data mobilisation, facilitating efficient integration and improved interoperability across diverse bioinformatics platforms for enhanced research and clinical insights.
Precision medicine aims at tailoring treatments to individual patient’s characteristics. In this regard, recognizing the significance of sex and gender becomes indispensable for meeting the distinct healthcare needs of diverse populations. To this end, continuing a trend of improving data quality observed since 2014, the European Genome-phenome Archive (EGA) established a policy in 2018 that mandates data providers to declare the sex of donor samples, aiming to enhance data accuracy and prevent imbalance in sex classification. We analyzed sex classification imbalance in human data from EGA and the U.S. counterpart, the database of genotypes and phenotypes (dbGaP). Our findings show a significant decrease in samples classified as unknown in EGA, potentially promoting better sex reporting during data collection. Based on our findings, we raise awareness of sample imbalance problems and provide a list of recommendations for enhancing biomedical research practices.
The European Genomic Data Infrastructure (GDI) project aims to establish a federated and secure platform facilitating access and analysis of genomic, phenotypic, and clinical data across Europe. This initiative operates through a network, where each node represents an European country responsible for implementing the required software stack to enable data sharing and processing within this infrastructure. The IMPaCT-Data Biomedical Cloud is the Spanish national implementation of the European Genomic Data Infrastructure (GDI). It is being established to provide a scalable and flexible analysis environment, enabling the integration, management and analysis of clinical, genomic and medical imaging data available within the Spanish National Precision Medicine Infrastructure associated with Science and Technology (IMPaCT).
Precision medicine aims at tailoring treatments to individual patient needs. In this context, artificial intelligence (AI)-based technologies are viewed as revolutionary since they have the capacity to identify key features that link genomic and phenotypic traits at the individual level. AI techniques therefore depend on the quantity and quality of patient data. When variables like sex, age, or race are ignored in sample records, it can result in biased predictions as they will not be considered in the training of the AI algorithm. To this end, the European Genome-phenome Archive (EGA) took action in 2018 and put into place a rule that requires data providers to declare the sex of donor samples uploaded into their repository to improve data quality and prevent the spread of biased results. In this work we quantified biases in sex classification over time in human data from studies deposited in EGA and the database of Genotypes and Phenotypes (dbGaP), which represents the EGA's equivalent in the USA. The main result is that the EGA policy is effective to fight sex classification biases because there are significantly less samples classified as unknown after 2018 in this repository than in dbGaP. Additionally, we qualitatively assessed public opinion on this issue. A survey addressed to users, creators, maintainers, and developers of biological databases revealed that specialized training and additional knowledge about diversity criteria are required. Based on our findings, we raise awareness of sample bias problems and provide a list of recommendations for enhancing biomedical research practices.
Artificial intelligence (AI) is transforming the field of medical imaging and has the potential to bring medicine from the era of ‘sick-care’ to the era of healthcare and prevention. The development of AI requires access to large, complete, and harmonized real-world datasets, representative of the population, and disease diversity. However, to date, efforts are fragmented, based on single–institution, size-limited, and annotation-limited datasets. Available public datasets ( e.g. , The Cancer Imaging Archive, TCIA, USA) are limited in scope, making model generalizability really difficult. In this direction, five European Union projects are currently working on the development of big data infrastructures that will enable European, ethically and General Data Protection Regulation-compliant, quality-controlled, cancer-related, medical imaging platforms, in which both large-scale data and AI algorithms will coexist. The vision is to create sustainable AI cloud-based platforms for the development, implementation, verification, and validation of trustable, usable, and reliable AI models for addressing specific unmet needs regarding cancer care provision. In this paper, we present an overview of the development efforts highlighting challenges and approaches selected providing valuable feedback to future attempts in the area. Key points • Artificial intelligence models for health imaging require access to large amounts of harmonized imaging data and metadata. • Main infrastructures adopted either collect centrally anonymized data or enable access to pseudonymized distributed data. • Developing a common data model for storing all relevant information is a challenge. • Trust of data providers in data sharing initiatives is essential. • An online European Union meta-tool-repository is a necessity minimizing effort duplication for the various projects in the area.
After approval of the GA4GH Beacon v2 standard, a first Beacon v2 Network prototype implementation was used to demonstrate a good level of interoperability. Based on the learnings from several new Beacon v2, an updated Beacon Network Aggregator has been implemented, connected with dedicated adjustments to the Beacon specification itself. To support the emerging GA4GH Beacon v2 Networks design, the Centre for Genomic Regulation (CRG) developed a new Beacon Network v2 User Interface which, in conjunction with the Barcelona Supercomputing Center (BSC) Beacon Network Aggregator, provides a complete federated Beacon v2 querying solution with improved user experience. Attention was put on the interoperability, as new implementations were tested. Although the reference implementation provides the Beacon v2 compatibility verification tool, substantial divergences in the implementation were detected. As a result, the Beacon Network v2 working group has prepared a set of guidelines for the Beacon v2 developers that should improve the interoperability within the Beacon Network. In summary, the current Beacon Network v2 project provides an important milestone to facilitate cooperation in terms of sharing genomic data and associated annotations. We expect that the project will have a big impact on how researchers, especially practitioners in medical genetics and cancer genomics, will approach the sharing and discovery of such data and utilize the internet to empower future discoveries. This current implementation should also serve as the stepping stone for further developments coupled with new data types becoming available in Beacon, e.g. clinical and medical imaging data.
With increasingly strict regulations for managing sensitive human omics and healthcare-derived data, solutions are required to ensure research and clinical data can be accessed across national borders. Safely sharing sensitive data is vital for data reuse to support biomedical research and personalised medicine programs. The Federated European Genome-phenome Archive (FEGA) provides a data sharing solution of infrastructure and governance frameworks to support discovery of and secure access to human data globally, while respecting national data protection regulations. Prompted by European initiatives such as the 1+ Million Genomes and European Health Data Space, the FEGA was officially launched in 2022 with a network of nodes in the UK, Spain, Norway, Sweden, Finland, and Germany. The next phase of FEGA will build upon these early successes to expand data sharing both within and outside of Europe. Accelerating this expansion can be achieved through collaboration with initiatives already working towards human data sharing, for example: the European Genome Data Infrastructure (GDI), Data Science for Health Discovery and Innovation in Africa (DS-I Africa), and Australian BioCommons. Ultimately, the FEGA vision is to build a global, interoperable discovery and access network of human data resources, to accelerate disease research and understand and improve human health.
The Global Alliance for Genomics and Health (GA4GH) is a standards-setting organization that is developing a suite of coordinated standards for genomics. The GA4GH Phenopacket Schema is a standard for sharing disease and phenotype information that characterizes an individual person or biosample. The Phenopacket Schema is flexible and can represent clinical data for any kind of human disease including rare disease, complex disease, and cancer. It also allows consortia or databases to apply additional constraints to ensure uniform data collection for specific goals. We present phenopacket-tools, an open-source Java library and command-line application for construction, conversion, and validation of phenopackets. Phenopacket-tools simplifies construction of phenopackets by providing concise builders, programmatic shortcuts, and predefined building blocks (ontology classes) for concepts such as anatomical organs, age of onset, biospecimen type, and clinical modifiers. Phenopacket-tools can be used to validate the syntax and semantics of phenopackets as well as to assess adherence to additional user-defined requirements. The documentation includes examples showing how to use the Java library and the command-line tool to create and validate phenopackets. We demonstrate how to create, convert, and validate phenopackets using the library or the command-line application. Source code, API documentation, comprehensive user guide and a tutorial can be found at https://github.com/phenopackets/phenopacket-tools. The library can be installed from the public Maven Central artifact repository and the application is available as a standalone archive. The phenopacket-tools library helps developers implement and standardize the collection and exchange of phenotypic and other clinical data for use in phenotype-driven genomic diagnostics, translational research, and precision medicine applications.
Jaap Heringa合作论文数Department of Computer Science, Vrije Universiteit Amsterdam/Department of Bioinformatics, Vrije Universiteit Amsterdam/ELIXIR/Dutch Techcentre for Life Sciences7