Scientific research increasingly depends on robust and scalable IT infrastructures to support complex computational workflows. With the proliferation of services provided by research infrastructures, NRENs, and commercial cloud providers, researchers must navigate a fragmented ecosystem of computing environments, balancing performance, cost, scalability, and accessibility. Hybrid cloud architectures offer a compelling solution by integrating multiple computing environments to enhance flexibility, resource efficiency, and access to specialised hardware. This paper provides a comprehensive overview of hybrid cloud deployment models, focusing on grid and cloud platforms (OpenPBS, SLURM, OpenStack, Kubernetes) and workflow management tools (Nextflow, Snakemake, CWL). We explore strategies for federated computing, multi-cloud orchestration, and workload scheduling, addressing key challenges such as interoperability, data security, reproducibility, and network performance. Drawing on implementations from life sciences, as coordinated by the ELIXIR Compute Platform and their integration into a wider EOSC context, we propose a roadmap for accelerating hybrid cloud adoption in research computing, emphasising governance frameworks and technical solutions that can drive sustainable and scalable infrastructure development.
Are you a sysadmin who would like to make cloud-based storage analytics solutions available to your institution, or would you like to make your institutional resources available to the ELIXIR community? Are you a developer who would like to integrate your tools and services into a growing federated cloud environment? Are you a researcher that is interested to learn how cloud computing may benefit your large-scale data analysis project at your home institution or across national and international institutions? If you are still here, the ELIXIR Cloud might be for you! We are developing building blocks for cloud-based data storage and analytics that interoperate through common authentication/authorization guidelines and dedicated open community standards defined by the Global Alliance for Genomics and Health (GA4GH), an international standard-setting organization bringing together opinion leaders in the life sciences and the biomedical sector. You can use these modules to build up your own cloud, for on premise, hybrid and multicloud setups. Or you can integrate with or make use of the ELIXIR Cloud, a federated analysis platform built on these modules and set up across multiple ELIXIR Nodes. Either way, the modular design of our solutions ensures a high degree of agnosticism with regard to underlying technologies (e.g., workflow engines, compute solutions) and extensibility. Here we would like to present to you the current status of the project, our roadmap for 2023, and how we may be able to help you with your own cloud infrastructure, integration and/or data analysis needs.
Research software is a critical component of computational research. Thus, being able to discover, understand and adequately utilize software is essential. All these activities rely on software metadata. However, these metadata are sparse, expensive to maintain, sometimes is not machine readable, and frequently inconsistent across different resources. The ELIXIR Research Software Ecosystem acts as a proxy to maintain and preserve high-quality metadata for describing research software, aggregating heterogeneous metadata resources from multiple providers. These providers often have different objectives when they are maintaining and aggregating metadata by themselves, as well as user communities and technical implementations. The goals of this project are to consolidate these metadata, and facilitate their integration, curation, and re-usability by the contributors and the wider research community.
The main goals and challenges for the life science communities in the Open Science framework are to increase reuse and sustainability of data resources, software tools, and workflows, especially in large-scale data-driven research and computational analyses. Here, we present key findings, procedures, effective measures and recommendations for generating and establishing sustainable life science resources based on the collaborative, cross-disciplinary work done within the EOSC-Life (European Open Science Cloud for Life Sciences) consortium. Bringing together 13 European life science research infrastructures, it has laid the foundation for an open, digital space to support biological and medical research. Using lessons learned from 27 selected projects, we describe the organisational, technical, financial and legal/ethical challenges that represent the main barriers to sustainability in the life sciences. We show how EOSC-Life provides a model for sustainable data management according to FAIR (findability, accessibility, interoperability, and reusability) principles, including solutions for sensitive- and industry-related resources, by means of cross-disciplinary training and best practices sharing. Finally, we illustrate how data harmonisation and collaborative work facilitate interoperability of tools, data, solutions and lead to a better understanding of concepts, semantics and functionalities in the life sciences.
Poster describing the ELIXIR Compute Platform work. Presented at the ELIXIR All Hands Meeting 2023 in Dublin, Ireland.
Research software is a critical component of computational research. Being able to discover, understand and adequately utilize software is essential. Many existing services facilitate these tasks, all of them relying heavily on software metadata. The continued upkeep of such complex and large sets of metadata comes at the cost of multiple efforts of curation, and the resulting metadata is often sparse and inconsistent. The ELIXIR Research Software Ecosystem (RSEc) aims to act as a proxy to maintain and preserve high-quality metadata for describing research software. These metadata are retrieved and synchronized with many major software-related services, within and beyond the ELIXIR Tools Platform. The EDAM ontology enables the semantic description of the scientific function of the described software. The RSEc central repository is a GitHub repository that aggregates software metadata mostly related to life sciences, and spanning the multiple aspects of software discovery, evaluation, deployment and execution. The aggregation of metadata in a centralized, open, and version-controlled repository enables the cross-linking of services, the validation and enrichment of software metadata, the development of new services and the analysis of these metadata.
Toxicology has been an active research field for many decades, with academic, industrial and government involvement. Modern omics and computational approaches are changing the field, from merely disease-specific observational models into target-specific predictive models. Traditionally, toxicology has strong links with other fields such as biology, chemistry, pharmacology and medicine. With the rise of synthetic and new engineered materials, alongside ongoing prioritisation needs in chemical risk assessment for existing chemicals, early predictive evaluations are becoming of utmost importance to both scientific and regulatory purposes. ELIXIR is an intergovernmental organisation that brings together life science resources from across Europe. To coordinate the linkage of various life science efforts around modern predictive toxicology, the establishment of a new ELIXIR Community is seen as instrumental. In the past few years, joint efforts, building on incidental overlap, have been piloted in the context of ELIXIR. For example, the EU-ToxRisk, diXa, HeCaToS, transQST, and the nanotoxicology community have worked with the ELIXIR TeSS, Bioschemas, and Compute Platforms and activities. In 2018, a core group of interested parties wrote a proposal, outlining a sketch of what this new ELIXIR Toxicology Community would look like. A recent workshop (held September 30th to October 1st, 2020) extended this into an ELIXIR Toxicology roadmap and a shortlist of limited investment-high gain collaborations to give body to this new community. This Whitepaper outlines the results of these efforts and defines our vision of the ELIXIR Toxicology Community and how it complements other ELIXIR activities.
The Global Alliance for Genomics and Health (GA4GH) aims to accelerate biomedical advances by enabling the responsible sharing of clinical and genomic data through both harmonized data aggregation and federated approaches. The decreasing cost of genomic sequencing (along with other genome-wide molecular assays) and increasing evidence of its clinical utility will soon drive the generation of sequence data from tens of millions of humans, with increasing levels of diversity. In this perspective, we present the GA4GH strategies for addressing the major challenges of this data revolution. We describe the GA4GH organization, which is fueled by the development efforts of eight Work Streams and informed by the needs of 24 Driver Projects and other key stakeholders. We present the GA4GH suite of secure, interoperable technical standards and policy frameworks and review the current status of standards, their relevance to key domains of research and clinical care, and future plans of GA4GH. Broad international participation in building, adopting, and deploying GA4GH standards and frameworks will catalyze an unprecedented effort in data sharing that will be critical to advancing genomic medicine and ensuring that all populations can access its benefits.
Intrinsically disordered proteins (IDPs) and intrinsically disordered regions (IDRs) are now recognised as major determinants in cellular regulation. This white paper presents a roadmap for future e-infrastructure developments in the field of IDP research within the ELIXIR framework. The goal of these developments is to drive the creation of high-quality tools and resources to support the identification, analysis and functional characterisation of IDPs. The roadmap is the result of a workshop titled “An intrinsically disordered protein user community proposal for ELIXIR” held at the University of Padua. The workshop, and further consultation with the members of the wider IDP community, identified the key priority areas for the roadmap including the development of standards for data annotation, storage and dissemination; integration of IDP data into the ELIXIR Core Data Resources; and the creation of benchmarking criteria for IDP-related software. Here, we discuss these areas of priority, how they can be implemented in cooperation with the ELIXIR platforms, and their connections to existing ELIXIR Communities and international consortia. The article provides a preliminary blueprint for an IDP Community in ELIXIR and is an appeal to identify and involve new stakeholders.
Intrinsically disordered proteins (IDPs) and intrinsically disordered regions (IDRs) are now recognised as major determinants in cellular regulation. This white paper presents a roadmap for future e-infrastructure developments in the field of IDP research within the ELIXIR framework. The goal of these developments is to drive the creation of high-quality tools and resources to support the identification, analysis and functional characterisation of IDPs. The roadmap is the result of a workshop titled "An intrinsically disordered protein user community proposal for ELIXIR" held at the University of Padua. The workshop, and further consultation with the members of the wider IDP community, identified the key priority areas for the roadmap including the development of standards for data annotation, storage and dissemination; integration of IDP data into the ELIXIR Core Data Resources; and the creation of benchmarking criteria for IDP-related software. Here, we discuss these areas of priority, how they can be implemented in cooperation with the ELIXIR platforms, and their connections to existing ELIXIR Communities and international consortia. The article provides a preliminary blueprint for an IDP Community in ELIXIR and is an appeal to identify and involve new stakeholders.
This paper defines the UK Infra-Red Telescope (UKIRT) Hemisphere Survey (UHS) and release of the remaining similar to 12 700 deg(2) of J-band survey data products. The UHS will provide continuous J- and K-band coverage in the Northern hemisphere from a declination of 0 degrees to 60 degrees by combining the existing Large Area Survey, Galactic Plane Survey and Galactic Clusters Survey conducted under the UKIRT Infra-red Deep Sky Survey (UKIDSS) programme with this new additional area not covered by UKIDSS. The released data include J-band imaging and source catalogues over the new area, which, together with UKIDSS, completes the J-band UHS coverage over the full similar to 17 900 deg(2) area. 98 per cent of the data in this release have passed quality control criteria. The remaining 2 per cent have been scheduled for re-observation. The median 5s point source sensitivity of the released data is 19.6mag (Vega). The median full width at half-maximum of the point spread function across the data set is 0.75 arcsec. In this paper, we outline the survey management, data acquisition, processing and calibration, quality control and archiving as well as summarizing the characteristics of the released data products. The data are initially available to a limited consortium with a world-wide release scheduled for 2018 August.
This research was based on observations obtained with XMM– Newton, an ESA science mission with instruments and contributions directly funded by ESA Member States and the National Aeronautics and Space Administration. This research was also based on observations made at the Anglo-Australian Telescope. MJP acknowledges financial support from the UK Science and Technology Facilities Council. FJC and SM acknowledge financial support through grant AYA2015-64346-C2-1-P (MINECO/FEDER). MTC acknowledges support by the Spanish Programma Nacional de Astronomia y Astrofisica under grant AYA2009-08059. MK acknowledges support by DFG grant KR 3338/3-1. The NASA/IPAC Extragalactic Database is operated by the Jet Propulsion Laboratory, California Institute of Technology, under contract with the National Aeronautics and Space Administration.
The availability of workflows for data publishing could have an enormous impact on researchers, research practices and publishing paradigms, as well as on funding strategies and career and research evaluations. We present the generic components of such workflows to provide a reference model for these stakeholders. The RDA-WDS Data Publishing Workflows group set out to study the current data-publishing workflow landscape across disciplines and institutions. A diverse set of workflows were examined to identify common components and standard practices, including basic self-publishing services, institutional data repositories, long-term projects, curated data repositories, and joint data journal and repository arrangements. The results of this examination have been used to derive a data-publishing reference model comprising generic components. From an assessment of the current data-publishing landscape, we highlight important gaps and challenges to consider, especially when dealing with more complex workflows and their integration into wider community frameworks. It is clear that the data-publishing landscape is varied and dynamic and that there are important gaps and challenges. The different components of a data-publishing system need to work, to the greatest extent possible, in a seamless and integrated way to support the evolution of commonly understood and utilized standards and—eventually—to increased reproducibility. We therefore advocate the implementation of existing standards for repositories and all parts of the data-publishing process, and the development of new standards where necessary. Effective and trustworthy data publishing should be embedded in documented workflows. As more research communities seek to publish the data associated with their research, they can build on one or more of the components identified in this reference model.
In disciplines such as biomedicine and social sciences, sharing and combining sensitive individual-level data is often prohibited by ethical-legal or governance constraints and other barriers such as the control of intellectual property or the huge sample sizes. DataSHIELD ( D ata A ggregation T hrough A nonymous S ummary-statistics from H armonised I ndividual-lev EL D atabases) is a distributed approach that allows the analysis of sensitive individual-level data from one study, and the co-analysis of such data from several studies simultaneously without physically pooling them or disclosing any data. Following initial proof of principle, a stable DataSHIELD platform has now been implemented in a number of epidemiological consortia. This paper reports three new applications of DataSHIELD including application to post-publication sensitive data analysis, text data analysis and privacy protected data visualisation. Expansion of DataSHIELD analytic functionality and application to additional data types demonstrate the broad applications of the software beyond biomedical sciences.
Software underpins the academic research process across disciplines. To be able to understand, use/reuse and preserve data, the software code that generated, analysed or presented the data will need to be retained and executed. An important part of this process is being able to persistently identify the software concerned. This paper discusses the reasons for doing so and introduces a model of software entities to enable better identification of what is being identified. The DataCite metadata schema provides a persistent identification scheme and we consider how this scheme can be applied to software. We then explore examples of persistent identification and reuse. The examples show the differences and similarities of software used in academic research, which has been written and reused at different scales. The key concepts of being able to identify what precisely is being used and provide a mechanism for appropriate credit are important to both of them. Â Â