The 21st century drastically transformed the way scientific research is carried out. All stages, from planning to results interpretation, are heavily dependant on specialised research software, the quality of which becomes an important issue during project implementation. Establishing generally-accepted software quality metrics is an important step in software adoption, continuous improvement, and sustainability. EVERSE is an EU-funded project focusing on promoting research software as a first-class citizen in the scientific research community, providing a framework for the quality assessment and evaluation of such software. The project provides a set of Research Software Quality Indicators covering different dimensions, which goes beyond the FAIR principles applied to Research Software. In this context, the project consists of the assembling ‘resqui’ workflow, which facilitates the evaluation of a list of indicators, and the ‘dashVERSE’ web-interface which visualizes the results. OpenEBench, the ELIXIR platform supporting community-driven scientific benchmarking activities and the technical monitoring of research software, has implemented the FAIRsoft indicators. FAIRsoft high- and low-level indicators focus on automatically measuring how FAIR a given software is. Indicators represent a community-effort in translating the original FAIR principles from data to software and then proposing concrete approaches to measure such principles.. Here we present the integration of the FAIRsoft indicators, as implemented in the OpenEBench Software Observatory, into the EVERSE framework. Such integration effort aims to reduce duplicated efforts, leveraging current implementations, and facilitating specialized knowledge exchange across scientific communities with a common goal: making research software of high quality and, therefore, contributing towards its long-term adoption and sustainability.
Genome analysis have become a regular procedure for disease diagnosis and research. Software frameworks and platforms employed for such analysis are quite complex and usually include sophisticated workflows executed in a cloud environment. Given the intrinsically sensitive nature of the processed data, providing solid security during the entire data lifetime is very important for a broad adoption of the platforms. The Global Alliance for Genomics and Health (GA4GH) has developed the Crypt4GH file encryption standard, which despite of being quite simple, provides a very powerful way for protecting sensitive data. Although the Crypt4GH standard was warmly received by the research community, standard support requires a broad adoption by tools developers. Some popular tools like SAMTools included Crypt4GH support natively, but for many other tools data must be decrypted before being processed. Here we present the Liquid Crypt4GH Java file system implementation which may provide a transparent (“liquid”) Crypt4GH support for Java-based tools that otherwise lacks such support. The tool may also be used in an opaque mode where tool developers would decide which files must be encrypted / decrypted. Liquid Crypt4GH is a pure Java Crypt4GH implementation based on the JSR-203 NIO2 Java File System API and acts as a proxy between Java application and a file system. This way Java applications use standard Java NIO API and only required the private or public Crypt4GH key for transparent encryption / decryption.
OpenEBench (https://openebench.bsc.es/) is the ELIXIR open-data collaborative platform to support community-driven scientific benchmarking. As part of the ELIXIR Tools Platform, OpenEBench is dedicated to advancing scientific benchmarking and technical monitoring practices of research software in Life Sciences. The open nature of the platform facilitates its use and adoption by different communities within ELIXIR and beyond. Currently, there are 12 active communities within the platform with the expectation to reach 20 in the near future, thanks partially to the engagement with different European projects. OpenEBench has captured an overall diversity of scientific benchmarking needs since its creation. This diversity is also present among the new communities, i.e. for long-term storage and results display (CAID), benchmarking workflow execution (LRGASP), periodic benchmarking events (QfO), or continuous benchmarking (CAMEO). Recently, there has been a noticeable increase of Artificial Intelligence (AI) software and models benchmarks and OpenEBench has swiftly responded to these demands by extending its backend and capabilities. Indeed, OpenEBench is part of future deployments across different projects, including cancer imaging benchmark for enhanced AI in oncology (EuCanImage), adoption and deployment of research software best practices (EOSC-EVERSE and ELIXIR STEERS), sex and gender biases in AI models for health applications (BAIHA), and benchmarking of multilingual natural language processing to standardise the structuring of cardiology reports across European regions (DataTools4Heart). In addition to this diversity, OpenEBench introduces a new feature called "Project Spaces" designed to facilitate collaboration by providing a web space where projects and communities can present their efforts, provide guidelines and share relevant information with anyone interested in engaging with them.
The Long-read RNA-Seq Genome Annotation Assessment Project (LRGASP) Consortium was formed to evaluate the effectiveness of long-read approaches for transcriptome analysis. The consortium generated over 427 million long-read sequences from cDNA and direct RNA datasets, encompassing human, mouse, and manatee species, using different protocols and sequencing platforms. These data were utilized by developers to address challenges in transcript isoform detection and quantification, as well as de novo transcript isoform identification. The study revealed that libraries with longer, more accurate sequences produce more accurate transcripts than those with increased read depth, whereas greater read depth improved quantification accuracy. In well-annotated genomes, tools based on reference sequences demonstrated the best performance. When aiming to detect rare and novel transcripts or when using reference-free approaches, incorporating additional orthogonal data and replicate samples are advised. This collaborative study offers a benchmark for current practices and provides direction for future method development in transcriptome analysis.
OpenEBench, the ELIXIR platform supporting scientific community benchmarking activities and the technical monitoring of research software in Life Sciences, has strived since its inception in 2017 to serve a broad audience of end-users. The 2024 update brings forth a variety of new features with the primary goal of enriching user engagement and fostering collaboration. The widgets deployed in this update represent an innovative approach to data visualization. These encapsulated packages of code empower the visualization of results effortlessly, grouping all the functionalities in a simple and visual layout. Additionally, while initially developed for OpenEBench, these widgets are designed in such a way that can be easily deployed and integrated in third-party websites. Moreover, the introduction of the OpenEBench intranet marks a significant shift towards strengthening internal collaboration and communication within users. In the 2024 release, it is possible to create and manage communities and events directly within the platform, automatizing those operations and reducing the interactions with the OpenEBench helpdesk. This initial release will capture users’ feedback, which will be used to improve this newly deployed functionality and as the basis to streamline other operations within the platform. The primary objective is empowering scientific communities to manage their own content while reducing error-prone manual operations. As these enhancements are adopted, we expect that OpenEBench will continue to serve as an essential resource for scientific communities benchmarking activities within and beyond ELIXIR and as a framework for monitoring the research software quality across the Life Sciences community.
The European Genomic Data Infrastructure (GDI) project aims to establish a federated and secure platform facilitating access and analysis of genomic, phenotypic, and clinical data across Europe. This initiative operates through a network, where each node represents an European country responsible for implementing the required software stack to enable data sharing and processing within this infrastructure. The IMPaCT-Data Biomedical Cloud is the Spanish national implementation of the European Genomic Data Infrastructure (GDI). It is being established to provide a scalable and flexible analysis environment, enabling the integration, management and analysis of clinical, genomic and medical imaging data available within the Spanish National Precision Medicine Infrastructure associated with Science and Technology (IMPaCT).
The Spanish Precision Medicine Infrastructure associated with Science and Technology (IMPaCT), aims to lay the foundations for impulsing precision medicine within the Spanish National Health System. IMPaCT revolves around three main pillars: Predictive Medicine, Data Science and Genomic Medicine. As part of the Data Science program, the IMPaCT-Data Biomedical Cloud is being established for providing a scalable and flexible analysis environment, enabling the integration, management and analysis of clinical, genomic and medical imaging data available within IMPaCT. The IMPaCT-Data Biomedical Cloud presents a federated computing environment that provides a range of platforms, including UseGalaxy.* and VRE, with an authentication system based on OIDC (Keycloak) that incorporates Life Sciences Log-in capabilities and is supplemented by an authorization system based on GA4GH Passports. IMPaCT-Data Biomedical Cloud main objective is to provide researchers with a robust and efficient computing environment, enabling them to collaborate, analyze data, and share results in a secure and scalable manner and is advancing to finally provide a complete computing environment with included streamlined authentication, fine-grained authorization controls, and flexible platform integration. As part of its continuos work, two videos have been recently published showing the capabilities of the current prototype-based implementation illustrating the analysis of distributed data and the use of GA4GH-based Passports and Visas mechanisms for accessing data across different systems including the use of different identity providers across this federated infrastructure.
Here we report the results of a project started at the BioHackathon Europe 2022. Its goals were to cross-compare and analyze the metadata centralized in the Tools Ecosystem, and linked to the EDAM ontology, as well as to explore methods for connecting tools used in registered Galaxy workflows (i.e. WorkflowHub entries) to the annotations available in bio.tools.
After approval of the GA4GH Beacon v2 standard, a first Beacon v2 Network prototype implementation was used to demonstrate a good level of interoperability. Based on the learnings from several new Beacon v2, an updated Beacon Network Aggregator has been implemented, connected with dedicated adjustments to the Beacon specification itself. To support the emerging GA4GH Beacon v2 Networks design, the Centre for Genomic Regulation (CRG) developed a new Beacon Network v2 User Interface which, in conjunction with the Barcelona Supercomputing Center (BSC) Beacon Network Aggregator, provides a complete federated Beacon v2 querying solution with improved user experience. Attention was put on the interoperability, as new implementations were tested. Although the reference implementation provides the Beacon v2 compatibility verification tool, substantial divergences in the implementation were detected. As a result, the Beacon Network v2 working group has prepared a set of guidelines for the Beacon v2 developers that should improve the interoperability within the Beacon Network. In summary, the current Beacon Network v2 project provides an important milestone to facilitate cooperation in terms of sharing genomic data and associated annotations. We expect that the project will have a big impact on how researchers, especially practitioners in medical genetics and cancer genomics, will approach the sharing and discovery of such data and utilize the internet to empower future discoveries. This current implementation should also serve as the stepping stone for further developments coupled with new data types becoming available in Beacon, e.g. clinical and medical imaging data.
The OpenEBench platform (https://openebench.bsc.es) has been part of the ELIXIR Tools Ecosystem platform since its inception. Despite its broad functionality, OpenEBench contributes to the ecosystem by monitoring bioinformatics tools gathered from different sources such as bio.tools, Bioconda, Galaxy and DebianMed. It also provides a set of FAIR for Research Software indicators that are periodically updated. The current model for the tool descriptions has been greatly influenced by the biotoolsSchema and has been stable for the past six years. Nevertheless, the Tools Ecosystem Platform is working towards a more standard description for software tools, which led to the plan of adopting a metadata standard with broad community support being at the moment two options: Bioschemas ”ComputationalTool” Profile and CodeMeta. Considering the great degree of overlap and ongoing discussions to achieve converge and full interoperability, the OpenEBench tools monitoring platform has started the adoption of the Bioschemas ”ComputationalTool” profile. Although JSON-LD provides great semantic data representation, it greatly expands the size of the containing data. The new approach keeps storing data in a JSON format and enrich data with JSON-LD artifacts (@context, @type, @id, etc.) on the fly while preparing the response to any request through the available API. This effort should contribute to the existing discussions in the ELIXR Tools Ecosystem platform on the need to adopt widely supported standards for describing research software and should serve as exemplar on the implications of adopting Bioschemas as a standard data model within a given platform.
With its first concepts going back to 2014, in 2018 the Global Alliance for Genomics and Health (GA4GH) approved Beacon v1 protocol as its standard for genomics data discovery. Warmly welcomed by the genomics community, many institutions implemented the protocol to provide a standardized way to query over their genomic data. The main advantage of these efforts emerged through the federated data discovery enabled by networks of beacons. Recently approved by the GA4GH, the completely redesigned Beacon v2 is a big step towards the adoption of Beacon API in clinical genomics and healthcare. The protocol changes require a corresponding redesign of the Beacon Network architecture and infrastructure in order to support new protocol features and datatypes. The Barcelona Supercomputing Center (BSC) present together with IT Center for Science Ltd. (CSC), Centre for Genomic Regulation (CRG) and Swiss Institute of Bioinformatics (SIB) the prototype implementation of the ELIXIR Beacon Network v2 service that enables querying individual beacons in the network, the aggregation of Beacon responses, supports ELIXIR AAI with open, registered and controlled access tiers and the handover to ELIXIR Core data resources. The service is jointly operated by BSC in Spain and CSC in Finland. We expect that other ELIXIR communities such as Federated Human Data, Rare Diseases, human Copy Number Variation, Proteomics, Plants, Cancer Data, and Health Data will benefit from this service either directly or through the experience from the ELIXIR Beacon Network v2 design and implementation.
The size and complexity of the biomaterials literature makes systematic data analysis an excruciating manual task. A practical solution is creating databases and information resources. Implant design and biomaterials research can greatly benefit from an open database for systematic data retrieval. Ontologies are pivotal to knowledge base creation, serving to represent and organize domain knowledge. To name but two examples, GO, the gene ontology, and CheBI, Chemical Entities of Biological Interest ontology and their associated databases are central resources to their respective research communities. The creation of the devices, experimental scaffolds, and biomaterials ontology (DEB), an open resource for organizing information about biomaterials, their design, manufacture, and biological testing, is described. It is developed using text analysis for identifying ontology terms from a biomaterials gold standard corpus, systematically curated to represent the domain's lexicon. Topics covered are validated by members of the biomaterials research community. The ontology may be used for searching terms, performing annotations for machine learning applications, standardized meta-data indexing, and other cross-disciplinary data exploitation. The input of the biomaterials community to this effort to create data-driven open-access research tools is encouraged and welcomed.
OpenEBench is the ELIXIR benchmarking and technical monitoring platform for bioinformatics tools, web servers and workflows. OpenEBench is part of the ELIXIR Tools platform and its development is led by the Barcelona Supercomputing Center (BSC) in collaboration with partners within ELIXIR and beyond. Scientific benchmarking helps determining the precision, recall and other metrics of bioinformatics resources in unbiased scenarios, which have been set up through reference databases, ad-hoc input and test data sets reflecting specifying scientific challenges. Chosen metrics allow to objectively evaluate the relative scientific performance of the different participating resources. It is even possible to understand what are the software potential biases, strengths and weaknesses and/or under which conditions do they perform better or worse. Unbiased and objective evaluations of bioinformatics resources are challenging to set-up and can only be effective when built and implemented around community driven efforts. Several communities from different scientific domains collaborate with OpenEBench in order to set-up, host and further develop their scientific efforts. Communities can focus on specific problems, e.g. Quest for Orthologs (QfO); or having a broader spectrum e.g. Spanish Network of Biomedical Research Centers on Rare Diseases (CIBERER); or covering different challenges on each of their editions, e.g. DREAM Challenges. Benchmarking efforts led by scientific communities might have a national scope e.g. CIBERER; or a global one e.g., Global Microbial Identifier Initiative (GMI). Most communities have similar needs in terms of reference data sets, metrics, and benchmarking results accessible within the community and beyond, independently of the scientific challenges tackled by each community and their geographical scope. Data sets should reflect existing challenges of the scientific community in terms of size, complexity, and content. Moreover, data sets are used for producing predictions by participants and to compute the performance of each participant when comparing their predictions against a previously agreed, often private, data sets that are referred many times as golden data sets. Metrics are used to measure the performance of individual participants and should reflect the common practices in the field. Finally, making results available to the community and beyond is as relevant as generating them. It is important that researchers can access these results at any time, and have the tools to assist them in understanding them. Moreover, associated data and metadata to any benchmarking efforts should fulfill the FAIR principles and be available in long-term repositories such as Zenodo and/or EUDAT with permanent digital identifiers for further use and re-use. Thus, OpenEBench has engaged with different communities offering assistance to bring their previously generated data and activities into the platform. Communities can make use of any of the three available levels in the OpenEBench architecture. However, how communities use the platform depends on their specific needs and resources. To ensure the long-term sustainability of OpenEBench, we are implementing a co-production model to accelerate the incorporation of new communities and the maintenance of the existing ones.
jmpegg library provides MPEG-G files encoding/decoding features. The library is written in a pure java language without external dependencies. The library also supports conversion between MPEG-G and well-known genomic formats.
ABSTRACT Multiscale Genomics (MuG) Virtual Research Environment (MuGVRE) is a cloud-based computational infrastructure created to support the deployment of software tools addressing the various levels of analysis in 3D/4D genomics. Integrated tools tackle needs ranging from high computationally demanding applications (e.g. molecular dynamics simulations) to high-throughput data analysis applications (like the processing of next generation sequencing). The MuG Infrastructure is based on openNebula cloud systems implemented at the Institute for research in Biomedicine, and the Barcelona Supercomputing Center, and has specific interfaces for users and developers. Interoperability of the tools included in MuGVRE is maintained through a rich set of metadata allowing the system to associate tools and data in a transparent manner. Execution scheduling is based in a traditional queueing system to handle demand peaks in applications of fixed needs, and an elastic and multi-scale programming model (pyCOMPSs, controlled by the PMES scheduler), for complex workflows requiring distributed or multi-scale executions schemes. MuGVRE is available at https://vre.multiscalegenomics.eu and documentation and general information at https://www.multiscalegenomics.eu . The infrastructure is open and freely accessible.
We understand benchmarking as the comparison of research software performance under controlled conditions. It encompasses both the technical performance, including software quality metrics, and the scientific performance in predefined challenges. Scientific communities play an important role here as they are responsible for defining reference datasets and metrics, pointing out the existing scientific challenges in their respective fields. Thus, in the context of ELIXIR-EXCELERATE project, we have developed the OpenEBench platform (https://openebench.bsc.es) aiming to provide a reference place to host technical and scientific performance for research software across the life sciences. OpenEBench provides an infrastructure where end-users can learn from different available software options and select the one best fitting their scientific needs. Bioinformatics software developers can find relevant datasets and meaningful scientific challenges to evaluate their own developments, and communities interested in a particular scientific domain can easily define which datasets and metrics are relevant for developers to work on, which in turn will allow the field to move ahead. A web application (https://openebench.bsc.es/html/scientific) allows users to browse through the benchmarking results from the different communities engaged (TCGA, QFO, CAMEO, GMI) in the platform which can be viewed using one of the visualization charts and transformed to table format, which is easier to interpret by non-expert users. On the technical side, the application uses several REST APIs, which can be used by other developers to upload and access the data for future studies or use cases. Another key feature is the Virtual Research Environment (https://openebench.bsc.es/submission), which includes the necessary mechanisms to import and execute benchmarking workflows on top of cloud computing infrastructures. OpenEBench follows the recommendations made by ELIXIR on the development of open source software making its code publicly available at https://github.com/inab/openebench-hub.
Advances in DNA sequencing technologies lead to an increase in the amount of sequenced genomic data, requiring new approaches for how this information is stored and processed. The most popular formats for storing genomic data - Sequence Alignment/Map Format (SAM) and its binary counterpart (BAM) do not provide an acceptable level of compression. The latter was improved in the CRAM format. Alongside the need for better compression, a modern genomic storage format should also integrate metadata representation, security strategies, and standardized accesses to improve interoperability. These features are not contemplated in the CRAM specification. Furthermore, there is no standardization body guaranteeing the perennity of CRAM. This motivated MPEG standardization group to start working on the new MPEG-G ISO/IEC 23092 Genomic Information Representation standard. Countries contributing to the MPEG-G standard are involved in the development via local research efforts. Spanish research project - Secure GENomic information COMpression (GENCOM) contributed with benchmarking of compression methods, file format definition, security and information metadata. GENCOM is a collaboration project between Polytechnical University of Catalonia, University of Barcelona and Barcelona Supercomputing Center. Here we present the MPEG-G format, going over the three main parts which composes it: file format, genomic representation and compression, metadata and security strategies. In addition to the participation in the implementation of the reference software, we developed pure Java implementation of the standard to be included in the HTSJdk library. This integration will allow MPEG-G file format usage in modern genome analysis pipelines that include tools like GATK or Picard. MPEG-G opens new uses cases which were not supported by SAM/BAM or CRAM. The ISO/IEC standardization body is also a guarantee for users that the format will be supported in the future.