EDAM is an ontology of concepts prevalent in the data analysis and data management in life sciences. EDAM is developed in a participatory and transparent fashion, within a broad and growing community of contributors from both inside and outside of ELIXIR. This development model, based on the contribution of a large number of scientific experts, therefore comes with its own set of challenges. To streamline and accelerate the evolution of EDAM, we have developed and integrated a set of tools that automate the quality control and release process for the ontology. In addition to ensuring the global consistency of EDAM, it enforces edition best practices both at the syntactic and semantic levels. These tools have been integrated in a continuous integration (CI) pipeline, automated using GitHub Actions in the source-code repository. EDAM contribution toolkit also include: EDAM Browser, a web interface that lets users explore EDAM and its usage graphically EDAM Popovers, a web-browser extension showing details of EDAM concepts selected in any website EDAMmap, a tool for mapping natural language text to EDAM ontology concepts These tools enable more efficient and higher-quality maintenance of EDAM, and rapid feedback to the contributors who are encouraged to suggest direct contributions.
Homepage : https://edamontology.org/ Source Code : https://github.com/edamontology/edamontology License : CC BY-SA 4.0 EDAM [ 1 ] is a domain ontology of data analysis and data management in bio- and other sciences, and science-based applications. It comprises concepts related to analysis, modelling, optimisation, and data life-cycle. Targetting usability by diverse end users, the structure of EDAM is relatively simple, divided into 4 main sections: Topic Operation Data (incl. Identifier ) Format EDAM is used in numerous resources, for example Bio.tools , Galaxy , Debian , or the ELIXIR Europe training portal TeSS . Thanks to the annotations with EDAM, computational tools , workflows , standards , data , and learning materials are easier to find, compare, choose, and integrate. EDAM contributes to open science by allowing semantic annotation of processed data, thus making the data more understandable, findable, and comparable. EDAM and its applications lower the barrier and effort for scientists and citizens alike, towards doing scientific research in a more open , reliable , and inclusive way. The main improvements in 2021 include: The addition of essential concepts of data management and open science Improved automated validation (CI) Improved contribution processes towards more inclusion and engagement with communities of scientific experts, software engineers, and volunteers Easier to contribute , and therefore new contributors and new collaborations . And welcoming more! News from the applications: Galaxy can now visualise the tools panel sorted by EDAM. The EDAM community brings together software engineers and science experts (academic, industrial, citizen), professionals and volunteers. It can be followed and contributed to via various channels. One option is GitHub , with dedicated repositories incl. edamontology and others. Other platforms for contributions to EDAM are especially the generic ontology browsers of the NCBO BioPortal ( EDAM , EDAM Bioimaging [ 2 ]) and WebProtégé ( EDAM , EDAM Bioimaging ; free registration required also for viewing). The main communication channel is Gitter . Along with the ontology, the community has developed a number of tools that enhance user experience with EDAM: EDAM Browser [ 3 ] is a lightweight and fast web-based ontology browser that provides a number of user-oriented features such as aggregated search across various EDAM-annotated resources ( e.g . Bio.tools , TeSS ), and suggesting changes to EDAM. EDAMmap is a tool for text mining EDAM concepts from articles. EDAM Popovers is a web-browser add-on / extension for showing details of EDAM concepts found in a website. Great e.g. for code and textual data on GitHub References: [ 1 ] Ison, J., Kalaš, M., Jonassen, I., Bolser, D., Uludag, M., McWilliam, H., Malone, J., Lopez, R., Pettifer, S. and Rice, P. ( 2013 ). EDAM: an ontology of bioinformatics operations, types of data and identifiers, topics and formats. Bioinformatics , 29 (10): 1325-1332. DOI: 10.1093/bioinformatics/btt113 Open access [ 2 ] Matúš Kalaš, Laure Plantard, Joakim Lindblad, Martin Jones, Nataša Sladoje, Moritz A. Kirschmann, Anatole Chessel, Leandro Scholz, Fabienne Rössler, Laura Nicolás Sáenz, Estibaliz Gómez de Mariscal, John Bogovic, Alexandre Dufour, Xavier Heiligenstein, Dominic Waithe, Marie-Charlotte Domart, Matthia Karreman, Raf Van de Plas, Robert Haase, David Hörl, Lassi Paavolainen, Ivana Vrhovac Madunić, Dean Karaica, Arrate Muñoz-Barrutia, Paula Sampaio, Daniel Sage, Sebastian Munck, Ofra Golani, Josh Moore, Florian Levet, Jon Ison, Alban Gaignard, Hervé Ménager, Chong Zhang, Kota Miura, Julien Colombelli, Perrine Paul-Gilloteaux, and welcoming new contributors! ( 2020 ) . EDAM-bioimaging: the ontology of bioimage informatics operations, topics, data, and formats (update 2020) [version 1; not peer reviewed]. F1000Research , 9 (ELIXIR):162 (Poster) DOI: 10.7490/f1000research.1117826.1 Open access [ 3 ] Brancotte, B., Blanchet, C. and Ménager, H. ( 2018 ). A reusable tree-based web-visualization to browse EDAM ontology, and contribute to it. J. Open Source Softw. , 3 (27): 698. DOI: 10.21105/joss.00698 Open access Acknowledgement: EDAM maintainers and interns were supported by ELIXIR Europe, Norway, and France. Dr. Melissa Black and Gloria Umutesi were awarded with complimentary registration at ISMB/ECCB 2021, funded by the BOSC 2021 sponsors . The list of co-authors includes the substantial contributors to EDAM in 2020-2021 (version 1.26), without the contributors specific to EDAM Biomaging, EDAM Browser, or EDAMmap.
Computational models have great potential to accelerate bioscience, bioengineering, and medicine. However, it remains challenging to reproduce and reuse simulations, in part, because the numerous formats and methods for simulating various subsystems and scales remain siloed by different software tools. For example, each tool must be executed through a distinct interface. To help investigators find and use simulation tools, we developed BioSimulators (https://biosimulators.org), a central registry of the capabilities of simulation tools and consistent Python, command-line and containerized interfaces to each version of each tool. The foundation of BioSimulators is standards, such as CellML, SBML, SED-ML and the COMBINE archive format, and validation tools for simulation projects and simulation tools that ensure these standards are used consistently. To help modelers find tools for particular projects, we have also used the registry to develop recommendation services. We anticipate that BioSimulators will help modelers exchange, reproduce, and combine simulations.
The bio.tools registry is a main catalogue of computational tools in the life sciences. More than 17 000 tools have been registered by the international bioinformatics community. The bio.tools metadata schema includes semantic annotations of tool functions, that is, formal descriptions of tools’ data types, formats, and operations with terms from the EDAM bioinformatics ontology. Such annotations enable the automated composition of tools into multistep pipelines or workflows. In this Technical Note, we revisit a previous case study on the automated composition of proteomics workflows. We use the same four workflow scenarios but instead of using a small set of tools with carefully handcrafted annotations, we explore workflows directly on bio.tools. We use the Automated Pipeline Explorer (APE), a reimplementation and extension of the workflow composition method previously used. Moving “into the wild” opens up an unprecedented wealth of tools and a huge number of alternative workflows. Automated composition tools can be used to explore this space of possibilities systematically. Inevitably, the mixed quality of semantic annotations in bio.tools leads to unintended or erroneous tool combinations. However, our results also show that additional control mechanisms (tool filters, configuration options, and workflow constraints) can effectively guide the exploration toward smaller sets of more meaningful workflows.
BACKGROUND:Life scientists routinely face massive and heterogeneous data analysis tasks and must find and access the most suitable databases or software in a jungle of web-accessible resources. The diversity of information used to describe life-scientific digital resources presents an obstacle to their utilization. Although several standardization efforts are emerging, no information schema has been sufficiently detailed to enable uniform semantic and syntactic description-and cataloguing-of bioinformatics resources. FINDINGS:Here we describe biotoolsSchema, a formalized information model that balances the needs of conciseness for rapid adoption against the provision of rich technical information and scientific context. biotoolsSchema results from a series of community-driven workshops and is deployed in the bio.tools registry, providing the scientific community with >17,000 machine-readable and human-understandable descriptions of software and other digital life-science resources. We compare our approach to related initiatives and provide alignments to foster interoperability and reusability. CONCLUSIONS:biotoolsSchema supports the formalized, rigorous, and consistent specification of the syntax and semantics of bioinformatics resources, and enables cataloguing efforts such as bio.tools that help scientists to find, comprehend, and compare resources. The use of biotoolsSchema in bio.tools promotes the FAIRness of research software, a key element of open and reproducible developments for data-intensive sciences.
Scientific data analyses often combine several computational tools in automated pipelines, or workflows. Thousands of such workflows have been used in the life sciences, though their composition has remained a cumbersome manual process due to a lack of standards for annotation, assembly, and implementation. Recent technological advances have returned the long-standing vision of automated workflow composition into focus. This article summarizes a recent Lorentz Center workshop dedicated to automated composition of workflows in the life sciences. We survey previous initiatives to automate the composition process, and discuss the current state of the art and future perspectives. We start by drawing the “big picture” of the scientific workflow development life cycle, before surveying and discussing current methods, technologies and practices for semantic domain modelling, automation in workflow development, and workflow assessment. Finally, we derive a roadmap of individual and community-based actions to work toward the vision of automated workflow development in the forthcoming years. A central outcome of the workshop is a general description of the workflow life cycle in six stages: 1) scientific question or hypothesis, 2) conceptual workflow, 3) abstract workflow, 4) concrete workflow, 5) production workflow, and 6) scientific results. The transitions between stages are facilitated by diverse tools and methods, usually incorporating domain knowledge in some form. Formal semantic domain modelling is hard and often a bottleneck for the application of semantic technologies. However, life science communities have made considerable progress here in recent years and are continuously improving, renewing interest in the application of semantic technologies for workflow exploration, composition and instantiation. Combined with systematic benchmarking with reference data and large-scale deployment of production-stage workflows, such technologies enable a more systematic process of workflow development than we know today. We believe that this can lead to more robust, reusable, and sustainable workflows in the future.
The FAIR Guiding Principles, published in 2016, aim to improve the findability, accessibility, interoperability and reusability of digital research objects for both humans and machines. Until now the FAIR principles have been mostly applied to research data. The ideas behind these principles are, however, also directly relevant to research software. Hence there is a distinct need to explore how the FAIR principles can be applied to software. In this work, we aim to summarize the current status of the debate around FAIR and software, as basis for the development of community-agreed principles for FAIR research software in the future. We discuss what makes software different from data with regard to the application of the FAIR principles, and which desired characteristics of research software go beyond FAIR. Then we present an analysis of where the existing principles can directly be applied to software, where they need to be adapted or reinterpreted, and where the definition of additional principles is required. Here interoperability has proven to be the most challenging principle, calling for particular attention in future discussions. Finally, we outline next steps on the way towards definite FAIR principles for research software.
The corpus of bioinformatics resources is huge and expanding rapidly, presenting life scientists with a growing challenge in selecting tools that fit the desired purpose. To address this, the European Infrastructure for Biological Information is supporting a systematic approach towards a comprehensive registry of tools and databases for all domains of bioinformatics, provided under a single portal (https://bio.tools). We describe here the practical means by which scientific communities, including individual developers and projects, through major service providers and research infrastructures, can describe their own bioinformatics resources and share these via bio.tools.
Pub2Tools is a Java command-line tool that looks through the scientific literature available in Europe PMC and constructs entry candidates for the bio.tools software registry from suitable publications. Pub2Tools automates a lot of the process needed for growing bio.tools, though its results still need some manual curation before they are of satisfactory quality. Running Pub2Tools once per month would result in hundreds of new metadata entries of tools, databases and services published in bioinformatics and life sciences journals. First, Pub2Tools gets a list of all publications available for the given period. After this initial list is narrowed down by combinations of keyphrases, the contents of the remaining publications are downloaded and the abstract of each publication is mined for the potential tool name. Names are assigned confidence scores, with low confidence publications often not being suitable for bio.tools at all. In addition to the tool name, web links matching the name are extracted from the abstract and full text of a publication and divided to the homepage and other link attributes of bio.tools. The content of publications and extracted links is also mined for software license and programming language information and phrases for the tool description attribute are automatically constructed. Good enough results are chosen for inclusion to bio.tools. In addition to finding new content for bio.tools, Pub2Tools can also be used to improve the current content when run on existing entries of bio.tools. For downloading publications and links, Pub2Tools uses a complementary tool developed by the authors, called PubFetcher. PubFetcher is capable of fetching publications from many other resources besides Europe PMC. And in case of web pages, it can extract metadata from standard sites like code repositories. Once the information about tools is obtained by Pub2Tools, another tool -- called EDAMmap -- is used for automatically adding EDAM ontology annotations to the new bio.tools entries. Thus, the full toolset can be used to identify, obtain and annotate public domain information about bioinformatics tools and databases in an interoperable manner. Pub2Tools is free and open-source software available from https://github.com/bio-tools/pub2tools, with the documentation at https://pub2tools.readthedocs.io/.
Project website: http://edamontology.org Source code: https://github.com/edamontology/edamontology License: CC BY-SA 4.0 EDAM is an ontology of well-established, familiar concepts that are prevalent within bioinformatics, and bioscientific data analysis in general [ 1 , 2 ]. The scope of EDAM includes types of data and data identifiers, data formats, operations, and topics. EDAM has a relatively simple structure, and comprises a set of concepts with terms, synonyms, definitions, relations, links, and some additional information (especially for data formats). EDAM is developed in a participatory and transparent fashion, within a growing international community of contributors. The development of EDAM is coordinated with the development and curation of tools registries ( e.g. bio.tools and BIII.eu ); registries of training materials ( e.g. TeSS ); with packaging of open-source bioinformatics software (especially Debian Med [ 3 ]); the Common Workflow Language [ 4 ]; and other related communities and initiatives. These include the developers’ community of Galaxy [ 5 ], and collaborations with specialised networks of experts, such as within the development of EDAM-bioimaging [ 6 ]. EDAM-bioimaging is an extension of EDAM towards bioimage informatics and machine learning, where a broad group of experts in bioimaging, image analysis, and deep learning has been contributing to the common effort. The comprehensive but concise inclusion of machine learning topics is one of the new additions in 2020.The latest release of EDAM at the time of publication was version 1.24 [ 7 ], and EDAM-bioimaging version alpha06 [ 8 ]. In summary, EDAM functions as common controlled vocabulary when publishing, sharing, and integrating information about bioinformatics tools, workflows, training materials, and other resources. In addition, EDAM is also useful when choosing terminology, for data provenance, and in text mining ( e.g. EDAMmap ). Slightly shorter versions of this abstract were reviewed by members of the corresponding committees of the listed conferences. [1] Jon Ison, Matúš Kalaš, Inge Jonassen, Dan Bolser, Mahmut Uludag, Hamish McWilliam, James Malone, Rodrigo Lopez, Steve Pettifer, Peter Rice. (2013) . EDAM: an ontology of bioinformatics operations, types of data and identifiers, topics and formats. Bioinformatics , 29 (10): 1325-1332. DOI: 10.1093/bioinformatics/btt113 Open access [2] Matúš Kalaš, Hervé Ménager, Jon Ison, Egon Willighagen, Björn Grüning (2017-2020) . edamontology/edamontology (All versions). Zenodo . DOI: 10.5281/zenodo.822690 Open access [3] Steffen Möller, Stuart W. Prescott, Lars Wirzenius, Petter Reinholdtsen, Brad Chapman, Pjotr Prins, Stian Soiland-Reyes, Fabian Klötzl, Andrea Bagnacani, Matúš Kalaš, Andreas Tille, Michael R. Crusoe (2017) . Robust Cross-Platform Workflows: How Technical and Scientific Communities Collaborate to Develop, Test and Share Best Practices for Data Analysis. Data Sci. Eng. , 2 : 232–244. DOI: 10.1007/s41019-017-0050-4 Open access [4] Peter Amstutz, Michael R. Crusoe, Nebojša Tijanić (editors), Brad Chapman, John Chilton, Michael Heuer, Andrey Kartashov, Dan Leehr, Hervé Ménager, Maya Nedeljkovich, Matt Scales, Stian Soiland-Reyes, Luka Stojanovic (2016) . Common Workflow Language, v1.0. Specification, Common Workflow Language working group . https://w3id.org/cwl/v1.0/ DOI: 10.6084/m9.figshare.3115156.v2 Open access [5] Hervé Ménager, Jon Ison, Matúš Kalaš, Veit Schwämmle, EDAM contributors (2017) . The EDAM ontology and its integration into Galaxy [version 1; not peer reviewed]. F1000Research , 6 (Galaxy):1032 (Poster). DOI: 10.7490/f1000research.1114336.1 Open access [6] Matúš Kalaš, Laure Plantard, Joakim Lindblad, Martin Jones, Nataša Sladoje, Moritz A. Kirschmann, Anatole Chessel, Leandro Scholz, Fabienne Rössler, Laura Nicolás Sáenz, Estibaliz Gómez de Mariscal, John Bogovic, Alexandre Dufour, Xavier Heiligenstein, Dominic Waithe, Marie-Charlotte Domart, Matthia Karreman, Raf Van de Plas, Robert Haase, David Hörl, Lassi Paavolainen, Ivana Vrhovac Madunić, Dean Karaica, Arrate Muñoz-Barrutia, Paula Sampaio, Daniel Sage, Sebastian Munck, Ofra Golani, Josh Moore, Florian Levet, Jon Ison, Alban Gaignard, Hervé Ménager, Chong Zhang, Kota Miura, Julien Colombelli, and Perrine Paul-Gilloteaux. We are welcoming new contributors! (2020) . EDAM-bioimaging: the ontology of bioimage informatics operations, topics, data, and formats (update 2020) [version 1; not peer reviewed]. F1000Research , 9 (ELIXIR,NEUBIAS):162 (Poster). DOI: 10.7490/f1000research.1117826.1 Open access [7] Matúš Kalaš, Hervé Ménager, Jon Ison, Egon Willighagen, Björn Grüning (2020) . edamontology/edamontology: EDAM 1.24 (Version 1.24). Zenodo . DOI: 10.5281/zenodo.3608238 Open access [8] Joakim Lindblad, Laure Plantard, Martin Jones, Nataša Sladoje, Marie-Charlotte Domart, Matthia Karreman, Estibaliz Gómez de Mariscal, Laura Nicolás Sáenz, Christos Kyprianou, María Arrate Muñoz-Barrutia, Raf Van de Plas, Xavier Heiligenstein, Ivana Vrhovac Madunić, Dean Karaica, Daniel Sage, Robert Haase, and all contributors to the previous versions, Matúš Kalaš (2020) . edamontology/edam-bioimaging: alpha06 (Version alpha06). Zenodo . DOI: 10.5281/zenodo.3695725 Open access
EDAM is a well-established ontology of operations, topics, types of data, and data formats that are used in bioinformatics and its neighbouring fields [ 1 , 2 ] . EDAM-bioimaging is an extension of EDAM dedicated to bioimage analysis, bioimage informatics, and bioimaging. It is being developed in collaboration between the ELIXIR research infrastructure and the NEUBIAS and COMULIS COST Actions, in close contact with the Euro-BioImaging research infrastructure and the Global BioImaging network. EDAM-bioimaging contains an inter-related hierarchy of concepts including bioimage analysis and related operations, bioimaging topics and technologies, and bioimage data and their formats. The modelled concepts enable interoperable descriptions of software, publications, data, workflows, and training materials, fostering open science and "reproducible" bioimage analysis. New developments in EDAM-bioimaging at the time of publication [ 3 ] include among others: A concise but relatively comprehensive ontology of Machine learning, Artificial intelligence, and Clustering (to the level relevant in particular in bioimaging, biosciences, and also scientific data analysis in general) Added and refined topics and synonyms within Sample preparation and Tomography, and finalised coverage of imaging techniques (all of these to the high-level extent that influences choices of downstream analysis, i.e. the scope of EDAM) EDAM-bioimaging continues being under active development, with a growing and diversifying community of contributors. It is used in BIII.eu , the registry of bioimage analysis tools, workflows, and training materials, and emerging also in descriptions of Debian Med packages available in Debian and Bio-Linux, and tools in bio.tools . Development of EDAM-bioimaging has been carried out in a successful open community manner, in a fruitful collaboration between numerous bioimaging experts and ontology developers. The last stable release at the time of poster publication is version alpha06 [ 3 ], and the live development version can be viewed and commented on WebProtégé (free registration required). New contributors are warmly welcome! [ 1 ] Ison, J., Kalaš, M., Jonassen, I., Bolser, D., Uludag, M., McWilliam, H., Malone, J., Lopez, R., Pettifer, S. and Rice, P. (2013). EDAM: an ontology of bioinformatics operations, types of data and identifiers, topics and formats. Bioinformatics , 29(10): 1325-1332. DOI: 10.1093/bioinformatics/btt113 Open Access [ 2 ] Kalaš, M., Ménager, H., Schwämmle, V., Ison, J. and EDAM Contributors (2017) . EDAM – the ontology of bioinformatics operations, types of data, topics, and data formats (2017 update) [version 1; not peer reviewed]. F1000Research , 6(ISCB Comm J):1181 (Poster) DOI: 10.7490/f1000research.1114459.1 Open Access [ 3 ] Matúš Kalaš, Laure Plantard, Martin Jones, Nataša Sladoje, Marie-Charlotte Domart, Matthia Karreman, Arrate Muñoz-Barrutia, Raf Van de Plas, Ivana Vrhovac Madunić, Dean Karaica, Laura Nicolás Sáenz, Estibaliz Gómez de Marisca, Daniel Sage, Robert Haase Joakim Lindblad, and all contributors to previous versions (2020). edamontology/edam-bioimaging: alpha06 (Version alpha06). Zenodo . DOI: 10.5281/zenodo. 3695725 Open Access
Proteomics is a highly dynamic field driven by frequent introduction of new technological approaches, leading to high demand for new software tools and the concurrent development of many methods for data analysis, processing, and storage. The rapidly changing landscape of proteomics software makes finding a tool fit for a particular purpose a significant challenge. The comparison of software and the selection of tools capable to perform a certain operation on a given type of data rely on their detailed annotation using well-defined descriptors. However, finding accurate information including tool input/output capabilities can be challenging and often heavily depends on manual curation efforts. This is further hampered by a rather low half-life of most of the tools, thus demanding the maintenance of a resource with updated information about the tools. We present here our approach to curate a collection of 189 software tools with detailed information about their functional capabilities. We furthermore describe our efforts to reach out to the proteomics community for their engagement, which further increased the catalog to >750 tools being about 70% of the estimated number of 1097 tools existing for proteomics data analysis.
JIB.tools 2.0 is a new approach to more closely embed the curation process in the publication process. This website hosts the tools, software applications, databases and workflow systems published in the Journal of Integrative Bioinformatics (JIB). As soon as a new tool-related publication is published in JIB, the tool is posted to JIB.tools and can afterwards be easily transferred to bio.tools, a large information repository of software tools, databases and services for bioinformatics and the life sciences. In this way, an easily-accessible list of tools is provided which were published in JIB a well as status information regarding the underlying service. With newer registries like bio.tools providing these information on a bigger scale, JIB.tools 2.0 closes the gap between journal publications and registry publication. (Reference: https://jib.tools).
Bioinformaticians and biologists rely increasingly upon workflows for the flexible utilization of the many life science tools that are needed to optimally convert data into knowledge. We outline a pan-European enterprise to provide a catalogue ( https://bio.tools ) of tools and databases that can be used in these workflows. bio.tools not only lists where to find resources, but also provides a wide variety of practical information.
EDAM is a well-established ontology of operations, topics, types of data, and data formats that are used in bioinformatics and its neighbouring fields [ 1 , 2 , 3 ] . EDAM-bioimaging is an extension of EDAM dedicated to bioimage analysis, bioimage informatics, and bioimaging. It is developed in collaboration between the NEUBIAS and ELIXIR Europe|ELIXIR-EXCELERATE projects, in contact with Euro-BioImaging and Global BioImaging. EDAM-bioimaging contains an inter-related hierarchy of concepts including bioimage analysis and related operations, bioimaging topics and technologies, and bioimage data and their formats. The modelled concepts enable interoperable descriptions of software, publications, data, and workflows, fostering reliable and transparent, "reproducible" bioimage analysis. EDAM-bioimaging is under active development, with a few alpha releases publicly available. It is used in BISE|biii.eu , the bioimaging tools and resources information portal, and emerging also in descriptions of Debian Med packages available in Debian and Bio-Linux. Development of EDAM-bioimaging has been carried out in a successful open community manner, in a fruitful collaboration between numerous bioimaging experts and ontology developers. The last stable release at the time of poster submission is version alpha05 [ 4 ], and the live development version can be viewed and commented on WebProtégé (free registration required). New contributors are warmly welcome! [ 1 ] Ison, J., Kalaš, M., Jonassen, I., Bolser, D., Uludag, M., McWilliam, H., Malone, J., Lopez, R., Pettifer, S. and Rice, P. (2013). EDAM: an ontology of bioinformatics operations, types of data and identifiers, topics and formats. Bioinformatics , 29(10): 1325-1332. DOI: 10.1093/bioinformatics/btt113 Open Access [ 2 ] Kalaš, M., Ménager, H., Schwämmle, V., Ison, J. and EDAM Contributors (2017) . EDAM – the ontology of bioinformatics operations, types of data, topics, and data formats (2017 update) [version 1; not peer reviewed]. F1000Research , 6(ISCB Comm J):1181 (Poster) DOI: 10.7490/f1000research.1114459.1 Open Access [ 3 ] Kalaš, M., Ménager, H., Ison, J. and Willighagen, E. (2018). edamontology/edamontology: EDAM 1.21 (Version 1.21). Zenodo . DOI: 10.5281/zenodo. 1325952 Open Access [ 4 ] Matúš Kalaš, Nataša Sladoje, Laure Plantard, Martin Jones, Leandro Aluisio Scholz, Joakim Lindblad, and contributors (2019). edamontology/edam-bioimaging: alpha05 (Version alpha05). Zenodo . DOI: 10.5281/zenodo.2557012 Open Access
Numerous software utilities operating on mass spectrometry (MS) data are described in the literature that provide specific operations as building blocks for the assembly of purposespecific workflows. Working out which tools and combinations are applicable or optimal is often hard: insufficient annotation of tool functions and interfaces impedes finding viable tool combinations, and potentially compatible tools may not, in practice, operate together. Thus researchers face difficulties in selecting practical and effective data analysis pipelines for a specific experimental design.
Background: Bioinformaticians routinely use multiple software tools and data sources in their day-to-day work and have been guided in their choices by a number of cataloguing initiatives. The ELIXIR Tools and Data Services Registry (bio.tools) aims to provide a central information point, independent of any specific scientific scope within bioinformatics or technological implementation. Meanwhile, efforts to integrate bioinformatics software in workbench and workflow environments have accelerated to enable the design, automation, and reproducibility of bioinformatics experiments. One such popular environment is the Galaxy framework, with currently more than 80 publicly available Galaxy servers around the world. In the context of a generic registry for bioinformatics software, such as bio.tools, Galaxy instances constitute a major source of valuable content. Yet there has been, to date, no convenient mechanism to register such services en masse. Findings: We present ReGaTE (Registration of Galaxy Tools in Elixir), a software utility that automates the process of registering the services available in a Galaxy instance. This utility uses the BioBlend application program interface to extract service metadata from a Galaxy server, enhance the metadata with the scientific information required by bio.tools, and push it to the registry. Conclusions: ReGaTE provides a fast and convenient way to publish Galaxy services in bio.tools. By doing so, service providers may increase the visibility of their services while enriching the software discovery function that bio.tools provides for its users. The source code of ReGaTE is freely available on Github at https://github.com/C3BI-pasteur-fr/ReGaTE.
Søren Brunak合作论文数Rigshospitalet;Novo Nordisk Foundation Center for Protein Research, University of Copenhagen;Department of Systems Biology, Technical University of Denmark4