Computational workflows represent major investments of effort and expertise. As first-class, publishable research objects of their own, they are key to sharing methodological know-how for reuse, reproducibility, and transparency. Thus, the application of the FAIR Principles to workflows is inevitable to enable them to be Findable, Accessible, Interoperable, and Reusable. Making workflows FAIR reduces duplication of effort, assists in the reuse of best practice approaches and community-supported standards, and ensures that workflows as digital objects can support reproducible, robust science. FAIR workflows draw from both FAIR data and software principles, and they help ensure and support data FAIRification. The FAIR Principles emphasize the association of persistent identifiers and machine-actionable metadata with workflows. Implementing the Principles requires a framework with appropriate programmatic protocols and an accompanying ecosystem of services, tools, policies, and best practices, as well the buy-in of existing workflow systems. The European EOSC-Life Workflow Collaboratory is an example of such a digital infrastructure for the Biosciences. It includes a metadata standards framework for describing workflows that is managed and used by dedicated new FAIR workflow services and programmatic APIs for interoperability and metadata access. It includes the WorkflowHub registry and LifeMonitor workflow testing service, and it incorporates existing workflow systems and packaging solutions. Here, we introduce the FAIR Principles for workflows and connect FAIR workflows with the FAIR ecosystems they inhabit with the EOSC-Life Collaboratory as a concrete example. We also introduce other community efforts that are easing the ways that workflows are shared and reused by others, and we discuss how the variations in different workflow settings impact their FAIR perspectives.
1 Department of Computer Science, The University of Manchester, Manchester, United Kingdom, 2 Melandra Limited, Stockport, United Kingdom, 3 Informatics Institute, University of Amsterdam, Amsterdam, The Netherlands, 4 Research IT, IT Services, University of Manchester, Manchester, United Kingdom, 5 Meise Botanic Garden, Meise, Belgium, 6 Department of Plant Biotechnology and Bioinformatics, Ghent University, Ghent, Belgium, 7 VIB Center for Plant Systems Biology, Ghent, Belgium, 8 Bioinformatics Group, Department of Computer Science, Albert-Ludwigs-University Freiburg, Freiburg, Germany, 9 Science for Life Laboratory (SciLifeLab), Department of Biochemistry and Biophysics, Stockholm University, Stockholm, Sweden
The WorkflowHub (workflowhub.eu) is a FAIR workflow registry sponsored by the European RI Cluster EOSC-Life and the European Research Infrastructure ELIXIR. It is workflow management system agnostic: workflows may remain in their native repositories in their native forms. As workflows are multi-component objects, including example and test data, they are packaged, registered, downloaded and exchanged as workflow centric Research Objects using the RO-Crate specification, making the Hub an implementation of FAIR Digital Object principles. A schema.org based Bioschemas profile describes the metadata about a workflow and encouraged use of the Common Workflow Language is provides a canonical description of the workflow itself. Workflow management systems such as Galaxy, Nextflow, and snakemake are working with the Hub to seamlessly and automatically support object packaging, registration and exchange. The WorkflowHub provides features such as community spaces, collections, versioning and snapshots, and contributor credit. It supports community registry standards and services such as GA4GH TRS and ELIXIR-AAI, and integrates with the LifeMonitor workflow testing service.
There are many hypotheses about where the Indigenous ancestors first settled in Australia tens of thousands of years ago, but evidence is scarce. Few archaeological sites date to these early times. Sea levels were much lower and Australia was connected to New Guinea and Tasmania in a land known as Sahul that was 30% bigger than Australia is today. Our latest research advances our knowledge about the most likely routes those early Australians travelled as they peopled this giant continent.
This report reviews the current state-of-the-art applied approaches on automated tools, services and workflows for extracting information from images of natural history specimens and their labels. We consider the potential for repurposing existing tools, including workflow management systems; and areas where more development is required. This paper was written as part of the SYNTHESYS+ project for software development teams and informatics teams working on new software-based approaches to improve mass digitisation of natural history specimens.
FAIRDOM (Findable, Accessible, Interoperable, Reusable Data, Operating procedures and Models) is an initiative to establish sustained data, model and process management service to the European Systems Biology community ( https://fair-dom.org/ ). With the free SEEK ( https://seek4science.org/ ) data management software FAIRDOM offers a data management platform for interdisciplinary projects to support the storage and exchange of data and models from research partners based on the FAIR principles. SEEK can be installed, run and further developed as its own instance, and over 50 organisations have done that including many in ELIXIR Nodes. It is also the platform used for the web-accessible public Commons platform FAIRDOMHub ( https://fairdomhub.org/ ), offering public information and password protected user collaboration spaces. The Hub is a managed service hosted at HITS (ELIXIR-DE) and currently supports 185 projects.
The microbial production of fine chemicals provides a promising biosustainable manufacturing solution that has led to the successful production of a growing catalog of natural products and high-value chemicals. However, development at industrial levels has been hindered by the large resource investments required. Here we present an integrated Design–Build-Test–Learn (DBTL) pipeline for the discovery and optimization of biosynthetic pathways, which is designed to be compound agnostic and automated throughout. We initially applied the pipeline for the production of the flavonoid (2 S )-pinocembrin in Escherichia coli , to demonstrate rapid iterative DBTL cycling with automation at every stage. In this case, application of two DBTL cycles successfully established a production pathway improved by 500-fold, with competitive titers up to 88 mg L −1 . The further application of the pipeline to optimize an alkaloids pathway demonstrates how it could facilitate the rapid optimization of microbial strains for production of any chemical compound of interest.
Synthetic biology is typified by developing novel genetic constructs from the assembly of reusable synthetic DNA parts, which contain one or more features such as promoters, ribosome binding sites, coding sequences and terminators. While repositories of such parts exist to promote their reuse, there is still a need to design novel parts from scratch. PartsGenie, freely available at http://parts.synbiochem.co.uk , is introduced to facilitate the computational design of such synthetic biology parts. PartsGenie has been designed to bridge the gap between optimisation tools for the design of novel parts, the representation of such parts in community-developed data standards such as Synthetic Biology Open Language (SBOL), and their sharing in journal-recommended data repositories. Consisting of a drag-and-drop web interface, a number of DNA optimisation algorithms, and an interface to the well-used data repository JBEI ICE, PartsGenie facilitates the design, optimisation and dissemination of reusable synthetic biology parts through a single, integrated application. PartsGenie can therefore be used as a single, stand-alone tool, or integrated into larger synthetic biology pipelines that are being developed in the SYNBIOCHEM centre and elsewhere.
FAIR Bioinformatics computation and data management: FAIRDOM and the Norwegian Digital Life initiative . Abstract from NETTAB 2018 Network Tools and Applications in Biology , Genova, Italy.
Studies of cumulative and long-term effects of human activities in the ocean are essential for developing realistic conservation targets. Here, we report the results of a recent national marine biodiversity inventory along the Swedish West coast between 2004 and 2009. The expedition revisited many historical localities that have been sampled with the same methods in the early twentieth century. We generated comparable datasets from our own investigation and the historical data to compare species richness, abundance, and geographic distribution of diversity. Our analysis indicates that the benthic ecosystems in the region have lost a large part of its original species richness over the last seven decades. We find evidence that especially rare species have disappeared. This process has caused a more homogenized community structure in the region and diminished historical biodiversity hotspots. We argue that the contemporary lack of rare species in the benthic ecosystems of the Kattegat and Skagerrak offers less opportunity to respond to environmental perturbations in the future and suggest improving the poor representation of rare species in the region. The study shows the value of biodiversity inventories as well as natural history collections in investigations of accumulated effects of anthropogenic activities and for re-establishing species-rich, productive, and resilient ecosystems.
In many disciplines, data are highly decentralized across thousands of online databases (repositories, registries, and knowledgebases). Wringing value from such databases depends on the discipline of data science and on the humble bricks and mortar that make integration possible; identifiers are a core component of this integration infrastructure. Drawing on our experience and on work by other groups, we outline 10 lessons we have learned about the identifier qualities and best practices that facilitate large-scale data integration. Specifically, we propose actions that identifier practitioners (database providers) should take in the design, provision and reuse of identifiers. We also outline the important considerations for those referencing identifiers in various circumstances, including by authors and data generators. While the importance and relevance of each lesson will vary by context, there is a need for increased awareness about how to avoid and manage common identifier problems, especially those related to persistence and web-accessibility/resolvability. We focus strongly on web-based identifiers in the life sciences; however, the principles are broadly relevant to other disciplines.
Biologists and biochemists have at their disposal a number of excellent, publicly available data resources such as UniProt, KEGG, and NCBI Taxonomy, which catalogue biological entities. Despite the usefulness of these resources, they remain fundamentally unconnected. While links may appear between entries across these databases, users are typically only able to follow such links by manual browsing or through specialised workflows. Although many of the resources provide web-service interfaces for computational access, performing federated queries across databases remains a non-trivial but essential activity in interdisciplinary systems and synthetic biology programmes. What is needed are integrated repositories to catalogue both biological entities and-crucially-the relationships between them. Such a resource should be extensible, such that newly discovered relationships-for example, those between novel, synthetic enzymes and non-natural products-can be added over time. With the introduction of graph databases, the barrier to the rapid generation, extension and querying of such a resource has been lowered considerably. With a particular focus on metabolic engineering as an illustrative application domain, biochem4j, freely available at http://biochem4j.org, is introduced to provide an integrated, queryable database that warehouses chemical, reaction, enzyme and taxonomic data from a range of reliable resources. The biochem4j framework establishes a starting point for the flexible integration and exploitation of an ever-wider range of biological data sources, from public databases to laboratory-specific experimental datasets, for the benefit of systems biologists, biosystems engineers and the wider community of molecular biologists and biological chemists.
The FAIRDOMHub is a repository for publishing FAIR (Findable, Accessible, Interoperable and Reusable) Data, Operating procedures and Models (https://fairdomhub.org/) for the Systems Biology community. It is a web-accessible repository for storing and sharing systems biology research assets. It enables researchers to organize, share and publish data, models and protocols, interlink them in the context of the systems biology investigations that produced them, and to interrogate them via API interfaces. By using the FAIRDOMHub, researchers can achieve more effective exchange with geographically distributed collaborators during projects, ensure results are sustained and preserved and generate reproducible publications that adhere to the FAIR guiding principles of data stewardship.
Ecological niche modelling (ENM) Components are a set of reusable workflow components specialized for performing ENM tasks within the Taverna workflow management system. Each component encapsulates specific functionality and can be combined with other components to facilitate the creation of larger and more complex workflows. One key distinguishing feature of ENM Components is that most tasks are performed remotely by calling web services, simplifying software setup and maintenance on the client side and allowing more powerful computing resources to be exploited. This paper presents the current set of ENM Components in the context of the Taverna family of tools for creating, publishing and sharing workflows. An example is included showing how the components can be used in a preliminary investigation of the effects of mixing different spatial resolutions in ENM experiments.
Katy Wolstencroft合作论文数The University of Manchester, School of Computer Science, Manchester, United Kingdom6