Participation in online activities has become highly popular in contemporary society. There is evidence that the collection of technologies captured under the banner of Web 2.0 is changing the way we do research, facilitating communications between scientists and encouraging collaboration through sharing data. Research 2.0 (or Science 2.0) is the term commonly used to describe a platform, powered by Web 2.0 tools to support scientific and other research, for providing open access to information, allowing peer review, and the pooling of collective intelligence. myExperiment and Open Wet Ware are two important examples of such a platform.
Automation in science is increasingly marked by the use of workflow technology. The sharing of workflows through repositories supports the verifiability, reproducibility and extensibility of computational experiments. However, the subsequent discovery of workflows remains a challenge, both from a sociological and technological viewpoint. Based on a survey with participants from 19 laboratories, we investigate the current practices in workflow sharing, re‐use and discovery among life scientists chiefly using the Taverna workflow management system. To address their perceived lack of effective workflow discovery tools, we go on to develop benchmarks for the evaluation of discovery tools, drawing on a series of practical exercises. We demonstrate the value of the benchmarks on two tools: one using graph matching and the other relying on text clustering. Copyright © 2009 John Wiley & Sons, Ltd.
A model of computation (MoC) is a formal abstraction of execution in a computer. There is a need for composing diverse MoCs in e-science. Kepler, which is based on Ptolemy II, is a scientific workflow environment that allows for MoC composition. This paper explains how MoCs are combined in Kepler and Ptolemy II and analyzes which combinations of MoCs are currently possible and useful. It demonstrates the approach by combining MoCs involving dataflow and finite state machines. The resulting classification should be relevant to other workflow environments wishing to combine multiple MoCs (available at http://ptolemy.org/heterogeneousMoCs).
The myExperiment virtual research environment supports the sharing of research objects used by scientists, such as scientific workflows. For researchers it is both a social infrastructure that encourages sharing and a platform for conducting research, through familiar user interfaces. For developers it provides an open, extensible and participative environment. We describe the design, implementation and deployment of myExperiment and suggest that its four capabilities - research objects, social model, open environment and actioning research - are necessary characteristics of an effective Virtual Research Environment for e-research and open science.
Automation in science is increasingly marked by the use of workflow technology. The sharing of workflows through publication mechanisms or repositories supports the verifiability, reproducibility and extensibility of computational experiments. However, the subsequent discovery of workflows remains a challenge, both from a technological and sociological viewpoint. We investigate current practices in workflow sharing, re-use and discovery amongst life scientists chiefly using the Taverna workflow management system. The study draws on two key sources: (i) a survey of researchers drawn from 19 research labs and (ii) an analysis of scientists’ behaviour on the myExperiment social network site, designed to encourage workflow exchange. The results reveal a multi-modal approach to workflow discovery, based on a mix of search on the content of the workflow and its situated context. We go on to develop a benchmark specifically for the evaluation of workflow discovery and to demonstrate it on two example approaches.
Much has been written on the promise of Web service discovery and (semi-) automated composition. In this discussion, the value to practitioners of discovering and reusing existing service compositions, captured in workflows, is mostly ignored. We present the case for workflows and workflow discovery in science and develop one discovery solution. Through a survey with 21 scientists and developers from the myGrid/Taverna workflow environment, workflow discovery requirements are elicited. Through a user experiment with 13 scientists, an attempt is made to build a benchmark for workflow ranking. Through the design and implementation of a workflow discovery tool, a mechanism for ranking workflow fragments is provided based on graph sub-isomorphism detection. The tool evaluation, drawing on a corpus of 89 public workflows and the results of the user experiment, finds that, for a simple showcase, the average human ranking can largely be reproduced.
Scientific workflows are becoming a valuable tool for scientists to capture and automate e‐Science procedures. Their success brings the opportunity to publish, share, reuse and re‐purpose this explicitly captured knowledge. Within the $^{my}$ Grid project, we have identified key resources that can be shared including complete workflows, fragments of workflows and constituent services. We have examined the alternative ways that these resources can be described by their authors (and subsequent users) and developed a unified descriptive model to support their later discovery. By basing this model on existing standards, we have been able to extend existing Web service and Semantic Web service infrastructure whilst still supporting the specific needs of the e‐Scientist. The $^{my}$ Grid components enable a workflow life‐cycle that extends beyond execution to include the discovery of previous relevant designs, the reuse of those designs and their subsequent publication. Experience with example groups of scientists indicates that this cycle is valuable. The growing number of workflows and services mean more work is needed to support the user in effective ranking of search results and to support the re‐purposing process. Copyright © 2006 John Wiley & Sons, Ltd.
A model of computation (MoC) is a formal abstraction of execution in a computer. There is a need for composing MoCs in e-science. Kepler, which is based on Ptolemy II, is a scientific workflow environment that allows for MoC composition. This paper explains how MoCs are combined in Kepler and Ptolemy II and analyzes which combinations of MoCs are currently possible and useful. It demonstrates the approach by combining MoCs involving dataflow and finite state machines. The resulting classification should be relevant to other workflow environments wishing to combine multiple MoCs.
Much has been written on the promise of Web service discovery and (semi-) automated composition. In this discussion, the value to practitioners of discovering and reusing existing service compositions, captured in workflows, is mostly ignored. This paper presents one solution to workflow discovery. Through a survey with 21 scientists and developers from the myGrid workflow environment, workflow discovery requirements are elicited. Through a user experiment with 13 scientists, an attempt is made to build a gold standard for workflow ranking. Through the design and implementation of a workflow discovery tool, a mechanism for ranking workflow fragments is provided based on graph sub-isomorphism matching. The tool evaluation, drawing on a corpus of 89 public workflows from bioinformatics and the results of the user experiment, finds that the average human ranking can largely be reproduced
Life sciences research is based on individuals, often with diverse skills, assembled into research groups. These groups use their specialist expertise to address scientific problems. The in silico experiments undertaken by these research groups can be represented as workflows involving the co-ordinated use of analysis programs and information repositories that may be globally distributed. With regards to Grid computing, the requirements relate to the sharing of analysis and information resources rather than sharing computational power. The myGrid project has developed the Taverna Workbench for the composition and execution of workflows for the life sciences community. This experience paper describes lessons learnt during the development of Taverna. A common theme is the importance of understanding how workflows fit into the scientists' experimental context. The lessons reflect an evolving understanding of life scientists' requirements on a workflow environment, which is relevant to other areas of data intensive and exploratory science. Copyright © 2005 John Wiley & Sons, Ltd.
To date on-line processes (i.e. workflows) built in e-Science have been the result of collaborative team efforts. As more of these workflows are built, scientists start sharing and reusing stand-alone compositions of services, or work-flow fragments. They repurpose an existing workflow or workflow fragment by finding one that is close enough to be the basis of a new workflow for a different purpose, and making small changes to it. Such a "workflow by example" approach complements the popular view in the Semantic Web Services literature that on-line processes are constructed automatically from scratch, and could help bootstrap the Web of Science. Based on a comparison of e-Science middleware projects, this paper identifies seven bottlenecks to scalable reuse and repurposing. We include some thoughts on the applicability of using OWL for two bottlenecks: workflow fragment discovery and the ranking of fragments.
1 Reuse and repurposing in e-Science Workflow techniques are an important part of in silico experimentation, potentially allowing a scientist to describe and enact their experimental processes in a structured, repeatable and verifiable way. The my Grid (www.mygrid.org.uk) workbench, a set of components to build workflows in bioinformatics, currently allows access to a thousand globally distributed services and a hundred work-flows, some of which orchestrate up to fifty services. Figure 2 shows the example of a my Grid workflow which gathers information about genetic sequences in support of research on Williams-Beuren syndrome [10]. Much of the research geared towards the construction of on-line processes (i.e. workflows) is led by a vision of automatic composition of services based on extensive formalisation (see for example www.daml.org/services/owl-s/pub-archive.html). Such research can be complemented with techniques that exploit those cases where existing workflows and fragments of workflows can be reused, thereby benefitting from hard-won human experience in composing services. A workflow fragment is a piece of an experimental description that is a coherent sub-workflow that makes sense to a domain specialist. Each fragment forms a useful resource in its own right and is identified and annotated at publication time. We distinguish between reuse, where workflows and workflow fragments created by one user might be used as is, and repurposing, where they are used as a starting point by others. The idea of repurposing is that a user looks for workflows that are close enough to the user's requirements so that these workflows can be fit to a new purpose. In Figure 1, we show the lifecycle of a repurposed workflow. 1. Before embarking on a new design the scientist consults a registry of existing workflows. Search facilities based on an ontology and a database repository identify any existing workflows that are relevant to them. 2. Workflows or their fragments are potentially edited; services are parame-terised or bound to end points but rarely altered. Other services, workflows or workflow fragments are sought, or new ones are created.
Abstract The first part of this thesis provides a comparative overview of UPML, DAML-S and F-X, three approaches to describe service capability on the Web. In our overview, we have special attention for the reasoning support offered by UPML, because to date it has not received much attention in the Web services literature. In the second part, we use our insights from this comparison to add functionality to F-X. In particular, we facilitate the use of domain ontologies in F-X, and we exploit such domain knowledge to support more flexible competence matching. We achieve this based on query relaxation techniques and a recent translation from a subset of Description Logics to Logic Programs. 1 Acknowledgements Many thanks to Dave Robertson and Stephen Pot-
Bioinformatics is a discipline that uses computational and mathematical techniques to store, manage, and analyze biological data in order to answer biological questions. Bioinformatics has over 850 databases [154] and numerous tools that work over those databases and local data to produce even more data themselves. In order to perform an analysis, a bioinformatician uses one or more of these resources to gather, filter, and transform data to answer a question. Thus, bioinformatics is an in silico science.
repurposing. In this abstract, we provide a partial answer by speculating on how, in future, experiments from the myGrid project could be cast ,within the wider setting of the experimental method. In a parallel All Hands
Duncan Hull合作论文数School of Computer Science, Manchester in the UK4
Pinar Alper合作论文数School of Computer Science
Kilburn Building
University of Manchester4
Daniele Turi合作论文数School of Computer Science The University of Manchester4
Katy Wolstencroft合作论文数The University of Manchester, School of Computer Science, Manchester, United Kingdom3
Simon Miles合作论文数Aerogility1