
The integration of grid, cloud and other e-infrastructures into the fields of biology, bioinformatics, biomedicine, and healthcare are crucial if optimum use is to be made of the latest high-performance and distributed computer technology in these areas. Science gateways are concerned with offering intuitive graphical user interfaces to applications, data, and tools on distributed computing infrastructures. This book presents the joint proceedings of the Tenth HealthGrid Conference and the Fourth International Workshop on Science Gateways for Life Sciences (IWSG-Life), held in Amsterdam, the Netherlands in May 2012. The HealthGrid conference promotes the exchange and debate of ideas, technologies and solutions likely to promote the integration of grids into biomedical research and health in the broadest sense. The IWSG-Life workshop series is a forum that brings together scientists from the field of life sciences, bioinformatics, and computer science to advance computational biology and chemistry in the context of science gateways. These events have been jointly organized to maximize the benefit from synergies and stimulate the forging of further links in joint research areas. HealthGrid Applications and Technologies Meet Science Gateways for Life Sciences is divided into three parts. Part I includes contributions accepted to the HealthGrid conference; Part II contains the papers about various aspects of the development and usage of science gateways for life sciences. The joint session is recorded in Part III, and addresses the topic of science gateways for biomedical research. This book will provide insights and new perspectives for all those involved in the research and use of infrastructures and technology for healthcare and life sciences.
In the proposed demonstration we will present DCV (Desktop Cloud Visualization): a unique technology that allows users to remote access 2D and 3D interactive applications over a standard network. This allows geographically dispersed doctors work collaboratively and to acquire anatomical or pathological images and visualize them for further investigations.
One of the important questions in biological evolution is to know if certain changes along protein coding genes have contributed to the adaptation of species. This problem is known to be biologically complex and computationally very expensive. It, therefore, requires efficient Grid or cluster solutions to overcome the computational challenge. We have developed a Grid-enabled tool (gcodeml) that relies on the PAML (codeml) package to help analyse large phylogenetic datasets on both Grids and computational clusters. Although we report on results for gcodeml, our approach is applicable and customisable to related problems in biology or other scientific domains.
This paper presents a study of the performance of federated queries implemented in a system that simulates the architecture proposed for the Scalable Architecture for Federated Translational Inquiries Network (SAFTINet). Performance tests were conducted using both physical hardware and virtual machines within the test laboratory of the Center for High Performance Computing at the University of Utah. Tests were performed on SAFTINet networks ranging from 4 to 32 nodes with databases containing synthetic data for several million patients. The results show that the caGrid FQE (Federated Query Engine) is capable and suitable for comparative effectiveness research (CER) federated queries given its nearly linear scalability as partner nodes increase in number. The results presented here are also important for the specification of the hardware required to run a CER grid.
Massively-parallel sequencing (MPS) technologies and their diverse applications in genomics and epigenomics research have yielded enormous new insights into the physiology and pathophysiology of the human genome. The biggest hurdle remains the magnitude and diversity of the datasets generated, compromising our ability to manage, organize, process and ultimately analyse data. The Wiki-based Automated Sequence Processor (WASP), developed at the Albert Einstein College of Medicine (hereafter Einstein), uniquely manages to tightly couple the sequencing platform, the sequencing assay, sample metadata and the automated workflows deployed on a heterogeneous high performance computing cluster infrastructure that yield sequenced, quality-controlled and 'mapped' sequence data, all within the one operating environment accessible by a web-based GUI interface. WASP at Einstein processes 4-6 TB of data per week and since its production cycle commenced it has processed ~ 1 PB of data overall and has revolutionized user interactivity with these new genomic technologies, who remain blissfully unaware of the data storage, management and most importantly processing services they request. The abstraction of such computational complexity for the user in effect makes WASP an ideal middleware solution, and an appropriate basis for the development of a grid-enabled resource - the Einstein Genome Gateway - as part of the Extreme Science and Engineering Discovery Environment (XSEDE) program. In this paper we discuss the existing WASP system, its proposed middleware role, and its planned interaction with XSEDE to form the Einstein Genome Gateway.
Computational neuroscience is a new field of research in which neurodegenerative diseases are studied with the aid of new imaging techniques and computation facilities. Researchers with different expertise collaborate in these studies. A study requires scalable computational and storage capacity and information management facilities to succeed. Many virtual laboratories are proposed and developed to facilitate these studies, however most of them cover only the parts related to the computational data processing. In this paper we describe and analyse the phases of the computational neuroscience studies including the actors, the tasks they perform, and the characteristics of each phase. Based on these we identify the required properties and functionalities of a virtual laboratory that supports the actors and their tasks throughout the complete study.
Production operation of large distributed computing infrastructures (DCI) still requires a lot of human intervention to reach acceptable quality of service. This may be achievable for scientific communities with solid IT support, but it remains a show-stopper for others. Some application execution environments are used to hide runtime technical issues from end users. But they mostly aim at fault-tolerance rather than incident resolution, and their operation still requires substantial manpower. A longer-term support activity is thus needed to ensure sustained quality of service for Virtual Organisations (VO). This paper describes how the biomed VO has addressed this challenge by setting up a technical support team. Its organisation, tooling, daily tasks, and procedures are described. Results are shown in terms of resource usage by end users, amount of reported incidents, and developed software tools. Based on our experience, we suggest ways to measure the impact of the technical support, perspectives to decrease its human cost and make it more community-specific.
Interactive visualization and correction of intermediate results are required in many medical image analysis pipelines. To allow certain interaction in the remote execution of compute- and data-intensive applications, new features of HTML5 are used. They allow for transparent integration of user interaction into Grid- or Cloud-enabled scientific workflows. Both 2D and 3D visualization and data manipulation can be performed through a scientific gateway without the need to install specific software or web browser plugins. The possibilities of web-based visualization are presented along the FreeSurfer-pipeline, a popular compute- and data-intensive software tool for quantitative neuroimaging.
Notwithstanding the benefits of distributed-computing infrastructures for empowering bioinformatics analysis tools with the needed computing and storage capability, the actual use of these infrastructures is still low. Learning curves and deployment difficulties have reduced the impact on the wide research community. This article presents a porting strategy of BLAST based on a multiplatform client and a service that provides the same interface as sequential BLAST, thus reducing learning curve and with minimal impact on their integration on existing workflows. The porting has been done using the execution and data access components from the EC project Venus-C and the Windows Azure infrastructure provided in this project. The results obtained demonstrate a low overhead on the global execution framework and reasonable speed-up and cost-efficiency with respect to a sequential version.
The discovery of knowledge from raw data is a multistage process, that typical requires collaboration between experts from disparate disciplines, and the application of a range of methods tailored to the research question. The aim of the eLab is to provide a web-based environment for health professionals and researchers to access health datasets, share knowledge and expertise and to apply methods for analysis and visualization of the results. The eLab is built around the core concept of the Research Object as the mechanism for preserving, reusing and disseminating the knowledge discovery process. The possible range of applications of the eLab is vast, and so the consideration of the trade off between specificity and generality is an important one, that is reflected in the requirements. The architecture and implementation of the eLab is described, and we report on the deployment of eLabs for applications in primary care, long-term conditions management, bariatric surgery and public health.
We report on the implementation of a software suite dedicated to the management and analysis of large scale RNAi High Content Screening (HCS). We describe the requirements identified amongst our different users, the supported data flow, and the implemented software. Our system is already supporting productively three different laboratories operating in distinct IT infrastructures. The system was already used to analyze hundreds of RNAi HCS plates.
Neuroimaging is a field that benefits from distributed computing infrastructures (DCIs) to perform data- and compute-intensive processing and analysis. Using grid workflow systems not only automates the processing pipelines, but also enables domain researchers to implement their expertise on how to best process neuroimaging data. To share this expertise and to promote collaborative research in neurosciences, ways to facilitate the exchange, re-use, and interoperability of workflow applications between different groups are required. The SHIWA project (SHaring Interoperable Workflows for large-scale scientific simulations on Available DCIs) is specifically addressing such use-cases, building a generic platform to facilitate workflow exchange and execution environments interoperability. The goal is to facilitate the dissemination and execution of workflows by diverse workflow management systems on multiple DCIs. This platform enables researchers to gain access to a variety of ready-to-use workflows, to reuse workflows developed by collaborators, to publish their own workflows to be used by others, and to use additional resources from external DCIs to run workflows. In this demonstration we present how the SHIWA platform is used to implement various usage scenarios in which workflow exchange supports collaboration in neuroscience. The SHIWA platform and the implemented solutions are presented from the user perspective, in this case the workflow developers and the neuroscientists. These workflow interoperability solutions aim to facilitate and enable more advanced and large-scale research in neuroscience. The demonstration will focus on usage scenarios currently employed to exchange, combine and interoperate neuroimaging workflows between Academic Medical Center, Amsterdam, the Charite Universittsmedizin, Berlin and the outGrid infrastructure. These workflows are developed for the analysis of neuroimaging data, in particular Magnetic Resonance Images (MRI) and Diffusion Tensor Imaging (DTI). Each group has ported workflows with complementary and overlapping functions to its Grid infrastructure using different workflow systems, so the goal is to combine and share them across the boundaries of the original DCIs. The following usage scenarios will be addressed in the demonstration: • Preparing the workflow for sharing with others (VO, workflow management system dependencies); • Using the SHIWA repository for publishing workflows to share a new workflow with other potential users; • Finding a workflow in the SHIWA repository and testing it with sample data; • Running the workflow found in the repository with own data using the SHIWA simulation platform; • Combining complementary workflows into a meta-workflow to implement additional functionality or combining different implementations of the same workflow to compute on different DCIs simultaneously The shown solution includes neuroimaging workflows using the GWES, MOTEUR, LONI pipeline and P-GRADE workflow engines submitting jobs to the German MediGRID, the Dutch BiG Grid, the European EGI, and the international outGrid infrastructures. Data to be accessed might be stored locally, on an iRODS data management system, in the LFC file catalog and on gridFTP-enabled sites. The user-interfaces are web-based and include a Liferay-based Grid portal and a P-GRADE workflow editor implemented as webstart application. The neuroimaging applications include self-developed tools for preprocessing DTI data and widely used methods from the ITK and the FSL toolboxes
European laws on privacy and data security are not explicit about the storage and processing of genetic data. Especially whole-genome data is identifying and contains a lot of personal information. Is processing of such data allowed in computing grids? To find out, we looked at legal precedents in related fields, current literature, and interviews with legal experts. We found that processing of genetic data is only allowed on distributed systems with specific security measures, both technical and organizational. Informed consent, although important, offers no substitute for such requirements.
The new science gateway MoSGrid (Molecular Simulation Grid) enables users to submit and process molecular simulation studies on a large scale. A conformational analysis of guanidine zinc complexes, which are active catalysts in the ring-opening polymerization of lactide, is presented as an example. Such a large-scale quantum chemical study is enabled by workflow technologies. Two times 40 conformers have been generated, for two guanidine zinc complexes. Their structures were optimized using Gaussian03 and the energies processed within the quantum chemistry portlet of the MoSGrid portal. All meta- and post-processing steps have been performed in this portlet. All workflow features are implemented via WS-PGRADE and submitted to UNICORE.
The Collaborative Computing Project for NMR (CCPN) has build a software framework consisting of the CCPN data model (with APIs) for NMR related data, the CcpNmr Analysis program and additional tools like CcpNmr FormatConverter. The open architecture allows for the integration of external software to extend the abilities of the CCPN framework with additional calculation methods. Recently, we have carried out the first steps for integrating our software Computer Simulation of Molecular Structures (COSMOS) into the CCPN framework. The COSMOS-NMR force field unites quantum chemical routines for the calculation of molecular properties with a molecular mechanics force field yielding the relative molecular energies. COSMOS-NMR allows introducing NMR parameters as constraints into molecular mechanics calculations. The resulting infrastructure will be made available for the NMR community. As a first application we have tested the evaluation of calculated protein structures using COSMOS-derived 13C Cα and Cβ chemical shifts. In this paper we give an overview of the methodology and a roadmap for future developments and applications.
In this paper we present the architecture of a framework for building Science Gateways supporting official standards both for user authentication and authorization and for middleware-independent job and data management. Two use cases of the customization of the Science Gateway framework for Semantic-Web-based life science applications are also described.