
In this paper, we introduce Apache Airavata, a software framework to compose, manage, execute, and monitor distributed applications and workflows on computational resources ranging from local resources to computational grids and clouds. Airavata builds on general concepts of service-oriented computing, distributed messaging, and workflow composition and orchestration. This paper discusses the architecture of Airavata and its modules, and illustrates how the software can be used as individual components or as an integrated solution to build science gateways or general-purpose distributed application and workflow management systems.
This paper introduces a gadget-based web 2.0 architecture for building scientific and educational tools. This architecture builds on ideas from both Google Web Toolkit (GWT) and Google gadget and adds AJAX functionality. The gadgets developed with GWT are hosted in the Open Gateway Computing Environments (OGCE) container. The community of scientific applications and science gateways developers face the challenges of rapid development of scientific and educational tools for their users. This web application development approach allows developers to create a gadget graphical user interface within hours and, thus, can potentially expand the variety of OpenSocial Gadgets. A gadget that has been developed using our gadget-based architecture for online Community Earth System Model (CESM) simulations is presented as a use case in this paper.
Science gateways broaden and simplify access to cyberinfrastructure (CI) by providing advanced interfaces to collaboration, analysis, data management, and other tools for students and researchers. As these science gateway interfaces to cyberinfrastructure grow in popularity, web portal developers adopt ad hoc approaches to the security challenges of authentication, authorization, and delegation. Science gateways integrate cyberinfrastructure resources on the researcher's behalf, i.e., accessing data, compute cycles, instruments, and other valuable resources. Resource access often requires use of the researcher's security credentials, in some cases exposing the researcher's long-lived password to potential compromise at the science gateway. There is no standard approach for a researcher to control and limit a science gateway's access to his or her resources. Thus, researchers are required to accept unnecessary risks when using science gateways. The "Distributed Web Security for Science Gateways" project is addressing these risks by providing authorization and delegation software for science gateways that complies with the Internet Engineering Task Force's standard OAuth protocol. The project is developing an OAuth server implementation and a set of client libraries and authentication modules to enable out of the box integration with common Web platforms, in coordination with gateways and cyberinfrastructure providers. In this paper, we introduce the project, including our planned software architecture.
This paper describes the NSF TeraGrid Science Gateways program, its formation, progress, lessons learned and current contributions over its seven-year life and new directions in the NSF XSEDE program. Early requirements analysis work with path-finding gateways that formed the underpinning of the program are described as are current projects and their unique contributions to the larger program. Future directions both within the XSEDE program and for gateways more generally are discussed.
Managing the growing volume of data being output by radio telescopes is a significant challenge faced by radio astronomers today. This challenge will only be further compounded with future telescopes such as the Square Kilometre Array (SKA), which will be the world's largest radio telescope when completed and produce data at unprecedented rates. This paper introduces the CyberSKA collaborative portal which is aimed at addressing the current and future needs of data-intensive radio astronomy. A wide variety of tools and services that have been developed and integrated with the CyberSKA portal, including a distributed data management system, a data access tool, remote visualization tools and a third party application interface are described. Current international usage of CyberSKA focusing on several different SKA Pathfinder survey projects and how they make use of the portal are also highlighted.
Science gateways enable researchers and students to use distributed scientific computing infrastructure (cyberinfrastructure) through Web browsers and Web-enabled desktop clients. This paper describes the use of the open source, open community Apache Rave project as the basis for developing science gateways. Building on Apache Shindig (for OpenSocial Gadgets) and Apache Wookie (for W3C Widgets), Rave provides an out-of-the box deployment that can be used to host reusable social Web components. Rave is based on the Spring MVC framework and so can also be extensively customized or extended with (for example) custom database back-ends and authentication modules. In this paper we consider Rave as a development platform for science gateways and discuss how the source code may be extended through three use cases that focus on gateway security requirements. A major consideration of this paper is how to design Rave as a development environment so that developers can make local customizations and extensions freely on both a rapidly changing code base (during Rave's initial development), and (later) between stable code bases during version upgrades. We conclude with a discussion of the implications of developing science gateways and other cyberinfrastructure software within the Apache Software Foundation and present its potential advantages.
The iPlant Collaborative is an NSF-funded cyberinfrastructure (CI) effort directed towards the plant sciences community. This paper enumerates the key concepts, middleware, tools, and extensions that create the unique capabilities of the iPlant Discovery Environment (DE) that provide access to our CI. The DE is a rich web-based application that brings flexible CI capabilities to a wide audience affiliated with the plant sciences, from computational biologists, bioinformaticians, applications developers, to bench biologists. The inherent interdisciplinary nature of plant sciences research produces diverse and complex data products that range from molecular sequences to satellite imagery as part of the discovery life cycle. With the constant creation of novel analysis algorithms, the advent and spread of large data repositories, and the need for collaborative data analysis, marshaling resources to effectively utilize these capabilities necessitates a highly flexible and scalable approach for implementing underlying CI. The iPlant infrastructure simultaneously supports multiple interdisciplinary projects providing essential features found in traditional science gateways as well as highly customized direct access to its underlying frameworks through use of APIs (Application Programming Interfaces). This allows the community to develop de novo applications. This approach allows us to serve broad community needs while providing flexible, secure, and creative utilization of our platform that is based on best practices and that leverages established computational resources.
The cloud platform complements traditional compute and storage infrastructures by introducing capabilities for efficiently provisioning resources in a self-service, on-demand manner. The new provisioning model promises to accelerate scientific discovery by improving access to customizable and task-specific computing resources. This paradigm is well-suited, especially for those applications tailored to leverage cloud-style of infrastructure capabilities. Adoption of the cloud model has been challenging for many domain scientists and scientific software developers due to the technical expertise required to effectively utilize this infrastructure. Some of the key limitations of cloud infrastructure are: limited integration with institutional authentication and authorization frameworks, lack of frameworks to enable domain-specific configurations for instances, and integration with scientific data repositories alongside existing computational clusters and grid deployments. Specifically designed to address some of these operational barriers towards adoptions by the plant sciences community, the iPlant Collaborative cloud platform, aptly named Atmosphere, is an open-source, robust, configurable gateway that extends established cloud infrastructure to meet the diverse computing needs for the plant science. Atmosphere manages the Virtual Machine (VM) lifecycle while maximizing the utilization of cloud resources for scientific workflows. Thus, Atmosphere allows researchers developing novel analytical tools to deploy them with ease while abstracting the underlying computing infrastructure, at the same time making it relatively easy for the users to access these tools via web browser. Atmosphere also provides a rich extensible Application Programming Interface (APIs) for integration and automation with other services. Since its launch, Atmosphere has seen a wide adoption by the plant sciences community for a broad array of applications that range from image processing to next generation sequence (NGS) analysis and can serve as a template for providing similar capabilities to other domains.
Today there are so much data being available from sources like sensors (RFIDs, Near Field Communication), web activities, transactions, social networks, etc. Making sense of this avalanche of data requires efficient and fast processing. Processing of high volume of events to derive higher-level information is a vital part of taking critical decisions, and Complex Event Processing (CEP) has become one of the most rapidly emerging fields in data processing. e-Science use-cases, business applications, financial trading applications, operational analytics applications and business activity monitoring applications are some use-cases that directly use CEP. This paper discusses different design decisions associated with CEP Engines, and proposes some approaches to improve CEP performance by using more stream processing style pipelines. Furthermore, the paper will discuss Siddhi, a CEP Engine that implements those suggestions. We present a performance study that exhibits that the resulting CEP Engine--Siddhi--has significantly improved performance. Primary contributions of this paper are performing a critical analysis of the CEP Engine design and identifying suggestions for improvements, implementing those improvements through Siddhi, and demonstrating the soundness of those suggestions through empirical evidence.
Understanding the evolutionary history of living organisms is a central problem in biology. Until recently the ability to infer evolutionary relationships was limited by the amount of DNA sequence data available, but new DNA sequencing technologies have largely removed this limitation. As a result, DNA sequence data are readily available or obtainable for a wide spectrum of organisms, thus creating an unprecedented opportunity to explore evolutionary relationships broadly and deeply across the Tree of Life. Unfortunately, the algorithms used to infer evolutionary relationships are NP-hard, so the dramatic increase in available DNA sequence data has created a commensurate increase in the need for access to powerful computational resources. Local laptop or desktop machines are no longer viable for analysis of the larger data sets available today, and progress in the field relies upon access to large, scalable high-performance computing resources. This paper describes development of the CIPRES Science Gateway, a web portal designed to provide researchers with transparent access to the fastest available community codes for inference of phylogenetic relationships, and implementation of these codes on scalable computational resources. Meeting the needs of the community has included developing infrastructure to provide access, working with the community to improve existing community codes, developing infrastructure to insure the portal is scalable to the entire systematics community, and adopting strategies that make the project sustainable by the community. The CIPRES Science Gateway has allowed more than 1800 unique users to run jobs that required 2.5 million Service Units since its release in December 2009. (A Service Unit is a CPU-hour at unit priority).
A key aspect of data cyberinfrastructure is the middleware for managing distributed scientific data with services for organizing, publishing and preserving data entities in a collaborative, data sharing environment. In this paper, we describe a virtual web-based file system for managing NMR (Nuclear Magnetic Resonance) data from multiple, distributed NMR spectrometers across the campus of UC San Diego. The environment provides an authenticated, web-based, single point of entry for users to access their data from any of the NMR systems. An intuitive graphical user interface allows users to view, annotate, search, refresh, upload, and download their own data files or files belonging to a group, based on user privileges. We provide details of the system, including the design of the NMR Portal, the central NMR data repository, and the data harvesting process.
The NERSC Web Toolkit (NEWT) brings High Performance Computing (HPC) to the web through easy to write web applications. Our work seeks to make HPC resources more accessible and useful to scientists who are more comfortable with the web than they are with command line interfaces. The effort required to get a fully functioning web application is decreasing, thanks to Web 2.0 standards and protocols such as AJAX, HTML5, JSON and REST. We believe HPC can speak the same language as the web, by leveraging these technologies to interface with existing grid technologies. NEWT presents computational and data resources through simple transactions against URIs. In this paper we describe our approach to building web applications for science using a RESTful web service. We present the NEWT web service and describe how it can be used to access HPC resources in a web browser environment using AJAX and JSON. We discuss our REST API for NEWT, and address specific challenges in integrating a heterogeneous collection of backend resources under a single web service. We provide examples of client side applications that leverage NEWT to access resources directly in the web browser. The goal of this effort is to create a model whereby HPC becomes easily accessible through the web, allowing users to interact with their scientific computing, data and applications entirely through such web interfaces.
Modern web pages are no longer static plain HTML files, and they provide rich content through scripting languages such as JavaScript and multimedia. However, communications between browser applications that use scripting languages remain challenge. As a solution to this problem, we propose BISSA, a communication model which provides a unified and time-decoupled communication platform based on tuple spaces for browser applications. BISSA consists of an in-browser tuple space and a scalable and distributed peer-to-peer global tuple space, and both can either act standalone or collaborate with each other. The in-browser tuple space provides a solid communication infrastructure for web gadgets. The integration with the peer-to-peer global tuple space further enforces this paradigm of communication, effectively allowing web gadgets to contribute or co-ordinate with the underlying computing infrastructure. This paper presents BISSA, its architecture and show how browser applications can use BISSA as an inter-gadget communication solution, storage platform for application generated data and as a middleware to develop web-based applications that brings the computation power of browser in to the grid.
To make research and development investments where they will have the most impact, it is critical to understand why some science and engineering gateway or portal projects change the way that science is conducted at a fundamental level in a given community. This paper provides some initial reflections on a June 2010 focus group with the goals of uncovering some of these characteristics of success and generating practical insights that draw on the strength of multidisciplinary perspectives. We identify five key tensions that are challenges to gateway sustainability, and we offer policy recommendations that could benefit future gateway projects.
Research teams collaborating across institutional, geographical and cultural boundaries are increasingly common. Funding agencies including the National Science Foundation (NSF) and National Institutes of Health (NIH) strongly encourage virtual organization building and collaboration across institutions and disciplines. A set of software tools that enables scientists to efficiently share information and resources with distributed collaborators is key to facilitate collaboration of research teams across institutional and geographical boundaries. Utilizing HUBzero, we have implemented a regional infrastructure supporting the High Performance Computing Consortium (HPC2) initiative within New York State. This hub consists of a web-portal infrastructure that allows investigators and research groups to easily create and share research materials as well as computational and storage resources among their members. In this paper, we describe the process of launching the hub and providing access to a regional grid consisting of three supercomputing sites with heterogeneous computing platforms and security/access policies. We address issues of capacity and security that arose as we road-tested this evolving platform.
Scientific data continues to grow in volume making the tasks of managing, accessing and sharing such data more challenging. Providing data to scientists via scientific gateways or collaborative portals can aid scientists in achieving these tasks. This paper presents a general data management system that has been built on top of Elgg, an open source social networking platform. The tool enables scientists to upload, browse, view and share a wide variety of scientific data, as well as define and evolve meta data standards in a collaborative manner. The data management system is currently being used as part of GeoChronos, a scientific gateway for Earth observation scientists, for creating and sharing collections of spectral and satellite data.
A chemist given a compound would be interested in knowing the experiments performed using the compound, journals containing the compound and also molecular properties of the compound. If there is a way to integrate this data, it would enhance the chemist's knowledge about a given compound. The Object Reuse and Exchange (ORE) specification may provide a solution to this problem. ORE is a model proposed by the digital libraries community to aggregate resources on the web. OREChem is a research project funded by Microsoft External Research that aims to apply and extend ORE to enable the integration of experimental, bibliographical and molecular properties data. OREChem targets crystallography as its primary application domain. This effort will design a prototypical, semantic-based eScience infrastructure for chemistry and chemical informatics. In this paper we describe how we have used REST as well as SOAP web services, TeraGrid cyberinfrastructure and Semantic Web technologies, such as RDF and triple stores, to facilitate the metadata integration.
Gateway computing environments face several challenges in providing robust, scalable, and sustainable capabilities to a wide range of users. Principles of encapsulation and cohesion have been applied in emerging trends of application framework development, where modular designs and abstraction layers allow these systems to remain flexible and agile as requirements evolve over time. Orbiter Commander is a modular and extensible application framework that leverages the Orbiter Federation Service Oriented Architecture to deliver fast and secure capabilities in an Eclipse RCP desktop application. Commander provides suites of modules that can be seamlessly delivered to end users on multiple platforms, enabling rapid component development through a flexible design and well-defined extension points. This paper presents our collaboration with the Spallation Neutron Source Neutron Experiment and Theory Hub (NExTHUB) and the Solenoidal Tracker at the at the RHIC (STAR) experiment, two suites of capabilities tailored to serve the needs of users at Oak Ridge National Laboratory and Brookhaven National Laboratory.
Science Gateways often execute computations on clusters managed by batch schedulers. These schedulers queue computations until resources are available and then execute them. Clusters are typically scarce resources so the amount of time a computation waits before it begins to execute can be significant. Gateways can potentially reduce the turn around time of their computations and improve their user experience if they use estimates of how long computations will wait in a queue before beginning to execute. This paper describes a service that provides such estimates as well as current and historical job information. An instance of this service has been deployed on TeraGrid and is available to TeraGrid science gateways.
My-Plant.org (My-Plant) is a social networking portal for the Plant Sciences community. As part of the iPlant Collaborative, My-Plant is charged with the goal of bringing together scientists, students, educators, and other interested parties by providing a new approach to connecting with others in the plant sciences thereby helping to spark new collaborations and communication among them. Many social networking sites exist where users can form groups and communicate, but the group structure is flat and has no inherent interconnectivity. My-Plant connects users via branches (called clades) of the phylogenetic tree of green plants, thus creating a unique, phylogenetically based social network structure. My-Plant users can join clades at any level of the tree and collaborate with other users interested in these clades. My-Plant is built upon the Drupal open source content management system. This paper discusses My-Plant in detail, its concept and contributions to the iPlant Collaborative, its implementation details, and its efforts to become an information hub for the plant sciences community. In addition, the paper focuses on some of the technology challenges, lessons learned, and future work of My-Plant.