
The CIPRES Science Gateway (CSG) provides researchers and educators with browser-based access to community codes for inference of phylogenetic relationships from DNA and protein sequence data. The CSG allows users to deploy jobs on the high-performance computers of the TeraGrid without requiring detailed knowledge of their complexities. Use of the CSG has grown rapidly; through March 2011 it had more than 2,200 users and enabled more than 180 peer-reviewed publications. The rapid growth in resource consumption was accommodated by deploying codes on Trestles, a new TeraGrid computer. Tools and policies were developed to insure efficient and effective resource use. This paper describes progress in managing the growth of this public cyberinfrastructure resource and reviews the domain science that it has enabled.
Program build information, such as compilers and libraries used, is vitally important in an auditing and benchmarking framework for HPC systems. We have developed a tool to automatically extract this information using signature-based detection, a common strategy employed by anti-virus software to search for known patterns of data within the program binaries. We formulate the patterns from various "features" embedded in the program binaries, and the experiment shows that our tool can successfully identify many different compilers, libraries, and their versions.
The neoGRID is under development in Quarry, a virtual hosting environment, for working with Taverna-based workflow utilizing grid computing. Taverna is a graphical workbench often used for biomedical informatics [1, and the references therein]. neoGRID is designed to offer a HPC-supported collaborative environment for the researchers from multidisciplinary scientific fields to gather data, integrate and analyze using teragrid resources. A significant number of resources including bioinformatics tools [2] has been deployed in TeraGrid, an NSF funded project. Besides, the myGrid team produced a suite of tools that are made available for analysis of proteins. In addition, their myExperiment site makes it easy to find, use and share scientific workflows and building scientific communities with common interests. This Cyberinfrastructure (CI)-supported neoGRID can be utilized for protein motifs analysis, particularly in the glycosyltransferase protein family and will be available to the Glyco-Community [3]. Workflows specifically designed for such analysis will be available in neoGRID. Moreover, a significant number of workflows are available in KEGG that are useful for protein analysis and drug discovery research. These and modification of these workflows in a collaborative environment will also be available for drug discovery research using members of glycosyltransferase enzyme family as target protein(s) [4]. Here, we demonstrate the usage of neoGRID for analysis of sialylmotifs of sialyltransferases. The presence of sialylmotifs is the cardinal feature of mammalian sialyltransferases [5], a group of enzymes that transfers sialic acid from CMP-NeuAc to the terminal carbohydrates group of various glycoproteins and glycolipids [6]. Sialic acid has been recognized as the key determinant of a diverse oligosaccharide structures involved in a large variety of biological events as diverse as animal cell-cell interaction to oncogenic transformation [7]. These conserved protein domains, sialylmotifs, have been shown to be involved in binding either the donor or acceptor substrates or both [8, 9], and a disulfide linkage between these two motifs has been shown to be essential for catalytic activity [10, 11]. Protein sequence [12] and structural analysis [11, 13] showed that mammalian sialyltransferase has no similarities with the bacterial enzymes, although a His residue serves as a catalytic center for both [11]. This provides a unique opportunity for discovery research on potential drug development. A thorough analysis of these motifs is now under study for drug discovery research using sialyltransferase as a target protein. Such analysis demanding high-performance computing power is available in neoGRID. The genome of protozoan parasite Toxoplasma gondii has been used here for bioinformatics analysis. T. gondii, which chronically infects roughly 30% of the world's population is typically asymptomatic but can cause life threatening disease in immune compromised individuals and birth defects of acquired during pregnancy [14]. Sequencing of this parasite genome of about 63 Mb in size with 14 chromosomes, has recently been completed [15, 16]. The availability of this sequence information from multiple isolates in a specially designed database (www.toxodb.org) that also include transcriptomic and proteomic data sets [16], has made this parasite an ideal candidate for in-depth bioinformatic interrogation. At this time, little is known about the parasite glycome or the encoded capacity for glycosylation. This deficit is in spite of the fact that the parasite form associated with chronic infection is heavily glycosylated [17]. Our initial studies have identified a broad spectrum of lectin reactivities [18]. Lectins recognize their specific glycan targets with high specificity and affinity [19]. We were particularly surprised to find evidence for sialylation of the parasite tissue cyst forms given that the genome does not appear to encode recognizable sialyltransferases based on the 'sialylmotif' [5]. This finding suggests that Toxoplasma may hijack the sialyltrasferase activities of the infected host cell or alternatively possess activities with entirely novel functional signatures. The integration of state of the art computational biology and bioinformatics with experimental validation provides a unique opportunity toward new discovery. By conducting a systematic survey of glycosyltransferase activities in silico we can establish both the presence and absence of specific activities in the Toxoplasma genome. Given the evolutionary antiquity of the parasitic protozoa we expect the combination of in silico analysis and experimental validation to offer new insights into the biology of these parasites. The availability of sequenced genomes of several related parasites housed at www.eupathdb.org [20] provides an ideal resource for expanding these studies to other parasites including Plasmodium species, the agent of malaria. In silico identification of potentially unique enzymatic activities could open doors toward the discovery of novel drugs to treat these often deadly infections.
Many new users of TeraGrid or other HPC resources are scientists or other domain experts by training and are not necessarily familiar with core principles, practices, and resources within the HPC community. As a result, they often make inefficient use of their own time and effort and of the computing resources as well. In this work, we present a training roadmap for new members of the HPC community. These new users benefit from a basic understanding of the the key concepts, technologies, and tools outlined in the roadmap, even if they are not required to develop immediate proficiency in all of them. In addition to providing a brief overview of a wide variety of topics, the roadmap includes links and pointers to numerous resources that users can refer to for their own training as the need arises.
We use deep adaptive mesh refinement simulations of isothermal self-gravitating supersonic turbulence to study the imprints of gravity on the mass density distribution in molecular clouds. The simulations show that the density distribution in self-gravitating clouds develops an extended power-law tail at high densities on top of the usual lognormal. We associate the origin of the tail with self-similar collapse solutions and predict the power index values in the range from -7/4 to -3/2 that agree with both simulations and observations of star-forming molecular clouds.
The top supercomputers typically have aggregate memories in excess of 100 TB, with simulations running on these systems producing datasets of comparable size. The size of these datasets and the speed with which they are produced define the minimum performance that modern analysis and visualization must achieve. We report on interactive visualizations of large simulations performed on Kraken at the National Institute for Computational Sciences using the parallel cosmology code Enzo, with grid sizes ranging from 10243 to 64003. In addition to the asynchronous rendering of over 570 timesteps of a 40963 simulation (150 TB in total), we developed the ability to stream the rendering result to multi-panel display walls, with full interactive control of the renderer(s).
The NSF DataONE [1] DataNet project and the NSF Tera-Grid [2] project have initiated a pilot collaboration to deploy and operate the DataONE Member Node software stack on TeraGrid infrastructure. The appealing feature of this collaboration is that it opens up the possibility to add large scale computing as an adjunct to DataONE data, metadata, and workflow manipulation and analysis tools. Additionally, DataONE data archive and curation services are exposed as an option for large scale computing and storage efforts such as TeraGrid/XSEDE. With this joint effort, DataONE also brings an open, persistent, robust, and secure method for accessing Earth sciences data collected by science communities such as The National Evolutionary Synthesis Center's Dryad [3], The Ecological Society of America's Ecological Archive [4], NASA's Distributed Active Archive Center at the Oak Ridge National Laboratory [5], the USGS's National Biological Information Infrastructure [6], the Fire Research & Management Exchange System [7], the Long Term Ecological Research Network [8], and the Knowledge Network for Biocomplexity [9]. Beginning with an April 1st, 2011, allocation, the DataONE Core Cyberinfrastructure Team has been working with the IU Quarry [10] virtual hosting service, and more generally with the TeraGrid data area, on this pilot implementation. The implementation includes multiple virtual servers in order to test different reference implementations of the common DataONE Member Node RESTful web-service functions [11]. These implementations include implementation as a Metacat server [12], as well as a Python Generic Member Node developed by DataONE [13]. The implementations will also mount TeraGrid-wide global storage services (DC-WAN [14] and Albedo [15]) and thus allow integration of input and output of large scale computational runs with wide area archival data and metadata services.
In this paper, we describe the formative activities related to HPC applications and TeraGrid at Texas Southern University (TSU). TSU is an urban MSI (minority serving institution) and HBCU with a very small student body in STEM areas (7% of the undergraduate population). We have a vibrant and diverse faculty. Limited resources compound the challenges faced by the HPC community at TSU. We faced several difficulties as well as successes in the building of HPC access and its utilization at TSU.
We describe how a new graduate course in scientific computing, taught during Fall 2010 at Louisiana State University, utilized TeraGrid resources to familiarize students with some of the real world issues that computational scientists regularly deal with in their work. The course was designed to provide a broad and practical introduction to scientific computing, creating the basic skills and experience to very quickly get involved in research projects involving modern cyberinfrastructure and complex real world scientific problems. As an integral part of the course, students had to utilize various TeraGrid resources, e.g., by deploying, using and extending scientific software within the national cyberinfrastructure.
This paper aims to describe methods that can be used to create new Student Cluster Competition teams from the standpoint of the team advisors. The purpose is to share these methods in order to create an easier path for organizing a successful team. These methods were gleaned from a survey of advisors that have formed teams in the last four years. Four advisors responded to the survey and those responses fit into five categories: (1) early preparation, (2) coursework specific to the competition, (3) close relationships with the hardware vendors, (4) concentration on the applications over the hardware, and (5) the need to encourage the team members to write papers about their experiences. In addition to these commonalities which may be best practices there are a few divergent but intriguing techniques that may also prove useful for potential advisors. Both will be discussed here and these methods can serve as a primer for anyone looking to start a new Student Cluster Competition team.
The usage and adoption of General Purpose GPUs (GPGPU) in HPC systems is increasing due to the unparalleled performance advantage of the GPUs and the ability to fulfill the ever-increasing demands for floating points operations. While the GPU can offload many of the application parallel computations, the system architecture of a GPU-CPU-InfiniBand server does require the CPU to initiate and manage memory transfers between remote GPUs via the high speed InfiniBand network. In this paper we introduce for the first time a new innovative technology - GPUDirect that enables Tesla GPUs to transfer data via InfiniBand without the involvement of the CPU or buffer copies, hence dramatically reducing the GPU communication time and increasing overall system performance and efficiency. We also explore for the first time the performance benefits of GPUDirect using Amber and LAMMPS applications.
Ongoing efforts by the Large Synoptic Survey Telescope (LSST) involve the study of asteroid search algorithms and their performance on both real and simulated data. Images of the night sky reveal large numbers of events caused by the reflection of sunlight from asteroids. Detections from consecutive nights can then be grouped together into tracks that potentially represent small portions of the asteroids' sky-plane motion. The analysis of these tracks is extremely time consuming and there is strong interest in the development of techniques that can eliminate unnecessary tracks, thereby rendering the problem more manageable. One such approach is to collectively examine sets of tracks and discard those that are subsets of others. Our implementation of a subset removal algorithm has proven to be fast and accurate on modest sized collections of tracks, but unfortunately has extremely large memory requirements for realistic data sets and cannot effectively use conventional high performance computing resources. We report our experience running the subset removal algorithm on the TeraGrid Appro Dash system, which uses the vSMP software developed by ScaleMP to aggregate memory from across multiple compute nodes to provide access to a large, logical shared memory space. Our results show that Dash is ideally suited for this algorithm and has performance comparable to or superior to that obtained on specialized, heavily demanded, large-memory systems such as the SGI Altix UV.
One of the foremost challenges in high performance computing (HPC) is promoting involvement of historically underrepresented undergraduate students. The Blue Waters Undergraduate Petascale Education Program, funded by the National Science Foundation Office of CyberInfrastructure, has been facing this challenge head-on while supporting undergraduate internship experiences that involve the application of HPC to problems in the sciences, engineering, or mathematics. This paper describes an evolving approach to the recruitment efforts for undergraduate interns and mentors that demonstrates the importance of formative assessment supporting proactive program changes leading to success in generating a large, diverse applicant pool including substantial numbers of qualified women and minority candidates.
The rapid acquisition of mutations conferring resistance to particular drugs remains a significant cause of anti-HIV treatment failure. Informatics based techniques give resistance scores to individual mutations which can be combined additively to assess the resistance levels of complete sequences. It is likely however, that the full picture is more complicated, with non-linear epistatic effects between combinations of mutations playing an important role in determining the level of viral resistance [1, 2]. Molecular dynamics is one simulation technique which offers the ability to derive quantitative (as well as qualitative) insight into the interplay of resistance-causing mutations. The dynamics of sequence specific models can be simulated and the free energy change associated with the drug binding calculated. The free energy change (or binding affinity) is the thermodynamic quantity which determines how tightly a drug will bind to its target. Hence, comparing values for mutant and wildtype systems allows the level of resistance of a particular sequence to be estimated. Validating such an approach is computationally demanding and the management of large numbers of simulations is a considerable administrative challenge. Here we present the use of tools based on the Simple API for Grid Applications (SAGA) to tackle this computational problem. This allows us to utilise high-end infrastructure, such as TeraGrid/XD, to provide extreme scales of throughput for highperformance simulations. This paper presents initial results and experience of using the TeraGrid in conjunction with DEISA, the European analogue of the US TeraGrid. High-throughput (high-performance) calculations are one of the few classes of computational problems that can easily exploit the computational power of widely distributed computing resources, thereby amassing more computational power than any of the individual resources could offer. In addition to the resource utilization rationale the need for cross-Grid capabilities in this case arises from shared and complementary scientific and technical skills found in this intercontinental collaboration.
This abstract describes the aggregation of TeraGrid and Open Science Grid to run the SCEC CyberShake application faster than on TeraGrid alone. Because the resources are distributed and data movement is required to use more than one resource, a careful analysis of the cost of data movement vs. the benefits of distributed computation has been done in order to best distribute the work across the resources.
Infrastructure cloud computing introduces a significant paradigm shift that has the potential to revolutionize how scientific computing is done. However, while it is actively adopted by a number of scientific communities, it is still lacking a well-developed and mature ecosystem that will allow the scientific community to better leverage the capabilities it offers. This paper introduces a specific addition to the infrastructure cloud ecosystem: the cloudinit.d program, a tool for launching, configuring, monitoring, and repairing a set of interdependent virtual machines in an infrastructure-as-a-service (IaaS) cloud or over a set of IaaS clouds. The cloudinit.d program was developed in the context of the Ocean Observatory Initiative (OOI) project to help it launch and maintain complex virtual platforms provisioned on demand on top of infrastructure clouds. Like the UNIX init.d program, cloudinit.d can launch specified groups of services and the VMs in which they run, at different run levels representing dependencies of the launched VMs. Once launched, cloudinit.d monitors the health of each running service to ensure that the overall application is operating properly. If a problem is detected in a service, cloudinit.d will restart only that service and any other service that failed that depended on it.
Scientific research creates substantially large volumes of data throughout the processes of discovery and analysis. Given the necessity for data sharing and data relocation, members of the scientific community are often faced with a productivity loss that correlates with the time cost incurred during the data transfer process. The GridFTP protocol was developed to improve this situation by addressing the performance, reliability, and security limitations of standard FTP and other commonly used data movement tools such as SCP. The Globus implementation of GridFTP is widely used to rapidly and reliably move data between geographically distributed systems. Traditionally, GridFTP performs well for datasets containing large files. When the data is partitioned into many small files, however, it suffers from lower transfer rates. Although the pipelining and concurrency solution in GridFTP provides improved transfer rates for datasets using lots-of-small-files, these solutions cannot be applied in environments that have strict firewall rules. In some cases, tarring the files in a dataset on the fly will help; in other cases, a checksum of the files after they are written to disk is desired. In this paper, we present the Globus XIO Pipe Open Driver which enables GridFTP to leverage the standard Unix tools to perform these tasks. We demonstrate the effectiveness of this functionality through several experiments.
The Ultrascan gateway provides a user friendly web interface for evaluation of experimental analytical ultracentrifuge data using the UltraScan modeling software. The analysis tasks are executed on the TeraGrid and campus computational resources. The gateway is highly successful in providing the service to end users and consistently listed among the top five gateway community account usage. This continued growth and challenges of sustainability needed additional support to revisit the job management architecture. In this paper we describe the enhancements to the Ultrascan gateway middleware infrastructure provided through the TeraGrid Advanced User Support program. The advanced support efforts primarily focused on a) expanding the TeraGrid resources incorporate new machines; b) upgrading UltraScan's job management interfaces to use GRAM5 in place of the deprecated WS-GRAM; c) providing realistic usage scenarios to the GRAM5 and INCA resource testing and monitoring teams; d) creating general-purpose, resource-specific, and UltraScan-specific error handling and fault tolerance strategies; and e) providing forward and backward compatibility for the job management system between UltraScan's version 2 (currently in production) and version 3 (expected to be released mid-2011).
Portals and gateways are increasingly offering users complex interfaces to interact with massive data sets. As dealing with big data becomes more commonplace, portal and gateway developers need to readdress how data is stored and rethink the supporting infrastructure that enables quick and simple access and analysis of data. It is becoming evident that traditional, relational databases are not always the most appropriate solution to allow users on-demand access to big data sets. In this study we show that using non-relational, "NoSQL" databases such as key-value stores and document stores can offer large benefits in performance, accessibility, and availability. We present a use case from the TeraGrid User Portal that demonstrates solutions for processing and auditing user job data efficiently in order to provide users rapid access to this data. One of the goals of TeraGrid User Portal is to offer users and PIs detailed job statistics such as service unit (SU) usage and job history via the user portal interface. While building a portal application to analyze batch job data records in the TeraGrid Central Database (TGCDB), we quickly ran into stumbling blocks. The TGCDB has over 17 million job records from December 2003 through March 2011. Between January 2011 and April 2011 alone, there are over 2.8 million job records. This data is growing at an ever-faster rate and will continue to grow as new computing resources become available. Even properly indexed tables took longer than ideal to query and still be responsive in a portal application. The current solution to this was to cache the jobs query results and access those cached results in the portal. This solved the issue with the speed of the query, but did not address the problem of dealing with this massive data set. We still needed the rich query interface that a database provides. In order to solve our issues we looked at a two different options. First, we tested moving the TGCDB to a newer, faster machine than the one it currently runs on to determine how much of the bottleneck was due to aging hardware. Second, we tested migrating the jobs data off of the relational PostgreSQL TGCDB and into a key-value store using Apache CouchDB instead of the flat file cache we had been using. CouchDB is a document-oriented database that is queried using MapReduce. CouchDB also offers specific benefits for portals and gateways, providing a RESTful JSON API that can be accessed using HTTP requests. Our initial tests have shown that moving the TGCDB to new hardware can provide a query speedup of 3.7x on average for the job queries we tested. Querying the same data using MapReduce queries to CouchDB gave an additional 8.24x speedup for a total of 30.6x speedup over the current TGCDB on average. The huge speedups offered by CouchDB come at the cost of additional disk usage. CouchDB maintains B-tree indices on the document store as well as any defined queries or âĂIJviewsâĂİ. These indices use a greater amount of disk than a relational database, but enables CouchDB to take full advantage of high-performance disks and file systems. We show that the increase in performance gained from using a data warehouse for certain large data sets can offer great benefits to building on-demand data analysis tools in portals and gateways. By identifying these large data sets such as the TeraGrid jobs data and migrating them to high performance data stores such as CouchDB we can make much more information readily available to users.
Traditionally, Hadoop is run on parallel machines with modest processing capabilities and large amounts of disk. A large shared-memory system may not, at first blush, seem a likely host for a Hadoop cluster, but when researchers asked to run Hadoop on our 16 TB shared memory system, Blacklight, we set out to determine how effectively Hadoop could work in this non-traditional environment. We will begin by discussing the technical hurdles involved with making the Hadoop platform run efficiently on Blacklight, including configuration management and dynamic port allocation. Then we will discuss the logistics of using RAM disk on Blacklight, where a portion of the system's memory is allocated on a per-job basis to be used in place of the local disk expected by Hadoop. To conclude, we'll share some test results comparing our shared-memory system running Hadoop to a traditional parallel cluster. We expect this work to increase user productivity by making fast memory-based I/O available to solve a large set of data-intensive problems, including those already utilizing, or easily adaptable to, the ubiquitous MapReduce framework.