High Performance Computing (HPC) Best Practice offers opportunities to implement lessons learned in areas such as computational chemistry and physics in genomics workflows, specifically Next-Generation Sequencing (NGS) workflows. In this study we will briefly describe how distributed-memory parallelism can be an important enhancement to the performance and resource utilization of NGS workflows. We will illustrate this point by showing results on the parallelization of the Inchworm module of the Trinity RNA-Seq pipeline for de novo transcriptome assembly. We show that these types of applications can scale to thousands of cores. Time scaling as well as memory scaling will be discussed at length using two RNA-Seq datasets, targeting the Mus musculus (mouse) and the Axolotl (Mexican salamander). Details about the efficient MPI communication and the impact on performance will also be shown. We hope to demonstrate that this type of parallelization approach can be extended to most types of bioinformatics workflows, with substantial benefits. The efficient, distributed-memory parallel implementation eliminates memory bottlenecks and dramatically accelerates NGS analysis. We further include a summary of programming paradigms available to the bioinformatics community, such as C++/MPI.
We introduce novel ideas involving aspect-oriented instrumentation, Multi-Faceted Program Monitoring, as well as novel techniques for a selective and detailed event-based application performance analysis, with an eye toward exascale. We give special attention to the spatial, temporal, and level-of-detail aspects of the three important phases of compile-time filtering, application execution, and runtime filtering. We use an event-based monitoring approach to allow selected and focused performance analysis.
Exemplifying collaborative software development between industry and academia to tackle computational challenges in manipulating large volumes of next-gen sequence data, leveraging advances in algorithm development and compute hardware, we describe our efforts to optimize the performance of the Trinity RNA-Seq de novo assembly software. Three versions of Trinity's Inchworm computationally intensive part (one that is based on the original OpenMP version, and two new versions that are based on MPI and on Fortran2008). New results of Inchworm’s parallel performance for various real-life problems (e.g., mouse, schizosaccharomyces pombe) are presented, as well as a detailed discussion of the MPI and PGAS computations scheme for Inchworm. Computations at Cray, Inc. MPI inchworm
State-of-the-art high performance computing (HPC) applications have to scale over an increasing number of processing elements, meanwhile application developers recently have to face the programming of special accelerator hardware. Although computing languages like CUDA and programming standards like OpenACC provide a fairly easy way to exploit the computational power of general purpose graphics processing units (GPGPUs), their programming is still challenging. Performance analysis is a vital procedure to efficiently use the available hardware and programming models. This paper presents the Vampir performance analysis capabilities by taking the example of a molecular dynamics code, which uses message passing (MPI), threading (OpenMP) and offloading to accelerators (OpenACC and CUDA). It is shown that the Vampir tool-set allows a holistic view on the combined usage of all commonly utilized programming paradigms in heterogeneous HPC applications.
De novo assembly of RNA-seq data enables researchers to study transcriptomes without the need for a genome sequence; this approach can be usefully applied, for instance, in research on 'non-model organisms' of ecological and evolutionary importance, cancer samples or the microbiome. In this protocol we describe the use of the Trinity platform for de novo transcriptome assembly from RNA-seq data in non-model organisms. We also present Trinity-supported companion utilities for downstream applications, including RSEM for transcript abundance estimation, R/Bioconductor packages for identifying differentially expressed transcripts across samples and approaches to identify protein-coding genes. In the procedure, we provide a workflow for genome-independent transcriptome analysis leveraging the Trinity platform. The software, documentation and demonstrations are freely available from http://trinityrnaseq.sourceforge.net. The run time of this protocol is highly dependent on the size and complexity of data to be analyzed. The example data set analyzed in the procedure detailed herein can be processed in less than 5 h.
Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the funding agencies that supported this work.
Digital instruments and simulations are creating an ever-increasing amount of data. The need for institutions to acquire these data and transfer them for analysis, visualization, and archiving is growing as well. In parallel, networking technology is evolving, but at a much slower rate than our ability to create and store data. Single fiber 100 Gbps networking solutions have recently been deployed as national infrastructure. This article describes our experiences with data movement and video conferencing across a networking testbed, using the first commercially available single fiber 100 Gbps technology. The testbed is unique in its ability to be configured for a total length of 60, 200, or 400 km, allowing for tests with varying network latency. We performed low-level TCP tests and were able to use more than 99.9% of the theoretical available bandwidth with minimal tuning efforts. We used the Lustre file system to simulate how end users would interact with a remote file system over such a high performance link. We were able to use 94.4% of the theoretical available bandwidth with a standard file system benchmark, essentially saturating the wide area network. Finally, we performed tests with H.323 video conferencing hardware and quality of service (QoS) settings, showing that the link can reliably carry a full high-definition stream. Overall, we demonstrated the practicality of 100 Gbps networking and Lustre as excellent tools for data management.
Using the first commercially available 100 Gbps Ethernet technology with a link of varying length, we have evaluated the performance of the Lustre file system and its networking layer under different latency scenarios. The results led us to a better understanding of the impact that the network latency has on Lustre's performance. In particular spanning Lustre's networking layer, striped small I/O, and the parallel creation of files inside a common directory. The main contribution of this work is the derivation of useful rules of thumbs to help users and system administrators predict the variation in Lustre's performance produced as a result of changes in the latency of the I/O network.
As part of the SCinet Research Sandbox at the Supercomputing 2011 conference, Indiana University (IU) demonstrated use of the Lustre high performance parallel file system over a dedicated 100 Gbps wide area network (WAN) spanning more than 3,500 km (2,175 mi). This demonstration functioned as a proof of concept and provided an opportunity to study Lustre's performance over a 100 Gbps WAN. To characterize the performance of the network and file system, low level iperf network tests, file system tests with the IOR benchmark, and a suite of real-world applications reading and writing to the file system were run over a latency of 50.5 ms. In this article we describe the configuration and constraints of the demonstration and outline key findings.
A molecular dynamics code simulating the diffusion in dense nuclear matter in white dwarf stars is analyzed in this collaboration between PTI (Indiana University) and ZIH (Technische Universität Dresden). The code is highly configurable allowing MPI, OpenMP, or hybrid runs and additional fine-tuning with a range of parameters. The first step in the code analysis is to identify the best performing parameter set of the serial version. This configuration represents the most promising candidate for further studies. Aim of the parallel analysis is then to measure the scalability limits of the different parallel code implementations and to detect bottlenecks possibly preventing higher parallel efficiency. This work has been done with the parallel analysis framework Vampir.
Das Hochleistungsrechenzentrum der TU Dresden wurde in Kooperation mit Alcatel-Lucent und T-Systems über eine 100-Gigabit-Strecke mit dem Rechenzentrum der TU Bergakedemie Freiberg verbunden. Die Umsetzung erfolgt erstmals mit kommerzieller Hardware. Eine weitere Besonderheit ist die Übertragung der Daten über eine einzige Wellenlänge. In dieser Arbeit werden nach der Beschreibung des 100-Gigabit-Testbeds erste Ergebnisse synthetischer Lasttests auf verschiedenen Ebenen des OSI-Referenzmodells vorgestellt und diskutiert. Die Vorgaben und der spezielle Aufbau des Testbeds machen eine Überarbeitung vorhandener Messmethoden notwendig. Die Entwicklung der Burst-Tests wird ebenfalls in dieser Arbeit beschrieben. Zusätzlich wird ein Ausblick auf weitere geplante Teilprojekte, sowie Einsatzmöglichkeiten der 100-Gigabit-Technik in Forschung und Industrie gegeben.
To process large amounts of data, in some fields of science hundreds or thousands of single jobs are submitted into a Grid. Monitoring the enormous numbers of jobs and their resource usage in such environments (like the LCG/gLite middleware) effectively becomes an important issue for the users. Current tools in LCG / gLite provide only limited value to the user as they are often simple command line applications only. Keeping an eye on large number of jobs can thus become quite painful.