SUMMARY Image segmentation is a very important step in the computerized analysis of digital images. The maxflow mincut approach has been successfully used to obtain minimum energy segmentations of images in many fields. Classical algorithms for maxflow in networks do not directly lend themselves to efficient parallel implementations on contemporary parallel processors. We present the results of an implementation of Goldberg–Tarjan preflow‐push algorithm on the Cray XMT‐2 massively multithreaded supercomputer. This machine has hardware support for 128 threads in each physical processor, a uniformly accessible shared memory of up to 4 TB and hardware synchronization for each 64 bit word. It is thus well‐suited to the parallelization of graph theoretic algorithms, such as preflow‐push. We describe the implementation of the preflow‐push code on the XMT‐2 and present the results of timing experiments on a series of synthetically generated as well as real images. Our results indicate very good performance on large images and pave the way for practical applications of this machine architecture for image analysis in a production setting. The largest images we have run are 32000 2 pixels in size, which are well beyond the largest previously reported in the literature.Copyright © 2013 John Wiley & Sons, Ltd.
We explore the comparative performance of the Cray XMT and XMT‐2 massively multithreaded supercomputers. We use benchmarks to evaluate memory accesses for various types of loops. We also compare the performance of these machines on matrix multiply and on three previously implemented dynamic programming algorithms. It is shown that the relative performance of these machines is dependent on the size (number of processors) of the configuration, as well as the size of the problem being evaluated. In particular, small configurations of the original XMT can sometimes show slightly better performance than larger configurations of the XMT‐2, for the same problem size. We note that, under heavy memory load, performance of loops can saturate well before the maximum number of processors available. This suggests that it may not always be useful to use the maximum number of processors for a specific run. We also show that manual restructuring of nested loops, including decreasing the parallelism, can result in major improvements in performance. The results in this paper indicate that careful exploration of the space of problem sizes, number of processors used, and choices of loop parallelization can yield substantial improvements in performance. These improvements can be very significant for production codes that run for extended periods of time. Copyright © 2012 John Wiley & Sons, Ltd.
This index covers all technical items - papers, correspondence, reviews, etc. - that appeared in this periodical during the year, and items from previous years that were commented upon or corrected in this year. Departments and other items may also be covered if they have been judged to have archival value. The Author Index contains the primary entry for each item, listed under the first author's name. The primary entry includes the co-authors' names, the title of the paper or other item, and its location, specified by the publication abbreviation, year, month, and inclusive pagination. The Subject Index contains entries describing the item under all appropriate subject headings, plus the first author's name, the publication abbreviation, month, and year, and inclusive pages. Note that the item title is found only under the primary entry in the Author Index.
Many viruses of interest, such as influenza A, have distinct segments in their genome. The evolution of these viruses involves mutation and reassortment, where segments are interchanged between viruses that coinfect a host. Phylogenetic trees can be constructed to investigate the mutation-driven evolution of individual viral segments. However, reassortment events among viral genomes are not well depicted in such bifurcating trees. We propose the concept of reassortment networks to analyze the evolution of segmented viruses. These are layered graphs in which the layers represent evolutionary stages such as a temporal series of seasons in which influenza viruses are isolated. Nodes represent viral isolates and reassortment events between pairs of isolates. Edges represent evolutionary steps, while weights on edges represent edit costs of reassortment and mutation events. Paths represent possible transformation series among viruses. The length of each path is the sum edit cost of the events required to transform one virus into another. In order to analyze tau stages of evolution of n viruses with segments of maximum length m, we first compute the pairwise distances between all corresponding segments of all viruses in O(m(2)n(2)) time using dynamic programming. The reassortment network, with O(tau n(2)) nodes, is then constructed using these distances. The ancestors and descendents of a specific virus can be traced via shortest paths in this network, which can be found in O(tau n(3)) time.
The use of space filling curves for proximity-improving mappings is well known and has found many useful applications in parallel computing. Such curves permit a linear array to be mapped onto a 2D (respectively, 3D) structure such that points that are distance d apart in the linear array are distance O (d 1/2 ) (O(d 1/3 )) apart in the 2D (3D) array and vice versa. We extend the concept of space filling curves to space filling surfaces and show how these surfaces lead to mappings from 2D to 3D so that points at distance d 1/2 on the 2D surface are mapped to points at distance O(d 1/3 ) in the 3D volume. Three classes of surfaces, associated respectively with the Peano curve, Sierpinski carpet, and the Hilbert curve, are presented. A methodology for using these surfaces to map from 2D to 3D is developed. These results permit efficient execution of 2D computations on processors interconnected in a 3D grid. The space filling surfaces proposed by us are the first such fractal objects to be formally defined and are thus also of intrinsic interest in the context of fractal geometry.
David Abramson, Monash University, Australia Enrique Alba, University of Malaga, Spain Srinivas Aluru, Iowa State University, USA Shahid H. Bokhari, University of Engineering & Technology, Pakistan Vincent Breton, CNRS/IN2P3, LPC Clermont-Ferrand, France Jean-Philippe Cassar, Polytech’Lille, France Amitava Data, University of Western Australia, Australia Youping Deng, The University of Southern Mississippi, USA Hans de Sterck, University of Waterloo, Canada Mario R. Guarracino, ICAR-CNR, Italy Ryoko Hayashi, Advanced Institute of Science and Technology (JAIST), Japan Matthew He, Nova Southeastern University, USA Alfons Hoekstra, University of Amsterdam, The Netherlands Lizy John, University of Texas at Austin Tamer Kahveci, University of Florida Sang-Tae Kim, Purdue University, USA Arun Krishnan, Bioinformatics Institute, Singapore Wenjun Li, UT Southwestern Medical Center, USA Yiming Li, National Chiao Tung University, Taiwan Shiyong Lu, Wayne State University, USA Robert L. Martino, National Institutes of Health, USA Michael Mascagni, Florida State University, USA Martin Middendorf, University of Leipzig, Germany Maria Mirto, University of Lecce, Italy Giri Narasimhan, Florida International University, USA Jun Ni, University of Iowa, USA Sergei Petoukhov, Russian Academy of Sciences, Russia Pascal Poullet, West French Indies University, France Youxing Qu, University of Georgia, USA Nagiza Samatova, Oak Ridge National Lab, USA Bertil Schmidt, Nanyang Technological University, Singapore Tony Solomonides, University of the West of England, UK El-Ghazali Talbi, LIFL, France Daming Wei, University of Aizu, Japan Tiffani Williams, Texas A & M University, USA C. M. Yang, Nankai University, China Yanqing Zhang, Georgia State University, USA
Mass storage systems (MSSs) play a key role in data-intensive parallel computing. Most contemporary MSSs are implemented as redundant arrays of independent/inexpensive disks (RAID) in which commodity disks are tied together with proprietary controller hardware. The performance of such systems can be difficult to predict because most internal details of the controller behavior are not public. We present a systematic method for empirically evaluating MSS performance by obtaining measurements on a series of RAID configurations of increasing size and complexity. We apply this methodology to a large MSS at Ohio Supercomputer Center that has 16 input/output processors, each connected to four 8 + 1 RAID5 units and provides 128 TB of storage (of which 116.8 TB are usable when formatted). Our methodology permits storage-system designers to evaluate empirically the performance of their systems with considerable confidence. Although we have carried out our experiments in the context of a specific system, our methodology is applicable to all large MSSs. The measurements obtained using our methods permit application programmers to be aware of the limits to the performance of their codes. Copyright © 2006 John Wiley & Sons, Ltd.
Program Committee Members David Abramson, Monash University, Australia Enrique Alba, University of Malaga, Spain Srinivas Aluru, Iowa State University, USA Shahid H. Bokhari, University of Engineering & Technology, Pakistan Vincent Breton, CNRS/IN2P3, LPC Clermont-Ferrand, France Kevin Burrage, University of Queensland, Australia Amitava Data, University of Western Australia, Australia Hans de Sterck, University of Waterloo, Canada Mario Rosario Guarracino, ICAR-CNR, Italy Ryoko Hayashi, Advanced Institute of Science and Technology (JAIST), Japan Matthew He, Nova Southeastern University, USA Alfons Hoekstra, University of Amsterdam, The Netherlands Chun-Hsi Huang, University of Connecticut, USA Lizy John, University of Texas, Austin, USA Tamer Kahveci, University of Florida, USA Sangtae Kim, Purdue University, USA Arun Krishnan, Bioinformatics Institute, Singapore Wenjun Li, UT Southwestern Medical Center, USA Yiming Li, National Chiao Tung University, Taiwan Robert L. Martino, National Institutes of Health, USA Michael Mascagni, Florida State University, USA Martin Middendorf, University of Leipzig, Germany Maria Mirto, University of Lecce, Italy Giri Narasimhan, Florida International University, USA Jun Ni, University of Iowa, USA Sergei Petoukhov, Russian Academy of Sciences, Russia Pascal Poulet, French West Indies University, France Youxing Qu, University of Georgia, USA Nagiza Samatova, Oak Ridge National Lab, USA Bertil Schmidt, Nanyang Technological University, Singapore Tony Solomonides, University of the West of England, UK El-Ghazali Talbi, LIFL, France
This chapter contains sections titled: Introduction Parallel Computer Architecture Bioinformatics Algorithms on the Cray MTA System Summary References
The standard algorithm for alignment of DNA sequences using dynamic programming has been implemented on the Cray MTA-2 (Multithreaded Architecture-2) at ENRI (Electronic Navigation Research Institute), Japan. Descriptions of several variants of this algorithm and their measured performance are provided. It is shown that the use of "Full/Empty" bits (a feature unique to the MTA) leads to implementations that provide almost perfect speedup for large problems on 1-8 processors. These results demonstrate the potential power of the MTA and emphasize its suitability for bioinformatic and dynamic programming applications.
Internet-based communication is assuming an increasingly important role in the developing world. It is thus crucial that students be exposed to contemporary networking equipment in a realistic setting, in order to connect theoretical material taught in lecture courses with the realities of physical hardware. To this end, a large computer networking laboratory has been set up to provide a realistic environment for teaching internetworking concepts. This laboratory provides university-level students with a testbed to experiment with fundamental issues of internetworking in a way that cannot be provided by simulators and to a degree of rigor not possible with the commonly available laboratory setups designed for technicians. We describe the motivations for setting up the laboratory, its network structure and equipment, and the type of experiments students conduct. The laboratory structure is influenced heavily by the limited funds at our disposal - a common problem in the developing world. Many of the problems we faced in setting up our equipment (such as the crucial impact of proper electrical grounding on system performance) are not ordinarily encountered in developed nations. Our experiences are thus likely to be of value to others in the developing world who are contemplating setting up experimental facilities for teaching networking.
MOTIVATION With the potential availability of nanopore devices that can sense the bases of translocating single-stranded DNA (ssDNA), it is likely that 'reads' of length approximately 10(5) will be available in large numbers and at high speed. We address the problem of complete DNA sequencing using such reads. We assume that approximately 10(2) copies of a DNA sequence are split into single strands that break into randomly sized pieces as they translocate the nanopore in arbitrary orientations. The nanopore senses and reports each individual base that passes through, but all information about orientation and complementarity of the ssDNA subsequences is lost. Random errors (both biological and transduction) in the reads create further complications. RESULTS We have developed an algorithm that addresses these issues. It can be considered an extreme variation of the well-known Eulerian path approach. It searches over a space of de Bruijn graphs until it finds one in which (a) the impact of errors is eliminated and (b) both possible orientations of the two ssDNA sequences can be identified separately and unambiguously. Our algorithm is able to correctly reconstruct real DNA sequences of the order of 10(6) bases (e.g. the bacterium Mycoplasma pneumoniae) from simulated erroneous reads on a modest workstation in about 1 h. We describe, and give measured timings of, a parallel implementation of this algorithm on the Cray Multithreaded Architecture (MTA-2) supercomputer, whose architecture is ideally suited to this 'unstructured' problem. Our parallel implementation is crucial to the problem of rapidly sequencing long DNA sequences and also to the situation where multiple nanopores are used to obtain a high-bandwidth stream of reads.
The Ohio Supercomputing Center (OSC) is in the process of augmenting its disk storage capacity by more than 500Terabytes. One of the key elements of the new storage units is a pool made up of 16 Intel P4 Xeon input/output processors, each connected to a dedicated IBM FAStT600 Storage Controller that, in turn, supports four 8+1 RAID5 units. We present an experimental evaluation of the sustained read rates that can be achieved with the FAStT600 units. We observe that the maximum sustained read rate of each FAStT600 is approximately 280MBytes/sec when multithreaded reads are used. We demonstrate through a series of experiments with smaller RAID5 configurations that it is unlikely that the sustained data rate can exceed this figure. Applications programmers implementing data-intensive applications should therefore employ multithreading and not expect more than 280MByte/sec sustained read performance. Sustained write performance is also analyzed. For 8+1 RAID5 systems this should, in theory, be89 times as fast as reads. Our experiments, however, reveal the write performance to be 1 3 to 1 2 times as fast.
The Cray MTA-2 (Multithreaded Architecture) is an unusual parallel supercomputer that promises ease of use and high performance. We describe our experience on the MTA-2 with a molecular dynamics code, SIMU-MD, that we are using to simulate the translocation of DNA through a nanopore in a silicon based ultrafast sequencer. Our sequencer is constructed using standard VLSI technology and consists of a nanopore surrounded by field effect transistors (FETs). We propose to use the FETs to sense variations in charge as a DNA molecule translocates through the pore and thus differentiate between the four building block nucleotides of DNA. We were able to port SIMU-MD, a serial C code, to the MTA with only a modest effort and with good performance. Our porting process needed neither a parallelism support platform nor attention to the intimate details of parallel programming and interprocessor communication, as would have been the case with more conventional supercomputers.
Abstract The Cray MTA, a multithreaded architecture, is a new parallel supercomputer installed at San Diego Supercomputer Center (SDSC). This machine,has an architecture quite different from those of other contemporary parallel machines. It has a flat, shared memory without locality and has hardware support for very fine-grained multithreading. The machine,and its associated parallelizing compiler promise great ease in scalable parallel computing. We report the results of a study, carried out in July‐September 1999, to evaluate the execution of EUL3D, a code that solves the Euler equations on an unstructured mesh, on the 8 processor MTA at SDSC. EUL3D captures the essential features of most unstructured mesh codes used in aerodynamic,research and development.
The impact of Linux on the developing world is, in many ways, even greater than on industrialized nations. The authors have used Linux productively in Pakistan in both academia and industry; here they detail what they learned in setting up an 84-seat Linux lab for undergraduate teaching, and how an ISP provider in Pakistan is using Linux to improve its services.
The Tera Multithreaded Architecture (MTA) is a new parallel supercomputer currently being installed at San Diego Supercomputing Center (SDSC). This machine has an architecture quite different from contemporary parallel machines. The computational processor is a custom design and the machine uses hardware to support very fine grained multithreading. The main memory is shared, hardware randomized and flat. These features make the machine highly suited to the execution of unstructured mesh problems, which are difficult to parallelize on other architectures. We report the results of a study carried out during July-August 1998 to evaluate the execution of EUL3D, a code that solves the Euler equations on an unstructured mesh, on the 2 processor Tera MTA at SDSC. Our investigation shows that parallelization of an unstructured code is extremely easy on the Tera. We were able to get an existing parallel code (designed for a shared memory machine), running on the Tera by changing only the compiler directives. Furthermore, a serial version of this code was compiled to run in parallel on the Tera by judicious use of directives to invoke the full/empty tag bits of the machine to obtain synchronization. This version achieves 212 and 406 Mflop/s on one and two processors respectively, and requires no attention to partitioning or placement of data issues that would be of paramount importance in other parallel architectures.
Partitioning is an important issue in a variety of applications. Two examples are domain decomposition for parallel computing and color image quantization. In the former we need to partition a computational task over many processors; in the latter we need to partition a high resolution color space into a small number of representative colors. In both cases, partitioning must be done in a manner that yields good results as defined by an application-specific metric. Binary dissection is a technique that has been widely used to partition non-uniform domains over parallel computers. It proceeds by recursively partitioning the given domain into two parts, such that each part has approximately equal computational load. The basic dissection algorithm does not consider the perimeter, surface area or aspect ratio of the two sub-regions generated at each step and can thus yield decompositions that have poor communication to computation ratios. We have developed and implemented several variants of the binary dissection approach that attempt to remedy this limitation, are faster than the basic algorithm, can be applied to a variety of problems, and are amenable to parallelization. We first present the Parametric Binary Dissection (PBD) algorithm, which takes into account volume and surface area when partitioning computational domains for use in parallel computing applications. We then consider another variant, the Fast Adaptive Dissection (FAD) algorithm, which provides rapid spatial partitioning for use in color image quantization. We describe the performance of PBD and FAD on representative problems and present ways of parallelizing the PBD algorithm on 2- or 3-d meshes and on hypercubes.
In the preceding three chapters we have seen how maximum flow and shortest path algorithms can be used to find the optimal assignment of a serial distributed program. Recent research has shown that a sum-bottleneck path algorithm can be employed to find the optimal assignment of the modules of a parallel or pipelined program in several types of distributed systems. This approach can also be used to find the optimal global assignment of a set of independent serial distributed programs over a single-host, multiple-satellite system. Since this technique can explicitly take concurrency into account, it represents a major development over the work presented in the preceding chapters.
Giri Narasimhan合作论文数School of Computing & Information Science
Florida International University1