The single source shortest path (SSSP) problem lacks parallel solutions which are fast and simultaneously work-efficient. We propose simple criteria which divide Dijkstra's sequential SSSP algorithm into a number of phases, such that the operations within a phase can be done in parallel. We give a PRAM algorithm based on these criteria and analyze its performance on random digraphs with random edge weights uniformly distributed in [0,1]. We use the G (n, d/n) model: the graph consists of n nodes and each edge is chosen with probability d/n. Our PRAM algorithm needs O(n 1/3 log n) log n) time and O (n log n+dn) work with high probability (whp). We also give extensions to external memory computation. Simulations show the applicability of our approach even on non-random graphs.
The construction of full-text indexes on very large text collections is nowadays a hot problem. The suffix array [16] is one of the most attractive full-text indexing data structures due to its simplicity, space efficiency and powerful/fast search operations supported. In this paper we analyze theoretically and experimentally, the I/O-complexity and the working space of six algorithms for constructing large suffix arrays. Additionally, we design a new external-memory algorithm that follows the basic philosophy underlying the algorithm in [13] but in a significantly different manner, thus combining its good practical qualities with efficient worst-case performances. At the best of our knowledge, this is the first study which provides a wide spectrum of possible approaches to the construction of suffix arrays in external memory, and thus it should be helpful to anyone who is interested in building full-text indexes on very large text collections.
Data to be processed has dramatically increased during the last years. Nowadays, external memory (mostly hard disks) has to be used to store this massive data. Algorithms and data structures that work on external memory have different properties and specialties that distinguish them from algorithms and data structures, developed for the RAM model. In this thesis, we first explain the functionality of external memory,which is realized by disk drives. We then introduce the most important theoretical I/O models. In the main part, we present the C++ class library LEDA-SM. Library LEDA-SM is an extension of the LEDA library towards external memory computation and consists of a collection of algorithms and data structures that are designed to work efficiently in external memory. In the last two chapters, we present new external memory data structures for external memory priority queues and new external memory construction algorithms for suffix arrays. These new proposals are theoretically analyzed and experimentally tested. All proposals are implemented using the LEDA-SM library. Their efficiency is evaluated by performing a large number of experiments. Die zu verarbeitenden Datenmengen sind in den letzten Jahren dramatisch gestiegen, so das Externspeicher (in Form von Festplatten) eingesetzt wird, um die Datenmengen zu speichern. Algorithmen und Datenstrukturen, die den Externspeicher benutzen, haben andere algorithmische Anforderungen als eine Vielzahl der bekannten Algorithmen und Datenstrukturen, die fur das RAM-Modell entwickelt wurden. Wir geben in dieser Arbeit erst einen Einblick in die Funktionsweise von Externspeicher anhand von Festplatten und erklaren die wichtigsten theoretischen Modelle, die zur Analyse von Algorithmen benutzt werden. Weiterhin stellen wir eine neu entwickelte C++ Klassenbibliothek namens LEDA-SM vor. LEDA-SM bietet eine Sammlung von speziellen Externspeicher Algorithmen und Datenstrukturen. Im zweiten Teil entwickeln wir neue Externspeicher-Prioritatswarteschlangen und neue Externspeicher- Konstruktionsalgorithmen fur Suffix Arrays. Unsere neuen Verfahren werden theoretisch analysiert, mit Hilfe von LEDA-SM implementiert und anschliesend experimentell getestet.
During the last years, many software libraries for in-core computation have been developed. Most internal memory algorithms perform very badly when used in an external memory setting. We introduce LEDA-SM that extends the LEDA-library [22] towards secondary memory computation. LEDA-SM uses I/O-efficient algorithms and data structures that do not suffer from the so called I/O bottleneck. LEDA is used for in-core computation. We explain the design of LEDA-SM and report on performance results.
A priority queue is a data structure that stores a set of items, each one consisting of a tuple which contains some (satellite) information plus a priority value (also called key) drawn from a totally ordered universe. A priority queue supports the following operations on the processed set: access_minimum (returns the item in the set having minimum key), delete_min (returns and deletes the item in the set having the minimum key) and insert (inserts a new item into the set). Priority queues (hereafter PQs) have numerous important applications: combinatorial optimization (e.g. Dijkstra’s shortest path algorithm [7]), time forward processing [5], job scheduling, event simulation and online sorting, just to cite a few. Many PQ implementations currently exist for small data sets fitting into the internal memory of the computer, e.g. k—ary heaps [23], Fibonacci heaps [10], radix heaps [1], and some of them are also publicly available to the programmers (see e.g. the LEDA library [15]). However, in large-scale event simulations or on instances of very large graph problems (as they recently occur in e.g. geographical information systems), the performance of these internal-memory PQs may significantly deteriorate, thus being a bottleneck for the overall application. In fact, as soon as parts of the PQs do not fit entirely into the internal memory of the computer, but reside in its external memory (e.g. in the hard disk), we may observe a heavy paging activity of the external-memory devices because the pattern of memory accesses is not tuned to exhibit any locality of reference. Due to the technological features of current disk systems [17], this situation may determine a slow down of 5 or 6 orders of magnitude in the final performance of each PQ-operation 1. Consequently, it is required to design PQs which take explicitly into account the physical properties of the disk systems in order to achieve efficient I/O-performances that allow these data structures to be plugged successfully in software libraries.
We show that the well-known random incremental construction of Clarkson and Shor18 can be adapted to provide efficient external-memory algorithms for some geometric problems. In particular, as the main result, we obtain an optimal randomized algorithm for the problem of computing the trapezoidal decomposition determined by a set of N line segments in the plane with K pairwise intersections, that requires [Formula: see text] expected disk accesses, where M is the size of the available internal memory and B is the size of the block transfer. The approach is sufficiently general to derive algorithms for other geometric problems: 3-d half-space intersections, 2-d and 3-d convex hulls, 2-d abstract Voronoi diagrams and batched planar point location; these algorithms require an optimal expected number of disk accesses and are simpler than the ones previously known. The results extend to an external-memory model with multiple disks.
Contents1 Introduction 31.1 License Terms . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 31.2 Overview . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 41.3 The External Memory Model of Computation . . . . . . . . . . . . . . . . . . . . . . . . . 42 Basics 62.1 User Defined Parameter Types . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 62.1.1 A starting example . . . . . . . . . . . . . . ...
Article Free Access Share on q-gram based database searching using a suffix array (QUASAR) Authors: Stefan Burkhardt MPI für Informatik, Im Stadtwald, 66123 Saarbrücken, Germany MPI für Informatik, Im Stadtwald, 66123 Saarbrücken, GermanyView Profile , Andreas Crauser MPI für Informatik, Im Stadtwald, 66123 Saarbrücken, Germany MPI für Informatik, Im Stadtwald, 66123 Saarbrücken, GermanyView Profile , Paolo Ferragina Dipartimento di Informatica, Università di Pisa, Corso Italia 56125 Pisa, Italy Dipartimento di Informatica, Università di Pisa, Corso Italia 56125 Pisa, ItalyView Profile , Hans-Peter Lenhof MPI für Informatik, Im Stadtwald, 66123 Saarbrücken, Germany MPI für Informatik, Im Stadtwald, 66123 Saarbrücken, GermanyView Profile , Eric Rivals Deutsches Krebsforschungszentrum, Abt. Theoretische Bioinformatik , INF 280, D-69120 Heidelberg, Germany Deutsches Krebsforschungszentrum, Abt. Theoretische Bioinformatik , INF 280, D-69120 Heidelberg, GermanyView Profile , Martin Vingron Deutsches Krebsforschungszentrum, Abt. Theoretische Bioinformatik , INF 280, D-69120 Heidelberg, Germany Deutsches Krebsforschungszentrum, Abt. Theoretische Bioinformatik , INF 280, D-69120 Heidelberg, GermanyView Profile Authors Info & Claims RECOMB '99: Proceedings of the third annual international conference on Computational molecular biologyApril 1999 Pages 77–83https://doi.org/10.1145/299432.299460Online:01 April 1999Publication History 71citation941DownloadsMetricsTotal Citations71Total Downloads941Last 12 Months13Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
We investigate the I/O-complexity of computing the trapezoidal decomposition deened by a set of N line segments in the plane. We present a randomized algorithm which solves optimally this problem requiring O(N B log M=B N B + K B) expected I/O operations, where K is the number of pairwise intersections, M is the size of available internal memory and B is the size of the block transfer. The proposed algorithm requires an optimal expected number of internal operations. As a by-product, the algorithm also solves the segment intersections problem requiring the same number of I/Os and internal operations.