The purpose of our work is to provide a method which exploits the parallel blockwise algorithmic approach used in the framework of high performance sparse direct solvers in order to develop robust preconditioners based on a parallel incomplete factorization. The idea is then to define an adaptive blockwise incomplete factorization that is much more accurate (and numerically more robust) than the scalar incomplete factorizations commonly used to precondition iterative solvers.
Solving large sparse symmetric positive definite systems of linear equations is a crucial and time-consuming step, arising in many scientific and engineering applications. The block partitioning and scheduling problem for sparse parallel factorization without pivoting is considered. There are two major aims to this study: the scalability of the parallel solver, and the compromise between memory overhead and efficiency. Parallel experiments on a large collection of irregular industrial problems validate our approach.
In this paper, we present the developments realized in the OURAGAN project around the parallelization of a MATLAB-like tool called SCILAB. These developments use high performance numerical libraries and different approaches based either on the duplication of SCILAB processes or on computational servers. This tool, SCILAB , allows users to perform high level operations on distributed matrices in a metacomputing environment. We also present performance results on different architectures. Key-words: SCILAB , Parallel libraries, Computational servers, CORBA, Data Redistribution.
This paper deals with task scheduling, where each task is one particular iteration of a DO loop with partial loop-carried dependencies. Independent iterations of such loops can be scheduled in an order different from the one of classical serial execution, so as to increase program performance.The approach that we present is based both on the use of a directive added to the High Performance Fortran (KPF2) language, which specifies the dependencies between iterations, and on inspector/executor support, implemented in the CoLUMBO library, which builds the task graph and schedules tasks associated with iterations. We validate our approach by showing results achieved on an IBM SP2 for a sparse Cholesky factorization algorithm applied to real problems. (C) 2000 Elsevier Science B.V. All rights reserved.
Over the last few years a number of changes occurred in the area of high performance computing. First, in the mid-1990’s a number of vendors have ”disappeared” from the high performance computing arena. This process resulted, among others, in a rapid convergence of HPC hardware architectures into few, relatively similar, developmental lines. More recently, we have witnessed the introduction of clusters of cheap and powerful PC’s which provide a more economical way of approaching medium size computational problems. Similarly, in the area of HPC software we also observe a slow convergence toward a few tools, which gained popularity and became de-facto standards for software writing. In addition, a number of software libraries have been developed and by now are in their n-th releases guaranteeing high quality and dependability. These are all signs of the maturing discipline, which is coming out of the initial stages of uncertainty into a stage of sustained growth. While these processes take place, a wide spread problem of the difficulty of moving toward parallel computing is also being recognized. The same way as in the past researchers tried to run their ”dusty-deck” codes on Cray computers and complained about the lack of performance, nowadays (already) vectorized codes have difficulty finding their way to parallel machines. The aim of this tutorial is to provide an overview of the recent developments and state of the art in the areas of high performance hardware, tools and environments, and libraries as related to the matrix algorithms. We will also attempt at summarize the most interesting current research projects that we deem ”worth watching.” The intended audience consists of anyone who is interested in the area of high performance computing. Since the aim of the tutorial is to provide an introduction and an overview of the field, only a minimal background in the computational sciences is required. THE PARALLEL COMPUTATION OF EIGENSYSTEMS Maurice Clint, Queen’s University of Belfast, UK Abstract With the rapid increase in computing power offered by high performance machines there has been a corresponding increase in the size of applications problems which are now tractable. In particular, in many important application areas, it is now commonplace to require the computation of (partial) eigensystems of matrices of order : often, however, these matrices are sparse. Since, in general, direct methods of eigensolution entail unacceptable fill–in and since, often, only partial eigensolutions are required iterative methods have now assumed major importance. Iterative methods employ the matrix of interest only as a multiplier, so sparsity is preserved. In addition the basic operations from which these methods are built are well suited to parallel implementation on a range of different architectures. In this tutorial some aspects of iterative and direct methods for the (partial) eigensolution of real symmetric matrices are addressed in the context of their efficient implementation on a range of high performance computers. In particular, a number of important features of the Lanczos method – including restarting and reorthogonalisation techniques and convergence monitoring – are discussed. The performances (on a number of machines with different architectures) of some recently developed variants of the Lanczos method are analysed and compared with that of a Lanczos routine from the ARPACK library. In addition, some salient features of alternative iterative approaches and direct methods are considered. The tutorial is based on recent work by the presenter, M. Szularz, J.S. Weston, K. Murphy, R.H. Perrott and others.With the rapid increase in computing power offered by high performance machines there has been a corresponding increase in the size of applications problems which are now tractable. In particular, in many important application areas, it is now commonplace to require the computation of (partial) eigensystems of matrices of order : often, however, these matrices are sparse. Since, in general, direct methods of eigensolution entail unacceptable fill–in and since, often, only partial eigensolutions are required iterative methods have now assumed major importance. Iterative methods employ the matrix of interest only as a multiplier, so sparsity is preserved. In addition the basic operations from which these methods are built are well suited to parallel implementation on a range of different architectures. In this tutorial some aspects of iterative and direct methods for the (partial) eigensolution of real symmetric matrices are addressed in the context of their efficient implementation on a range of high performance computers. In particular, a number of important features of the Lanczos method – including restarting and reorthogonalisation techniques and convergence monitoring – are discussed. The performances (on a number of machines with different architectures) of some recently developed variants of the Lanczos method are analysed and compared with that of a Lanczos routine from the ARPACK library. In addition, some salient features of alternative iterative approaches and direct methods are considered. The tutorial is based on recent work by the presenter, M. Szularz, J.S. Weston, K. Murphy, R.H. Perrott and others.
In this paper, we present the HPFIT project whose aim is to provide a set of interactive tools integrated in a single environment to help users to parallelize scientific applications to be run on distributed memory parallel computers. HPFIT is built around a restructuring tool called TransTOOL which includes an editor, a parser, a dependence analysis tool and an optimization kernel. Moreover, we provide a clean interface to help developers of tools around High Performance Fortran to integrate their software within our tool.
The HPFIT project has the aim to provide a set of interactive tools integrated in a single environment to help users to parallelize scientific applications to be run on distributed memory parallel computers. In this paper about the HPFIT project we present a data structure visualization tool called Visit, and HPF extensions for irregular problems.
In this paper, we consider the problem of data partitioning for block sparse Cholesky factorization on distributed memory MIMD computers. We propose a preprocessing algorithm which computes and distributes a column block partition based on an initial partition induced by a nested dissection ordering. This preprocessing algorithm works by optimizing load balancing under precedence constraints and communication traffic. It can be performed in linear time and space complexities.
Serge Chaumette合作论文数Computer Science Department;Universit?? Bordeaux 17
Pierre Ramet合作论文数INRIA Bordeaux Sud-Ouest Universitée Bordeaux 11
Pascal Hénon合作论文数INRIA associated team PhYleAS1
Olivier Beaumont合作论文数LaBRI - Laboratoire Bordelais de Recherche en Informatique,;Projet INRIA C??page;Universit?? Bordeaux 11