In this paper we present a new incomplete factorization of a square matrix into triangular factors in which we get standard $LU$ or $LDL^T$ factors (direct factors) and their inverses (inverse factors) at the same time. Algorithmically, we derive this method from the approach based on the Sherman–Morrison formula [R. Bru, J. Cerdán, J. Marín, and J. Mas, SIAM J. Sci. Comput., 25 (2003), pp. 701–715]. In contrast to the robust incomplete decomposition (RIF) algorithm [M. Benzi and M. Tůma, Numer. Linear Algebra Appl., 10 (2003), pp. 385–400] the direct and inverse factors here directly influence each other throughout the computation. Consequently, the algorithm to compute the approximate factors may mutually balance dropping in the factors and control their conditioning in this way. For the symmetric positive definite case, we derive the theory and present an algorithm for computing the incomplete $LDL^T$ factorization, and we discuss experimental results. We call this new approximate $LDL^T$ factorization the balanced incomplete factorization (BIF). Our experimental results confirm that this factorization is very robust and may be useful in solving difficult ill conditioned problems by preconditioned iterative methods. Moreover, the internal coupling of the computation of direct and inverse factors results in much shorter setup times (times to compute approximate decomposition) than RIF, a method of a similar and very high level of robustness. We also derive and present the theory for the general nonsymmetric case, but do not discuss its implementation.
In the paper a mixed-hybrid approximation of the potential fluid flow problem based on prismatic discretization of the domain is presented. Trilateral prismatic elements with vertical faces and nonparallel bases suitable for the modelling of real geological circumstances are considered. The set of linearly independent vector basis functions is defined and existence and uniqueness of the approximate solution from the resulting symmetric indefinite system are examined. Possible approaches to the solution of the discretized system are discussed.
A two-level additive Schwarz preconditioning scheme for so lving Ciarlet-Raviart, Hermann-Miyoshi, and Hellan-Hermann-Johnson mixed metho d equations for the biharmonic Dirichlet problem is presented. Using suitably defined mesh-dependent forms, a unified approach, with ties to the work of Brenner for n nconforming methods, is provided. In particular, optimal preconditioning o f a Schur complement formulation for these equations is proved on polygonal domains w thout slits, provided the overlap between subdomains is sufficiently large.
We consider matrix-free solver environments where information about the underlying matrix is available only through matrix vector computations which do not have access to a fully assembled matrix. We introduce the notion of partial matrix estimation for constructing good algebraic preconditioners used in Krylov iterative methods in such matrix-free environments, and formulate three new graph coloring problems for partial matrix estimation. Numerical experiments utilizing one of these formulations demonstrate the viability of this approach.
Some details of arithmetic of two representatives of computers (a superscalar workstation and a vector uniprocessor) available in the Czech Republic for scientific computing are described. Consequently, their efficiency and precision on a set of linear algebraic tasks solved by different solvers is compared.
In the paper we deal with some computational aspects of the Generalized minimal residual method (GMRES) for solving systems of linear algebraic equations. The key question of the paper is the importance of the orthogonality of computed vectors and its influence on the rate of convergence, numerical stability and accuracy of different implementations of the method. Practical impact on the efficiency in the parallel computer environment is considered.
This paper describes a technique for constructing robust preconditioners for the CGLS method applied to the solution of large and sparse least squares problems. The algorithm computes an incomplete LDLT factorization of the normal equations matrix without the need to form the normal matrix itself. The preconditioner is reliable (pivot breakdowns cannot occur) and has low intermediate storage requirements. Numerical experiments illustrating the performance of the preconditioner are presented. A comparison with incomplete QR preconditioners is also included.
We describe a novel technique for computing a sparse incomplete factorization of a general symmetric positive definite matrix A . The factorization is not based on the Cholesky algorithm (or Gaussian elimination), but on A ‐orthogonalization. Thus, the incomplete factorization always exists and can be computed without any diagonal modification. When used in conjunction with the conjugate gradient algorithm, the new preconditioner results in a reliable solver for highly ill‐conditioned linear systems. Comparisons with other incomplete factorization techniques using challenging linear systems from structural analysis and solid mechanics problems are presented. Copyright © 2003 John Wiley & Sons, Ltd.
Sparse linear systems Kx = b are considered, where K is a specially structured symmetric indefinite matrix. These systems arise frequently, e.g., from mixed finite element discretizations of PDE problems. The LDL T factorization of K with diagonal D and unit lower triangular L is known to exist for natural ordering of K , but the resulting triangular factors can be rather dense. On the other hand, for a given permutation matrix P , the LDL T factorization of P TKP may not exist.In this paper a new way to obtain a fill-in minimizing permutation based on initial fill-in minimizing ordering is introduced. For an important subclass of matrices arising from mixed and hybrid finite element discretizations, the existence of the LDL T factorization of the permuted matrix is proved. Experimental results on practical problems indicate that the amount of computational savings can be substantial when compared with the approach based on Schur complement.
We consider the parallel computation of the stationary probability distribution vector of ergodic Markov chains with large state spaces by preconditioned Krylov subspace methods. The parallel preconditioner is obtained as an explicit approximation, in factorized form, of a particular generalized inverse of the generator matrix of the Markov process. Graph partitioning is used to parallelize the whole algorithm, resulting in a two-level method.Conditions that guarantee the existence of the preconditioner are given, and the results of a parallel implementation are presented. Our results indicate that this method is well suited for problems in which the generator matrix can be explicitly formed and stored.
Mixed-hybrid finite element approximation of the potential fluid flow problem leads to the solution of a large symmetric indefinite system for the velocity and potential head vector components. Such discretization gives rise to a very accurate approximation of the continuity equation in every element, and for low-order discretizations, the structural properties of the discrete matrix blocks allow cheap block elimination of the positive-definite diagonal block and subsequent reduction to the Schur complement system for the pressure and Lagrangian vector components. This system is then frequently solved by the iterative conjugate gradient-type method. Whereas this approach is well known, considerably less attention has been paid to the numerical stability aspects of such transformation. It was shown in [5] that block LU factorization can be unstable even when the system matrix is symmetric positive definite. In this paper we examine this type of conditional stability for a particular application in the underground water flow modelling. We show that the actual error of the computed approximate solution depends not only on the user-defined tolerance in the conjugate gradient process but also on the spectral properties of the corresponding matrix blocks eliminated during the Schur complement reduction. It is often observed that although the backward error of the approximate solution in the iterative part is reduced to the level of machine accuracy, the total residual norm after the back-substitution process remains at certain accuracy level. We give a bound for this maximal attainable accuracy and illustrate our theoretical results on a model example.
We present a variant of the AINV factorized sparse approximate inverse algorithm which is applicable to any symmetric positive definite matrix. The new preconditioner is breakdown-free and, when used in conjunction with the conjugate gradient method, results in a reliable solver for highly ill-conditioned linear systems. We also investigate an alternative approach to a stable approximate inverse algorithm, based on the idea of diagonally compensated reduction of matrix entries. The results of numerical tests on challenging linear systems arising from finite element modeling of elasticity and diffusion problems are presented.
The influence of reorderings on the performance of factorized sparse approximate inverse preconditioners is considered. Some theoretical results on the effect of orderings on the fill-in and decay behavior of the inverse factors of a sparse matrix are presented. It is shown experimentally that certain reorderings, like minimum degree and nested dissection, can be very beneficial. The benefit consists of a reduction in the storage and time required for constructing the preconditioner, and of faster convergence of the preconditioned iteration in many cases of practical interest.
The mixed-hybrid finite element discretization of Darcy's law and continuity equation describing the potential fluid flow problem in porous media leads to a symmetric indefinite linear system for the pressure and velocity vector components. As a method of solution the reduction to three Schur complement systems based on successive block elimination is considered. The first and second Schur complement matrices are formed eliminating the velocity and pressure variables, respectively, and the third Schur complement matrix is obtained by elimination of a part of Lagrange multipliers that come from the hybridization of a mixed method. The structural properties of these consecutive Schur complement matrices in terms of the discretization parameters are studied in detail. Based on these results the computational complexity of a direct solution method is estimated and compared to the computational cost of the iterative conjugategradient method applied to Schur complement systems. It is shown that due to special block structure the spectral properties of successive Schur complement matrices do not deteriorate and the approach based on the block elimination and subsequent iterative solution is well justified. Theoretical results are illustrated by numerical experiments.
Over the last few years a number of changes occurred in the area of high performance computing. First, in the mid-1990’s a number of vendors have ”disappeared” from the high performance computing arena. This process resulted, among others, in a rapid convergence of HPC hardware architectures into few, relatively similar, developmental lines. More recently, we have witnessed the introduction of clusters of cheap and powerful PC’s which provide a more economical way of approaching medium size computational problems. Similarly, in the area of HPC software we also observe a slow convergence toward a few tools, which gained popularity and became de-facto standards for software writing. In addition, a number of software libraries have been developed and by now are in their n-th releases guaranteeing high quality and dependability. These are all signs of the maturing discipline, which is coming out of the initial stages of uncertainty into a stage of sustained growth. While these processes take place, a wide spread problem of the difficulty of moving toward parallel computing is also being recognized. The same way as in the past researchers tried to run their ”dusty-deck” codes on Cray computers and complained about the lack of performance, nowadays (already) vectorized codes have difficulty finding their way to parallel machines. The aim of this tutorial is to provide an overview of the recent developments and state of the art in the areas of high performance hardware, tools and environments, and libraries as related to the matrix algorithms. We will also attempt at summarize the most interesting current research projects that we deem ”worth watching.” The intended audience consists of anyone who is interested in the area of high performance computing. Since the aim of the tutorial is to provide an introduction and an overview of the field, only a minimal background in the computational sciences is required. THE PARALLEL COMPUTATION OF EIGENSYSTEMS Maurice Clint, Queen’s University of Belfast, UK Abstract With the rapid increase in computing power offered by high performance machines there has been a corresponding increase in the size of applications problems which are now tractable. In particular, in many important application areas, it is now commonplace to require the computation of (partial) eigensystems of matrices of order : often, however, these matrices are sparse. Since, in general, direct methods of eigensolution entail unacceptable fill–in and since, often, only partial eigensolutions are required iterative methods have now assumed major importance. Iterative methods employ the matrix of interest only as a multiplier, so sparsity is preserved. In addition the basic operations from which these methods are built are well suited to parallel implementation on a range of different architectures. In this tutorial some aspects of iterative and direct methods for the (partial) eigensolution of real symmetric matrices are addressed in the context of their efficient implementation on a range of high performance computers. In particular, a number of important features of the Lanczos method – including restarting and reorthogonalisation techniques and convergence monitoring – are discussed. The performances (on a number of machines with different architectures) of some recently developed variants of the Lanczos method are analysed and compared with that of a Lanczos routine from the ARPACK library. In addition, some salient features of alternative iterative approaches and direct methods are considered. The tutorial is based on recent work by the presenter, M. Szularz, J.S. Weston, K. Murphy, R.H. Perrott and others.With the rapid increase in computing power offered by high performance machines there has been a corresponding increase in the size of applications problems which are now tractable. In particular, in many important application areas, it is now commonplace to require the computation of (partial) eigensystems of matrices of order : often, however, these matrices are sparse. Since, in general, direct methods of eigensolution entail unacceptable fill–in and since, often, only partial eigensolutions are required iterative methods have now assumed major importance. Iterative methods employ the matrix of interest only as a multiplier, so sparsity is preserved. In addition the basic operations from which these methods are built are well suited to parallel implementation on a range of different architectures. In this tutorial some aspects of iterative and direct methods for the (partial) eigensolution of real symmetric matrices are addressed in the context of their efficient implementation on a range of high performance computers. In particular, a number of important features of the Lanczos method – including restarting and reorthogonalisation techniques and convergence monitoring – are discussed. The performances (on a number of machines with different architectures) of some recently developed variants of the Lanczos method are analysed and compared with that of a Lanczos routine from the ARPACK library. In addition, some salient features of alternative iterative approaches and direct methods are considered. The tutorial is based on recent work by the presenter, M. Szularz, J.S. Weston, K. Murphy, R.H. Perrott and others.
Standard preconditioners, like incomplete factorizations, perform well when the coefficient matrix is diagonally dominant, but often fail on general sparse matrices. We experiment with nonsymmetric permutations and scalings aimed at placing large entries on the diagonal in the context of preconditioning for general sparse matrices. The permutations and scalings are those developed by Olschowka and Neumaier [Linear Algebra Appl., 240 (1996), pp. 131--151] and by Duff and Koster [SIAM J. Matrix Anal. Appl., 20 (1999), pp. 889--901; Tech. report Ral-Tr-99-030, Rutherford Appleton Laboratory, Chilton, UK, 1999]. We target highly indefinite, nonsymmetric problems that cause difficulties for preconditioned iterative solvers. Our numerical experiments indicate that the reliability and performance of preconditioned iterative solvers are greatly enhanced by such preprocessing.
A number of recently proposed preconditioning techniques based on sparse approximate inverses are considered. A description of the preconditioners is given, and the results of an experimental comparison performed on one processor of a Cray C98 vector computer using sparse matrices from a variety of applications are presented. A comparison with more standard preconditioning techniques, such as incomplete factorizations, is also included. Robustness, convergence rates, and implementation issues are discussed.
This paper is concerned with a new approach to preconditioning for large, sparse linear systems. A procedure for computing an incomplete factorization of the inverse of a nonsymmetric matrix is developed, and the resulting factorized sparse approximate inverse is used as an explicit preconditioner for conjugate gradient-type methods. Some theoretical properties of the preconditioner are discussed, and numerical experiments on test matrices from the Harwell-Boeing collection and from Tim Davis's collection are presented. Our results indicate that the new preconditioner is cheaper to construct than other approximate inverse preconditioners. Furthermore, the new technique insures convergence rates of the preconditioned iteration which are comparable with those obtained with standard implicit preconditioners.
In the paper the potential fluic flow problem in porous media using Darcy's law and the continuity equation is solved. Mixed-hybrid finite element formulation based on general trilateral prismatic elements is considered. Spectral properties of resulting symmetric indefinite system of linear equations are examined. Minimal residual method for the solution of systems with a symmetric indefinite matrix is applied. The rate of convergence and the asymptotic convergence factor which depend on the eigenvalue distribution of the system matrix are estimated.
A method for computing a sparse incomplete factorization of the inverse of a symmetric positive definite matrix A is developed, and the resulting factorized sparse approximate inverse is used as an explicit preconditioner for conjugate gradient calculations. It is proved that in exact arithmetic the preconditioner is well defined if A is an H-matrix. The results of numerical experiments are presented.
Olivier Beaumont合作论文数LaBRI - Laboratoire Bordelais de Recherche en Informatique,;Projet INRIA C??page;Universit?? Bordeaux 11