For over two decades the dominant means for enabling portable performance of computational science and engineering applications on parallel processing architectures has been the bulk-synchronous parallel programming (BSP) model. Code developers, motivated by performance considerations to minimize the number of messages transmitted, have typically pursued a strategy of aggregating message data into fewer, larger messages. Emerging and future high-performance architectures, especially those seen as targeting Exascale capabilities, provide motivation and capabilities for revisiting this approach. In this paper we explore alternative configurations within the context of a large-scale complex multi-physics application and a proxy that represents its behavior, presenting results that demonstrate some important advantages as the number of processors increases in scale.
The computing community is in the midst of a disruptive architectural change. The advent of manycore and heterogeneous computing nodes forces us to reconsider every aspect of the system software and application stack. To address this challenge there is a broad spectrum of approaches, which we roughly classify as either revolutionary or evolutionary. With the former, the entire code base is re-written, perhaps using a new programming language or execution model. The latter, which is the focus of this work, seeks a piecewise path of effective incremental change. The end effect of our approach will be revolutionary in that the control structure of the application will be markedly different in order to utilize single-instruction multiple-data/thread (SIMD/SIMT), manycore and heterogeneous nodes, but the physics code fragments will be remarkably similar. Our approach is guided by a set of mission driven applications and their proxies, focused on balancing performance potential with the realities of existing application code bases. Although the specifics of this process have not yet converged, we find that there are several important steps that developers of scientific and engineering application programs can take to prepare for making effective use of these challenging platforms. Aiding an evolutionary approach is the recognition that the performance potential of the architectures is, in a meaningful sense, an extension of existing capabilities: vectorization, threading, and a re-visiting of node interconnect capabilities. Therefore, as architectures, programming models, and programming mechanisms continue to evolve, the preparations described herein will provide significant performance benefits on existing and emerging architectures.
The Edinburgh Concurrent Supercomputer Project is built around a Meiko Computing Surface, with presently some 400 floating-point transputers and 1.6 Gbytes of memory. The first part of this paper gives a brief overview of the project's origins and status. In the second part we review the results of applications in high energy physics, including lattice gauge theory and Monte Carlo event simulation.
We investigate the low mass properties of our quenched QCD data for the range of beta values: 5.70, 6.00, 6.15 and 6.30. The chiral condensate, the ground-state meson spectrum and the scaling behaviour of various quantities are examined. A criterion for assessing at what value of the quark mass finite lattice effects become intolerable is suggested.
We report results of quenched hadron mass calculations using 163 × 24 lattices, extending our earlier work to higher values of β. We continue to use staggered fermions and local lattice hadron operators. At β = 6.15, there is good flavour symmetry in the meson sector and indications that finite-size effects are small for quark masses between 0.01 and 0.16 in lattice units. Although there is some indication of crossover from the heavy- to the light-quark regime in the nucleon-to-rho mass ratio, finite-size effects prevent us making measurements at quark masses for which the pion is significantly less than half the mass of the rho. At 0β = 6.3 the finite-size effects appear to be substantially worse.
The Computing Surface is a flexible, extensible computing system which is constructed from a choice of intelligent building blocks. A variety of building blocks are available with a range of computational or peripheral capabilities
The remarkable processing capabilities of the nervous system must derive at least in part from the large numbers of neurons participating (roughly 1010), since the timescales involved are of the order of milliseconds, rather than the nanoseconds of modern computers. We summarise common features of the neural network models which attempt to capture this behaviour and describe the many levels of parallelism which they exhibit. A range of models has been implemented on the SIMD (ICL Distributed Array Processor) and MIMD (Meiko Computing Surface) hardware at Edinburgh. Examples include: (i) training algorithms in the context of the Hopfield net, with specific application to the storage of words and text with content-addressable memory; (ii) the back-propagation training algorithm for the multi-layer perception; (iii) image restoration with Hopfield and Tank analogue neurons, and (iv) the Durbin and Willshaw elastic net, as applied to the travelling salesman problem.
We describe methods for solving large sparse systems of linear equations on computers with limited fast memory, or high ratios of processor speed to bandwidth between main and fast memory. Our algorithms are designed for calculating quark propagators, columns of the inverse of the fermion matrix in Lattice Quantum Chromodynamics, but are more generally applicable. We compare their rates of convergence and the balance between CPU time and I/O overhead. We present a block-iterative algorithm which when implemented on the DAP is 5 times as efficient as the Conjugate Gradient Method for 163 × 24 lattices (complex linear systems of size approximately 3 × 105).
Hadron masses are calculated in quenched quantum chromodynamics (QCD) using pure gauge configurations on 164 lattices, periodically extended in time where necessary. We use the Susskind formulation of lattice fermions and construct the propagators for hadrons created by local lattice operators. The inversion of the fermion matrix is accomplished using either the even/odd partitioned conjugate gradient algorithm, or block successive over-relaxation, depending on lattice size. Low statistics results at β = 5.7 confirm the high value of the nucleon-to-rho mass ratio found on smaller lattices. Together with the absence of flavour symmetry in the meson spectrum, this leads us to conclude that continuum behaviour is not observed at this coupling. Higher statistics results at β = 6.0 give a nucleon-to-rho mass ratio of approximately 1.5 and satisfy flavour symmetry in the meson sector to within 10%. We observe two baryon timeslice propagators which become increasingly different as the quark mass is decreased. This effect appears to be exaggerated by our choice of antiperiodic boundary conditions in space and is interpreted as a finite size effect.
The distant source method (DSM) computes quark propagators on large lattices via two or more calculations on smaller lattices. Here it is tested in quenched QCD. Systematic errors are shown to be acceptably small. However, the DSM is not competitive with other algorithms for generating quark propagators from scratch; its usefulness lies in permitting the extension of an existing set of propagators in time.
The remarkable processing capabilities of the nervous system must derive from the large numbers of neurons participating (roughly 1010), since the time-scales involved are of the order of a millisecond, rather than the nanoseconds of modern computers. The neural network models which attempt to capture this behaviour are inherently parallel.
I discuss the Hybrid Monte Carlo algorithm for performing lattice gauge theory calculations. This is a large step method which has none of the discrete step size errors usually associated with the Molecular Dynamics, Langevin, or Hybrid algorithms. The method allows the inclusion of dynamical fermion fields in a straightforward way.
Modern microprocessor technology makes it possible to design array processors to run specific algorithms. We discuss performance estimation for such systems. We use the Occam model of concurrency and plan to study the performance of Transputer arrays.
It is shown that the success of acceleration for abelian gauge field dynamics need not depend on any choice of gauge. The discussion leads to proposing a particular scheme for acceleration in non-abelian theories which is also gauge independent.
We present the results of a high statistics study of the chiral condensate in quenched lattice QCD on an 84 lattice at β = 5.4, 5.5, 5.6, 5.7, 5.8 and 6.0. We see clear evidence for deviation from asymptotic scaling in the range of β considered. Our results are in agreement with the behaviour anticipated from recent Monte Carlo renormalisation group studies of the β-function. We find indications of a common scaling behaviour for the condensate, the string tension and the deconfinement temperature.