
This paper presents parallel algorithms for MPEG video compression by using the PVM library. Because of the huge amount of computation, a sequential software encoder is slow at the coding of reasonably sized images (e.g. 0.1-0.3 frames per second). In order to speed up the process, we developed two kinds of parallel algorithm, which run on a cluster of workstations connected to the network The first method encodes slices in parallel, while the second one encodes equally sized Groups Of Pictures (GOP) in parallel. The communication between the machines is facilitated by the PVM package. The reference sequential algorithm can encode a SIF sized sequence at 0.2 frames per second on a SUN Spare 2 using the GOP pattern of IBBP. The slice level parallelization can take this rate up to nearly 0.7 frames per second with the use of 6 Spare 2s connected via Ethernet. The degradation of the performance is a consequence of the low bandwidth of the network and the relatively high communication activity. The GOP level parallelization provides an encoding rate of nearly 1.0 frame per second at the same conditions. In this case, the communication is reduced significantly, however, the load balancing may be slightly worse if the size of the GOP is large (greater than or equal to 10). Our parallel algorithms are entirely portable and produce bitstreams fully conforming the MPEG-1 standard. The methods can be applied to MPEG-2 encoding as well. The system performs nearly as well as the parallel Berkeley encoder by using the second algorithm. However, it produces a video stream of better quality because the encoding is done correctly.
In this paper we compare the different programming paradigms available on the Cray T3D for the implementation of a 3D prototype of an Atmospheric Chemistry/Transport Model. We discuss the amount of work needed to convert existing codes to the T3D and the portability of the resulting codes. Tests show that the scalability with respect to the model size and the number of processors is linear except for the data parallel implementation.
Sharing information and resources between homogenous computers is already operational on computers networks. However, the development of communication technologies increases heterogeneity of the machines which are involved in network communication and this heterogeneity is a real problem for the specification of distributed applications. To propose solutions to this problem, International Standards Organization (ISO) offered, in 1993, a document, the International Draft Proposal of Remote Procedure Call (RPC) mechanism, in order to share power computing resources between heterogeneous computers on a wide or local area network. RPC is one of the most simple and useful mean of communication used to share information and resources. The ISO RPC offers the same functions as the RPC for homogenous environment. The ISO model proposes the use of ROSE, Presentation, Session upper layers to provide a solution for the heterogeneity.Due to the simplicity of RPC mechanism, it would be interesting to use it to specify industrial distributed applications. For this, it is necessary to verify that ISO RPC suits to the time constraints of the production. In this intention, we have developed a RPC program, according to the 1993 RPC draft document and we have measured its performances, in a context which is currently one of the most used in industry: the ISO stack for upper layers communication protocols, with TCP/IP environment on Ethernet network.Two kinds of measures have been realized; data transfer capacity which is the measure of the real data flow the RPC can give, and latency time which is a measure of the duration of a simple call on the remote computer without any execution of program.The analysis of ISO RPC performances have been made in comparison with the most famous and widespread RPC within homogenous environment, the SUN RPC system. From the features of distributed applications used in industry and an analysis of the performances obtained with our software, we point out the lacunas of the ISO RPC model and we propose improvements of this model for industrial implementation.
A ''Divide and Conquer'' (D&C) algorithm enriched with a dynamic reconfiguration is discussed. The algorithm has been designed to run on a tree structure embedded in a hypercube. Application on symbolic normal form computations (on a 64-node nCUBE2 computer) showed that under certain conditions ''D&C with Reconfiguration'' may perform considerably better than pure D&C.
Time-accurate simulation of turbulent flames in high Reynolds number flows is a challenging task since both fluid dynamics and combustion must be modeled accurately. To numerically simulate this phenomenon, very large computer resources (both time and memory) are required. Although current vector supercomputers are capable of providing adequate resources for simulations of this nature, the high cost and their limited availability, makes practical use of such machines less than satisfactory. At the same time, the explicit time integration algorithms used in unsteady flow simulations often possess a very high degree of parallelism, making them very amenable to efficient implementation on large-scale parallel computers. Under these circumstances, distributed memory parallel computers offer an excellent near-term solution for greatly increased computational speed and memory, at a cost that may render the unsteady simulations of the type discussed above more feasible and affordable.This paper discusses the study of unsteady turbulent flames using a simulation algorithm that is capable of retaining high parallel efficiency on distributed memory parallel architectures. Numerical studies are carried out using large-eddy simulation (LES). In LES, the scales larger than the grid are computed using a time- and space-accurate scheme, while the unresolved small scales are modeled using eddy viscosity based subgrid models. This is acceptable for the moment/energy closure since the small scales primarily provide a dissipative mechanism for the energy transferred from the large scales. However, for combustion to occur, the species must first undergo mixing at the small scales and then come into molecular contact. Therefore, global models cannot be used. Recently, a new model for turbulent combustion was developed, in which the combustion is modeled, within the subgrid (small-scales) using a methodology that simulates the mixing and the molecular transport and the chemical kinetics within each LES grid cell. Finite-rate kinetics can be included without any closure and this approach actually provides a means to predict the turbulent rates and the turbulent flame speed. The subgrid combustion model requires resolution of the local time scales associated with small-scale mixing, molecular diffusion and chemical kinetics and, therefore, within each grid cell, a significant amount of computations must be carried out before the large-scale (LES resolved) effects are incorporated. Therefore, this approach is uniquely suited for parallel processing and has been implemented on various systems such as: Intel Paragon, IBM SP-2, Cray T3D and SGI Power Challenge (PC) using the system independent Message Passing Interface (MPI) compiler. In this paper, timing data on these machines is reported along with some characteristic results.
Massively parallel processors require new formalisms to effectively control and utilise such platforms. Declarative languages, with their property of implicit parallelism, represent one such formalism. However, a difficulty with such languages is providing efficient and timely implementations on new architectures. This paper is concerned with the support of a computational model for a parallel functional language, Hope(+) [6], on top of a portable run-time system, Nexus [2,3].
The rapidly evolving changes in wireless communication technologies increase the importance of radio-monitoring and surveillance. In this paper, we present a scalable distributed model for radio-reconnaissance. The HOPPER system is designed as a net of geographically distributed bearing units connected to a message-passing network of computers and commanded by a decentralized command unit. The model has been implemented in portable simulation prototypes, used to develop, test and analyze distributed identification methods based on structure recognition and pattern matching. We present a method that constructs in parallel synthesized pattern signatures by segmentation and that locates emitters by direction area intersections.
In this paper, a short overview of project 3, “Advanced Flux Modelling”, in the NOWESP project of the MAST II programme is given. The computational requirements of large scale 3-D flow and transport models of the North-West European Shelf are considered. The possibilities and implications of massively parallel processors for the simulation of these models are discussed. We also present some of the results and experience obtained in parallelising the shallow water flow and transport models within the NOWESP project.
The finite element package DIANA is being ported to a distributed environment for several reasons - parallel execution after domain decomposition, but also GUI and concurrent engineering come into reach. in this paper we first explain how we look at parallel programming in general, thereby justifying the use of distributed data structures for the realization of a distributed environment. Then we discuss the design decisions taken during the development of a pilot version. We finish with an evaluation of the pilot version.
Traditionally, atmospheric models have avoided the use of fully compressible governing equations since the length of a time step in such a model would be limited by the acoustic modes. Tanguay et al 11] have shown that the semi-implicit technique of Robert, originally applied to the primitive equations in large-scale models of the atmosphere, can be successfully applied to integrate the fully compressible non-hydrostatic equations. More recently Robert 10] applied a semi-implicit, semi-Lagrangian integration scheme to the Euler equations in order to simulate the convection of warm air in a dry isen-tropic atmosphere and demonstrated that this scheme works well at convective scales in the atmosphere The bubble convection model of Robert 10] serves as a test problem for compressible non-hydrostatic atmospheric codes. We describe a distributed-memory implementation of the bubble convection problem and discuss parallel implementation issues related to both advection and the Helmholtz solver. The accuracy, stability, and performance of semi-Lagrangian advection are determined by the trajectory integration and interpolation algorithms. More eecient trajectory and piecewise interpolation algorithms have now been incorporated into the model. For the Helmholtz solver, the choice of an eecient parallel preconditioner is crucial. The portable, distributed-memory implementation of the bubble convection problem will serve as a prototype for the development of a parallel mesoscale atmospheric model.
An important obstacle for an industrial break through of parallel computing is the complexity of parallel programming and of porting existing simulation packages to parallel computers. In the current paper a rigorous approach is worked out which simplifies the port of a package and parallel programming in general. The approach offers flexibility in modifying or extending a simulation package by identifying stages or levels of abstraction each containing a formulation of the (class of) problem(s) to be simulated. Flexibility, 'high level of abstraction', prevention of side-effects and other advantages will be discussed and illustrated.
A global snapshot scheme, modifying the distributed scheme in [4] is presented. The modified scheme is based on the echo algorithm [2], and uses a concurrent, hierarchical approach. A correctness proof is outlined, and an analysis of the message complexity given.
The partitioning of 3D topologically rectangular finite element grids onto 2D mesh connected arrays of processors is considered. A simple rectilinear grid partitioning is found to be most effective when the span of the processor mesh is an integer divisor of the problem grid size. Otherwise, a simulated annealing procedure may be used to improve load balancing. The details and tuning parameters of this annealing procedure are given.
A hybrid of the consistent and independent checkpointing schemes is presented. Coordination between nodes is less than normally required in consistent checkpointing schemes.
The acceptance of parallel computing in workstation clusters has increased in the past years. One important reason for this is the cost-efficiency of workstation clusters as an alternative to specialized distributed-memory parallel computer systems. A potential bottleneck for distributed-memory architectures is the interconnection network between the processing elements. This is the main disadvantage of clusters which arises due to the local area network (LAN) connecting the workstations. A LAN does not reach the low latency, high bandwidth and capacity of a specialized interconnection network used in distributed-memory parallel computer architectures.In this paper, a new workstation cluster interconnection network architecture which increases the communication performance of clusters is presented and evaluated. in contrast to traditional workstation clusters, the Concurrent Network Architecture (CNA) improves the interconnection network of a workstation cluster with the introduction of a structured network consisting of several parallel and independent LAN communication channels.The basic idea of independent LAN channels within a CNA can be achieved by a common technology such as Ethernet or FastEthernet. Hereby, all components of a realization of a CNA workstation cluster are based on standards and protocols such as the internet protocols (TCP/IP, UDP/IP). Furthermore, the CNA is able to migrate to new LAN technology as well as to integrate advances in protocol standards, resulting in a higher communication performance. Due to the use of common standards, the communication performance of a CNA cluster can be improved compared to a traditional cluster (using a single, shared LAN channel) while still maintaining the cost-efficiency of workstation clusters compared to specialized parallel computer architectures.An exemplary implementation of the CNA is built on Ethernet (10Base-2) as a cost-efficient LAN standard. A message-passing system was designed for the CNA in order to support an application programmer to utilize the structured Ethernet channels. The CNA and its message-passing system is entirely based on the use of standard hardware components and other common standards such as UNIX, the TCP/IP protocols and Ethernet.
Over the last three years, we at the NOAA Forecast Systems Laboratory (FSL) have been developing a numerical weather prediction model specifically designed to run well on Massively Parallel Processors (MPPs). In addition, our model uses a new quasi-nonhydrostatic well-posed meteorological formulation based on the ''approximate system'' of Browning and Kreiss [1]. After extensively testing the model - called the Quasi-NonHydrostatic meteorological model (QNH)-we have recently added the capability of inputing weather observations and are doing short-term forecasts. The model can output the forecast as three-dimensional graphics, rendered using the Advanced Visualization System (AVS) running on a graphics workstation, as well as the usual two-dimensional weather map contour plots. QNH is coded using the Scalable Modeling System (SMS) parallel software developed at FSL, which provides portability as well as high-performance on most MPPs, and contains portable and efficient parallel I/O.