It is an established fact that the network topology can have an impact on the performance of scientific parallel applications. However, little work has been done to design an easy to use solution inside a communication library supporting a parallel programming model where the complexities of making the application performance network topology agnostic is hidden from the end user. Similarly, the rapid improvements in networking technology and speed are resulting in many commodity clusters becoming heterogeneous, with respect to networking speed. For example, switches and adapters belonging to different generations (SDR - 8 Gbps, DDR - 16 Gbps and QDR - 36 Gbps speeds in InfiniBand) are integrated into a single system. This leads to an additional challenge to make the communication library aware of the performance implications of heterogeneous link speeds. Accordingly, the communication library can perform optimizations taking link speed into account. In this paper, we propose a framework to automatically detect the topology and speed of an InfiniBand network and make it available to users through an easy to use interface. We also make design changes inside the MPI library to dynamically query this topology detection service and to form a topology model of the underlying network. We have redesigned the broadcast algorithm to take into account this network topology information and dynamically adapt the communication pattern to best fit the characteristics of the underlying network. To the best of our knowledge, this is the first such work for InfiniBand clusters. Our experimental results show that, for large homogeneous systems and large message sizes, we get up to 14% improvement in the latency of the broadcast operation using our proposed network topology-aware scheme over the default scheme at the micro-benchmark level. At the application level, the proposed framework delivers up to 8% improvement in total application run-time especially as job size scales up. The proposed network speed-aware algorithms are able to attain micro-benchmark performance on the heterogeneous SDR-DDR InfiniBand cluster to perform on par with runs on the DDR only portion of the cluster for small to medium sized messages. We also demonstrate that the network speed aware algorithms perform 70% to 100% better than the naive algorithms when both are run on the heterogeneous SDR-DDR InfiniBand cluster.
A domain decomposition approach and finite element formulation are developed for parallel distributed simulation of surface tension driven viscous flow and transport processes. The scheme is implemented in a finite element code MGFLO and is used to study performance on CRAY T3E and SGI Origin 2000 and 3000 parallel supercomputers as well as several PC clusters. The parallel algorithm implementation is briefly discussed, and scaled speedup studies are presented. Representative simulation results for surfactant and thermocapillary driven surface tension flows are also presented.
A domain decomposition strategy and parallel gradient-type iterative solution scheme have been developed and implemented for computation of complex 3D viscous flow problems involving heat transfer and surface tension effects. Special attention has been paid to the kernels for the computationally intensive matrix-vector products and dot products, to memory management, and to overlapping communication and computation. Details of these implementation issues are described together with associated performance and scalability studies. Representative Rayleigh- Bénard and microgravity Marangoni flow calculations on the Cray T3D are presented, and performance results verifying a sustained rate in excess of 16 gigaflops on 512 nodes of the T3D have been obtained. The work is currently being extended to the T3E and we have begun carrying out further performance benchmarks and scalability studies on this platform. Preliminary performance studies have recently been carried out and sustained rates above 50 gigaflops and 100 gigaflops have been achieved on the 512 node T3E-600 and 1024 node T3E-900 configurations respectively.
We are developing and implementing techniques for hierarchic visualization of very large scale data sets that may arise as a result of computer simulations of engineering applications or from experimental measurements. The basic approach is motivated by the ideas used in adaptive mesh refinement and coarsening strategies, multigrid solution schemes and wavelet multiresolution techniques. Of particular interest are large scale simulations on parallel supercomputers. The work exploits tree-traversal algorithms and data structures for adaptive refinement/coarsening to enable selective hierarchic extraction and compression of solution data for visualization locally or at a remote site. Error or feature indicators provide a theoretical framework to guide adaptive mesh schemes and provide a mechanism for selectively screening the simulation results. These indicators can also be used to construct so-called "intelligent agents" to help direct the visualization. A prototype "data handler" that incorporates some of these primitives is under development and testing. Representative examples involving adaptively refined triangular meshes in 2D and meshes of "quadrilateral brick" elements in 3D are investigated. These illustrate, for instance, how adaptive refinement may generate a nonuniform quadtree or octree based on an analysis error indicator and then how selective hierarchic visualization using a different feature indicator can be applied on the same tree. Supporting timing studies for remote visualization of hierarchic data using nested meshes or multiresolution approaches are also presented. The extension of these schemes to parallel distributed data sets using mesh partitioning strategies is also considered.
In this study we describe parallel iterative and direct solution strategies for viscous flow computation using domain decomposition. Both stream function vorticity and primitive variable formulations are considered. The basic ideas associated with the element-by-element approach and subdomain approach can also be extended to other cell or control-volume approximate formulations. Numerical results and parallel performance studies are presented for computations with parallel sparse solution based on p − multilevel schemes and multiple front elimination schemes.
The orientation tensor L is introduced to construct a modified Leslie-Ericksen model for the viscous, incompressible flow of anisotropic suspensions (including electric field effects). This is then utilized to develop a weak variational formulation and finite element scheme for computing the flow and orientation fields. Numerical results are presented for exploratory test problems.
AbstractA double‐transform technique provides a semi‐analytic solution in the form of a series expansion for unsteady axisymmetric Stokes flow in the entrance region of a semi‐infinite rigid cylindrical tube. This in turn offers an appropriate bench‐mark problem for evaluating the quality of numerical approximations. To illustrate this, periodic axial flow in a circular cylinder is considered. Some aspects of the bench‐mark problem that are of interest include the reverse flow in the wall layers, the accuracy of the approximate method in different flow regimes and the mesh grading. This bench‐mark problem and the numerical study provide some insight into practical issues pertinent to the approximate solution of unsteady and periodic flows.
AbstractA finite element formulation and analysis is developed to study coupled heat transfer and viscous flow in a weld pool. The thermal effects generate not only buoyancy forces but also a variation in the surface tension which acts to drive the viscous flow in the molten weld pool. A moving phase boundary separates molten and solid material. Numerical experiments reveal the nature of the highly convective flow in the weld pool and the associated thermal profiles. The relative importance of buoyancy, surface tension, phase change, convection, etc. are examined. We also consider the sensitivity of the solution to the finite element mesh and related non‐linear numerical instabilities. Of particular interest is the coupling of the thermal and viscous flow fields for the case when radial flow is inward or outward.
AbstractVector and parallel algorithms for finite‐element analysis using the element‐by‐element (EBE) data structure are developed. The algorithms are based on the EBE approach in conjunction with gradient‐type iterative solution. The essential idea is to exploit the independent dense local‐element matrix‐vector calculations and reconfigure them to take advantage of vector or parallel processing capabilities. The ideas also are well suited to finite‐element adaptive refinement computations. Specific algorithmic details related to the implementations are given, together with results of speed‐up performance studies conducted on the vector and parallel architectures. The basic vector and parallel versions of the matrix‐vector product for the EBE scheme are very straightforward modifications of the EBE algorithm, requiring little coding change.
AbstractThe 8‐node (serendipity) velocity basis with C° bilinear pressure is a popular element but has been observed to yield poor pressures. We present some details of numerical experiments that indicate the local nature of the error and the effects of mesh refinement, increasing Reynolds number and regularity of the data. This leads to a strategy for appropriately modifying the data near the corners that is effective in improving the computed pressure approximation.
In this work we consider the problem of solving large-scale 3-dimensional coupled flow and transport processes. The treatment includes algorithms and methodology and performance studies for parallel simulation on supercomputer s and on PC clusters. It also concerns phenomenological flow and transport calculations and issues related to reliability of computations, grid effects and validation studies as the complexity of the problem increases. The flow models considered cover both Navier Stokes and generalized Newtonian fluids such as Powell-Eyring and Williamson models. We consider coupling to a variety of transport processes dealing, in particular, with coupled heat and fluid flow including thermocapillary surface tension but also including studies of species transport and coupling to other fields such as electric fields in electrorheological flow problems. The coupled fluid flow and heat transfer applications will also include Nusselt number calculations for a benchmark problem in the computational heat transfer community where a physical experiment has been especially designed for validation of computations. Some recent theoretical results on the error analysis of specific classes of generalized Newtonian fluids will be summarized and comparison simulations presented for a familiar test geometry. Comparative studies including performance at the processor level and scaled speedup performance studies on PC parallel clusters of varying size will be included. Relevant aspects of the algorithms and the parallel implementation will be summarized and some aspects of the software engineering implementation will be briefly sketched. The work on PC clusters will also include some recent progress in the general area of GRID computing in the sense of building a capability that includes Computational Networks of parallel cluster architectures that are geographically remote. In addition to the coupled viscous flow and transport processes, we also consider the Elder benchmark for flow for the evolution of a plume from a surface contaminant. This problem has led to conflicting numerical results in the literature and we demonstrate related grid sensitivity to help resolve this question. Some of the work described will deal with 3-dimensional adaptive simulations and the use of software libraries to facilitate inclusion of adaptive techniques in support of these applications.