
Advances in microprocessor technology, power management and network communication have altered the course of development of multiprocessor architectures in order to bring higher level of processing. The introduction of multi-core technology has boosted computing power provided by high-speed network of workstations and SMPs, providing large computational power at an affordable cost using solely commodity components. In this paper, it is presented a tool for integration of several clusters in a single High-Performance System based on MPI standard. The Gateway Process is responsible for MPI process communication channels control and message forwarding, through the use of a protocol that guarantees message ordering and sender/receiver synchronization. It is implemented to support system scalability, offering resources for point to point and collective operations. Results of experimental tests show that the proposed tool is practical and efficient.
The boundary-value problem for polarized-radiation transfer equation in layered medium with Fresnel matching conditions at the boundaries of the medium partition is considered. Parallel numerical algorithm in the MPI environment based on recursive modification of Monte-Carlo method for solving the boundary-value problem is proposed and proved.
This paper deals with the parallelization of free-surface three-dimensional oceanographic model. The model is based on full nonlinear "primitive" equations of the ocean. Generalized vertical coordinate system is applied for better resolution of main features of the simulated basin. The numerical model is conservative. Numerical integration procedure is based on time-splitting method with Robert-Asselin filtering. Numerical model code is implemented for running on cluster computers. The parallelization is achieved using domain decomposition method and standard MPI to ensure portability of the code.
Resource performance prediction is the basis of dynamic load balance in distributed computing. A model for resource performance prediction named ARPP is introduced and carried out. ARPP model monitors key parameters of resources and estimates the directions using ant algorithm. The implement and analysis of ARPP is based on GridSim simulator and the process of astronomical image mosaicking application. The experiment result shows the efficiency of the model and the determination of optimized parameters.
The paper focuses on a numerical method for detecting, visualizing and monitoring abnormal cell growth using large-scale mathematical simulations. The discretization of multi-dimensional partial differential equation (PDE) is based on finite difference method. The predictor system depending on users input data via a user interface, generating the initial and boundary condition generated from parabolic or elliptic type of PDE. The processing large sparse matrixes are based on multiprocessor computer systems for abnormal growth visualization. The multi-dimensional abnormal cell has produced the numerical analysis and understanding results at the target area for the potential improvement of detection and monitoring the growth. The development of the prediction system is the combinations of the parallel algorithms, open source software on Linux environment and distributed multiprocessor system. The paper ends with a concluding remark on the parallel performance evaluations and numerical analysis in reducing the execution time, communication cost and computational complexity.
Two new cellular-automata models of the diffusion process are proposed. They are based on integer states of cells instead of Boolean ones in the known models: asynchronous naive diffusion by Toffolli and block-synchronous Margolus diffusion. Computing experiments have been carried out with these models; they demonstrate a good correlation with this physical phenomenon. The main advantages of the proposed models are (i) low automata noise and (ii) variable diffusion rate.
Cloud computing opens new opportunities for application providers because with the policy "add as needed and pay as used" they can economize the cost for computing resources. In cloud environments, issues such as resource allocation and dynamic resource provisioning based on users' QoS constraints are yet to be addressed for interactive workflow applications. This paper develops an effective load metric, remaining tasks, for interactive workflow applications. Based on this metric load dispatching and dynamic resource provisioning approaches are proposed which outperform existing methods under a series of simulation evaluations. Experimental results show that the proposed approaches offer application providers better maintenance of QoS-satisfied response time under time-varying workload, at the minimum cost of resource usage.
With the large amount of data which is produced by scientific experiments and simulations, the data replication has become an important issue in data grid. In order to shorten the file accessing time, multiple copies of file and stores can be created in the appropriate location. In this paper, we present a dynamical maintenance service of replication to maintain the data in grid environments. Based on the Bayesian Networks (BN), the optimization strategy was proposed in this study; named Implicit Dynamic Maintenance Service with Bayesian Network (IDMSBN).
Organization of high performance execution of fragmented programs met the problem of choice of acceptable way of their execution. The possibilities of execution optimization on the stages of fragmented program development, compilation and execution are considered. The methods and algorithms of optimizations are suggested to be included both in fragmented programming language and in run-time system.
Clusters became the de-facto standard in modern high-performance computing. At present it is rather often when a single organization has a few clusters and wants to connect them into a multicluster to benefit from reduced task waiting time and increased total available processing power. This paper studies one of possible approaches to the problem of uniting the computing resources which addresses exactly the case of owning clusters by a single proprietor. Such approach was implemented in Metacluster system of Nizhni Novgorod State University, Russia. Current state of Metacluster was reviewed. Key features, component-based architecture and main functions implementation details were described.
The paper describes investigations of some aspects of creating stochastic wave models of dynamic acoustic noise. The problem about sound radiation by homogeneous and stationary pressure fluctuations on a surface of layered waveguide is considered. The method of statistical modeling based on the randomization of spectral density of surface sources is used. It allows calculating random realizations of sound pressure and particle velocity that are exact solution of the equations of linear acoustics. The software for calculation of statistical characteristics of surface noise in a distributed computing environment is developed. Some examples of its applications are given.
With the advancement of network and techniques of clusters, joining clusters to construct a wide parallel system becomes a trend. Irregular array redistribution employs generalized blocks to help utilize the resource while executing scientific application on such platforms. Research for irregular array redistribution is focused on scheduling heuristics because communication cost could be saved if this operation follows an efficient schedule. In this paper, a two-step communication cost modification (T2CM) and a synchronization delay-aware scheduling heuristic (SDSH) are proposed to normalize the communication cost and reduce transmission delay in algorithm level. The performance evaluations show the contributions of proposed method for irregular array redistribution.
Resource discovery in distributed computing systems is a critical issue to find and retrieve distributed resources rapidly. In general, most of previous proposed strategies focus on developing keyword searching approaches with preserving system scalability. In this paper, we propose a cluster-based hybrid overlay, which supports efficient keyword searching with the highly churn rate. The cluster-based hybrid overlay groups the nodes with the same attributes to form unstructured attribute-groups, and then clusters these attribute-groups with similar attributes to form attribute-clusters. Our proposed hybrid overlay could provide efficient multi-attribute and range-query searches with load balancing in large-scale P2P networks. Experimental results show that the proposed overlay performs well.
A technique of parallel computing in simulating the deposition of diamond-like carbon thin film by molecular dynamics is proposed. The Tersoff potential which is a multi-body potential is adopted here in determining inter-atomic forces. The deposition of carbon thin film on diamond substrates and silicon substrates under different incident kinetic energies and different substrate temperatures are investigated. The multiprocessor of workstation computer containing 8 cores used for simulating the deposition is based on MPICH2 which is an implementation of message passing interface. The results show that the percentages of deposited sp3 carbon atoms differ from 6.1% to 34.8% depending on the type of substrate, incident kinetic energy and substrate temperature.
The code generator in a compiler attempts to match a subject tree against a collection of tree-shaped patterns for generating instructions. Tree-pattern matching may be considered as a generalization of string parsing. We propose a new generalized LR (GLR) parser, which extends the LR parser stack with a parser cactus. GLR explores all plausible parsing steps to find the least-cost matching. GLR is fast due to two properties: (1) duplicate parsing steps are eliminated and (2) partial parse trees that will not lead to a least-cost matching are discarded as early as possible.
The paper describes the basic features of the developed computer-aided system of design, modeling, and electronic simulation of integral panel manufacturing. Results of system application for three-dimensional stress analysis computations and simulations of panel shaping under various thermomechanical and speed conditions are demonstrated. As is seen from computations, obtaining these solutions without parallelization can last for weeks; therefore, urgent problems can be hardly solved. These computations can be accelerated by using parallelization algorithms. In modern operational systems, the execution thread capability of generating another thread allows building multithreaded programs of recursively called subroutines, recursive calla being substituted by creation of a thread. Based on these activities, a computer-aided design system was designed for manufacturing structural panel elements with complex curvature and engraving. Application of the system for intricate manufacturing processes (blank and die tooling for shaping purposes based on material creeping) assists in eliminating the reject symptoms for wing panels of modern aircraft, considerably reduces the production costs, and improves the product quality.
In this paper, we present a fast scalable method to reduce the computation time of genetic algorithms for traveling salesman problem, called the Parallel Pattern Reduction Enhanced Genetic Algorithm (PPREGA). The general idea behind the proposed algorithm is twofold: (1) Eliminate the redundant computations of GA on its convergence process by pattern reduction and (2) Minimize the completion time of GA by parallel computing. Our simulation result shows that the proposed algorithm can significantly reduce not only the computation time but also the maximum completion time of GA. Moreover, our simulation result shows further that the loss of the quality of the end result is small.
In this paper we introduce a software system which allows to carry out and visualize computational experiments for studying and researching the parallel algorithms of solving complicated computational problems in imitation mode on one single sequential computer. User can “assemble” a parallel computational system of cluster type that consists of multiprocessor and multicore nodes connected with the network, set up the problem to be solved, carry out the parallel solving algorithm, collect and analyze the results of computational experiments. To estimate the execution time of parallel method on current hardware system we use the sophisticated models. For every implemented parallel method we proved the theoretical estimations of the execution time by comparing the real time of the execution on the NNSU high performance cluster with the time, that can be calculated using the model.
Discrete simulation method of physicochemical kinetic processes is proposed and investigated. The method is based on formal representation of classical Von-Neumann’s Cellular Automaton (CA) extension, which allow all kind of discrete alphabets, probabilistic transition functions, and asynchronous mode of operation. Some techniques for simple CA composition are given for simulating complex processes. Transformation of asynchronous CA into block-synchronous type is used to provide high efficiency of parallel implementation.