
In this work, we introduce resources co-allocation algorithms for parallel jobs execution in distributed computing with non-dedicated and heterogeneous hosts. Complex distributed computing systems often operate under conditions of the resources availability uncertainty. Imprecise estimations of jobs execution runtime, unplanned maintenance works and other global and local events do not allow to consider accurate schedules of the resources utilization. On the other hand, an efficient job-flow execution in compliance with QoS constraints requires reliable mechanisms for advanced resources allocation and reservation. The novelty of the proposed resources allocation approach is in a general procedure efficiently selecting computing nodes according to the resources availability criteria. Special knapsack and greedy algorithms are implemented and compared in a market-based computing simulation model.
Distributed computing systems are widely used for the execution of loosely coupled many-task applications, such as parameter sweeps, workflows, distributed optimization. These applications consist of a potentially large number of computational tasks that can be executed more or less independently. Since the application users often have an access to multiple computing resources, it is important to provide a convenient and efficient environment for execution of applications across the user-defined heterogeneous resource pools. The paper discusses the related challenges and presents an approach for solving them based on Everest, a web-based distributed computing platform. The presented solution supports reliable and efficient execution of many-task applications, while taking into account resource performance, adapting to queuing delays and providing a mechanism for communication between tasks.
A high-resolution mesoscale meteorological model for forecasting and studying weather events and surface air quality in an urbanized area or a large industrial or transportation hub is presented. An effective semi-implicit second-order finite volume method with parallel implementation on multiprocessor computing system was developed to solve the equations of the model. The results of testing the parallel program on the supercomputer Cyberia of Tomsk State University demonstrated its high efficiency. The approach developed was successfully applied to predicting heavy precipitation events and an urban heat island effect.
In this paper we present the design of data processing workflow for scientific experiments, which require complicated multi-step analysis procedure. We test it on datasets from Single Particle Imaging (SPI) experiments. The workflow is based on microservice architecture, Docker containers and Kubernetes platform. For workflow setup and management we use REANA software which is compatible with Kubernetes ochestrator and supports standard Common Workflow Language (CWL) to describe complex computing jobs. Our approach allows easy construction of workflows of diverse architecture for a wide range of applications. It allows integration of heterogeneous software in a uniform way as well as easy modification or replacement of workflow components. In the same time it allows easy scaling of computations in a cloud infrastructure. We show the applicability of the designed scheme and estimate the overhead of the platform middleware.
The paper describes the construction of the coupled ocean-ice Global model INMIO-CICE-CMF2.0 for predicting the state of water and ice in the Arctic Ocean with high spatial resolution (0.1°). The 3D ocean model INMIO is developed at the Institute of Numerical Mathematics and Institute of Oceanology, Russian Academy of Sciences. The sea-ice model CICE (Community Ice CodE) is developed Los Alamos National Laboratory. The models are fully coupled at each four-time steps using own software named Compact Modeling Framework (CMF ver.2.0). Outputs are the surface variables sea level and ice conditions (concentration, thickness, velocity, convergence, strength, etc.) and 3-dimensional maps of current, temperature and salinity. The main aim of the research is developed the algorithm to find the optimal processor configuration of the coupled ocean-ice model for increase performance and optimizing computer resources usage. This is nontrivial problem because these two models are not uniformed and are used for numerical experiments with thousands of processors cores on Supercomputers with shared memory using MPI technology. The theoretical and practical aspects of the problem are discussed.
The problem of classifying a large set of texture images using deep neural networks is considered. To reduce the learning time of the neural network, it is proposed to use a desktop grid system. The deep neural network architecture is selected and its implementation is described. A method for organizing deep learning based on the data separation approach and synchronous parameter updates during distributed learning is presented. The features of deep learning on a desktop grid system are discussed and the results of a computational experiment are presented.
The binary32 and binary64 floating-point formats provide good performance on current hardware, but also introduce a rounding error in almost every arithmetic operation. Consequently, the accumulation of rounding errors in large computations can cause accuracy issues. One way to prevent these issues is to use multiple-precision floating-point arithmetic. This preprint, submitted to Russian Supercomputing Days 2020, presents a new library of basic linear algebra operations with multiple precision for graphics processing units. The library is written in CUDA C/C++ and uses the residue number system to represent multiple-precision significands of floating-point numbers. The supported data types, memory layout, and main features of the library are considered. Experimental results are presented showing the performance of the library.
The classical problem of oscillations of liquid droplets is a good test for the applicability of computer simulation. We discuss the details of our approach to a simulation scheme based on the Boltzmann lattice equation. We show the results of modeling induced vibrations in a chain of three drops in a closed tube. In the initial position, the central drop has formed as an ellipsoid, out of the spherical equilibrium form. The excitation of vibrations in the left and right droplets depends on the viscosity of the surrounding fluid and the surface tension. Droplets are moving out of the initial position as well. We discuss the limits of the applicability of our model for the study of such a problem. We will also show the dynamics of the simulated process.
In recent years digital rock physics technology is regarded as a promising tool which can supplement traditional laboratory techniques. This technology is based on numerical experiment with direct resolution of pore space of a rock sample, which is obtained with computed microtomography. Necessity of high resolution leads to a high dimension of discrete settings (106–109 numerical cells). The work is devoted to an application of different partitioning algorithms to the problem of flow simulation within geometry of rock sample pore space. Simulation of a single-phase fluid flow within pore space of a sandstone sample with voxel representation is used to compare the partitions obtained by various methods using parallel partitioning tools ParMETIS, Zoltan, and GridSpiderPar. Average time spent on interprocess exchange during 200 time steps of the considered parallel simulation was compared when the grid was distributed over the cores in accordance with various partitions. The obtained results demonstrate advantages of some algorithms and reveal the criteria, important for the problem.
It is known that the nucleation of stars in the universe occurs, as a rule, in areas of the interstellar medium, where nebulae - gas and dust clouds consisting of molecular hydrogen, can collide with each other, changing their state during dynamic interaction with gravitational and magnetic fields. The gravitational-turbulent description of these processes is quite common in explaining the reasons for the creation of pre-stellar regions as a consequence of the collision of molecular clouds. We adhere to this approach with some simplification, assuming that the main factor in the dynamic transformations of gas formations is the influence on the collision process of mainly the kinetic energy of the clouds, separating these effects from the effects of gravitational collapse and from the influence of magnetic fields. To simulate gas-dynamic processes of different scales, a parallel numerical code was developed using grids with improved resolution, which was used in a numerical experiment on high-performance computers. The simulation showed that sharp changes in the distribution of matter in the shock core of new formations can be triggered by the Kelvin-Helmholtz instability and nonlinear thin-shell instability, which lead to sharp perturbations of the gas density in the resulting clumps, outer shells and gas filaments, with density fluctuations in the external interstellar medium. A predictive analysis of the appearance of possible pre-stellar zones during the evolution of new cloud formations is given.
Rapidly developing new technologies, especially in the field of modern aircraft, stimulate great interest in creating high-energy materials for various purposes. Recently, modern computer technologies have been playing an increasingly important role in creating new materials with determined properties. Focus of this work is on computer design of new compounds that have not yet been synthesized – high-enthalpy derivatives of heterocycles, such as tetrazine, furazan, furoxan, triazole, etc. using quantum chemical calculation methods. For experimentally studied substances C2N6O4, C2N6O3, C2N8O4, the calculated values of enthalpy Δf 298 o (g) are 8–15% higher than the experimental values, which is significantly less than the spread of experimental values for these compounds. The enthalpy of formation of the studied gaseous molecules was calculated using the atomization method. The simulation was performed within the GAUSSIAN 09 software package using the B3LYP hybrid density functional with the basis 6-311+G(2d,p) and the combination of methods G4 and G4(MP2). Tasks have high computational complexity and the calculation time it takes from several hours to months on multi-node supercomputer configurations.
The intellectual structure of scientific discipline consists of a set of interacting topics. The evolution of these topics is the subject of special attention because it reflects the actual interest of researchers and stakeholders. This paper analyzes issues of High-Performance Computing (HPC) on the base of the formal topic modeling technique. Analyzing the abstracts of 7661 publications referenced in Web of Science in 2005–2019, we identified seven topics that concern different aspects of HPC science. The central theme is the Large Scale Applications focused on practical and scientific problems solved using HPC. It is closely linked with Parallel Algorithms that should effectively exploit the thousands of processing cores, Parallel Software for heterogeneous distributed systems, and Interconnected systems that study the integration of HPC facilities in systems of larger size. These topics are relatively stable both in terms of popularity (number of publications) and impact (number of citations). The single topic, which popularity and impact continuously grow in the last 15 years, is Energy efficiency since power consumption is a critical issue of exascale systems. We also found that the topic of Heterogeneous systems dedicated mainly to GPU usage declines after the peak of interest in 2010–2015. The results obtained shed light on the structure of HPC science and supplement the known publications that declare research direction towards exascale performance.
The work is dedicated to domain-decomposition parallel iterative linear system solver. An algebraic multilevel preconditioner is used on a subdomain. The cornerstone of the method is the dual-threshold second-order Crout incomplete LU factorization. The Crout version of LU factorization permits condition estimation for inverse factors ‖ L^-1‖ and ‖ U^-1‖ . The factorization of k-th row and column of the matrix is postponed upon the growth of the estimated condition. Following factorization, the Schur complement is computed for the postponed part and the factorization continues in the multilevel fashion until the complete matrix is factorized. Before each level factorization begins, the matrix is first permuted by finding maximum transversal and reverse Cuthill-McKee permutations and then rescaled into I-dominant matrix. The performance of the method on a coupled multiphysics problem of blood coagulation is demonstrated.
Huge computer resources needed to promptly compute the 24-h global weather forecast dictate the necessity to optimize the numerical algorithms of the model and their parallel implementation. We present some experience gained while implementing the new high-resolution version of the SL-AV global atmosphere model for numerical weather prediction at parallel systems with many thousands of processor cores. Unlike our previous scalability studies, we need to minimize the elapsed time of the forecast at given processor cores number which is currently about 4000. The results of optimizations are shown for two Roshydromet high-performance computer systems.
INMOST (Integrated Numerical Modeling Object-oriented Supercomputing Technologies) is an open-source platform for fast development of efficient and flexible parallel multi-physics models. In this paper we review capabilities of the platform and present two INMOST-based applications for parallel simulations of multi-phase flow in porous media and clot formation in blood flow. For a more detailed description we refer to [1].
In order to evaluate efficiency of some global optimization method or compare efficiency of different methods, it is necessary to select a set of test problems, define comparison measures, and, finally, choose a way of visual presentation of the computational results. In this paper, a wide set of test optimization problems is considered including a new global constrained optimization problem generator (GCGen). Main performance measures and comparative criteria of efficiency are presented. The ways of visual presentation of computational results are suggested.
One of the most effective ways to enhance students’ learning in programming disciplines is carrying out practical assignments covering all key topics on their own. Ideally, a student deals with an individual set of tasks. While the relevance of such an approach is self-evident, its practical implementation is no easy matter. The main difficulty encountered by teachers is providing a high-quality assessment of the solutions developed by students on a variety of problems. Therefore, it is extremely important to automate the evaluation process. The present article is devoted to the description of SoftGrader – automated checking system, developed in UNN, which allows performing mass testing of student program correctness, both sequential and parallel programs.
The paper covers the research and numerical implementation of an graph model of natural and technogenic factors’ interaction in shallow water productivity. Based on it, the analysis of pulse propagation in computing environment from the vertices is performed in the context of research situation of valuable fish degradation of the Azov Sea that are subject to excessive commercial fishing withdrawal. The model takes into account the convective transport, microturbulent diffusion, taxis, catch, and the influence of spatial distribution of salinity, temperature and nutrients on changes in plankton and fish concentrations. Discrete analogue of proposed model problem of water ecology, included in software complex, were developed using schemes of second order of accuracy taking into account the partial filling of computational cells. The adaptive modified alternately triangular method was used for solving the system of grid equations of large dimension, arising at model discretization. Effective parallel algorithms were developed for numerical implementation of biological kinetics problem and oriented on multiprocessor computer system and NVIDIA Tesla K80 GPU with the data storage format modification. Due to it, the production processes of biocenose populations of shallow water were analyzed in real and accelerated time.
The article describes computational experiments aimed to enumerating the number of orthogonal diagonal Latin squares for general and special types of orthogonality. General type of orthogonality can be verified using Euler-Parker method, corresponding the number of main classes of orthogonal diagonal Latin squares, the number of normalized orthogonal diagonal Latin squares and total number of orthogonal diagonal Latin squares of general type form previously unknown numerical series A330391, A305570 and A305571 (calculated up to order 8) has been added to OEIS. Self-orthogonal (SODLS), doubly self-orthogonal (DSODLS) and extended self-orthogonal diagonal Latin squares (ESODLS) form the set of special types of orthogonality. For each of these types corresponding numerical series was calculated and published in OEIS with numbers A329685, A287761, A287762 (SODLS, up to order 10), A333366, A333367, A333671 (DSODLS, up to order 10) and A309210, A309598, A309599 (ESODLS, up to order 8). Values for orders 1–8 were obtained by analyzing the complete lists of canonical forms of the main classes of orthogonal DLSs obtained by the authors by exhaustive search. Values for order 9 were derived from the SODLS list of order 9 provided by Harry White. Values for order 10 were obtained by analyzing the list of SOLS of order 10, available online (van Vuuren et al.). The values obtained confirm the similar values for SODLS and DSODLS obtained previously by Francis Gaspalou and partially published by Harry White. For some of the obtained numerical values previously unknown mathematical relations are empirically established.
The paper proposes an approach to semi-automatic program parallelization in SAPFOR (System FOR Automated Parallelization). SAPFOR proposes opportunities to perform user-guided source-to-source program transformations and to reveal implicit parallelism in sequential programs. The LLVM compiler infrastructure is used to examine a program and Clang is used to perform source-to-source program transformation. This paper highlights benefits of IR-level (Intermediate Representation) program analysis which allows us to apply low-level program transformations to investigate properties of the original program. To exploit program parallelism SAPFOR relies on DVMH which is a directive-based programming model. We use subset of C-DVMH language which allows us to run parallel program on GPU as well on multiprocessors. Evaluation of presented approach has been performed using the C version of the NAS Parallel Benchmarks.