
The advent of interconnected machines laid the foundation for utilizing distributed computing resources. Coinciding with rapid advancements in computing technologies and significant hardware innovations at the turn of the century, the field of science also experienced exponential growth. As supercomputers remained limited in availability and usage, distributed computing leveraged the potential of idle and dedicated resources, bridging the gap between scientists and computationally intensive projects. This paper provides a review of the evolutionary journey of grid computing in scientific applications, starting from the advancements in network connection technologies and concept of metacomputing and progressing to the current developments integrating cloud technologies with large-scale grids. The paper aims to outline the key milestones, advancements, and challenges encountered throughout this evolution, highlighting the potential of grid computing in enabling scientific breakthroughs and addressing future research directions. The most popular middleware systems are considered, as well as a description of scientific grid systems that existed in the past and are still in operation today is given. At the end of the article, we examined two of the most significant scientific discoveries that became possible largely thanks to grid technologies.
Full-atomistic modeling of the deposition of TiO2, SiO2 and TiO2–SiO2 films is performed using parallel calculations. The dependence of film density on the deposition angle and deposition energy is studied. Simulation of post-deposition annealing of film structures is also carried out. Mechanical stresses in TiO2–SiO2 films, arising due to differences in the properties of silicon dioxide and titanium dioxide, are calculated. It is found that the film density decreases with decreasing deposition energy and increasing deposition angle. The use of surfacing annealing leads to an increase in film thickness. In two-layer TiO2–SiO2 films, the stresses are compressive. Particular attention is paid to reducing computational costs when simulating large atomistic clusters, consisting of hundreds of thousands of atoms. Reducing the parameter that determines the calculation of the electrostatic part of interatomic energy significantly reduces the simulation time. At the same time, in this case, the accuracy of determining the electrostatic energy in the reciprocal space decreases, which should be taken into account during modeling.
The article examines potential applications of WAVEWATCH III (WW3), the thirdgeneration wind-wave model. This study delves into the implementation of hybrid parallelization (MPI-OpenMP) and the development of multiple-cell grids tailored for the Azov Sea region. It elucidates fundamental equations of the model, their discretization, and software execution. The multiple-cell grid strategy employs high-resolution cells within the region of interest, gradually increasing cell density in other areas to optimize memory consumption. A 6-level multiple-cell grid was specifically crafted for the Azov Sea, with an algorithm outlined for its generation incorporating two refinement methods. This algorithm enables the creation of refined multiple-cell grids near shorelines at varying levels, along with the capability to refine grid structures in arbitrary zones. Additionally, the article presents hybrid parallelization techniques for the wave spectral component (MPI-OpenMP), assessing scalability in both MPI and hybrid deployments. The WW3 model offers a multigrid option facilitating parallel operation of subdomains akin to domain decomposition, while ensuring parallelization of each subnet via the component decomposition method.
Global models of the Earth's gravitational field, built from data collected by space geodetic missions, play a very important role in studying global processes across Earth's various geospheres. The paper is devoted to the development of a program for the recovery of the Earth's gravity field parameters. This program will enable in the future to process the results of measurements from the Russian satellite constellation and to build gravity field models of different spatial and temporal resolution. The recovery of the gravity field from satellite measurements is a rather resourceconsuming computational process, and parallel computations are crucial for its optimization. This paper describes the mathematical model, the algorithm and the results of parallelization, as well as presents the results of the gravity field recovery using parallel computations working with real measurement data. The static model of the Earth's gravitational field MSU2024-1 was built using the GRACE-FO mission data for the whole year 2021. The model is decomposed to degrees and orders of n = m = 120 and presented in terms of geoid heights. We also compared the EGF recovery on a monthly interval using the GRACE-FO data obtained in this work with the CSR, GFZ, and JPL temporal models built at other world centers.
Computer aided structural based approach was used to find inhibitors of SARS-CoV-2 nsp16 (2'-O-methyltransferase). Docking based virtual screening of three libraries, Enamine Coronavirus Library, Enamine Nucleoside Mimetics Library, and Chemdiv Nucleoside Analogue Library, was performed. In total, 39350 3D-structures of low molecular weight ligands were docked into a model of nsp16 prepared using the structure of 6WKQ complex from the Protein Data Bank. Docking was performed by the SOL docking program. For the best SOL scored ligands, the protein-ligand binding enthalpy was calculated using the PM7 semiempirical quantum-chemical method with the COSMO implicit solvent model. The most promising eleven compounds were purchased and their inhibitory activity against the recombinant viral nsp16 protein was measured using MST assay with Monolith NT.115. As a result, two compounds, Z195979162 and Z1333277068, from Enamine Coronavirus Library demonstrated dissociation constants Kd for nsp16/nsp10 complex equal to 2.0 and 5.0 μM. The relative stability of these ligands in their docked positions in the nsp16 S-adenosylmethionine (SAM) binding site was confirmed in the molecular dynamics simulations along 70 ns trajectories. Z195979162 and Z1333277068 compounds belong to two chemical classes: 1,4-disubstituted tetrahydropyridines and derivatives of pyrazole-5-carboxamide, respectively, and can be good starting points for further hit optimization in the field of nsp16 inhibitors design.
The features of a statically typed functional-dataflow model of parallel computation and its mapping to the statically typed language of functional-dataflow parallel programming Smile are considered. To provide support for architecture-independent parallel programming, we used: a functional style, an implicit managing of calculations on data readiness, structured data objects that provide representation of various types of parallelism. A distinctive feature of the approach is the inclusion in the model of special asynchronous data objects that can generate events on partial filling. These data objects are stream and swarm. Each of these data objects has its own specifics to control by parallel calculations. A stream is used to process data of the same type that arrives sequentially and asynchronously at random intervals. A swarm is used to describe independent data of the same type or different types, on which it is possible to perform massive parallel operations. The use of streams and swarms in various situations as well as their mapping into each other and other program objects are shown. An analysis is made of the possibilities of transforming the formed language constructs into programming languages used in writing programs for modern parallel architectures.
One of the promising directions for improving hybrid reconfigurable high-performance computer platforms operating in the mode of collaborative applied computing centers is their inclusion as an active component in the machine learning ecosystem, which opens up new opportunities to enhance the actual outperformance of solving various application tasks by intellectualizing the management of available computing resources. The task scheduler operation is crucial in improving the efficiency of hybrid supercomputer platforms, which combine dozens of processor blocks with different architectures, including specialized graphics and reconfigurable accelerators. To form an optimal order of jobs in the HPC queue, the article proposes to apply deep survival machine learning models, which increase the accuracy of the estimated time of the tasks successful execution and the required amount of computing resources. The main peculiarity of the machine learning models is that they are trained on censored heterogeneous data collected from previous periods of task execution observations using a multi-agent scheduler. In order to ensure high accuracy, the random survival forest is used as a part of the machine learning model which provides survival and hazard functions in the framework of the survival analysis. A specific weighted clustering procedure is proposed to divide tasks in accordance with their execution times as well as the feature vectors. Various numerical experiments with actual data illustrate the outperformance of the presented approach.
c (cid:13) The Authors 2023. This paper is published with open access at SuperFri.org In this paper, we introduce the design of an advanced high-performance interconnect architecture for supercomputers. In the first part of the paper, we consider the first generation high-performance Angara interconnect (Angara G1). The Angara interconnect is based on the router ASIC, which supports a 4D torus topology, a deterministic and an adaptive routing, and has the hardware support of the RDMA technology. The interface with a processor unit is PCI Express. The Angara G1 interconnect has an extremely low communication latency of 850 ns using the MPI library, as well as a link bandwidth of 75 Gbps. In the paper, we present the scalability performance results of the considered application problems on the supercomputers equipped with the Angara G1 interconnect. In the second part of the paper, using research results and experience we present the architecture of the advanced interconnect for supercomputers (G2). The G2 architecture supports 6D torus topology, the advanced deterministic and zone adaptive routing algorithms, and a low-level interconnect operations including acknowledgments and notifications. G2 includes support for exceptions, performance counters, and SR-IOV virtualization. A G2 hardware is planned in the form factor of a 32-port switch with the QSFP-DD connectors and a two-port low profile PCI Express adapter. The switches can be combined to 4D torus topology. We show the performance evaluation of an experimental FPGA prototype, which confirm the possibility of implementing the proposed advanced high performance interconnect architecture.
In the paper we consider the high-level synthesis toolchain for transformation of programs written in C (the standard ISO/IEC 9899:1999) into configuration files of field programmable gate arrays (FPGAs) used in multichip reconfigurable computer systems. Unlike most academic (DWARV, BAMBU, LEGUP) and commercial (CatapultC, Vivado HLS, Vivado Vitis) high-level synthesis tools, “Theseus” uses the original methodology of transformation (porting) sequential calculations into a parallel-pipeline configuration of FPGA hardware. For a sequential program, an information graph is created and transformed into the maximally parallel structure, which is then ported to a specified configuration of the reconfigurable computer system using formal methods of reduction of performance and hardware costs without marking the source text with auxiliary parallelization directives. The distinctive feature of the approach is a significantly smaller number of analyzed variants in comparison to parallelizing compilers. Due to this, it is possible to reduce the porting time of sequential programs in the synthesis of solutions for reconfigurable computer systems with a set of FPGA chips interconnected by a spatial communication system. In the paper we show the results of porting a number of application tasks to the architecture of various reconfigurable computer systems using the proposed “Theseus” toolchain.
The article is devoted to the analysis of the current state of the supercomputer industry and the prospects for its development. In terms of methodological approach and tools, the work is a continuation of a series of similar analytical reviews by the authors. The main source of information for analysis is the archive of editions (releases) of the world ranking of the five hundred most productive supercomputers in the world. The novelty of this work lies not only in updating the information, taking into account the latest editions of the Top500 list, but also in focusing on the following circumstance: the global supercomputer industry is undergoing a radical restructuring – transition from the "petascale era" to the "exascale era". Technological trends and features of solutions for the most productive systems in the world in recent years are given. The pace of development of supercomputer technologies, development trends are discussed: hybrid architectures, interconnect technologies and changes in the positions of supercomputer manufacturing companies. Based on the results of the analysis, reliable forecasts were made for the coming years about the general appearance of exascale systems.
In the paper we consider a promising universal reconfigurable supercomputer, the computing nodes of which are reconfigurable computing device Arcturus. It was developed at the Supercomputers and Neurocomputers Research Center (Taganrog) and based on modern Xilinx FPGAs of UltraScale+ family of HBM-series. The purpose is to achieve the highest computational layout density, to ensure balanced power supply and cooling, as well as the implementation of powerful data exchange configuration. The supercomputer can have up to 1.5 thousand FPGAs of a single computing field. It has extensive information exchange capabilities between FPGAs within the device and between devices to solve tightly coupled problems. Differential lines with multi-gigabit transceivers connected to them are used as the main connections between FPGAs. It provides exchange at the velocity up to 25 Gbit/s. Information interaction between the RCD is performed through optical channels with the capacity up to 4.5 Tbit/s. Immersion technology is used for cooling components of computing system. It provides the removal of the total heat output up to 20 kW. The developed power supply configuration is based on the input constant voltage of 380 V and provides stable power supply to the components. Owning to the implementation of time-consuming algorithms for various scientific and technical problems with the high real performance, it is possible to widespread use of the Arcturus supercomputer. Scaling of computing nodes will allow designing an entire computing circuit of a supercomputer with the performance up to several tens of Petaflops.
The paper is devoted to a numerical study of detonation initiation in the gas mixture in the channel with a profiled end. Initiation occurs as a result of the reflection of the shock wave of relatively low intensity from the end of the channel. Numerical calculations are carried out using unstructured triangular grids. The numerical algorithm is parallelized by the computational domain decomposition method using the METIS library. The exchange of grid function values between computing cores is performed using the MPI library. Numerical calculations are conducted on grids with different numbers of triangular cells. Detonation initiation patterns are obtained, which correspond to each other. The differences are mainly related to the degree of resolution of the elements of the gas mixture flow in the channel.
The widespread development of unmanned aerial vehicles and light propeller-driven aircraft poses the task of reducing the community noise of such vehicles. To solve this problem, tools are needed to calculate the noise of such devices. The paper presents the results of the numerical simulation of the noise of an AV-2 propeller mounted on an AN-2 light propeller-driven aircraft. The authors use the acoustic-vortex method to solve the problem of aeroacoustic modeling of propeller noise in the presence of an incoming flow. The paper shows a good agreement of computed data with the in-flight experiment results and the calculation by the semi-empirical method. For the flight mode with an airspeed of 180 km/h, the deviation of the numerical simulation results from the experimental data does not exceed 2 dB.
Polyatomic gas cloud expansion into vacuum under pulsed laser ablation is studied on the basis of one-dimensional model kinetic equation. To account for the influence of internal energy on the vapor-gas cloud parameters (density, temperature, and velocity) a model kinetic equation of the BGK-type is applied, which considers the energy exchange using a two-temperature model. In this approach, the collision integral is approximated by the sum of two terms corresponding to elastic (translational relaxation) and inelastic (rotational relaxation) collisions. Note that the vibrational energy is not taken into account in this paper. The independence of the differential part of transfer equation from the rotational energy makes it possible to reduce the equation to a system with two functions. A comparison of the gas flow macroparameters obtained by the kinetic equations and the DSMC is carried out. The influence of the number of rotational degrees of freedom on the evolution of average temperature and on the gas parameters is shown. The calculations are performed on non-uniform grids in the phase space with global dynamic velocity mesh adaptation to suppress the “ray effect”.
Excitation of a laminar gas micro-jet by acoustic impact and by pulse-periodic heat source was simulated using the FlowVision software package in 2D formulation at normal conditions. Heat source imitates an electrical discharge. Air jet was formed by channel with inner size of 0.7 mm with the Poiseuille velocity profile at inlet boundary, the maximum profile velocity was varied in a range of 2.5–10 m/s. Influence of heat source frequency and power on the large-scale vortex formation was described. In the case of a jet with a speed of 5 m/s, the natural oscillations of the jet in response to a single pulse had a frequency fres = 1380 Hz, so excitation of the jet was possible at close frequencies of 1190 Hz and 1500 Hz. At the same time, at a frequency of 1000 Hz (approximately equal to 2/3 fres), every second impulse acted in antiphase and the oscillations developed poorly. Dependence of flow structure from the jet velocity was obtained. The results obtained show the possibility of exciting a micro-jet using low-power electrical discharges such as spark, DBD or corona.
The article discusses a parallel algorithm of solving linear algebraic equations systems for symmetric sparse matrices, which allows to split a large task into many small subtasks, thereby both increasing performance and reducing memory consumption. It is based on a method of simultaneous calculation of intermediate values during matrix factorization with maintaining load balancing on processors so that when the final result of the left parts of the factorization is obtained, the right parts of the factorization do not depend on them. This approach allows the initial stiffness matrix to be represented as a product of a large number of simple matrixes and solve a system of linear algebraic equations in the form of a sequence of solutions by substitution. To reduce the filling of sparse factorization matrices, an approximate minimum degree method was used, which, in addition to being one of the most efficient and fastest ones existing at the moment, allows the developed algorithm to distribute the load of calculations more evenly. The developed method is implemented in APM Ltd. software products for systems with shared memory, but it can also be performed for distributed memory systems.
A three-dimensional model of the multiphase flow based on the Eulerian–Eulerian approach was implemented using the FlowVision CFD package and, on this basis, a numerical algorithm for study of evaporation of liquid fuel was developed. The created high-performance complex for the carrier and dispersed phases interaction simulation was validated against the well-studied experimental problem of the evaporation and mixing of kerosene emerging from a flat pre-filming airblast atomizer for gas turbine combustors. In this work, the carrier phase is supposed to be air and kerosene vapors, and the dispersed phase is selected as liquid kerosene. Based on the calculated kerosene evaporation drops distributions, an important parameter that characterizes the spray fineness, Sauter mean diameter, is determined. Numerically calculated in the developed model the evaporation rate and Sauter mean diameter of fuel droplets agreed well with the experimental data. In famous works, the air temperature and pressure varied during the experiments. At the same time, in comparison with the calculated data, a stronger influence on the kerosene evaporation was obtained by air temperature than pressure. The dependence on pressure can be seen in the case of taking into account the corresponding changes in the liquid fuel properties. It is also noted that the initial fuel temperature is an important parameter for evaporation. This can be seen in the results of the kerosene evaporation numerical simulation carried out in this study.
The modern development of complex Earth system models forces developers to take into account not only the computational efficiency, but also the performance of the data input and output. This work evaluates the data output performance of the INM RAS Earth system model and optimizes its weak points. The output operations were found to be surprisingly slow on the Cray XC40-LC supercomputer compared to the results obtained on the INM RAS cluster. To identify the bottleneck, the computational time, the distributed data gathering time, and the file system output time were measured separately. The distributed data gathering time was the cause of the slowdown on the Cray XC40-LC, so optimizations were made to the gathering routines without any additional rework of the existing output code. The optimizations resulted in a significant reduction in the overall model running time on the Cray XC40-LC, while the gathering time itself was reduced by a factor of 102–103. The results highlight the importance of optimizing the output performance in Earth system models.
Stationary disturbances in a free supersonic flow caused by two-dimensional roughness in a turbulent boundary layer on a wall were investigated using numerical simulation in the FlowVision software package. Calculations were performed for cases where the roughness heights were significantly smaller than the thickness of the boundary layer. In order to reduce the numerical oscillations, the computational grid has been adapted to the perturbation fronts by means. It was found that a disturbance in the form of a small amplitude N-wave is formed in the flow above the boundary layer. The effect of roughness height on disturbances formed in a free flow was studied. It was found that as the roughness height increases, there is an observed increase in the amplitude and spatial scale of the disturbance. It was also found that the gradient of the flow parameters between the disturbance fronts remains practically unchanged for all roughness heights considered. The numerical results were verified with experimental data. A strong agreement was achieved between the simulation and experimental result.