This work explores the performance of single- and multi-GPU computing on state-of-the-art NVIDIA- and AMD-based server-class hardware using various programming interfaces to accelerate a real-world scientific application for solidification modeling based on the phase-field method. The main computations of this memory-bound application correspond to 20 stencils computed across grid nodes. We investigate the application's scalability for two basic schemes of organizing computation: without and with hiding data transfers behind computation, combined with using either peer-to-peer inter-GPU data transfers through NVIDIA NVLink and AMD Infinity interconnects or communication over the PCIe and main memory. Among the studied programming interfaces is CUDA, HIP, and OpenMP Accelerator Model. While the first two are designed to write the codes for a specific hardware platform, OpenMP enables code portability between NVIDIA and AMD GPUs. The resulting performance is experimentally assessed on computing platforms containing NVIDIA V100 (up to 8 GPUs) and A100 (one GPU), as well as AMD MI210 (one device) and MI250 (up to 8 logical GPUs).
SCADvanceXP is an industrial network intrusion detection system that scans and monitors data exchange between engineering stations, field divides, controllers, supervisory control and data acquisition (SCADA), and other elements of the operational technology network in detail. SCADvanceXP has the potential to detect advanced attacks on industrial infrastructures with the use of rulebased, signature-based, and behavioural detection methods, which are supported by sophisticated machine and deep learning models. As a system developed in Poland, it addresses the needs of industry in that region of Europe. The goal of this work was to assess SCADvanceXP’s potential to detect common industrial threats. In order to check SCADvanceXP’s potential, an effort was undertaken to evaluate its functionality on major industrial threats. For that purpose, twelve malware strains interfering with industrial systems were described. Later, the SCADvanceXP functionality was overlapped on malware behavioural and detection markers, pointing out exact mechanisms in SCADvanceXP that would detect analysed threats. The results show that SCADvanceXP is able to detect a wide range of attacks on industrial networks. SCADvanceXP’s rich functionality is able to provide a high standard of security. However, if a threat is affecting systems not directly connected with industrial networks, SCADvanceXP will not be able to detect it. SCADvanceXP only monitors industrial systems; hence, corporate networks must be protected by a different solution to provide the required level of security. Nonetheless, SCADvanceXP is dedicated to operating within industrial networks and does not have access to regular IT networks. It can be concluded that SCADvanceXP is a specialist tool providing desired security for industrial networks.
The Natural History Collections of Adam Mickiewicz University (AMUNATCOLL) in Poznań contain over 2.2 million specimens. Until recently, access to the collections was limited to specialists and was challenging because of the analogue data files. Therefore, this paper presents a new approach to data sharing called the Scientific, Educational, Public, and Practical Use (SEPP) Model. Since the stakeholder group is broad, the SEPP Model assumes the following key points: full open access to the digitized collections, the structure of metadata in accordance with certain standards, and a versatile tool set for data mining or statistical and spatial analysis. The SEPP Model was implemented in the AMUNATCOLL IT system, which consists of a web portal equipped with a wide set of explorative functionalities tailored to different user groups: scientists, students, officials, and nature enthusiasts. An integral part of the system is a mobile application designed for field surveys, enabling users to conduct studies comparing their own field data and AMUNATCOLL data. The AMUNATCOLL IT database contains digital data on specimens, biological samples, bibliographic sources, and multimedia nature documents. The metadata structure was developed in accordance with ABCD 2.06 and Darwin Core standards.
Abstract This paper describes a project aimed at digitizing and openly sharing the natural history collections (AMUNATCOLL) of the Faculty of Biology at Adam Mickiewicz University in Poznań (Poland). The result of this project is a database (including 2.2 million records) of plant, fungal and animal specimens, which is available online via the AMUNATCOLL portal and on the Global Biodiversity Information Facility website. This article presents selected aspects of the “life cycle” of this project, with a particular focus on its preparatory phase.
The phase-field (PF) method is a powerful tool for solving interfacial problems in materials science. This paper's primary goal is to assess the impact of various state-of-the-art C/C++ compilers available for novel AMD EPYC Rome and Milan multi-core processors on the performance of real-world scientific codes corresponding to the solidification modeling application using the PF method and generalized finite difference scheme for solving governing PDEs. Among the studied compilers are AOCC, Clang, GCC, Intel compiler, and PGI. Besides performance, the numerical accuracy of the simulation is verified since various optimizations used by the compilers can cause differences in the simulation results. For the objectivity of the assessment, we study different application versions with the increasing complexity of memory access patterns and control structures. All tested codes are executed on a dual-socket server equipped with 64-core AMD EPYC Rome 7742 CPUs. This assessment confirms the advantage of the Intel compiler over other compilers. In particular, while for the static intensity only GCC lags noticeably behind other compilers, in the case of the dynamic intensity, the clear winner is the Intel compiler, which allows increasing the performance up to about 1.3 times over the AMD compiler. Moreover, we reveal the sensitivity of the application version and compiler to selecting a different number of NUMA domains and enabling/disabling the NUMA balancing option in EPYC processors. By comparing the performance gain achieved by various compilers due to vectorization of the application codes, we show that the Intel compiler's performance advantage results primarily from considerably better utilization of capabilities of AVX2 vector instructions. Finally, we show the importance of choosing a compiler when comparing the performance of AMD EPYC and Intel Xeon processors.
The synergy between Artificial Intelligence and the Edge Computing paradigm promises to transfer decision-making processes to the periphery of sensor networks without the involvement of central data servers. For this reason, we recently witnessed an impetuous development of devices that integrate sensors and computing resources in a single board to process data directly on the collection place. Due to the particular context where they are used, the main feature of these boards is the reduced energy consumption, even if they do not exhibit absolute computing powers comparable to modern high-end CPUs. Among the most popular Artificial Intelligence techniques, clustering algorithms are practical tools for discovering correlations or affinities within data collected in large datasets, but a parallel implementation is an essential requirement because of their high computational cost. Therefore, in the present work, we investigate how to implement clustering algorithms on parallel and low-energy devices for edge computing environments. In particular, we present the experiments related to two devices with different features: the quad-core UDOO X86 Advanced+ board and the GPU-based NVIDIA Jetson Nano board, evaluating them from the performance and the energy consumption points of view. The experiments show that they realize a more favorable trade-off between these two requirements than other high-end computing devices.
This work summarizes the results of a set of executions completed on three fat-tree network supercomputers: Stampede at TACC (USA), Helios at IFERC (Japan) and Eagle at PSNC (Poland). Three MPI-based, communication-intensive scientific applications compiled for CPUs have been executed under weak-scaling tests: the molecular dynamics solver LAMMPS; the finite element-based mini-kernel miniFE of NERSC (USA); and the three-dimensional fast Fourier transform mini-kernel bigFFT of LLNL (USA). The design of the experiments focuses on the sensitivity of the applications to rather different patterns of task location, to assess the impact on the cluster performance. The accomplished weak-scaling tests stress the effect of the MPI-based application mappings (concentrated vs. distributed patterns of MPI tasks over the nodes) on the cluster. Results reveal that highly distributed task patterns may imply a much larger execution time in scale, when several hundreds or thousands of MPI tasks are involved in the experiments. Such a characterization serves users to carry out further, more efficient executions. Also researchers may use these experiments to improve their scalability simulators. In addition, these results are useful from the clusters administration standpoint since tasks mapping has an impact on the cluster throughput.
The mission of the PRACE Research Infrastructure is to enable high impact scientific discovery and engineering research and development across all disciplines to enhance European competitiveness for the benefit of society. The paper presents the current state of the computational infrastructure, which is unique on the European level and procedures helping with the seamless access to the European HPC infrastructure: supercomputers, applications, storage. Several Polish computational grants have been called in the paper in order to present examples of the Polish domestic scientific engagement.
This investigation summarizes a set of executions completed on the supercomputers Stampede at TACC (USA), Helios at IFERC (Japan), and Eagle at PSNC (Poland), with the molecular dynamics solver LAMMPS, compiled for CPUs. A communication-intensive benchmark based on long-distance interactions tackled by the Fast Fourier Transform operator has been selected to test its sensitivity to rather different patterns of tasks location, hence to identify the best way to accomplish further simulations for this family of problems. Weak-scaling tests show that the attained execution time of LAMMPS is closely linked to the cluster topology and this is revealed by the varying time-execution observed in scale up to thousands of MPI tasks involved in the tests. It is noticeable that two clusters exhibit time saving (up to 61
The work undertaken in this paper was done in the Centre of Excellence for Global Systems Science (CoeGSS) – an interdisciplinary project funded by the European Commission. CoeGSS project provides a computer-aided decision support in the face of global challenges (e.g. development of energy, water and food supply systems, urbanisation processes and growth of the cities, pandemic control, etc.) and tries to bring together HPC and global systems science. This paper presents a proposition of GSS benchmark which evaluates HPC architectures with respect to GSS applications and seeks for the best HPC system for typical GSS software environments. The outcome of the analysis is defining a benchmark which represents the average GSS environment and its challenges in a good way: spread of smoking habits and development of tobacco industry, development of green cars market and global urbanisation processes. Results of the tests that have been run on a number of recently appeared HPC platforms allow comparing processors’ architectures with respect to different applications using execution times, TDPs3 and TCOs4 as the basic metrics for ranking HPC architectures. Finally, we believe that our analysis of the results conveys a valuable information to the broadened GSS audience which might help to determine the hardware demands for their specific applications, as well as to the HPC community which requires a mature benchmark set reflecting requirements and traits of the GSS applications. Our work can be considered as a step into direction of development of such mature benchmark.
The work undertaken in this paper was done in the Centre of Excellence for Global Systems Science (CoeGSS), an interdisciplinary project, funded by the European Commission. The project provides decision-support in the face of global challenges. It brings together HPC and global systems science. This paper presents a proposition of GSS benchmark with the aim to find the most suitable HPC architecture and the best HPC system which allows to run GSS applications effectively. The GSS provides evidence about global systems challenges, e.g. the network structure of the world economy, energy, water and food supply systems, the global financial system or the global city system, and the scientific community. The outcome of the analysis is defining a benchmark which represents the GSS environment in the best way. Three exemplary challenges were defined as pilot applications: Health Habits, Green Growth and Global Urbanisation extended with additional applications from GSS ecosystem: Iterative proportional fitting (IPF), Data rastering - a preprocessing process converting all vectorial representations of georeferenced data into raster files to be later used as simulation input, Weather Research and Forecasting (WRF) model, CMAQ/CCTM (Community Air Multiscale Quality Modelling System/The CMAQ Chemistry-Transport Mode), CM1 (Cloud Modelling), ABMS (Agent-based Modelling and Simulation), OpenSWPC (An Open-source Seismic Wave Propagation Code). The above list seems to be quite rich and reflects the real GSS world as much as possible, having in mind, for example the real-world applications availability. Additionally, the authors tested new HPC platforms based on Intel® Xeon® Gold 6140, AMD EpycTM, ARM Hi1616 and IBM Power8+. Due to the hardware availability, the testbed consisted of a limited number of nodes. This restricted the ability to provide full tests of scalability for given applications. However, this small number of available computational units (cores) can provide valuable outcome including architecture comparison for different applications based on execution times, TDPs1 and TCO2. These are the basic metrics used for providing a ranking of HPC architectures. Finally, this document is thought to be valuable information for the GSS community for future purposes and analysis to determine their specific demands as well as - in general - to help develop a mature final benchmark set reflecting the GSS environment requirements and specialty. As none of the existing benchmarks is dedicated to the GSS community, the authors decided to create one by calling it a GSS benchmark to serve and help GSS users in their future work.
W dobie zagrożeń asymetrycznych cyberbezpieczeństwo infrastruktury krytycznej staje się poważną kwestią, a jednocześnie wyzwaniem dla twórców systemów zabezpieczeń. W niniejszym artykule przedstawiono czynniki eskalujące poziom trudności detekcji zaawansowanych zagrożeń, a także, na przykładzie dwóch projektów naukowo-badawczych, opisano realizowane przez Poznańskie Centrum Superkomputerowo-Sieciowe (PCSS) prace podejmujące to wyzwanie. Na przykładzie krajowego projektu SCADvance opisano zastosowanie algorytmów uczenia maszynowego do wykrywania zagrożeń w protokołach sieci przemysłowych. Wskazano również na rolę, jaką środowisko naukowe jest w stanie odegrać w tworzeniu innowacyjnych systemów zabezpieczeń infrastruktury krytycznej, a także na konieczność zastosowania rozwiązań tej klasy dla właściwej ochrony wrażliwych sieci teleinformatycznych.
In this paper an overview of the problem of cybersecurity monitoring and analytics in HPC centers is performed from two intersecting points of view: challenges of assuring the necessary security level of HPC infrastructures themselves as well as new, not available earlier, opportunities to effectively analyze large volumes of heterogeneous data, facilitated by using large HPC clusters together with scalable analytic software. A major part of this paper is devoted to the most relevant methodologies and solutions that can be used by security analytics in order to at least partially face the challenge of analyzing large volumes of data potentially related with cyber-security events, in real-time or quasi-real-time. Particular solutions are considered in the context of their applicability in an HPC infrastructure. Relying on the results of experiments conducted within the SECOR project we have shown an approach of further development of the prepared architecture in HPC environment – within the confines of another R&D project, PROTECTIVE.
The article describes the PL-Grid computing infrastructure built as a result of the first PL-Grid-based project "Polish Infrastructure for Supporting Computational Science in the European Research Space", then enlarged, upgraded and improved during the PLGrid Plus project. Five Polish supercomputing sites joined forces to create a cross-country integrated computing service for our scientists. These sites were Cyfronet in Kraków, ICM in Warsaw, PSNC in Poznań, TASK in Gdańsk and WCSS in Wrocław. PL-Grid Infrastructure enables Polish scientists carrying out scientific research based on the simulations and large-scale calculations using the computing clusters as well as it provides convenient access to distributed computing resources.
Luděk Matyska合作论文数Institute of Computer Science4
Cezary Mazurek合作论文数Poznan Supercomputing and Networking Center3