The confrontation of complex Earth system model (ESM) codes with novel supercomputing architectures poses challenges to efficient modeling and job submission strategies. The modular setup of these models naturally fits a modular supercomputing architecture (MSA), which tightly integrates heterogeneous hardware resources into a larger and more flexible high-performance computing (HPC) system. While parts of the ESM codes can easily take advantage of the increased parallelism and communication capabilities of modern GPUs, others lag behind due to the long development cycles or are better suited to run on classical CPUs due to their communication and memory usage patterns. To better cope with these imbalances between the development of the model components, we performed benchmark campaigns on the Jülich Wizard for European Leadership Science (JUWELS) modular HPC system. We enabled the weather and climate model Icosahedral Nonhydrostatic (ICON) to run in a coupled atmosphere–ocean setup, where the ocean and the model I/O is running on the CPU Cluster, while the atmosphere is simulated simultaneously on the GPUs of JUWELS Booster (ICON-MSA). Both atmosphere and ocean are running globally with a resolution of 5 km. In our test case, an optimal configuration in terms of model performance (core hours per simulation day) was found for the combination of 84 GPU nodes on the JUWELS Booster module to simulate the atmosphere and 80 CPU nodes on the JUWELS Cluster module, of which 63 nodes were used for the ocean simulation and the remaining 17 nodes were reserved for I/O. With this configuration the waiting times of the coupler were minimized. Compared to a simulation performed on CPUs only, the MSA approach reduces energy consumption by 45 % with comparable runtimes. ICON-MSA is able to scale up to a significant portion of the JUWELS system, making best use of the available computing resources. A maximum throughput of 170 simulation days per day (SDPD) was achieved when running ICON on 335 JUWELS Booster nodes and 268 Cluster nodes.
Compute nodes in modern HPC systems are growing in size and their hardware has become ever more diverse. Still, many HPC centers allocate the resources of full nodes exclusively to avoid contention, despite the associated risk of underutilization. This paper describes a thorough resource utilization study of CPU and GPU compute and memory capacity, and interconnect bandwidth on JUWELS, a mature leadership-class modular supercomputer, with the aim of identifying opportunities for improving utilization through advanced scheduling and node sharing. Separate analysis of CPU-only and GPU-accelerated nodes finds that CPU compute usage is already close to optimal for the CPU-only nodes, whereas there is plenty of scope for co-scheduling CPU-based jobs on GPU-accelerated nodes. Memory capacity and node-level interconnect bandwidth are sufficient to provision co-scheduled jobs. We analyze multiple one-month datasets to validate robustness of conclusions over time and compare with previous studies on other systems to establish generalizability of results.
High Performance Computing (HPC) systems are among the most energy-intensive scientific facilities, with electric power consumption reaching and often exceeding 20 Megawatts per installation. Unlike other major scientific infrastructures such as particle accelerators or high-intensity light sources, which are few around the world, the number and size of supercomputers are continuously increasing. Even if every new system generation is more energy efficient than the previous one, the overall growth in size of the HPC infrastructure, driven by a rising demand for computational capacity across all scientific disciplines, and especially by Artificial Intelligence (AI) workloads, rapidly drives up the energy demand. This challenge is particularly significant for HPC centers in Germany, where high electricity costs, stringent national energy policies, and a strong commitment to environmental sustainability are key factors. This paper describes various state-of-the-art strategies and innovations employed to enhance the energy efficiency of HPC systems within the national context. Case studies from leading German HPC facilities illustrate the implementation of novel heterogeneous hardware architectures, advanced monitoring infrastructures, high-temperature cooling solutions, energy-aware scheduling, and dynamic power management, among other optimisations. By reviewing best practices and ongoing research, this paper aims to share valuable insight with the global HPC community, motivating the pursuit of more sustainable and energy-efficient HPC architectures and operations.
<p>Significant progress has been made in recent years to develop km-scale versions of global Earth System Models (ESM), combining the chance of replacing uncertain model parameterizations by direct treatment and the improved representation of orographic and land surface features (Sch&#228;r et al., 2020, Hohenegger et al., 2022). However, adapting climate codes to new hardware and at the same time keeping the performance portability, still remains a major issue. Given the long development cycles, the various maturity of ESM modules and their large code bases, it is not expected that all code parts can be brought to the same level of exascale readiness in the near future. Instead, short term model adaptation strategies need to focus on software abilities as well as hardware availability. Moreover, energy use efficiency is of growing importance on both sides, supercomputer providers and scientific projects employing climate simulations.</p> <p>Here, we present results from first simulations of the coupled atmosphere-ocean modelling system ICON-v2.6.6-rc on the supercomputing system JUWELS at the J&#252;lich Supercomputing Centre (JSC) with a global resolution of 5 km, using significant parts of the HPC system. While the atmosphere part of ICON (ICON-A) is capable of running on GPUs, model I/O currently performs better on a CPU cluster and the ocean module (ICON-O) has not been ported to modern accelerators yet. Thus, we make use of the modular supercomputing architecture (MSA) of JUWELS and its novel batch job options for the coupled ICON model with ICON-A running on the NVIDIA A100 GPUs of JUWELS Booster, while ICON-O and the model I/O are running simultaneously on the CPUs of the JUWELS Cluster partition. As expected, ICON performance is limited by ICON-A. Thus we chose the performance-optimal Booster-node configuration for ICON-A considering also memory requirements (84 nodes) and adapted ICON-O configuration to achieve minimum waiting times for simultaneous time step execution and data exchange (63 cluster nodes).&#160; We compared runtime and energy efficiency to cluster-only simulations (on up to 760 cluster nodes) and found only small improvements in runtime for the MSA case, but energy consumption is already reduced by 26% without further improvements in vector length applied with ICON. When switching to even higher ICON resolutions, cluster-only simulations are not fitting to most of current HPC systems and upcoming exascale systems will rely to a large extent on GPU acceleration. Thus exploiting MSA capabilities is an important step towards performance portable and energy efficient use of km-scale climate models.</p> <p>References:</p> <p>Hohenegger et al., ICON-Sapphire: simulating the components of the Earth System and their interactions at kilometer and subkilometer scales, https://doi.org/10.5194/gmd-2022-171, in review, 2022.</p> <p>Sch&#228;r et al., Kilometer-Scale Climate Models: Prospects and Challenges, https://doi.org/10.1175/BAMS-D-18-0167.1, 2020.</p> <p>&#160;</p>
Moments of the quark density, helicity, and transversity distributions are calculated in unquenched lattice QCD. Calculations of proton matrix elements of operators corresponding to these moments through the operator product expansion have been performed on 163 × 32 lattices for Wilson fermions at β = 5.6 using configurations from the SESAM collaboration and at β = 5.5 using configurations from SCRI. One-loop perturbative renormalization corrections are included. At quark masses accessible in present calculations, there is no statistically significant difference between quenched and full QCD results, indicating that the contributions of quark-antiquark excitations from the Dirac Sea are small. Close agreement between calculations with cooled configurations containing essentially only instantons and the full gluon configurations indicates that quark zero modes associated with instantons play ∗present address DESY/Zeuthen
Electroweak form factors provide a precise experimental probe of the quark substructure of hadrons, so it is of great interest to calculate them in QCD from first principles. Since the only known way to solve QCD is numerical solution on a discrete lattice, this work addresses the decomposition of matrix elements of the electromagnetic current into form factors consistent with cubic lattice symmetry and the relation of these form form factors to the familiar ones associated with continuum Lorentz symmetry. The breaking of continuum symmetries leads, in general, to more form factors on the lattice than in the continuum. In a practical Monte Carlo calculation on the lattice, it is advantageous to calculate these lattice form factors since all momenta related by a lattice symmetry can be combined to obtain the highest statistical accuracy. Therefore, it is necessary to determine the appropriate combination of these form factors that approaches the physical form factors in the continuum limit. In this work, we consider the electromagnetic form factors of the nucleon. The matrix element of the electromagnetic current in the continuum is
The DEEP projects have developed a variety of hardware and software technologies aiming at improving the efficiency and usability of next generation high-performance computers. They evolve around an innovative concept for heterogeneous systems: the Cluster-Booster architecture. In it, a general purpose cluster is tightly coupled to a manycore system (the Booster). This modular way of integrating heterogeneous components enables applications to freely choose the kind of computing resources on which it runs most efficiently. Codes might even be partitioned to map specific requirements of code-parts onto the best suited hardware. This paper presents for the first time measurements done by a real world scientific application demonstrating the performance gain achieved with this kind of code-partition approach.
The J¨ulich Supercomputing Centre is internationally recognised by two major features: the ex- cellent high-performance computing services that it provides to the scientific community, and its ability to drive the development of cutting-edge supercomputing technology. The latter is perfectly illustrated with the supercomputer evolution experienced at JSC, with three unique architecture approaches developed in the centre in slightly over a decade, namely the Dual Supercomputer approach, the Cluster-Booster concept and the Modular Supercomputer architecture. They represent an evolution in the way HPC systems are built and operated, aiming at offering real-world applications the best possible computing platform.
The recently completed research project DEEP-ER has developed a variety of hardware and software technologies to improve the I/O capabilities of next generation high-performance computers, and to enable applications recovering from the larger hardware failure rates expected on these machines. The heterogeneous Cluster-Booster architecture --first introduced in the predecessor DEEP project-- has been extended by a multi-level memory hierarchy employing non-volatile and network-attached memory devices. Based on this hardware infrastructure, an I/O and resiliency software stack has been implemented combining and extending well established libraries and software tools, and sticking to standard user-interfaces. Real-world scientific codes have tested the projects' developments and demonstrated the improvements achieved without compromising the portability of the applications.
Heading towards exascale, the challenges for process management with respect to flexibility and efficiency grow accordingly. Running more than one application simultaneously on a node can be the solution for better resource utilization. However, this approach of co-scheduling can also be the way to go for gaining a degree of flexibility with respect to process management that can enable some kind of interactivity even in the domain of high-performance computing. This chapter gives an introduction into such co-scheduling policies for running multiple MPI sessions concurrently and interactively within a single user allocation. The chapter initially introduces a taxonomy for classifying the different characteristics of such a flexible process management, and discusses actual manifestations thereof during the course of the reading. In doing so, real world examples are motivated and presented by means of ParaStation MPI, a high-performance MPI library supplemented by a complete framework comprising a scalable and dynamic process manager. In particular, four scheduling policies, implemented in ParaStation MPI, are detailed and evaluated by applying a benchmarking tool that has especially been developed for measuring interactivity and dynamicity metrics of job schedulers and process managers for high-performance computing.
Heading towards exascale, the challenges for process management with respect to flexibility and efficiency grow accordingly. Running more than one application simultaneously on a node can be the solution for better resource utilization. However, this approach of co-scheduling can also be the way to go for gaining a degree of flexibility with respect to process management that can enable some kind of interactivity even in the domain of high-performance computing. This chapter gives an introduction into such co-scheduling policies for running multiple MPI sessions concurrently and interactively within a single user allocation. The chapter initially introduces a taxonomy for classifying the different characteristics of such a flexible process management, and discusses actual manifestations thereof during the course of the reading. In doing so, real world examples are motivated and presented by means of ParaStation MPI, a high-performance MPI library supplemented by a complete framework comprising a scalable and dynamic process manager. In particular, four scheduling policies, implemented in ParaStation MPI, are detailed and evaluated by applying a benchmarking tool that has especially been developed for measuring interactivity and dynamicity metrics of job schedulers and process managers for high-performance computing.
SummaryHomogeneous cluster architectures, which used to dominate high‐performance computing (HPC), are challenged today by heterogeneous approaches utilizing accelerator or co‐processor devices. The DEEP (Dynamical Exascale Entry Platform) project is implementing a novel architecture for HPC, in which a standard HPC cluster is directly connected to a so‐called ‘Booster’: a cluster of many‐core processors. By these means heterogeneity is organized differently as in today's standard approach, where accelerators are added to each node of the cluster. In order to adapt application codes to this Cluster‐Booster architecture as seamless as possible, DEEP has developed a complete programming environment. It integrates the offloading functionality given by the Message Passing Interface standard with an abstraction layer based on the task‐based OmpSs programming paradigm. This paper presents the DEEP project with an emphasis on the DEEP programming environment. Copyright © 2015 John Wiley & Sons, Ltd.
Homogeneous cluster architectures dominating high-performance computing (HPC) today are challenged, in particular when thinking about reaching Exascale by the end of the decade, by heterogeneous approaches utilizing accelerator elements. The DEEP (Dynamical Exascale Entry Platform) project aims for implementing a novel architecture for high-performance computing consisting of two components - a standard HPC Cluster and a cluster of many-core processors called Booster. In order to make the adaptation of application codes to this Cluster-Booster architecture as seamless as possible, DEEP provides a complete programming environment. It integrates the offloading functionality given by the MPI standard with an abstraction layer based on the task-based OmpSs programming paradigm. This paper presents the DEEP project with an emphasis on the DEEP programming environment.
The High Order Method Modeling Environment (HOMME) and the modified version of The Parallel Ocean Program (POPperf) are two important applications for atmospheric and weather research. With an emphasis on efficiency, portability, maintainability and most importantly, scalability, HOMME and POPperf have been successfully deployed over the years on a wide variety of highperformance systems, such as Cray and Blue-Gene. With the increased adoption of HPC commodity clusters based on the high-speed InfiniBand network, understanding HOMME and POPperf scalability and optimization options available for clusters at scale are crucial for utilizing the capability of InfiniBand-based systems to serve as cost effective high-performance and highly scalable solutions. Our results identify HOMME and POPperf scaling capabilities on one of the world’s largest InfiniBand networks and demonstrate the critical elements for optimizing these applications at scale.
Cluster computers are dominating high performance computing (HPC) today. The success of this architecture is based on the fact that it proffits from the improvements provided by mainstream computing well known under the label of Moore's law. But trying to get to Exascale within this decade might require additional endeavors beyond surfing this technology wave. In order to find possible directions for the future we review Amdahl's and Gustafson's thoughts on scalability. Based on this analysis we propose an advance architecture combining a Cluster with a so called Booster element comprising of accelerators interconnected by a high performance fabric. We argue that this architecture provides significant advantages compared to today's accelerated clusters and might pave the way for clusters into the era of Exascale computing. The DEEP project has been presented aiming for an implementation of this concept. Six applications from fields having the potential to exploit Exascale systems will be ported to DEEP.We analyze one application in detail and explore the consequences of the constraints of the DEEP systems on its scalability.