This study introduces a novel framework designed to enhance the performance, scalability, and portability of the kilometer-scale E3SM Land Model (km-ELM) within the E3SM modeling infrastructure. By seamlessly integrating cutting-edge data tools, we address existing challenges such as slow performance, limited scalability, and difficulties in software integration in current data-driven ELM simulation over large geographic areas. Our innovative approach leverages the KiloCraft data toolkit to generate unified inputs for simulations ranging from a single-cite case, to a 72,083-cell regional case to a continental configuration encompassing 21.6 million land grid cells at a 1 km & times; 1 km resolution. We conduct extensive strong- and weak-scaling experiments on three state-of-the-art supercomputers, utilizing up to 100,800 CPU cores across 2400 compute nodes to evaluate end-to-end metrics including wall-clock time, simulation-years-per-day (SYPD), initialization costs, and I/O throughput. Our results reveal the land (LND) component's efficient scaling, demonstrating near-ideal weak scaling and strong-scaling parallel efficiencies reaching up to 87% at 50,400 cores. We confirm portability and reproducibility through bitwise-equivalent outputs across different machines using identical inputs over supported machines. Notably, at extreme scales, we identify I/O as a critical bottleneck and that leads to effective solution with the SCORPIO/ADIOS stack. Collectively, these findings validate the deployment of km-ELM at a continental scale with high parallel efficiency and provide essential guidance on configuration, decomposition, and I/O settings for optimized kilometer-scale land simulations in E3SM. This work emphasizes the innovative design and practical solutions that enhance the operational capabilities of km-ELM, focusing on software performance and scalability while leaving detailed scientific evaluations of simulated land processes for future investigations.
As high-performance computing (HPC) applications move to higher-resolution model grids and increase the frequency of the simulation output, achieving good I/O performance is critical to overall application performance. HPC applications also need to run on several different platforms for different users and simulations. Achieving good portable I/O performance is a challenging problem, especially considering the wide range of evolving file systems and computer architectures. Application I/O libraries such as ADIOS2, HDF5, PnetCDF, and netCDF have tried to solve this issue by providing high-level I/O abstractions and hiding the complexity of achieving portable I/O performance across computer architectures within the libraries. Nevertheless, applications often require further performance tuning, optimizations, and the ability to choose among these libraries and output file formats to achieve good portable I/O performance at scale across multiple computer architectures and file systems. In this paper we describe the design and implementation of the SCORPIO library, which was created to alleviate this problem for Earth system models. We discuss the features of the library that we leverage to achieve good I/O performance for production configurations of the DOE E3SM Earth system model.
This paper presents advancements in scaling the ultrahigh-resolution E3SM Land Model (uELM) for deployment on leadership-class supercomputers, addressing the increased demand for km-scale Earth system modeling. By focusing on km-scale ELM simulations, we enhance predictive capabilities for climate interactions, facilitating improved responses to climate change impacts on energy systems, agriculture, and water resources. Our approach leverages innovative software architecture optimizations, sophisticated data handling techniques, and advanced parallel processing, achieving strong scalability on two leadership supercomputers (2400 nodes (105,600 cores) on Summit, and 1200 nodes (76,800 cores) on Frontier). Results from extensive scalability assessments on the Summit and Frontier also demonstrate outstanding I/O performance (close to 400 GB/s write throughput) and the model's ability to efficiently handle increasing computational demands. This study not only establishes uELM's capability for high-resolution simulations over vast geographical domains, but also sets a foundation for future Earth system modeling breakthroughs.
The development of a kilometer-scale E3SM Land Model (km-scale ELM) is an integral part of the E3SM project, which seeks to advance energy-related Earth system science research with state-of-the-art modeling and simulation capabilities on exascale computing systems. Through the utilization of high-fidelity data products, such as atmospheric forcing and soil properties, the km-scale ELM plays a critical role in accurately modeling geographical characteristics and extreme weather occurrences. The model is vital for enhancing our comprehension and prediction of climate patterns, as well as their effects on ecosystems and human activities. This study showcases the first set of full-capability, km-scale ELM simulations over various computational domains, including simulations encompassing 21.6 million land gridcells, reflecting approximately 21.5 million square kilometers of North America at a 1 km x 1 km resolution. We present the largest km-scale ELM simulation using up to 100,800 CPU cores across 2,400 nodes. This continental-scale simulation is 300 times larger than any previous studies, and the computational resources used are about 400 times larger than those used in prior efforts. Both strong and weak scaling tests have been conducted, revealing exceptional performance efficiency and resource utilization. The km-scale ELM uses the common E3SM modeling infrastructure and a general data toolkit known as KiloCraft. Consequently, it can be readily adapted for both fully-coupled E3SM simulations and data-driven simulations over specific areas, ranging from a single gridcell to the entire North America.
We present an efficient and performance portable implementation of the Simple Cloud Resolving E3SM Atmosphere Model (SCREAM). SCREAM is a full featured atmospheric global circulation model with a nonhydrostatic dynamical core and state-of-the-art parameterizations for microphysics, moist turbulence and radiation. It has been written from scratch in C++ with the Kokkos library used to abstract the on-node execution model for both CPUs and GPUs. SCREAM is one of only a few global atmosphere models to be ported to GPUs. As far as we know, SCREAM is the first such model to run on both AMD GPUs and NVIDIA GPUs, as well as the first to run on nearly an entire Exascale system (Frontier). On Frontier, we obtained a record setting performance of 1.26 simulated years per day for a realistic cloud resolving simulation.
2 (E3SMv2) is a significant evolution from its predecessor E3SMv1, resulting in a model that is nearly twice as fast and with a simulated climate that is improved in many metrics.We describe the physical climate model in its lower horizontal resolution configuration consisting of 110 km atmosphere, 165 km land, 0.5°river routing model, and an ocean and sea ice with mesh spacing varying between 60 km in the mid-latitudes and 30 km at the equator and poles.The model performance is evaluated by means of a standard set of Coupled Model Intercomparison Project Phase 6 (CMIP6) Diagnosis, Evaluation, and Characterization of Klima (DECK) simulations augmented with historical simulations as well as simulations to evaluate impact of different forcing agents.The simulated climate is generally realistic, with notable improvements in clouds and precipitation compared to E3SMv1.E3SMv1 suffered from an excessively high equilibrium climate sensitivity (ECS) of 5.3 K.In E3SMv2, ECS is reduced to 4.0 K which is now within the plausible range based on a recent World Climate Research Programme (WCRP) assessment.However, E3SMv2 significantly underestimates the global mean temperature in the second half of the historical record.An analysis of single-forcing simulations indicates that correcting the historical temperature bias would require a substantial reduction in the magnitude of the aerosol-related forcing.
This work documents version two of the Department of Energy's Energy Exascale Earth System Model (E3SM). E3SMv2 is a significant evolution from its predecessor E3SMv1, resulting in a model that is nearly twice as fast and with a simulated climate that is improved in many metrics. We describe the physical climate model in its lower horizontal resolution configuration consisting of 110 km atmosphere, 165 km land, 0.5° river routing model, and an ocean and sea ice with mesh spacing varying between 60 km in the mid‐latitudes and 30 km at the equator and poles. The model performance is evaluated with Coupled Model Intercomparison Project Phase 6 Diagnosis, Evaluation, and Characterization of Klima simulations augmented with historical simulations as well as simulations to evaluate impacts of different forcing agents. The simulated climate has many realistic features of the climate system, with notable improvements in clouds and precipitation compared to E3SMv1. E3SMv1 suffered from an excessively high equilibrium climate sensitivity (ECS) of 5.3 K. In E3SMv2, ECS is reduced to 4.0 K which is now within the plausible range based on a recent World Climate Research Program assessment. However, a number of important biases remain including a weak Atlantic Meridional Overturning Circulation, deficiencies in the characteristics and spectral distribution of tropical atmospheric variability, and a significant underestimation of the observed warming in the second half of the historical period. An analysis of single‐forcing simulations indicates that correcting the historical temperature bias would require a substantial reduction in the magnitude of the aerosol‐related forcing.
Earth science high-performance applications often require extensive analysis of their output in order to complete the scien- tific goals or produce a visual image or animation. Often this analysis cannot be done in situ because it requires calculating time-series statistics from state sampled over the entire length of the run or analyzing the relationship between similar time series from previous simulations or observations. Many of the tools used for this postprocessing are not themselves high- performance applications, but the new Parallel Gridded Analysis Library (ParGAL) provides high-performance data-parallel versions of several common analysis algorithms for data from a structured or unstructured grid simulation. The library builds on several scalable systems, including the Mesh Oriented DataBase (MOAB), a library for representing mesh data that sup- ports structured, unstructured finite element, and polyhedral grids; the Parallel-NetCDF (PNetCDF) library; and Intrepid, an extensible library for computing operators (such as gradient, curl, and divergence) acting on discretized fields. We have used ParGAL to implement a parallel version of the NCAR Command Language (NCL) a scripting language widely used in the climate community for analysis and visualization. The data-parallel algorithms in ParGAL/ParNCL are both higher performing and more flexible than their serial counterparts.
Earth science high-performance applications often require extensive analysis of their output in order to complete the scientific goals or produce a visual image or animation. Often this analysis cannot be done in situ because it requires calculating time-series statistics from state sampled over the entire length of the run or analyzing the relationship between similar time series from previous simulations or observations. Many of the tools used for this postprocessing are not themselves highperformance applications, but the new Parallel Gridded Analysis Library (ParGAL) provides high-performance data-parallel versions of several common analysis algorithms for data from a structured or unstructured grid simulation. The library builds on several scalable systems, including the Mesh Oriented DataBase (MOAB), a library for representing mesh data that supports structured, unstructured finite element, and polyhedral grids; the Parallel-NetCDF (PNetCDF) library; and Intrepid, an extensible library for computing operators (such as gradient, curl, and divergence) acting on discretized fields. We have used ParGAL to implement a parallel version of the NCAR Command Language (NCL) a scripting language widely used in the climate community for analysis and visualization. The data-parallel algorithms in ParGAL/ParNCL are both higher performing and more flexible than their serial counterparts.
Many fields that employ computation require extensive analysis of the output from a petascale simulation of a grid(or mesh)-based application in order to complete their scientific goals or produce a visual image or animation. Often this analysis cannot be done in-situ because it requires calculating time-series statistics from state sampled over the entire length of the run or analyzing the relationship between similar time series from previous simulations or observations. The programs that perform this analysis are not nearly as flexible in their choice of grids or as high-performing as the primary applications. We will describe a new Parallel Gridded Analysis Library (ParGAL) that performs data-parallel versions of several common analysis algorithms on data from a structured or unstructured grid simulation. The library builds on several scalable systems starting with the Mesh Oriented DataBase (MOAB). MOAB is a library for representing mesh data that supports structured, unstructured finite element and polyhedral grids and also supports parallel operations on those grids including loading to and from disk using parallel I/O. We are using the Parallel-NetCDF (PNetCDF) library to perform parallel I/O operations between the popular NetCDF format and ParGAL. Finally, we also make use of Intrepid, an extensible library for computing operators (such as gradient, curl, divergence, etc.) acting on discretized fields. The design and performance of ParGAL will be described and an example of its application to climate compared to a widely used tool is given.
Climate models are both outputting larger and larger amounts of data and are doing it on more sophisticated numerical grids. The tools climate scientists have used to analyze climate output, an essential component of climate modeling, are single threaded and assume rectangular structured grids in their analysis algorithms. We are bringing both task- and data-parallelism to the analysis of climate model output. We have created a new data-parallel library, the Parallel Gridded Analysis Library (ParGAL) which can read in data using parallel I/O, store the data on a compete representation of the structured or unstructured mesh and perform sophisticated analysis on the data in parallel. ParGAL has been used to create a parallel version of a script-based analysis and visualization package. Finally, we have also taken current workflows and employed task-based parallelism to decrease the total execution time.
Parallel programming models on large-scale systems require a scalable system for managing the processes that make up the execution of a parallel program. The process-management system must be able to launch millions of processes quickly when starting a parallel program and must provide mechanisms for the processes to exchange the information needed to enable them communicate with each other. MPICH2 and its derivatives achieve this functionality through a carefully defined interface, called PMI, that allows different process managers to interact with the MPI library in a standardized way. In this paper, we describe the features and capabilities of PMI. We describe both PMI-1, the current generation of PMI used in MPICH2 and all its derivatives, as well as PMI-2, the second-generation of PMI that eliminates various shortcomings in PMI-1. Together with the interface itself, we also describe a reference implementation for both PMI-1 and PMI-2 in a new processmanagement framework within MPICH2, called Hydra, and compare their performance in running MPI jobs with thousands of processes.
Commercial HPC applications are often run on clusters that use the Microsoft Windows operating system and need an MPI implementation that runs efficiently in the Windows environment. The MPI developer community, however, is more familiar with the issues involved in implementing MPI in a Unix environment. In this paper, we discuss some of the differences in implementing MPI on Windows and Unix, particularly with respect to issues such as asynchronous progress, process management, shared-memory access, and threads. We describe how we implement MPICH2 on Windows and exploit these Windows-specific features while still maintaining large parts of the code common with the Unix version. We also present performance results comparing the performance of MPICH2 on Unix and Windows on the same hardware. For zero-byte MPI messages, we measured excellent shared-memory latencies of 240 and 275 nanoseconds on Unix and Windows, respectively.
Pavan Balaji合作论文数Mathematics and Computer Science Division,;Argonne National Laboratory4