Power is a fundamental constraint as supercomputing advances to exascale. Efficient operation within strict power budgets requires application-aware power management based on a detailed understanding of application-level power behavior. This work analyzes the Energy Exascale Earth System Model (E3SM) atmosphere component, SCREAM, on Perlmutter (NERSC) and Frontier (OLCF). We characterize power variation across inputs, concurrency levels, and power caps, evaluate the energy impact of code optimizations, and attribute energy within the code using a newly developed GPU energy model. Results show that SCREAM’s peak power remains stable during its core execution phase and decreases gradually as concurrency increases. Power capping experiments reveal a performance–energy "sweet spot". On Perlmutter, limiting GPU power to 50% of thermal design power (TDP) achieves up to 15% energy savings with a 7% performance penalty. On Frontier, a 40% TDP cap yields up to 10% energy savings with less than 10% performance loss. Code optimizations reduce SCREAM energy by shortening run time without increasing power. Modeling reveals a critical insight: data movement accounts for approximately 70% of SCREAM’s GPU energy. This fundamentally shifts the optimization focus from FLOPS to data transfer reduction for this class of applications, offering the most impactful strategy for improving energy efficiency. This work establishes a foundation for practical, application-aware power management at exascale.
This paper presents the development of a new entropy-based feature selection method for identifying and quantifying impacts. Here, impacts are defined as statistically significant differences in spatio-temporal fields when comparing datasets with and without an external forcing in an Earth system model. Temporal feature selection is performed by first computing the cross-fuzzy entropy to quantify similarity of patterns between two datasets and then applying changepoint detection to identify regions of statistically constant entropy. The method is used to capture temperate north surface cooling from a 9-member simulation ensemble of the Mt. Pinatubo volcanic eruption, which injected 10 Tg of SO2 into the stratosphere. The results estimate a mean difference decrease in near surface air temperature of −0.560 K with a 99% confidence interval between −0.864 K and −0.257 K between April and November of 1992, one year following the eruption. A sensitivity analysis with decreasing SO2 injection revealed that the impact is statistically significant at 5 Tg but not at 3 Tg. Using identified features, a dependency graph model based on a 9-day lag had significantly fewer nodes than a graph based on monthly means. This demonstrates our method’s ability to perform dimension reduction while still uncovering source-to-impact pathways.
We propose an approach for characterizing source-impact pathways, the interactions of a set of variables in space-time due to an external forcing, in climate models using in-situ analyses that circumvent computationally expensive read/write operations. This approach makes use of a lightweight open-source software library we developed known as CLDERA-Tools. We describe how CLDERA-Tools is linked with the U.S. Department of Energy’s Energy Exascale Earth System Model (E3SM) in a minimally invasive way for in-situ extraction of quantities of interested and associated statistics. Subsequently, these quantities are used to represent source-impact pathways with time-dependent directed acyclic graphs (DAGs). The utility of CLDERA-Tools is demonstrated by using the data it extracts in-situ to compute a spatially resolved DAG from an idealized configuration of the atmosphere with a parameterized representation of a volcanic eruption known as HSW-V.
High-resolution simulations of polar ice sheets play a crucial role in the ongoing effort to develop more accurate and reliable Earth system models for probabilistic sea-level projections. These simulations often require a massive amount of memory and computation from large supercomputing clusters to provide sufficient accuracy and resolution; therefore, it has become essential to ensure performance on these platforms. Many of today’s supercomputers contain a diverse set of computing architectures and require specific programming interfaces in order to obtain optimal efficiency. In an effort to avoid architecture-specific programming and maintain productivity across platforms, the ice-sheet modeling code known as MPAS-Albany Land Ice (MALI) uses high-level abstractions to integrate Trilinos libraries and the Kokkos programming model for performance portable code across a variety of different architectures. In this article, we analyze the performance portable features of MALI via a performance analysis on current CPU-based and GPU-based supercomputers. The analysis highlights not only the performance portable improvements made in finite element assembly and multigrid preconditioning within MALI with speedups between 1.26 and 1.82x across CPU and GPU architectures but also identifies the need to further improve performance in software coupling and preconditioning on GPUs. We perform a weak scalability study and show that simulations on GPU-based machines perform 1.24–1.92x faster when utilizing the GPUs. The best performance is found in finite element assembly, which achieved a speedup of up to 8.65x and a weak scaling efficiency of 82.6% with GPUs. We additionally describe an automated performance testing framework developed for this code base using a changepoint detection method. The framework is used to make actionable decisions about performance within MALI. We provide several concrete examples of scenarios in which the framework has identified performance regressions, improvements, and algorithm differences over the course of 2 years of development.
We present an efficient and performance portable implementation of the Simple Cloud Resolving E3SM Atmosphere Model (SCREAM). SCREAM is a full featured atmospheric global circulation model with a nonhydrostatic dynamical core and state-of-the-art parameterizations for microphysics, moist turbulence and radiation. It has been written from scratch in C++ with the Kokkos library used to abstract the on-node execution model for both CPUs and GPUs. SCREAM is one of only a few global atmosphere models to be ported to GPUs. As far as we know, SCREAM is the first such model to run on both AMD GPUs and NVIDIA GPUs, as well as the first to run on nearly an entire Exascale system (Frontier). On Frontier, we obtained a record setting performance of 1.26 simulated years per day for a realistic cloud resolving simulation.
Earth and Space Science Open Archive This work has been accepted for publication in Journal of Advances in Modeling Earth Systems (JAMES). Version of RecordESSOAr is a venue for early communication or feedback before peer review. Data may be preliminary. Learn more about preprints. preprintOpen AccessYou are viewing an older version [v1]Go to new versionConvection-Permitting Simulations with the E3SM Global Atmosphere ModelAuthorsPeter MartinCaldwelliDChristopher RyutaroTeraiiDBenjamin RHillmanNoel D.KeeniDPeter ABogenschutzWuyinLinHassanBeydouniDMark ATayloriDLucaBertagnaiDAndrewBradleyThomas CClevengeriDAaron SheffieldDonahueiDChrisEldredJames GFoucarJean-ChristopheGolaziDOksanaGubaRobert LJacobJeffJohnsoniDJagadishKrishnaWeiranLiuiDKyle GPresselAndrew G.SalingeriDBalwinderSinghAndrewSteyerPaulUllrichiDDanqingWuXingqiuYuanJacobShpundHsi-YenMaiDCharles SuttonZenderiDSee all authors Peter Martin CaldwelliDCorresponding Author• Submitting AuthorLawrence Livermore National Laboratory (DOE)iDhttps://orcid.org/0000-0001-8604-0844view email addressThe email was not providedcopy email addressChristopher Ryutaro TeraiiDUniversity of California - IrvineiDhttps://orcid.org/0000-0002-2433-0472view email addressThe email was not providedcopy email addressBenjamin R HillmanSandia National Laboratoriesview email addressThe email was not providedcopy email addressNoel D. KeeniDLawrence Berkeley National Laboratory (DOE)iDhttps://orcid.org/0000-0003-3607-3554view email addressThe email was not providedcopy email addressPeter A BogenschutzLawrence Livermore National Laboratoryview email addressThe email was not providedcopy email addressWuyin LinBrookhaven National Laboratoryview email addressThe email was not providedcopy email addressHassan BeydouniDKarlsruhe Institute of TechnologyiDhttps://orcid.org/0000-0003-4094-8173view email addressThe email was not providedcopy email addressMark A TayloriDSandia National LaboratoriesiDhttps://orcid.org/0000-0002-9267-2554view email addressThe email was not providedcopy email addressLuca BertagnaiDUnknowniDhttps://orcid.org/0000-0002-6171-3202view email addressThe email was not providedcopy email addressAndrew BradleySandia National Laboratoryview email addressThe email was not providedcopy email addressThomas C ClevengeriDSandia National LabiDhttps://orcid.org/0000-0002-3340-2482view email addressThe email was not providedcopy email addressAaron Sheffield DonahueiDLawrence Livermore National LaboratoryiDhttps://orcid.org/0000-0002-4710-753Xview email addressThe email was not providedcopy email addressChris EldredLAGA, University of Parisview email addressThe email was not providedcopy email addressJames G FoucarSandia National Laboratory (DOE)view email addressThe email was not providedcopy email addressJean-Christophe GolaziDLawrence Livermore National Laboratory (DOE)iDhttps://orcid.org/0000-0003-1616-5435view email addressThe email was not providedcopy email addressOksana GubaSandia National Laboratoriesview email addressThe email was not providedcopy email addressRobert L JacobArgonne Notional Laboratoryview email addressThe email was not providedcopy email addressJeff JohnsoniDCohere LLCiDhttps://orcid.org/0000-0002-0265-5241view email addressThe email was not providedcopy email addressJagadish KrishnaUnknownview email addressThe email was not providedcopy email addressWeiran LiuiDUC DavisiDhttps://orcid.org/0000-0002-7559-4726view email addressThe email was not providedcopy email addressKyle G PresselPacific Northwest National Laboratoryview email addressThe email was not providedcopy email addressAndrew G. SalingeriDSandia National LaboratoryiDhttps://orcid.org/0000-0003-4692-6813view email addressThe email was not providedcopy email addressBalwinder SinghPacific Northwest National Laboratory (DOE)view email addressThe email was not providedcopy email addressAndrew SteyerSandia National Laboratoryview email addressThe email was not providedcopy email addressPaul UllrichiDUniversity of California DavisiDhttps://orcid.org/0000-0003-4118-4590view email addressThe email was not providedcopy email addressDanqing WuArgonne National Labview email addressThe email was not providedcopy email addressXingqiu YuanArgonne National Labview email addressThe email was not providedcopy email addressJacob ShpundUnknownview email addressThe email was not providedcopy email addressHsi-Yen MaiDLLNLiDhttps://orcid.org/0000-0002-9628-1278view email addressThe email was not providedcopy email addressCharles Sutton ZenderiDUniversity of California, IrvineiDhttps://orcid.org/0000-0003-0129-8024view email addressThe email was not providedcopy email address
Focal Area(s): Primary focal areas are: predictive modeling through the use of AI-derived model components; advanced methods including network design/optimization/deep learning. Science Challenge: The atmospheric, ocean and ice dynamics components of the Energy Exascale Earth System Model (E3SM) are governed by Partial Differential Equations (PDEs) and significant efforts have been made during the last decades to develop such computational models. Here we propose to fundamentally improve these PDE-based codes by enhancing them with Machine Learning (ML) sub-models for complex, poorly understood physical processes in the context of ice sheet modeling. We propose to train these models with a novel approach that allows the assimilation of the different sources of data available (direct/indirect observations and possibly simulation data), improving on existing simplified models. We also highlight computational challenges originating from the coexistence of PDE-based and ML-based models.
We present an effort to port the nonhydrostatic atmosphere dynamical core of the Energy Exascale Earth System Model (E3SM) to efficiently run on a variety of architectures, including conventional CPU, many-core CPU, and GPU. We specifically target cloud-resolving resolutions of 3 km and 1 km. To express on-node parallelism we use the C++ library Kokkos, which allows us to achieve a performance portable code in a largely architecture-independent way. Our C++ implementation is at least as fast as the original Fortran implementation on IBM Power9 and Intel Knights Landing processors, proving that the code refactor did not compromise the efficiency on CPU architectures. On the other hand, when using the GPUs, our implementation is able to achieve 0.97 Simulated Years Per Day, running on the full Summit supercomputer. To the best of our knowledge, this is the most achieved to date by any global atmosphere dynamical core running at such resolutions.
Earth and Space Science Open Archive posterOpen AccessYou are viewing the latest version by default [v1]Performance-Portability Results for the Non-Hydrostatic Atmosphere Dycore of E3SM at Cloud-Resolving Resolutions.AuthorsLucaBertagnaiDOksanaGubaMarkTayloriDJamesFoucarAndrewBradleyAndrewSalingeriDSee all authors Luca BertagnaiDCorresponding Author• Submitting AuthorSandia National LaboratoriesiDhttps://orcid.org/0000-0002-6171-3202view email addressThe email was not providedcopy email addressOksana GubaSandia National Laboratoriesview email addressThe email was not providedcopy email addressMark TayloriDSandia National LaboratoriesiDhttps://orcid.org/0000-0002-9267-2554view email addressThe email was not providedcopy email addressJames FoucarSandia National Laboratoriesview email addressThe email was not providedcopy email addressAndrew BradleySandia National Laboratoriesview email addressThe email was not providedcopy email addressAndrew SalingeriDSandia National LaboratoriesiDhttps://orcid.org/0000-0003-4692-6813view email addressThe email was not providedcopy email address
We present an effort to port the nonhydrostatic atmosphere dynamical core of the Energy Exascale Earth System Model (E3SM) to efficiently run on a variety of architectures, including conventional CPU, many -core CPU, and GPU. We specifically target cloud -resolving resolutions of 3 km and 1 km, To express on -node parallelism we use the C++ library Kokkos, which allows us to achieve a performance portable code in a largely architecture -independent way. Our C++ implementation is at least as fast as the original Fortran implementation on IBM Power9 and Intel Knights Landing processors, proving that the code refactor did not compromise the efficiency on CPU architectures, On the other hand, when using the GPUs, our implementation is able to achieve 0,97 Simulated Years Per Day, running on the full Summit supercomputer. To the best of our knowledge, this is the most achieved to date by any global atmosphere dynamical core running at such resolutions.