As supercomputers continue to grow in size and power consumption, the ability to understand application run time performance and energy/power usage is of growing importance. In the future it may be necessary for systems to operate under power caps imposed by facilities due to external influences such as renewable power generation (e.g. solar) impacting energy availability at an affordable cost during certain times of day and more frequent natural disasters due to climate change causing power transmission disruptions. However, current energy consumption and power characterization metrics like energy-delay products do not express all of the characteristics of an application that are relevant to understanding the impact of power capping on performance. In this study, we (1) characterize the useful features of a metric that effectively captures time-to-solution and power performance dynamics; and (2) design and evaluate a novel set of application power efficiency (APE) metrics that accounts for time-to-completion, power utilization and power variability. Power utilization quantifies a system's trapped capacity (unused power budget) and power variability quantifies an application's variation of power consumption over time. We demonstrate how APE can be used to quantify application characteristics which, in turn, enables reasoning about the impacts of power capping on an application.
Power will be a first-class operating constraint for Exascale computing. In order to manage power consumption of systems, measurement and control methods need to be developed. While several approaches have been developed by hardware manufacturers, they are vendor-specific and in some cases implementation-specific interfaces. Integrating all of the individual device level measurement and control functionality in a single system is a difficult task that requires system specific code. Sandia National Laboratories, in collaboration with many industry and academic partners, has developed a Power API specification, consisting of a broad range of interfaces spanning from low-level hardware to platform management and accounting. In order for many of the interfaces to be useful, especially at large scale, measurement data must be collected and control directives must be distributed in a scalable manner. This paper details the challenges of providing large scale power measurement and control and the scalable collection and control distribution architecture that is being integrated into the Power API reference implementation.
The motivation for power and energy measurement and control capabilities for High Performance Computing (HPC) systems is now well accepted by the community. While technology providers have begun to deliver some capabilities in this area, interfaces to expose these features are vendor specific. The need for a standard way to leverage these emerging capabilities, now and in the future is clear. To address this need, the Department of Energy funded an effort to produce a Power application programming interface (API) specification for High Performance Computing systems with the goal of contributing this API to the community as a proposed standard for power measurement and control. In addition to the open publication of this standard an Advanced Power Management Non-recurring Engineering project has been initiated with Cray Inc. with the intention of advancing capabilities in this area and delivering them on a leadership class platform. We will detail the collaboration established between the Alliance for Computing at Extreme Scale (Sandia Laboratories and Los Alamos Laboratory) and Cray and the portions of the Power API that have been selected for the first production implementation of the standard. Keywords-power monitoring; power control; energy efficiency; power measurement;
Power API-the result of collaboration among national laboratories, universities, and major vendors-provides a range of standardized power management functions, from application-level control and measurement to facility-level accounting, including real-time and historical statistics gathering. Support is already available for Intel and AMD CPUs and standalone measurement devices.
Operations management of the New Mexico Alliance for Computing at Extreme Scale (ACES) (a collaboration between Los Alamos National Laboratory and Sandia National Laboratories) Trinity platform will rely on data from a variety of sources including System Environment Data Collections (SEDC); node level information, such as high speed network (HSN) performance counters and high fidelity energy measurements; scheduler/resource manager; and plant environmental facilities. The water-cooled Cray XC platform requires a cohesive way to manage both the facility infrastructure and the platform due to several critical dependencies. We present preliminary results from analysis of integrated data on the Trinity Application Readiness Testbed (ART) systems as they pertain to enabling advanced operational analysis through the understanding of operational behaviors, relationships, and outliers. Keywords-High Performance Computing; Monitoring
As high-performance computing systems continue to grow in size and complexity, energy efficiency and reliability have emerged as first-order concerns. Researchers have shown that data movement is a significant contributing factor to power consumption on these systems. Additionally, rollback/recovery protocols like checkpoint/restart can generate large volumes of data traffic exacerbating the energy and power concerns. In this work, we show that a coarse-grained model can be used effectively to speculate about the energy footprints of rollback/recovery protocols. Using our validated model, we evaluate the energy footprint of checkpoint compression, a method that incurs higher computational demand to reduce data volumes and data traffic. Specifically, we show that while checkpoint compression leads to more frequent checkpoints (as per the optimal checkpoint frequency) and increases per checkpoint energy cost, compression still yields a decrease in total application energy consumption due to the overall runtime decrease.
Accuracy of component based power measuring devices forms a necessary basis for research in the area of power-efficient and power-aware computing. The accuracy of these devices must be quantified within a reasonable tolerance. This study focuses on PowerInsight, an out- of-band embedded measuring device which takes readings of power rails on compute nodes within a HPC system in realtime. We quantify how well the device performs in comparison to a digital oscilloscope as well as PowerMon2. We show that the accuracy is within a 6% deviation on measurements under reasonable load.
The challenge of balancing between power and performance is now well established. While research in this area is well underway, the ability to measure power and energy in situ has remained an obstacle. This problem is magnified in the field of High Performance Computing (HPC). To meet this challenge, a device called PowerInsight has been designed to accomplish component level power and energy instrumentation of commodity hardware. PowerInsight was designed by Penguin Computing, in close cooperation with Sandia National Laboratories, to further power and energy research in HPC and other areas. This paper documents the design and development of PowerInsight, hardware and software. Validation of the functionality of PowerInsight was done during design and development as well as experimentally after integrating the first PowerInsight devices into a commodity cluster. This paper only begins to show the wide range of impact this level of power and energy instrumentation can have on a range of architectural and application research and analysis topics.