
Fine-grained power estimation in multicore Systems on Chips (SoCs) is crucial for efficient thermal management. BPI (Blind Power Identification) is a recent approach that determines the power consumption of different cores and the thermal model of the chip using only thermal sensor measurements and total power consumption. BPI relies on steady-state thermal data along with a naive initialization in its Non-negative Matrix Factorization (NMF) process, which negatively impacts the power estimation accuracy of BPI. This paper proposes a two-fold approach to reduce these impacts on BPI. First, this paper introduces an innovative approach for NMF initializing, i.e., density-oriented spatial clustering to identify centroid data points of active cores as initial values. This enhances BPI accuracy by focusing on dense regions in the dataset and excluding outlier data points. Second, it proposes the utilization of steady-state temperature data points to enhance the power estimation accuracy by leveraging the physical relationship between temperature and power consumption. Our extensive simulations of real-world cases demonstrate that our approach enhances BPI accuracy in estimating the power per core with no performance cost. For instance, in a four-core processor, the proposed approach reduces the error rate by 76% compared to BPI and by 24% compared to the state of the art in the literature, namely, Blind Power Identification Steady State (BPISS). The results underline the potential of integrating advanced clustering techniques in thermal model identification, paving the way for more accurate and reliable thermal management in multicores and SoCs.
Self-consumption is developing worldwide to increase renewable electricity consumption and reduce electricity bills. It can be carried out individually or collectively (grouping several entities), but is generally restricted geographically. In datacenters, one may use load shifting to benefit from the self-consumption tariffs, at the cost of increasing energy consumption. An alternative would be to extend the self-consumption rules to wider perimeters. This paper proposes a comparative study on several aspects influencing the collective self-consumption (CSC) of an Edge infrastructure, including spatial load shifting, temporal load shifting and extending the current rules to encompass wider geographical boundaries. Spatial shifting under the current CSC scheme is found more cost-effective (3.9% of cost reduction) with a negligible increase in energy consumption (0.19%), compared to the revised definition of collective self-consumption which leads to 3.7% cost reduction and no increase in energy consumption. Moreover, allowing up to 10% of the user tasks to be shifted in time can further increase the self-consumption rate by 1.3%.
Systolic arrays are commonly used for running deep neural networks (DNNs) at the edge, where latency and energy efficiency requirements are stringent. Monolithic 3D (Mono3D) is an emerging 3D integration technology that offers ultra-high vertical interconnect density among processing and memory layers. The bandwidth benefits provided by Mono3D can help meet the growing latency and energy efficiency demands for DNNs. This paper presents a novel implementation for weight stationary (WS) dataflow in Mono3D systolic arrays, called WS-Mono3D. WS-Mono3D utilizes multiple resistive RAM layers and SRAM with high-density vertical interconnects to multicast inputs and performs high-bandwidth weight pre-loading while maintaining the same order of multiply-and-accumulate operations as in native WS dataflow. Consequently, WS-Mono3D eliminates input and weight forwarding cycles, and, thus, provides up to a 40% reduction in energy-delay-product (EDP) over the native WS implementation in 2D with iso-configuration. The paper also demonstrates the impact of temperature on energy efficiency benefits in WS-Mono3D.
Datacenters could greatly improve their sustainability by shifting their power usage across time in response to electricity’s carbon intensity. However, despite decades of research, modulating demand in response to grid signals has not been deployed in wide practice. We review diverse studies and real-world practices to understand the reasons for the gap between concept and reality, identifying several significant challenges. Demand response frameworks (i) are often too complex, (ii) break abstraction layers in datacenter management, (iii) place too much emphasis on batch processing jobs, (iv) lack dynamic strategies, and (v) provide insufficient incentives for datacenter operators and users. Overcoming these challenges is essential to making datacenter demand response a reality, which would lead to sustainable, efficient computing.
Datacenters have been a significant source of global carbon emissions. Researchers have developed tools to quantify the emissions of datacenters, aiming to enable environmentally sustainable datacenter development. However, these tools include only some of the hardware components; focus primarily on the design stage; and do not account for the inherent uncertainties of hardware and software characteristics. To this end, we propose an uncertainty-aware dynamic unified carbon modeling tool, dubbed U-DUCT, which, to the best of our knowledge, is the first attempt to comprehensively understand the datacenter carbon emissions by addressing these limitations. Specifically, U-DUCT considers both well-studied computer servers and the rarely included storage and switches, which have a huge impact on carbon emissions. Moreover, we shift the focus of U-DUCT to the run-time stage and investigate the effect of uncertainties on both embodied and operational carbon estimations. By using U-DUCT to account for uncertainty, experiments on a small-scale cluster running real-world application workloads reveal new potential for reducing carbon emissions from computer servers, storage racks, and switches.
Dynamic adaptive video streaming over HTTP (DASH), the de-facto standard in video streaming, requires significant CPU energy for transcoding. For carbon efficiency, it is essential to adhere to a low energy budget when using non-renewable energy sources. However, this can reduce the available bitrate versions, negatively impacting overall video quality. To tackle this trade-off, we propose a new deep reinforcement learning (DRL)-based scheme that limits energy consumption while enhancing video quality on transcoding servers. The scheme leverages a learning model that accounts for variable transcoding times and dynamic popularity changes, calculating the expected video quality, which is returned as a reward to the agent for each action when each bitrate version is transcoded. This allows the agent to decide on the transcoding of each bitrate version, ensuring the energy budget threshold is met while maximizing video quality. Experimental results show that the proposed scheme improves video quality between 2.7% and 18.3% (average, 10.9%) under various energy budgets.
This paper presents a comprehensive analysis of code generated by Large Language Models (LLMs) in terms of energy consumption and speed, moving beyond traditional accuracy metrics. We evaluate eight state-of-the-art models, including ChatGPT-3.5, ChatGPT-4o, CodeGemma:7b, WizardCoder:33b, Llama3:8b, Phind-CodeLlama:34b-v2, Nous-Hermes2:10.7b, and Mistral:7b, using RuntimeRatio and EnergyRatio metrics on LeetCode 1 questions. To evaluate LLMs’ understanding on energy efficient concepts, we propose a three-step prompting methodology to test models on coding with runtime and energy efficiency in mind. In an overwhelming majority of cases, LLMs perform worse than the human solution in both code runtime and energy consumption. In cases where models do perform well, it is often because they are aware of the question online, and produce responses similar or identical to famous solutions. Our comparative analysis of the efficient prompt and chain of thought prompt methods to the base prompt emphasize that models are inconsistent with producing quicker and more efficient code when directly asked to assemble energy efficient code. Our findings highlight that LLMs currently lack understanding in energy efficiency in terms of code generation, pushing a need for efforts to add green computing ideas and energy efficiency for future training and fine-tuning.
Data centers have been relying on renewable energy integration coupled with energy efficient specialized processing units and accelerators to increase sustainability. Unfortunately, the carbon generated from manufacturing these systems is becoming increasingly relevant due to these energy decarbonization and efficiency improvements. Furthermore, it is less clear how to mitigate this aspect of embodied carbon. As workloads continue to evolve over each hardware generation we explore the tradeoffs of fabricating new application-tuned hardware compared with more general solutions such as Field Programmable Gate Arrays (FPGAs). We also explore how REFRESH FPGAs can amortize embodied carbon investments from previous generations to meet the requirements of future generations workloads.
Bit-truncation has demonstrated great potential to enable run-time quality-power adaptive data storage and thus enhance the power efficiency of data-intensive applications such as videos and deep learning. However, existing bit-truncation memories are custom designed for a specific working condition of the target application. In this paper, we present a novel bit-truncation memory with full truncation flexibility, and it can truncate any number of bits for optimal tradeoff between quality requirements of applications and power savings. Our experiments show that the proposed memory can support three different video applications (including luminance-aware, content-aware, and region-of-interest-aware) with enhanced power efficiency (up to 50.03% power savings) as compared to state-of-the art solutions. Also, the proposed memory achieves up to 66.56% and 63.29% power savings for baseline and lightweight deep learning models respectively, with a low implementation cost (2.42%).
Handling heavy computing tasks in real-time systems requires substantial power, leading to significant heat generation. This can cause serious thermal problems like rising temperatures, high thermal gradients, and hot spots. These issues can reduce chip performance, accelerate device aging, and cause premature failure. Thermal-Aware Scheduling (TAS) optimizes heat management to maintain a safe temperature. In this work, we evaluate RT-TAS, an advanced TAS algorithm using a steady-state thermal model, and compare it with our proposed dynamic physics-informed POD-based TAS algorithm. POD-TAS utilizes a reduced order thermal model with a dynamic predictive approach, allowing precise temperature control during task scheduling. This reduces peak temperatures and thermal variations. Our tests on a multi-core processor demonstrate that POD-TAS significantly improves thermal performance. It reduces the average spatial thermal variance by 6.18%, the variance of the maximum CPU temperature over time by 59.68%, and the variance of the average temperature by 13.05% due to its ability to maintain the chip temperature within a set threshold.
This paper proposes SRC, a novel framework for efficient and reliable inference on battery-free smart Internet of Things (IoT) devices. SRC supports various configurations that follow reactive configuration while using the innovative state machine and a safe threshold mechanism to proactively halt operations, reducing store/load operations by up to 75%. It strategically stores essential convolutional neural network (CNN) data (layer, kernel, etc.) to optimize input/output feature map management. This reactive design allows seamless task resumption across power cycles, ensuring continuity in unpredictable energy environments. Experiments show significant gains, with SRC achieving on average similar to 81.85% reduction in read/write operations and approximately 57.18% improvement in sensing compared to conventional reactive methods based on the intermittent Energy Trace 1.
Education technology (EdTech) is an important tool for streamlining and improving course administration and teaching. Many modern EdTech tools rely on cloud services to host containerized applications. While this is convenient, it is also costly in terms of both dollars and carbon emissions.We propose the alternative approach of hosting containerized EdTech applications on local clusters of upcycled Android devices. We perform an evaluation of the Google Pixel Fold for handling educational workloads. Our findings suggest that such repurposed device could effectively bridge the gap between mobile and traditional computing platforms in education, open new avenues for accessible educational computing environments.
Cloud platforms' rapid growth raises significant concerns about their electricity consumption and resulting carbon emissions. Power capping is a known technique for limiting the power consumption of data centers where workloads are hosted. Today's data center computer clusters co-locate latency-sensitive web and throughput-oriented batch workloads. When power capping is necessary, throttling only the batch tasks without restricting latency-sensitive web workloads is ideal because guaranteeing low response time for latency-sensitive workloads is a must due to Service-Level Objectives (SLOs) requirements. This paper proposes PADS, a hardware-agnostic workload-aware power capping system. Due to not relying on any hardware mechanism such as RAPL and DVFS, it can keep the power consumption of clusters equipped with heterogeneous architectures such as x86 and ARM below the enforced power limit while minimizing the impact on latency-sensitive tasks. It uses an application-performance model of both latency-sensitive and batch workloads to ensure power safety with controllable performance. Our power capping technique uses diagonal scaling and relies on using the control group feature of the Linux kernel. Our results indicate that PADS is highly effective in reducing power while respecting the tail latency requirement of the latency-sensitive workload. Furthermore, compared to state-of-the-art solutions, PADS demonstrates lower P95 latency, accompanied by a 90% higher effectiveness in respecting power limits.
In recent years, Machine learning (ML) methods have emerged as a promising approach for mitigating the challenges with software power measurement and estimation. While ML-driven methods have been effective in advancing the state of the art in energy-efficient computing, their predictive abilities are often limited by the quality and the size of the training datasets. This paper presents a diffusion-based generative method for synthesizing software energy data to create larger representative datasets for power modeling and related tasks. We demonstrate the effectiveness of our methods in the context of energy-efficient graph analytics. We use an augmented dataset to train a regression model and a classifier to predict the GPU power consumption of graph applications running on different classes of input graphs. Experimental evaluation shows the accuracy of the predictive models can be improved by as much as 11% when trained on the augmented dataset than over the original dataset.
Large language models (LLMs) have been widely used for their ability to handle complex natural language tasks with high accuracy. The lifecycle of LLMs development and deployment encompasses both training and serving phases. Although training takes months and consumes significant amounts of energy, recent studies show that the energy consumption of LLM serving has now surpassed that of training, leading to significant environmental impacts, especially in terms of carbon footprints. While much prior work has focused on improving LLM performance, the specific challenge of reducing the carbon footprint of LLM serving has been largely overlooked. This paper identifies key challenges and outlines research directions for making LLM serving more sustainable, aiming to inspire further environmentally responsible advancements in the field.
Power consumption of data centers is rapidly becoming more prominent as the demand for computation increases. Next-generation systems are expected to require significantly more power, making it essential to design them to operate under power constraints to achieve sustainability goals. Power utilities are responsible for constantly providing reliable power and this task becomes harder as a higher amount of renewables are integrated into the grid. Demand response (DR) programs are promising solutions to maintain grid reliability by exploiting the flexibility of power consumers, such as data centers.While prior research explores data center participation in DR, real-world examples are still limited due to the risks of operating under power constraints, such as quality-of-service (QoS) violations. To provide greater flexibility in power consumption while improving data centers’ ability to meet QoS targets, we propose Conductor, a novel framework that coordinates the participation of multiple data centers in DR, increasing their resilience to operate under power constraints without requiring any inter-data-center workload migration. Conductor assigns dynamic power targets to data centers based on their real-time QoS information and mitigates the risks of joining DR programs by recovering the QoS violations of jobs while achieving up to 78% better tracking of the power targets compared to individual data center DR participation.
Although highly energy efficient, adiabatic and reversible systems suffer from performance drawbacks inherent to the physical operations that make them so efficient. Superscalar processors provide high performance through out-of-order speculative work of which an effective branch predictor is a key component in those performance gains. In the context of reversibility, a branch predictor is a design focal point because any fully reversible system must also be able to predict branch outcomes when in reverse mode. Taking advantage of Temporal Streaming techniques, this paper introduces several reversible branch predictor implementations which enable reversible and out-of-order instruction execution. These first-of-their-kind designs allow for a superscalar architecture that would maintain both a high level of performance and a high level of energy efficiency with the ability to un-compute obsolete data stored in memory. Testing our designs using the SimpleScalar out-of-order simulator, we estimate possible additional savings of 24 fJ per MB of data recovered at room temperature and at reverse prediction rates 2.27% higher than the forward. This work opens new avenues for designing and developing what we are calling Fully Adiabatic, Reversible, and Superscalar (FARS) Processor Architectures and is the first of many adaptations of conventional superscalar components to a reversible system.
As the global population continues to urbanize, city residents are increasingly exposed to extreme heat due to the urban heat island phenomenon, which poses a serious threat to human health. Understanding heat exposure in urban areas is challenging due to the heterogeneity of urban form and land uses, which create micro-climates that expose individuals to a wide variation of temperatures in outdoor environments. To address this issue, we have deployed a sensor network throughout the city of Bloomington, Indiana and on the campus of Indiana University. The environmental sensor network includes air temperature, relative humidity, and soil moisture sensors in different urban-forms such as along streets, among densely clustered trees, in parking lots, and in community gardens. The sensor network captures a rich set of data related to local climate, heat exposure, and vegetative heat stress. Local environmental monitoring is an important area of research that enables researchers to more precisely predict and quantify an individuals’ exposure to extreme heat in urban environments. In this paper, we describe the application of our environmental sensor network, the Healthy Cities Sensor Network (HCSN) and how it can be utilized to increase climate resilience for local communities.
The growth of computing continues to place power demands on an ever-stressed energy grid. Embedded systems are a part of that demand. Focusing on the architecture of the System-On-Chip (SoC) is an area of interest for power optimization. Intellectual Property (IP) blocks that provide support functionality for SoC designs are often licensed from vendors who provide little or no visibility into their inner workings, requiring reliance on vendor-supplied metrics. This work proposes TurbOS, a small footprint open-source operating system (OS) for reducing required resources within IP blocks. TurbOS utilizes the Turbo9 microprocessor, and is just 20% the size of the popular FreeRTOS operating system, while providing the same functionality. For certain applications, the number of IP blocks can be reduced without loss of functionality.
The introduction of tensor core (TC) in NVIDIA GPUs has accelerated neural network computations. A TC is a set of storage and arithmetic units dedicated to handling matrix-multiply-and-accumulate (MMA) operations. While TC reduces the runtime of convolutional neural networks (CNNs), it increases power consumption. In particular, register files within TCs consume a significant fraction of leakage power as GPU designers steadily have increased the size of register files to boost performance. In this work, we propose a value-based approach to reduce leakage power in register files. We observe that register bits exhibit a strong bias towards zeros. We exploit this bit-level bias property and propose a low-power SRAM (LPS) cell that draws significantly less leakage current than regular SRAM cells. In the preferred state, the leakage power is smaller by as much as approximately 49x. We also propose LPS+ which increases the sparsity rate in SRAM cells by selectively inverting register bits. To reduce leakage power further, we propose precision-aware LPS+ (PLSP+) which exploits the error resiliency property of CNNs and drops low-order bits in network values. We evaluate our proposed techniques using state-of-the-art CNNs and show that leakage power is reduced by 77.3% with a negligible impact on accuracy.