
With the emergence of both power and performance as primary design constraints, energy efficiency has become the new design criteria. A platform with heterogeneous-ISA processors can provide multiple power-performance execution points needed for a varied mix of workloads. We argue that a new system software architecture is needed to obtain maximum energy efficiency on such heterogeneous-ISA platforms. We present our system software, a replicated-kernel operating system and a compiler framework, and quantify the advantages of such a system software on ARM-x86 using simulations. Based on our experimental observations, we propose a scheduling approach which considers system and application runtime characteristics along with platform profiles to maximize energy efficiency.
Datacenters demand big memory servers for big data. For blade servers, which disaggregate memory across multiple blades, we derive technology and architectural models to estimate communication delay and energy. These models permit new case studies in refusal scheduling to mitigate NUMA and improve the energy efficiency of data movement. Preliminary results show that our model helps researchers coordinate NUMA mitigation and queueing dynamics. We find that judiciously permitting NUMA reduces queueing time, benefiting throughput, latency and energy efficiency for datacenter workloads like Spark. These findings highlight blade servers' strengths and opportunities when building distributed shared memory machines for data analytics.
In this paper we aim to optimize power consumption of mobile games without compromising user experience. We study the behavior of 40 mobile games on a smartphone and identify two power-inefficient issues: 1) fixed high frame rate that consumes a high power but brings negligible extra benefits to user experience when the screen content does not change rapidly or stays nearly static, and 2) high overdraw rate---the same pixels are drawn for multiple times and thus wastes energy. We report the measurement results of our study and explore possible solutions to mitigate these two issues. In particular, for the first issue, we have implemented a prototype to enable dynamic frame rate scaling that is able to reduce the frame rate to save power based on how fast the game content changes. A lower frame rate is used when the game content does not change fast and thus user-perceived experience is retained. Preliminary experimental results show that our approach is promising.
Legacy archival workloads have a typical write-once-read-never pattern, which fits well for tape based archival systems. With the emergence of newer applications like Facebook, Yahoo! Flickr, Apple iTunes, demand for a new class of archives has risen, where archived data continues to get accessed, albeit at lesser frequency and relaxed latency requirements. We call these types of archival storage systems as active archives. However, keeping archived data on always spinning storage media to fulfill occasional read requests is not practical due to significant power costs. Using spin-down disks, having better latency characteristics as compared to tapes, for active archives can save significant power. In this paper, we present a two-tier architecture for active archives comprising of online and offline disks, and provide an access-aware intelligent data layout mechanism to bring power efficiency. We validate the proposed mechanism with real-world archival traces. Our results indicate that the proposed clustering and optimized data layout algorithms save upto 78% power over random placement.
Recently, single-ISA heterogemeous multi-core processors (SI-HMP) draw attention, pursuing optimal power-performance scaling. Leveraging differently optimized heterogeneous cores, SI-HMP can dynamically tune performance with minimal additional power consumption, or it can find maximum performance core combination with respect to a given power budget. However, the little-to-big, or big-to-little core switching has hidden costs. To properly scale up/down the power-performance, we should carefully analyze the actual performance gain, considering the multi-core processing model and inter-cluster communication. This paper reveals that there are some good and bad cases for core switching, and presents a possible way to achieve good power-performance scaling through big-little switching.
Semiconductor device engineers are hard-pressed to relate observed device-level properties of potential CMOS replacements to computation performance. We address this challenge by developing a model linking device properties to algorithm parallelism, total computational work, and degree of voltage and frequency scaling. We then use the model to provide insight into how device properties influence execution time, average power dissipation, and overall energy usage of parallel algorithms executing in the presence of hardware concurrency. The model facilitates studying tradeoffs: It lets researchers formulate joint energy-delay metrics that account for device properties. We support our analysis with data from a dozen large digital circuit designs, and we validate the models we present using performance and power measurements of a parallel algorithm executing on a state-of-the-art low-power multicore processor.
Battery State of Charge (SOC) estimation is a fundamental component of today's smartphones that affects the internal processes and observable behavior of the devices. This article systematically investigates and analyzes the SOC estimation techniques in smartphones. First, we discover that the voltage curve of a given smartphone implicitly captures the usable capacity of the battery while charging the mobile device. Second, we observe that today's SOC estimation techniques do not model battery capacity loss sufficiently to accurately capture the usable capacity. Finally, we report findings based on battery analytics of 2077 devices that validate the relationship between battery voltage and the usable capacity of a device. The presented results enable the development of more accurate battery gauges and metering solutions thus resulting in better power-saving decisions, recommendations for the users, and most importantly more reliable system.
In this paper we introduce ViRUS: Virtual function Replacement Under Stress. ViRUS allows the runtime system to switch between blocks of code that perform equivalent functionality at different Quality-of-Service levels when the system is under stress -- be it in the form of scarce energy resources, temperature emergencies, or various sources of environmental and process variability -- with the ultimate goal of energy efficiency. We demonstrate ViRUS with a framework for transparent function replacement in shared libraries and a polymorphic version of the standard C math library in Linux. Case studies show how ViRUS can tradeoff upwards of 4% degradation in application quality for a band of upwards of 50% savings in energy consumption.
Mobile apps connect with counterpart cloud services and pursue data transfer over wireless networks. The amount of mobile data transferred will only increase as app data needs and usage expand. Energy expenditure during periods of data transfer constitutes a significant portion of a mobile's battery usage. At the same time, due to the evolving nature of wireless technologies, there is a proliferation of networks (such as public Wifi hotspots, access points, and cellular technologies) that exhibit diverse energy and performance characteristics. As a user moves around (even within the same logical network) the hardware serving data transfer varies, often offering opportunities for data transfer with distinct bandwidth and latency capabilities. Apps and data services in smartphones however are oblivious to this diversity and the energy impact of using one opportunity vs. others. In this paper, we first present a study of Wifi characteristics in two domains, shopping malls and within an enterprise campus, demonstrating the fine-grained diversity of network opportunities in two commonplace scenarios. Next, we describe a system-level framework and set of interfaces that enable energy efficient app syncs by leveraging the right opportunities. Preliminary results using our approach show up to three times lower energy costs for popular mobile apps using media uploads.
Prior battery-aware systems research has focused on discharge power management in order to maximize the usable battery lifetime of a device. In order to achieve the vision of perpetual mobile device operation, we propose that software also needs to carefully consider the process of battery charging. This is because the power consumed by the system when plugged in can influence the rate of battery charging, and hence, the availability of the system to the user. We characterize the charging process of a Nexus 4 smartphone and analyze the charging behaviors of anonymous Nexus 4 users using the Device Analyzer dataset. We find that there is potential for software schedulers to increase device availability by distributing tasks across the charging period. We estimate that approximately 53% of the users we examined could benefit from up to 18.9% improvement in net energy gained by the battery while charging. Accordingly, we propose new threads of research in charging-aware power management and deferrable task scheduling that could improve overall availability for a significant portion of smartphone users.
Prior work has shown the benefits of Energy Storage Devices (ESDs), such as batteries, to smoothen/flatten power draws in Datacenters, for reducing demand during peak tariffs (for op-ex savings) and under-provisioning the power infrastructure (for cap-ex savings). Until now, all prior studies for such smoothening, referred to as Demand Response, have considered re-purposing existing UPS unit batteries for demand response. It is not clear if such dual usage - handling power outages and demand response - is the most effective option since the needs (energy and/or power), mandates (best effort vs. hard stipulations), costs, availability and health degradation considerations could be very different. In this paper, we study the design space of choices for provisioning ESDs for these dual purposes - separate ESDs for each purpose, common pool of ESDs for both purposes, and softreservations in this pool with possible re-purposing dynamically based on demand. Our evaluations show that: (i) provisioning lead-acid batteries for a peak "power" load needed to handle power outages already comes with sufficient energy capacity that is more than adequate to automatically supply the energy needs for demand response; (ii) this makes it economically attractive to use the same UPS batteries, originally intended for Power Outages, for Demand Response as well, despite any consequent health degradation (due to repeated discharges); (iii) the ability to handle the needs during a power outage is not compromised despite the dual-purposing of these UPS batteries; and (iv) the non-orthogonality of the power and energy capacities of these batteries (i.e. provisioning for the high power needs during an outage automatically comes with a lot of energy capacity) suggests the possibility of having different Energy Storage Technologies for the two purposes and we show that a heterogeneous/hybrid option using Ultra-capacitors or Flywheels for Power Backup and batteries for Demand Response is a more cost-effective option.
The emerging industry trend of ever-increasing display density on mobile devices has dramatically increased workload placed on a mobile GPU's. Because mobile GPU power consumption increases almost linearly with workload, increasing the display density directly decreases battery life of a device. While this tradeoff is acceptable if user experience is improved, display densities beyond that which the human eye can perceive would result in decreased device battery life for no perceptible gain. Further, the workload imposed by such high density displays may invalidate the previous requirement that the interface always run at high frame rates. In this paper, we show that the display on some modern devices already exceeds human perceptive capabilities, a feature which can be exploited at runtime in order to reduce GPU power consumption. Our proposed method includes both resolution and refresh rate scaling, which are shown to reduce mobile GPU power consumption by up to 33% and 38%, respectively. We conclude with a discussion of how such a system may be implemented on existing devices.
To curtail data centers' huge cooling power consumption and water demand (for cooling), air-side economizer has been increasingly adopted to cool down servers. Recently, another sustainability practice, rainwater harvesting, has also seen a growing adoption in data centers, potentially leading to water self-sufficiency without connecting to water utilities to supply cooling water. Nonetheless, various factors, e.g., unpredictable rain falls and limitations on water harvesting area, make water self-sufficiency challenging. In this paper, we present a first-of-its-kind study to evaluate whether it is feasible to achieve water self-sufficiency in data centers. We find that although water self-sufficiency depends on noncontrollable factors such as weather; improving power proportionality (via power management) and increasing water tank size will increase the feasibility, relieving the requirement on water harvesting area and making water self-sufficiency a reality in different locations.
Power and energy efficiency are major concerns in future supercomputing systems. We expect that applications will be constrained to operate under a power budget and achieving the expected levels of performance will be challenging. Understanding how power is consumed by an application throughout its different phases will be necessary to shift power to those resources on the critical path. In this paper, we identify opportunities for shifting power between components for a representative kernel of explicit hydrodynamics codes. Based on a linear regression model, we dynamically throttle the memory system in regions with low memory bandwidth requirements on an energy-efficient supercomputer. Our results show that we can save a significant amount of power that could be used on resources on the critical path and, thus, maximize performance under the operating power budget.
The recent availability of modern big. LITTLE ARM chips featuring heterogeneous processor cores has enabled a practical investigation of the heterogeneous chip concept and its implications for performance and energy. The ODROID XU+E platform produced by Hardkernel integrates hardware power monitors for the individual chip clusters, GPU and Memory. In this paper we investigate the detailed energy characteristics of the hardware components and applications to highlight the platform's potential for energy-aware resource management of the platform. We measure network, storage and CPU energy consumption as well as parallel application performance. We find that the platform offers interesting sweet-spots for fine tuning of resource scheduling, but also poses new challenges, for example for quantifying the cost of switching between clusters.
Despite that OLED screen has been increasingly adopted in smartphones to save power; screen is still one of the most energy-consuming modules in smartphones. Techniques such as local dimming are proposed to further reduce the power consumption of OLED screen, but it is hard to decide which part of the screen could be dimmed, and it often results in compromised user experience. Intuitively, when a user interacts with a smartphone via the touch screen, the screen areas are covered by the user's fingers and even some of the neighboring areas could be safely dimmed. Thus, in this paper, we propose FingerShadow, a new technique which does local dimming for the screen areas covered by user fingers to save more power, without compromising the user visual experience. We have studied 10 users' touch interaction behaviors and found that on average 11.14% of the screen were covered by fingers. For these 10 users, we estimate that FingerShadow can achieve 5.07%-22.32% power saving, averaging 12.96%, with negligible overhead. We discuss the challenges and future research work to implement Finger-Shadow in existing smartphone systems.
Nowadays, GPUs are widely used to accelerate many high performance computing applications. Energy conservation of such computing systems has become an important research topic. Dynamic voltage/frequency scaling (DVFS) is proved to be an appealing method for saving energy for traditional computing centers. However, there is still a lack of firsthand study on the effectiveness of GPU DVFS. This paper presents a thorough measurement study that aims to explore how GPU DVFS affects the system energy consumption. We conduct experiments on a real GPU platform with 37 benchmark applications. Our results show that GPU voltage/frequency scaling is an effective approach to conserving energy. For example, by scaling down the GPU core voltage and frequency, we have achieved an average of 19.28% energy reduction compared with the default setting, while giving up no more than 4% of performance. For all tested GPU applications, core voltage scaling is significantly effective to reduce system energy consumption. Meanwhile the effects of scaling core frequency and memory frequency depend on the characteristics of GPU applications.
Recently, there has been a surge of interests on developing techniques and architectures for prefetching ads to potentially reduce the smartphone energy drain by 3G/4G radios from fetching ads. Despite the development of prefetching techniques, it remains unclear (1) how much smartphone energy do ads consume in popular apps in dominant app markets, and (2) out of which, what portion can we realistically save from prefetching? We present a measurement study of the energy drain of top 100 free apps in Google Play, totaling more than 2.2 B downloads, to re-examine the above two motivational questions for ads energy research. We found the upper bound energy savings from prefetching ads is low: out of the top 100 apps, only 57 apps display ads, which incur on average 3.2% total energy on ads 3G tails. We further show the already-low upper bound ads energy saving is hard to achieve by ads prefetching as different apps exhibit very different ads behavior.
In order to minimize the container server power consumption, a new cooling system that incorporates fan-less servers and freshair cooling is proposed. In a conventional container data center, the required air flow for sever cooling is supplied by both server built-in fans and container facility fans. Therefore, this work has been carried out on fan-less servers to reduce power consumption. Although fanless servers are expected to reduce power consumption, facility fans have to provide excessive air to secure a safe operation of servers. In order to achieve optimized air-flow from facility fans to cool fan-less servers, a power saving control system incorporating the IT system and cooling facilities is proposed. Here, facility fans are controlled based on server information such as CPU temperature, rack position and so on. Through this study, we suggest that the minimum point in total power consumption of the container server with no performance penalty existed by the trade-off relationship between the power consumption changes of servers and of facility fans with CPU temperature. This enables us to operate the server system with minimized power consumption depending on the air temperature. To verify the energy-saving effect of this technology, a prototype container server with the proposed system was constructed. As a result, 22.8% energy saving was achieved with this new system, compared with the conventional container servers with built-in fans.
More than 90% of consumer computers use integrated graphics processors. In these processors, the CPU and the GPU share the same physical memory. Due to high density, good power efficiency, and low cost, integrated graphics processors are promising candidates for next-generation micro-servers and, hence, data-center workloads. While discrete graphics processors have been extensively studied, there is very little work on characterizing integrated GPUs. This paper is a step towards understanding the power and performance of integrated GPUs. Our results reveal many architectural caveats that programmers need to be aware of to exploit integrated GPUs: memory contention between the CPU and GPU, workload dependent energy efficiency, and data transfer tradeoffs.