We present Avatar, a data center environmental advisory system for raised floor data centers. Using limited information such as inlet and reference temperatures for IT equipment and basic floor plan geometry, Avatar produces recommendation to adjust the operation of computer room air conditioners (CRACs) and the configuration of vent tiles in a data center so as to reducing excess provisioning of cooling and to remove hot spots. Avatar reduces operating expenses by cooling the same load with less energy. Avatar reduces capital expenses by recovering stranded cooling capacity that would otherwise have to be replaced.
In this paper, we present the Daffy data model and messaging framework for data centers. The model is a hybrid model, combining physical, structural, geometrical, and logical modeling techniques. The messaging scheme has excellent performance and allows for the loose coupling of the various framework components. The framework bridges the gap between facilities and IT domains and enables the holistic, cross-domain management of data centers. The framework supports rich visualization, cross-domain queries and sophisticated cross-domain autonomic control systems.
In recent years, climate change, depletion of conventional energy sources and rising energy costs have led to an increased focus on sustainability. Within the Information Technology (IT) sector, data centers are significant energy consumers. The first steps towards reducing power consumption in data centers are to monitor it and to determine the heavy hitters.Unfortunately, fine-grain power information is often not readily available within data center environments. In this paper, we conduct an exploratory analysis of aggregate power data in a data center. We collect data from the power infrastructure of a data center in Palo Alto, CA, as well as from a data center in Bangalore, India. We examine the data in increasing detail, and reveal the opportunities and challenges for disaggregating data center power consumption data.
Efficient and reliable operation of today's data centers, which host IT equipment with ever-increasing power density, relies heavily on the cooling system to meet the thermal management needs of the IT equipment with minimal environmental footprint. The dynamic IT workload, together with the spatial variance of cooling efficiencies, creates both temporal and spatial non-uniformities within the data centers. Most data centers use zonal cooling actuators, such as computer room air conditioners (CRAC), to alleviate the local "hot spots". Without proper localized cooling actuation mechanisms, the cooling capacity is usually over-provisioned that leads to waste of energy. To address this problem, we introduce adaptive vent tiles (AVT) for local cooling adjustment, and develop a holistic multivariable model based on the mass and energy balance principles to capture the effects of both zonal and local cooling actuation on the inlet temperatures of the racks that host the IT equipment. A model predictive controller is then proposed to minimize the total cooling power while meeting the thermal requirements of the racks. The zonal and local cooling actuation is coordinated in such a unified framework for the provisioning, transport and distribution of the cooling resources in the data centers. The proposed holistic cooling approach is validated in a production data center. Experimental results indicate that up to 36% of CRAC units blower power can be saved, compared with the state of the art control solution.
In this paper, we describe an integrated design and management approach to creating a sustainable IT ecosystem: a physical infrastructure where information technology has been seamlessly interwoven to improve environmental efficiency while achieving lower cost. Specifically, we describe five principles to achieve such integration: ecosystem-scale life-cycle design; scalable and configurable resource microgrids; pervasive sensing; knowledge discovery and visualization; and autonomous control. Application of the approach is demonstrated for the case study of an urban water infrastructure, and we find that the proposed approach could potentially enable reduction of life-cycle energy use by over 15%.
Data centers contain IT, power and cooling infrastructures, each of which is typically managed independently. In this paper, we propose a holistic approach that couples the management of IT, power and cooling infrastructures to improve the efficiency of data center operations. Our approach considers application performance management, dynamic workload migration/consolidation, and power and cooling control to “right-provision” computing, power and cooling resources for a given workload. We have implemented a prototype of this for virtualized environments and conducted experiments in a production data center. Our experimental results demonstrate that the integrated solution is practical and can reduce energy consumption of servers by 35% and cooling by 15%, without degrading application performance.
Local airflow distribution in data center environments has historically been accomplished through ventilation tiles distributed over a raised floor air distribution plenum. The tiles are initially configured upon the commissioning of the facility and, as IT equipment configuration changes with time, the tiles are adjusted accordingly. However, tile adjustment is a manual process that is error-prone and often non-intuitive. Tile flow rates are a strong function of under floor plenum pressure distribution which is subject to change as tile layouts are reconfigured. Thermal models are often developed to assist with layout changes, but these models can be time-consuming to generate and require skilled users to achieve accurate results. This paper presents an adaptive vent tile (AVT) for use in raised floor data centers that can adapt to the needs of nearby IT equipment. We present a multi-input-multi-output (MIMO) AVT controller that automatically and dynamically adjusts a multiplicity of AVT openings in coordination such that thermal management requirements are met with minimum use of airflow. We describe the development of dynamic models and algorithm design of the MIMO controller. The controller was evaluated with a set of AVT units in a production data center environment. Results show that the controller can optimize local airflow distribution, provide fine-grained rack intake temperature control and respond to disturbances in a manner that is not achievable through static distribution of tiles.
In data centers with raised floor architecture, the floor tiles are typically perforated, delivering the cold air from the plenum to the inlets of equipment located in racks. The environment of these data centers is dynamic in that the workload and power dissipation fluctuate considerably over both short-term and long-term time scales. As such, airflow requirements vary continuously. However, due to labor costs and lack of expertise, the tiles are adjusted infrequently, and many data centers are grossly over provisioned for airflow in general and/or lack sufficient airflow delivery in certain local areas. This wastes energy and reduces data center thermal capacity. We have previously introduced Kratos, an Adaptive Vent Tile (AVT) technology that addresses this problem by automatically adjusting mechanical louvers mounted to the tiles in response to the needs of nearby IT equipment. Our initial results were limited to a 3-tile test bed that allowed us to prove concept but did not provide for scalability. This paper extends the previous work by expanding the size of the test bed to 28 tiles and 29 racks located in multiple thermal zones. We present experimental modeling results on the MIMO (Multi-Input Multi-Output) system and provide insights on the external behavior of the system through CFD (Computational Fluid Dynamic) analysis. We develop an MPC (Model-based Predictive Control) controller to maintain the temperatures of racks below the thresholds through vent tile tuning. Experimental results show that the controller can maintain the temperature below the thresholds while reducing overall cooling air requirements.
Energy consumption and energy prices are two critical concerns for data center operators. In 2005, servers alone accounted for 1.2% of energy consumption in the U.S. [1]. By 2012, the EPA predicts energy consumption in data centers will double from 2007 levels. Energy prices are also increasing. In 2006, they rose by 10% [1]. Energy prices are expected to rise even higher, due to regulatory and social concerns over green house gas emissions from energy production.In this work, we demonstrate how data center energy consumption can be reduced and thermal capacity improved by systematically considering the interactions between data center infrastructure elements (e.g., cooling and servers). Specifically, we quantify the benefits of the automated migration of workloads running in Virtual Machines (VMs). VMs are migrated from a set of servers located in an inefficient thermal zone to servers in a more efficient zone. Our experimental results show that our technique reduced data center energy usage by 30% and improved available thermal capacity by 22%."
To accommodate the dynamic environment within raised floor data centers, cooling capacity is tuned during operation through zonal control means, e.g., active management of air conditioning resources. However, due to the spatial variance of cooling efficiency and time-varying cooling demand within zones, zonal adjustments alone are not able to maximize the thermal capacity of data centers. Without making local adjustments to the physical structure, such as altering vent tile openings, a data center can suffer significant reduction in thermal capacity and cooling efficiency, and such that facility lifespan. In this paper, we present active cooling technologies using both local and zonal actuators that improve overall cooling efficiency. Experimental evaluation in a data center shows that the integrated controller can adapt to changes to the system under control, significantly improve the controllability of the temperatures and reduce the energy consumption of the cooling facility.
The environmental impact of data centers is significant and is growing rapidly. However, there are many opportunities for greater efficiency through integrated design and management of data center components. To that end, we propose a sustainable data center that replaces conventional services in the physical infrastructures with more environmentally friendly IT services. We have identified five principles for achieving this vision: data center scale lifecycle design, flexible and configurable building blocks, pervasive sensing, knowledge discovery and visualization, and autonomous control. We describe these principles and present specific use cases for their application. Successful implementation of the sustainable data center vision will require multi-disciplinary collaboration across various research and industry communities.
We present two new cooperative caching algorithms that allow a cluster of file system clients to cache chunks of files instead of directly accessing them from origin file servers. The first algorithm, called C-LRU (Cooperative-LRU), is based on the simple D-LRU (Distributed-LRU) algorithm, but moves a chunk's position closer to the tail of its local LRU list when the number of copies of the chunk increases. The second algorithm, called RobinHood, is based on the N-Chance algorithm, but targets chunks cached at many clients for replacement when forwarding a singlet to a peer. We evaluate these algorithms on a variety of workloads, including several publicly available traces, and find that the new algorithms significantly outperform their predecessors.
Distributed systems are notoriously difficult to implement and debug. One important tool for understanding the behavior of distributed systems is tracing. Unfortunately, effective tracing for modern distributed systems faces several challenges. First, many interesting behaviors in distributed systems only occur rarely, or at full production scale. Hence we need tracing mechanisms which impose minimal overhead, in order to allow always-on tracing of production instances. Second, for high-speed systems, messages can be delivered in significantly less time than the error of traditional time synchronization techniques such as network time protocol (NTP), necessitating time adjustment techniques with much higher precision. Third, distributed systems today may generate millions of events per second systemwide, resulting in traces consisting of billions of events. Such large traces can overwhelm existing trace analysis tools. These challenges make effective tracing difficult.We present techniques that address these three challenges. Our contributions include 1) a low-overhead tracing mechanism, which allows tracing of large systems without impacting their behavior or performance (0.14 mu s/event), 2) a post hoc technique for producing highly accurate time synchronization across hosts (within 10 mu s, compared to between 100 mu s to 2 ms for NTP), and 3) incremental data processing techniques which facilitate analyzing traces containing billions of trace points on desktop systems. We have successfully applied these techniques to two distributed systems, a cooperative caching system and a distributed storage system, and from our experience, we believe our techniques are applicable to other distributed systems.
Next generation data centers must be designed to meet Service Level Agreements (SLAs) for application performance while reducing costs and environmental impact. Traditional design approaches are manually intensive and must integrate thousands of components at multiple granularities, often with conflicting goals. We propose an Automated Data Center Synthesizer to design Sustainable Data Centers that meet SLA goals, minimize carbon emissions and embedded exergy, are optimally efficient and deliver significantly reduced Total Cost of Ownership (TCO). The paper concludes with a use case study that employs the synthesizer process flow to design an optimal data center to deliver a set of services for a hypothetical city using state of the art sustainable technologies.