
This chapter chronicles the 25-year evolution of Warehouse-Scale Computing (WSC), emphasizing the journey from its initial niche application in web search to its current status as the backbone of hyperscale applications and cloud platforms. We present our historical retrospective in five epochs, each highlighting key developments in WSC. In the second part of the chapter, we distill key lessons from this journey. We conclude with a discussion of future opportunities including the next wave of scale, new accelerators, software-defined hardware, and the room for optimization at higher levels of the software stack. We also highlight the importance of agility and modularity, the continued emphasis on trusted computing (reliability, privacy, security, sovereignty), sustainability, and the disruptive potential of AI for WSC design.
This chapter starts with a high-level characterization of data center energy before moving to a discussion of power, energy, and sustainability in warehouse-scale computers (WSCs). We discuss the economic and environmental importance of energy efficiency, and define metrics like PUE (Power Usage Effectiveness) and SPUE (Server PUE) that help assess and improve data center and server infrastructure efficiency. We also explore the concept of energy-proportional computing and techniques for enhancing the energy efficiency of computing resources. We then discuss data center power provisioning strategies, including oversubscription and energy storage, and conclude with a discussion of environmental sustainability and a high-level characterization of data center energy usage of the internet and cryptocurrency mining.
This chapter analyzes performance and costs of Warehouse Scale Computers (WSCs). It highlights key performance metrics and discusses system balance within WSCs. It then examines fleetwide workload characterization and performance modeling, drawing from Google case studies. The second part of the chapter investigates WSC capital and operational expenditures, presenting a Total Cost of Ownership (TCO) model for evaluating tradeoffs between on-premises solutions and public cloud offerings. Despite offering more functionality, flexible scaling, and integrations compared to on-premises IT, public cloud is more cost effective unless on-premises lifetime utilization is very high.
This chapter presents our selection of 25 top papers for 25 years of Warehouse-Scale Computing (WSC) at Google. The papers cover a range of topics crucial to WSC, including early search architecture, cluster management, file systems, data processing, distributed storage, profiling, networking, energy efficiency, and specialized hardware like TPUs and VCUs. These diverse papers present a nice view on the evolution of WSC from initial prototypes to sophisticated, globally-distributed systems, highlighting innovations that have significantly impacted the industry. Each paper is briefly summarized, with both a discussion of the key contributions of the paper and a note on why we selected this for our list.
This chapter examines the hardware building blocks essential to warehouse-scale computers (WSCs), detailing system design principles and key tradeoffs between scale-out and scale-up architectures. It explores WSC server and rack design, including the evolution from monolithic to disaggregated systems. Special emphasis is placed on the increasing importance of specialized accelerators, with in-depth analyses of machine learning accelerators like TPUs and GPUs, as well as video acceleration with VCUs. The discussion extends to networking, covering host networking, smartNICs, cluster networking, and WANs, and highlights innovations such as spine-less networking and optical switching. The chapter further delves into the varied storage needs of WSCs, discussing HDD and SSD characteristics and the trend toward disaggregated storage designs. Finally, it concludes with an overview of cross-cutting issues and a historical review of WSC design for servers, storage, and networking over the past twenty-five years.
This chapter focuses on Software-Defined Infrastructure, using software to optimize hardware efficiency in response to a slowing Moore’s Law. It covers software-defined servers, including compiler optimizations, hardware-aware scheduling, management of disaggregated resources, and software-defined memory. It also discusses software-defined networking, a pioneering area for these concepts, highlighting traffic and topology engineering, DDoS mitigation, bandwidth enforcement, and network-aware scheduling. Furthermore, it explores the application of ’software-defined’ approaches to accelerators and data center optimizations for power and fleet management. Finally, it concludes with machine learning’s role in optimizing control loops and future opportunities in software-defined infrastructure.
This chapter discusses warehouse-scale computers (WSCs), encompassing reliability, availability, security, and privacy. We discuss fault tolerance at scale, emphasizing the need to design systems that can handle both hardware and software failures. We explore various types of faults, their root causes, and strategies for predicting, localizing, repairing, and tolerating them. The discussion then transitions to security, addressing key aspects such as data center physical security, roots of trust, encryption, confidential computing, CPU vulnerabilities, and the principles of secure service deployment.