The IBM POWER8 (TM) architecture introduces many novel features that improve the overall performance and energy management of systems based on this new platform. To cover a wide range of workloads, the design team focused on not just the traditional server workloads like transaction processing and enterprise resource planning, but also on emerging workloads such as big data, analytics, and cloud. Designed for openness, the POWER8 platform enables other technology providers to innovate freely with OpenPower (TM). In designing the most critical POWER8 processor features, the development team relied on state-of-the-art performance and energy models for design optimization and trade-off studies, and for projecting the system-level performance for the most important workloads. In this paper, we describe these models in detail and spotlight how the models were used to influence the design of specific POWER8 features.
The IBM POWER8™ processor was designed for high performance on traditional server workloads as well as big data, analytics, and cloud workloads. In this paper, we describe key performance features of the IBM POWER8 processor. These include hardware assists that allow the POWER8 processor to automatically adapt to changing workloads by dynamically monitoring and tuning itself, enhancements to hardware instrumentation for performance monitoring, and performance improvements for encryption, virtualization, and I/O. We also describe the performance characteristics of a wide variety of applications, and we present the results of these applications running on POWER8 processor-based systems compared with previous generations of IBM Power Systems™.
This article consists of a collection of slides from the author's conference presentation on the special features, system design and architectures, processing capabilities, and targeted markets for IBM's POWER8 family of processor products.
In this paper, we describe the key performance enhancements in IBM POWER7 (R) microarchitecture and its memory hierarchy, including performance modeling and verification methodology. We also describe the performance characteristics of server applications, including Standard Performance Evaluation Corporation (SPEC) central processing unit, SAP Sales and Distribution, SPECjbb, online transaction processing workloads, and high-performance computing applications running on POWER7 processor-based systems compared with other systems.
This article consists of a collection of slides from the author's conference presentation. Some of the specific areas/topics discussed include: Power Dissipation and Efficiency Basics; POWER4 vs. POWERS; POWERS vs. POWER6; Roadrunner and Blue Gene System Efficiency; Conclusion; BACKUP: Looking Ahead: A Few Key Research Issues.
The newly released CPU2006 benchmarks are long and have large data access footprint. In this paper we study the behavior of CPU2006 benchmarks on the newly released Intel's Woodcrest processor based on the Core microarchitecture. CPU2000 benchmarks, the predecessors of CPU2006 benchmarks, are also characterized to see if they both stress the system in the same way. Specifically, we compare the differences between the ability of SPEC CPU2000 and CPU2006 to stress areas traditionally shown to impact CPI such as branch prediction, first and second level caches and new unique features of the Woodcrest processor. The recently released Core microarchitecture based processors have many new features that help to increase the performance per watt rating. However, the impact of these features on various workloads has not been thoroughly studied. We use our results to analyze the impact of new feature called "Macro-fusion" on the SPEC Benchmarks. Macro-fusion reduces the run time and hence improves absolute performance. We found that although floating point benchmarks do not gain a lot from macro-fusion, it has a significant impact on a majority of the integer benchmarks.
Workload characterization has become an integral part of the design of future servers since their characteristics can guide the developers to understand the workload requirements and how the underlying architecture would optimize the performance of the intended workload. In this paper, we give an overview of the POWER5 architecture. We also introduce the POWER5 performance monitor facilities and performance events that lead to the construction of a CPI (cycles per instruction) breakdown model. For our study, we characterize four different groups of workloads: commercial, HPC, memory, and scientific. Using the data obtained from the POWER5 performance counters, we breakdown the CPI stack into a base component, when the processor is completing work and a stall component when the processor is not completing instructions. The stall component can be further divided into cycles when the pipeline was empty and cycles when the pipeline was not empty but completion is stalled. With this model, we enumerate the number of processing cycles, i.e., a fraction of the CPI, a workload spent while progressing through the core resources and the incurred penalty upon encountering those resource usage inhibitors. The results show the CPI breakdown for each workload, identify where each workload spends its processing cycles and the associated CPI cost when accessing the core resources.
The authors compared three popular Internet server benchmarks with a suite of CPU-intensive benchmarks to evaluate the impact of front-end and middle-tier servers on modern microprocessor architectures.
Java has become fairly popular on commercial servers in recent years. However, the behavior of Java server applications has not been studied extensively. We characterize two Java server benchmarks, SPECjbb2000 and VolanoMark 2.1.2, on two IBM PowerPC architectures, the RS64-III and the POWER3-II, and compare them to more traditional workloads as represented by selected benchmarks from SPECint2000. We find that our Java server benchmarks have generally the same characteristics on both platforms: in particular, high instruction cache, ITLB, and BTAC (Branch Target Address Cache) miss rates. These benchmarks also exhibit high L2 miss rates due mostly to data loads. Instruction cache and L2 misses are seen to be the primary contributors to CPI.
The phenomenal growth of the World Wide Web has resulted in the emergence and popularity of several information technology related computer applications. These applications are typically executed on computer systems that contain state-of-the-art superscalar microprocessors. Superscalar microprocessors can fetch, decode, and execute multiple instructions in each clock cycle. They contain multiple functional units and generally employ large caches. Most of these superscalar processors execute instructions in an order different from the instruction sequence that is fed to them. In order to finish the job as soon as possible, they look further down into the instruction stream and execute instructions from places where sequential execution flow has not reached yet. With the aid of sophisticated branch predictors, they identify the potential path of program flow in order to find instructions to execute in advance. At times, the predictions are wrong and the processor nullifies the extra work that it speculatively performed. Most of the microprocessors that are executing today's internet workloads were designed before the advent of these emerging workloads. The SPEC CPU benchmarks (See sidebar on SPEC CPU benchmarks) have been used widely in performance evaluation during the last 12 years, but they are different in functionality from the emerging commercial applications. Whether the difference in functionality results in key differences in the exploitation of architectural and microarchitectural features of the processor, is the subject of this article. In order to answer this question one has to identify appropriate workloads and obtain performance metrics that indicate the execution characteristics. Emerging workloads contain several software packages, interfaces and standards that were triggered by the proliferation of web servers and the arrival of electronic commerce. An end-to-end e-business transaction typically involves a dozen or more different software layers, including the front end/portal, shopping carts, network communication, credit card or electronic check transaction, security software layers, etc. Literally all enterprises including airlines, banks, stock brokerage firms, and most consumer product vendors nowadays use their web servers to deal with
Java has, in recent years, become fairly popular as a platform for commercial servers. However, the behavior of Java server applications has not been studied extensively. We characterize two multithreaded Java server benchmarks, SPECjbb2000 and VolanoMark 2.1.2, on two IBM PowerPC architectures, the RS64-111 and the POWER3-11, and compare them to more traditional workloads as represented by selected benchmarks from SPECint2000. We find that our Java server benchmarks have generally the same characteristics on both platforms: in particular, high instruction cache, ITLB, and BTAC (Branch Target Address Cache) miss rates. These benchmarks also exhibit high L2 miss rates due mostly to loads. As one would expect, instruction cache and L2 misses are primary contributors to CPI. Also, the proportion of zero dispatch cycles is high, indicating the difficulty in exploiting ILP for these workloads.
Praveen Seshadri合作论文数Cornell University;Computer Science Department 1