Soft Error Resiliency (SER) is a major concern for Petascale high performance computing (HPC) systems. In designing Blue Gene/Q (BG/Q) [8], many mechanisms were deployed to target SER including extensive use of Silicon-On-Insulator (SOI), radiation-hardened latches [7,13], detection and correction in on-chip arrays, and very low radiation packaging materials. On the other hand, it is well known that application behavior has major impacts on the masking (or “derating” factor) in system SER calculations. The principal goal of this project is to understand the interaction between BG/Q hardware and high-performance applications when it comes to SER by performing and evaluating a chip irradiation experiment.
Fault injection through accelerated irradiation is an effective way to evaluate the overall soft error resiliency of microprocessors. In this work, we report on irradiation experiments on a Blue Gene/Q (BG/Q) compute processor chip running selected applications. Blue Gene/Q is the third generation of IBM's massively parallel, energy efficient Blue Gene series of supercomputers. In the experiments, we found 69 code fails. Out of these, 26 code fails are relevant for the calculation of the mean-time-between-failures (MTBF) for a 20 PetaFLOP, 96 rack system running a comparable workload mix. The expected MTBF for check-stops due to cosmic radiation and alpha particles from chip packaging materials is calculated to be 51 days for sea-level at New York City running the application mix studied. If the most vulnerable application is run exclusively, the projected MTBF is 35 days. These are outstanding results for a machine of this magnitude. The beaming experiment and projected MTBF validate the necessity to include autonomous hardware detection and recovery at the cost of design effort, silicon area and power.
Stream processing systems are designed to support applications that use real time data. Examples of streaming applications include security agencies processing data from communications media, battlefield management systems for military operations, consumer fraud detection based on online transactions, and automated trading based on financial market data. Many stream processing applications are faced with the challenge of increasingly large volumes of data and the requirement to deliver low-latency responses predicated by analysis of that data. In this paper, we assess the applicability of the Blue Gene architecture for stream computing applications. This work is part of a larger effort to demonstrate the efficacy of using a Blue Gene for streaming applications. Blue Gene supercomputers provide a high-bandwidth low-latency network connecting a set of I/O and compute nodes. We examine Blue Genepsilas suitability for stream computing applications by assessing its messaging capability for typical stream computing messaging workloads. In particular, this paper presents results from micro-benchmarks we used to evaluate the raw performance of Blue Gene/P (Blue Gene/P) supercomputer under loads produced by high volumes of streaming data. We measure the performance of data streams that originate outside the supercomputer, are directed through the I/O nodes to the compute nodes and then terminate outside. Our performance experiments demonstrate that the Blue Gene/P hardware delivers low-latency and high-throughput capability in a manner usable by streaming applications.