The 65 nm cell broadband enginetrade (cell BE) is a multi-core SoC, implemented in a high performance SOI technology featuring a separate dual power supply for SRAM arrays to improve stability and performance using an elevated voltage. A new method is shown to analyze the SRAM cell under application conditions which was used to tune the cell for stability, write-ability and performance. An improve...
Methods for estimating JPEGcompressed image quality without the original image are investigated. Estimation of peak signalto-noise (PSNR) is calculated by utilizing the Laplacian nature of the image transform coefficients along with the quantization step size used to generate the image. Estimation of the Laplacian lambda () values are shown as well as improvements to the PSNR estimation that are achieved when compensating for large quantization step sizes. PSNR estimation using the ITU-T J.240 standard is also implemented.
The 65nm CELL Broadband Enginetrade design features a dual power supply, which enhances SRAM stability and performance using an elevated array-specific power supply, while reducing the logic power consumption. Hardware measurements demonstrate low-voltage operation and reduced scatter of the minimum operating voltage. The chip operates at 6GHz at 1.3V and is fabricated in a 65nm CMOS SOI technology.
The floating point unit in the synergistic processor element of a CELL processor is a fully-pipelined 4-way SIMD unit designed to accelerate media and data streaming. It supports 32-bit single-precision floating point and 16-bit integer operands with two different latencies, optimizing the performance of critical single-precision multiply-add operations. It employs fine-grained clock gating for power saving. Architecture, logic, circuits and integration are co-designed to meet the performance, power, and area goals.
The authors describe the low-power design of the synergistic processor element (SPE) of the cell processor developed by Sony, Toshiba and IBM. CMOS static gates implement most of the logic, and dynamic circuits are used in critical areas. Tight coupling of the instruction set architecture, microarchitecture, and physical implementation achieves a compact, power-efficient design.
The floating-point unit in the synergistic processor element of the 1st generation multi-core CELL processor is described. The FPU supports 4-way SIMD single precision and integer operations and 2-way SIMD double precision operations. The design required a high-frequency, low latency, power and area efficiency with primary application to the multimedia streaming workloads, such as 3D graphics. The FPU has 3 different latencies, optimizing the performance critical single precision FMA operations, which are executed with a 6-cycle latency at an 11FO4 cycle time. The latency includes the global forwarding of the result. These challenging performance, power, and area goals were achieved through the co-design of architecture and implementation with optimizations at all levels of the design. This paper focuses on the logical and algorithmic aspects of the FPU we developed, to achieve these goals.
THE SYNERGISTIC PROCESSOR ELEMENT IS A NEW ARCHITECTURE ORIENTED FOR MULTIMEDIA AND STREAMING PROCESSING. IN THIS ARCHITECTURE, THE MEMORY IS NOT A CACHE BUT A PRIVATE OR SCRATCH PAD MEMORY. SUCH A MEMORY IS SIMPLE AND NEEDS TO BE HIGH-FREQUENCY AND LARGE SPACE IN LOW-POWER. THIS DESIGN USES AN 11 FAN-OUT OF FOUR (11 F04), Six-CYCLE, FULLY PIPELINED, EMBEDDED 256-KBYTE SRAM FOR THIS PURPOSE. THE DESIGN'S MEMORY IS NOT ONE HARD MACRO, BUT A GROUP OF CUSTOM MACROS PHYSICALLY DISTRIBUTED TO OPTIMIZE THE PIPELINE.
A 32b 4-way SIMD dual-issue synergistic processor element of a CELL processor is developed with 20.9 million transistors in 14.8mm/sup 2/ using a 90nm SOI technology. CMOS static gates implement the majority of the logic. Dynamic circuits are used in critical areas, occupying 19% of the nonSRAM area. ISA, microarchitecture and physical implementation are tightly coupled to achieve a compact and power efficient design. Correct operation has been observed up to 5.6GHz at 1.4V supply and 56/spl deg/C.
A 4-way SIMD streaming processor of a cell processor is developed in a 90nm SOI technology. CMOS static gates implement the majority of the logic. Dynamic circuits are used in critical areas, occupying 19% of the non-SRAM area. ISA, microarchitecture, and physical implementation are co-optimized to achieve a compact and power efficient design
Christian Jacobi, Ii合作论文数IBM Development Germany2