Accurate power modeling is crucial for energy-efficient CPU design and runtime management. An ideal power modeling framework needs to be accurate yet fast, achieve high temporal resolution (ideally cycle-accurate) yet with low runtime computational overheads, and easily extensible to diverse designs through automation. Simultaneously satisfying such conflicting objectives is challenging and largely unattained despite significant prior research. In this paper, we propose APOLLO, an automated per-cycle power modeling framework that serves as the basis for both a design-time power estimator and a low-overhead runtime on-chip power meter (OPM). APOLLO uses the minimax concave penalty (MCP)-based feature selection algorithm to automatically select less than 0.05% of RTL signals as power proxies. The power estimation achieves R-2 > 0.95 on Arm Neoverse N1 [3] and R-2 > 0.94 on Arm Cortex-A77 [2] microprocessors, respectively. When integrated with an emulator-assisted flow, APOLLO finishes per-cycle power estimation on millions-of-cycles benchmark in minutes for million-gate industrial CPU designs. Furthermore, the power model is synthesized and integrated into the microprocessor implementation as a runtime OPM. APOLLO's accuracy further improves when coarse-grained temporal resolution is preferred. To our best knowledge, this is the first runtime OPM that simultaneously achieves percycle temporal resolution and < 1% area/power overhead without compromising accuracy, which is validated on high-performance, out-of-order industrial CPU designs.
The increasing infant mortality and wear out failure rates observed in very deep sub micron silicon technologies is now a major problem for the design of future high-density SoCs. Emerging architectures based on Multi-Processor SoCs (MPSoCs) give the opportunity to exploit the natural redundancy to control the system performance in presence of failures. In this paper we evaluate the impact of a distributed fault-handler strategy on the system in term of cost and feasibility. We also discuss strategies for applying this technique to a distributed MPSoC architecture.
This paper proposes a novel strategy for optimizing resources in Multi-Processor Systems-on-Chip (MPSoC). The approach is based on using control-loop feedback mechanism to maximize the efficiency on exploiting available resources such as CPU time, operating frequency, etc. Each Processing Element (PE) in the architecture is equipped with a frequency scaling module responsible for tuning the frequency of processors at run-time according to the application requirements. Results show the system's capability of adapting to disturbing conditions. For validation purposes we have implemented a multi-threaded MJPEG decoder together with an ADPCM audio decoder and a FIR.
The increasing failure rates observed in very deep sub micron silicon technologies pose a major problem to the design of future high-density SoCs. Emerging new architecture based on Multiprocessor SoC (MPSoC) gives the opportunity to exploit the natural redundancy with replicated spare processor in order to maintain the system performance in presence of failures. Based on the assumption that a transient loss of functionality can be tolerated, we study the feasibility and propose a cost-effective dependable hardware/software method which self-substitutes faulty processors with spare processors in a distributed manner. It guarantees the integrity, improves the availability and eases the maintainability of the MPSoC at system-level.
In this paper we propose a strategy for better exploiting Multi-Processor Systems-on-Chip resources utilization by means of using a control-loop feedback mechanism. We apply the proposed techniques in a purely distributed memory MPSoC architecture that is composed of a frequency scaling module responsible for tuning the frequency of processors at run-time. Results show very promising in terms of adaptation capabilities for system with dynamic workload. Performance results demonstrate the effectiveness of the proposed approach when workload requirements for applications may vary, affecting the overall performance of the system. For validating the proposed approach we have implemented a multi-thread MJPEG decoder application and created an architecture model with/without perturbations in the system.
The increasing failure rates observed in very deep sub micron silicon technologies pose a major problem to the design of future high-density SoCs. While hardening techniques originated from critical application areas (automotive, avionics) exist, they usually incur a cost overhead that renders them inadequate for consumer market segments. Thus we present a concept, an implementation and an evaluation of a scalable software-hardware detection, isolation and recovery method. The method exploits the natural redundancy that exists in MPSoCs for enhancing their reliability. Based on the assumption that a transient loss of functionality can be tolerated, the proposed scheme relies on a hardware/software framework that makes it possible to diagnose and to isolate faulty processors in a distributed manner. It guarantees the integrity, improves the availability and eases the maintainability of the MPSoC at system-level.
Detecting and reacting to the various failure occurrences on multiprocessor system on chips (MP-SoC) will be challenging in the next few years. We propose two reliable architectural solutions: a watchdog for streaming architecture and a diagnostic and reconfiguration technique based on software tasks. These two methods incur a cost overhead, that can be tuned for signal and image processing algorithms.