In this paper we introduce a fast discrete event-driven simulation methodology, called KnightSim, that is intended for use in the development of future computer architectural simulations. KnightSim extends an older event-driven simulation library by (1) incorporating corrections to functional issues that were introduced by the recent additions of stack protection, pointer mangling, and source fortification in the Linux software stack, (2) incorporating optimizations to the event engine, and (3) introducing a novel parallel implementation. KnightSim implements events as independently executable x86 "KnightSim Contexts". KnightSim Contexts comprise a mechanism for fast context execution and automatically model occupancy and contention, which readily lends itself to use in computer architectural simulations. We present the implementation methodologies of KnightSim and Parallel KnightSim with a detailed performance analysis. Our performance analysis makes direct comparisons between KnightSim, Parallel KnightSim, and the discrete event-driven simulation engines found in three different mainstream computer architectural simulators. Our results show that on average KnightSim achieves speedups of 2.8 to 11.9 over the other discrete event-driven simulation engines. Our results also show that on average Parallel KnightSim can achieve speedups over KnightSim of 1.89, 3.33, 5.84, and 9.24 for 2, 4, 8, and 16 threaded executions respectively.
We introduce M2S-CGM a detailed architectural simulator that models the interactions between CPUs and GPUs operating in coherent heterogeneous compute environments. M2S-CGM extends an existing and established x86 CPU model and Southern Islands GPU model, adds a new custom-built memory system model and switching fabric called CGM, and incorporates a well-known SDRAM model. The CGM memory system simulator provides configurable entire system simulation and can support a range of non-coherent and coherent CPU-GPU configurations. M2S-CGM supports the runtime for OpenCL-based benchmarks in addition to traditional multithreaded CPU benchmarks and can run benchmarks from established heterogeneous benchmark collections. This allows us to experiment with different coherent CPU-GPU configurations and propose effective future improvements in these systems. We present the makeup of M2S-CGM's software architectural design, provide a validation of the simulator, and provide coherent CPU-GPU execution results. Our validation results show average differences between our physical test system and M2S-CGM, of 10.4%, 22%, and 6.4% for 2 threaded, 4 threaded, and heterogeneous benchmark runs respectively. Our coherent CPU-GPU experimental results show an average speedup of 2.8 for our benchmarks over the baseline noncoherent system.
Power management is becoming very important in data centers. Cloud computing is also one of the newest promising data center techniques which is appealing to many big companies. As cloud computing is different from current data centers in terms of power management due to a dynamic structure and property for its online service. Power budgeting, in terms of its important role in power management, provides powerful solutions for cloud computing with dynamic capabilities. To be specific, existing methods for data centers are based on power distribution units (PDU) divided by fixed locations on physical levels. However, it is not suitable for cloud with the dynamic property. We propose a power management design based at the logical level which uses a distribution tree with classified power capping by different service or workload types. By setting multiple trees, we can differentiate and analyze the effect of workload types and Service Level Agreements (SLAs) in terms of power characteristics.
— Researchers are studying the reconfiguration capabilities of current Static Random Access Memory (SRAM) based Field Programmable Gate Arrays (FPGAs). Researchers are working on improving the sustainability, survivability, and availability of these FPGAs. This concept is called Evolvable Hardware. The goal of researchers is to develop and employ efficient techniques which provide an autonomous capability for self healing after the FPGA has experienced a permanent physical fault. The techniques in development are targeted towards systems where human interaction after a permanent system fault is typically not possible i.e. satellites and deep space probes. This paper focuses on the techniques in development for fault detection, isolation, diagnosis, and repair within the FPGA. Also after identification and description of the techniques currently used by researchers a theoretical application of these techniques is presented. In the theoretical application a description of how to employ these techniques together, as a system, to produce an efficient means of fault detection and isolation is demonstrated.
Mark Heinrich合作论文数School of Electrical Engineering and Computer Science (EECS)2