This paper introduces a novel design automation methodology for charge recovery logic (CRL). The proposed methodology combines a novel logic compression algorithm with automatic schematic generation to automate the design process of CRL, enabling power and performance simulations for a large number and variety of CRL circuits. As a measure of the effectiveness of the proposed design flow, automated implementations of CRL equivalents of the LGSynth’91 combinational benchmark circuits are compared with their CMOS counterparts. The results demonstrate a trade-off in power for area: Automatically generated CRL circuits dissipate 51.3% less power on average compared to CMOS equivalents, occupying 54.9% larger area.
This paper presents the integration of resonant clocking to multi-die architectures to synchronize individual chiplets connected through an active silicon interposer. The proposed inter-chiplet synchronization through the active silicon interposer rotary oscillator array (ASI-ROA) provides a unitary clock domain to the multiple die (i.e. multiple chiplets) in the package with a very low design overhead. System performance analysis is performed with parasitics-extracted, post-layout simulation models of two different sizes of representative heterogeneous multi-die architectures, each with varying number of RISC-V cores per die. Each RISC-V core of the multi-die package belongs to the unitary clock domain, designed with ASI-ROA to operate at a frequency of 2 GHz. The proposed architecture is investigated for robustness in frequency and skew across the multi-die system (MDS) with SPICE based simulations of post layout models, demonstrating variations of only 80 MHz for a 2 GHz target frequency. The power savings are upto 41% for the overall MDS, compared to an equivalent implementation with a contemporary ADPLL used to synchronize the multiple chiplets over the active interposer. The average clock skew of the completely resonant architecture presented in this work is 8.2 ps.
A novel clock generation and distribution network is proposed for multi-die architectures connected through an active silicon interposer. The proposed clock network generates and distributes a resonant clock through the active silicon interposer between dies, with each die served through resonant local clock trees. The proposed active silicon interposer rotary oscillator array (AI-ROA) serves to establish a unitary clock domain, providing constant phase and magnitude clock sources to the multiple die (i.e. multiple chiplets) in the package. Analysis is performed with multiple ARM CORTEX M0 cores per die of a homogeneous multi-die package architecture. Each M0 core of the multi-die package belongs to the unitary clock domain, designed with AI-ROA to operate at a frequency of 1 GHz. The multiple die are designed in the 28 nm technology node and the active interposer is designed in the 65 nm technology node. SPICE based simulations of post-layout models provides analysis and evaluation of the proposed architecture for performance metrics under process, voltage, and temperature variations. In particular, performance metrics are reported for 1) power consumption in comparison to PLL based architectures designed and synthesized with an industrial tool, 2) robustness against process variations, and 3) clock skew across the cores throughout the multiple die.
In the typical application-specified integrated circuit (ASIC) design flow, reliability-driven performance loss is computed, in part, with switching activity files. However, for ASIC designs of multicore processors, the typical switching activity files lack multithreaded software workload information. An accurate switching activity for a multicore design can be generated using a logic simulator. However, the logic simulator process suffers from long runtimes when dealing with real workloads. This paper analyzes the effects of scaling multithreaded workloads and proposes Custard, a hardware methodology for lifetime improvement of multicore processors by obtaining multithreaded switching activity signatures in a short period of time using a performance simulator (gem5), logic simulator (VCS), and thermal simulator (HotSpot). Custard is particularly important for multicore, Internet of Things processors as the runtime feedback-based reliability mechanisms used on current multicore processors incur area and power overhead that could be prohibitive for smaller form factors and power budgets. Experiments are performed with Custard using real workloads on an OpenSPARC T1 design with two, four, and eight cores that are fully synthesized and routed. The default-sized T1 core is improved to have a reliability increase of 4.1x, with 0.08% and 1.57% increase on average in cell area and switching power, respectively.
Building clock trees for tight skew constraints of clock delivery networks is standard in the industry. Tight slew constraints of high-performance designs require post-processing techniques to satisfy slew constraints after clock tree synthesis (CTS). Post-processing adversely impacts the power dissipation. This paper proposes slew merging region CTS (SMRcts); a novel algorithm to satisfy bounded slew and skew constraints simultaneously during synthesis. Experimental results performed on International Symposium on Physical Design (ISPD) 2010 benchmarks using a 20-nm FinFET technology show an average reduction of 15% power over a bounded skew approach. Comparison to the ISPD 2010 CTS contest solutions in the literature shows SMRcts producing a 51% improvement in a utility metric. Scalability of SMRcts is demonstrated on ISPD 2013 benchmarks with up to 100k sinks.
The emergence of Network-on-Chip (NoCs) as a scalable interconnection infrastructure for Chip Multi-Processors (CMPs) creates the need to analyze the degradation and lifetime repercussions incurred by network traffic from exa-scale and distributed workloads. Reliability methods and redundancies are in place to circumvent defective parts or to roll back to a previous safe state. These reliability techniques are reactive in nature and do not focus effort in avoiding degradation which may result in full system failure. Current methods for improving degradation and lifetime focus on design and runtime optimizations. This paper proposes a workload-aware routing algorithm, which complements known techniques and optimizes the overall network traffic balance. The end product is a methodology that utilizes workload signatures and priority-based routing to improve port and router utilization across the network. The proposed WAR algorithm stays within approximate to 2% of the latency and average hop count of a NoC using XY routing. Simulation results show an improvement in network traffic balance upto approximate to 17% (average improvement of approximate to 8.6%) with WAR on NoCs for lifetime improvement.
Network-on-chip (NoC) routers have a non-uniform utilization based on the number of links, location, and the workload running. Uneven utilization can lead to reliability issues, namely Negative Bias Temperature Instability (NBTI), that results in a reduced lifetime for gates. To address this, a physical design-based solution using workload signatures and cell sizing is proposed to improve the lifetime of NoCs. Using real workloads in conjunction with logic simulation and physical design, this paper analyzes the utilization of each port of NoC routers and, based on a desired lifetime, resizes the physical design. Results using SPLASH2 benchmarks show that the NoC router lifetime can be increased by 3.4× over the default sized router with an average increase of 5.58% and 5.27% for area and power, respectively.
Binary clock tree (BCT) synthesis fundamentally depends on the quality of the process of merging pairs. Selecting the optimal merging nodes is computationally expensive, even using heuristic methods. This paper presents an automated synthesis approach based on genetic algorithms (GA), that reduces the search effort for feasible node pair selection. Insights and best practices are presented for using GA processes for generating BCTs that evolve over time. BCTs synthesized with this GA approach are demonstrated experimentally with HSPICE simulations. Furthermore, the impact of utilizing a human-in-the-loop in this GA process for merging pair selection is analyzed methodically. The outcome is a best-practices approach towards automating the synthesis of BCTs based on the proposed GA approach.
In the conventional ASIC design flow, slew constraints are imposed on clock sinks and clock buffers uniformly. The slew constraint has a significant affect not only on the timing but also on power in high performance designs. This paper investigates relaxing tight slew constraints for the reduction of low-impact buffers in clock trees. This is motivated by the observation that buffers in a clock tree that directly drive clock sinks have a high impact on meeting timing constraints. Buffers that drive other buffers in the clock tree do not directly impact the sink timing within some bound. These buffers can be considered low-impact, but still consume significant power. In order to reduce this power, the slew design constraint can be relaxed for low-impact buffers. Benchmarks from the ISPD2010 benchmark suite are synthesized to demonstrate the impact of a "slew-down" at the non-sink clock buffers of a clock tree with power and timing results from HSPICE. Results using a 20nm PTM FinFET technology at 4GHz shows that power savings of up to 50% (~10% on average) can be achieved compared to the minimum slew constraint while satisfying the same global slew constraint by a methodical slew-down process.
Next-generation, multi-core Internet of Things (IoT) processors use advanced technology nodes that have increased reliability issues such as Negative Bias Temperature Instability (NBTI). Run-time feedback reliability mechanisms used on current multi-core processors incur too much area and power overhead leading to limited design space for IoT applications. In the typical ASIC design flow, switching activity files are used in order to model the reliability-driven performance loss. However, for multi-core ASIC design, the typical switching activity files lack multi-threaded software workload information thus lacking representation of modern workloads. The exact switching activity for a multi-core design can be established using a logic simulator, which captures the scheduling and execution order of multi-threaded workloads but suffer from long run times when dealing with real workloads. This paper proposes a method to obtain multi-threaded switching activity signatures in a short period of time using a performance (gem5) and a logic simulator (VCS) for IoT applications. These switching activity files are compiled into a workload signature and through this work, maintain compatibility with design tools. This workload signature is used in the standard ASIC design flow for lifetime improvement by mitigating the reliability issues such as NBTI at design-time. Experiments are performed using real workloads on an OpenSPARC T1 core. The default-sized T1 core is improved to have a reliability increase of 4.1×, leading to a desired reliability time of 10 years.
Resonant rotary clocking is a low-power clocking technology for multi-phase clock generation in GHz frequency range. In this paper, Rotary Traveling Wave Oscillators (RTWOs) are analyzed under process variations and negative bias temperature instability (NBTI) at the 90nm technology node. The analysis is focused on 1) variations in the physical geometries of the rotary ring, 2) inter and intra-die transistor variations, 3) power supply fluctuation and 4) NBTI. Monte-Carlo based analysis are performed to study the effects of process variations and transistor aging on the operating frequency and power consumption of the rotary ring at a temperature of 110° C. SPICE based simulations show natural robustness against process variations, and NBTI.
Wide wire sizes are often used in clock trees to improve timing characteristics and reduce electromigration effects. Recent research suggests the attractiveness of wide wires is affected by the forbidden pitch issues in the lithography of sub-20nm technologies. Parallel wiring is a recently proposed technique to get around these lithography issues in the routing stage of the ASIC design flow. Routing multiple minimum sized wires instead of wide wires is advantageous for i) improved metal density, and ii) is opportunistic, as the timing characteristics are dictated by the number of parallel wires. This paper proposes Wire Type clock tree synthesis, WT-CTS, the first CTS methodology to adapt parallel wiring types. WT-CTS utilizes the opportunity of changing timing characteristics through parallel wires to implement an incremental delay matching technique while the clock tree is being synthesized. A 20nm PTM technology is used in HSPICE simulation after ISCAS89 benchmarks are synthesized using WT-CTS. Results show up to a 92% reduction in skew and 12% power reduction while only increasing routing resources by 8% when compared to the minimum width design.
Simulation is employed extensively to perform exploration of design spaces by computer designers. Contemporary simulation environments are now increasingly complex comprising of support for multiple cores and full operating systems. Resource use between simulation environments vary widely because of these different system contexts and the fact that multi-threaded applications have intrinsic non-determinism. In addition, more recent simulation environments use Dynamic Binary Instrumentation (DBI) traces collected on the system context (OS, library, threading API) of the host system. Methodologies that have been employed to validate and compare simulation frameworks are usually limited to comparing CPI and cache statistics and do not provide a detailed function-level breakdown or understanding of the source of mismatches. In this work, we attempt to identify and quantify the true sources of mismatch between a DBI framework and a full system simulation framework. We use memory traces of multithreaded applications that have been annotated with function call information to allow for a breakdown of the source of mismatch within an application. To the best of our knowledge, this level of detail in comparison has not been attempted before, especially with traces of multi-threaded applications. In this study, we find that the sources of mismatch come mainly from threading mechanisms/threading API function calls, Library/System function calls and User Space condition synchronization. Based on the results of the study, we identify specific functions in each category of mismatch. We then propose a few ways to close the gap and enable more reliable simulation for design space exploration.
It is formidable to embed iterative simulations into the clock tree synthesis process to verify the skew and slew constraints. Instead, accurate and simple timing models for clock buffers are traditionally used so as to perform clock tree synthesis with sufficient accuracy. Two-pole RC and/or piecewise linear models accurately models the gate delay without a waveform dependency for a wide range of waveform properties. However, they unnecessarily complicate the problem for the time modeling of clock buffers where, unlike logic gates, the input and output waveform properties are similar. Look-up table-based approaches are traditionally used in order to obtain the clock buffer timing with inputs being the input slew and the output capacitance. However, the effective capacitance estimation of the highly resistive wires of sub-45nm technologies is a challenge, making it hard to identify the output capacitance. Also, the multiple or dynamically-scaled voltage levels of the current designs necessitate a costly LUT-based pre-characterization process. In this work, a timing estimation scheme for clock buffers is proposed which models both the delay and the slew as linear equations, bypassing the costly LUT characterization process. The experimental results performed with SAED 32nm buffer library show that the proposed timing model can achieve a maximum absolute value error of ≈5ps to ≈10ps for the buffer timing compared to SPICE simulations. Furthermore, the proposed timing model provides an error from 0.2% to 4.6% at different timing constraints and operating voltage levels, when used for insertion delay computation.
Mineo Kaneko合作论文数School of Information Science1