A fully synthesizable DCiM on-die accelerator with 128 input channels, 32 output channels, and 2 weight sets, supporting zero-point-quantized INT8xINT8 computation, is fabricated in Intel 18A technology and occupies 0.0856mm(2). The DCiM accelerator operates at 2.62GHz for a supply voltage of 1.1V, and 25 degrees C, achieving a peak area-efficiency of 250TOPS/mm(2). Robust operation of DCiM is maintained down to 400mV, 25 degrees C, where it delivers a peak energy-efficiency of 147TOPS/W at 25% input activity, with a power consumption of 6.1mW.
A fully synthesizable DCiM on-die accelerator with 128 input channels, 32 output channels, and 2 weight sets, supporting zero-point-quantized INT8×INT8 computation, is fabricated in Intel 18A technology and occupies $0.0856 \text{mm}^{2}$. The DCiM accelerator operates at 2.62 GHz for a supply voltage of 1.1 V, and 25 °C, achieving a peak area-efficiency of $250 \text{TOPS} / \text{mm}^{2}$. Robust operation of DCiM is maintained down to $400 \text{mV}, 25^{\circ} \mathrm{C}$, where it delivers a peak energy-efficiency of 147TOPS/W at 25% input activity, with a power consumption of 6.1 mW.
A side-channel-attack (SCA) resistant AES engine with multiplicative-masked Sboxes is fabricated in Intel 4 CMOS, achieving 1.8× lower area overhead compared to conventional additive-masked implementations. Balanced dual-rail detector circuits pre-empt zero-value attacks while providing a 34,000× increase in side-channel-attack resistance, with a measured minimum-traces-to-disclose (MTD) of 850M encryption traces.
A fully configurable and synthesiz able timing characterization test-bench for memory IPs enables high resolution clk2q, setup, hold and cycle-time delay measurements. The timing test-bench features distributed regional capture FFs, mesh based low-skew clock and setup difference measurement across regional capture FFs to minimize error, multiple data/input delay generators to handle timing permutations across memory inputs, automated relative placement/pre-routing for matched layout and XORed clock delay generators to create multiple edges for measuring read after write delay/cycle time.
Power and electromagnetic (EM) side-channel attacks (SCA) exploit data-dependent power consumption from cryptographic engines to extract embedded secret keys. While series-connected voltage regulators [1], [2] and arithmetic countermeasures like heterogenous Galois-field arithmetic [3] provide acceptable levels of side-channel leakage suppression, they cannot defend against determined adversaries. Random additive masking [4] on the other hand, provides a provably-secure solution [5] that disrupts first-order correlations between measured power/EM signatures and secret keys, while incurring $2\times$ overhead in area and power consumption. In this paper, we demonstrate a reconfigurable AES accelerator fabricated in Intel 4 CMOS process with minimum-time-to-disclosure (MTD) $> 1\text{B power}/\text{EM}$ traces in on-demand SCA-resistant mode, while providing a $2.2\times$ boost in encryption performance during a dual-core mode of operation (Fig. 34.4.1). When coupled with side-channel attack detection techniques [6], [7], this approach allows the user to operate at $> 2\times$ AES throughput during the safe mode of operation in trusted environments, with the ability to quickly trade-off throughput for a higher level of SCA-resistance when the onset of an attack is detected. In the blind-bulk mode of operation, the accelerator randomly switches at a user-specified rate between SCA-resistant and dual-core modes while encrypting bulk data, providing $1.14-\text{to}-1.6\times$ boost in encryption throughput with measured MTD $> 50\mathrm{M}$ traces.
A double-buffered, 4kb standard-cell-based register file with measured 2.4GHz operation at 0.65V, 100°C and scalable performance to 3.7GHz at 0.8V is fabricated in a leading-edge CMOS node. Double-buffering, Gray-coded read/write addressing, super-multi-bit macro standard cells, and mixed-frequency clocking enable a measured peak energy-efficiency of 5.73TOPS/W at 0.5V, 0°C with 14%/11% register file read/write power savings and 28% area reduction over conventional ping-pong design.
A side-channel attack (SCA) hardened AES-128 and RSA crypto-processor in 14-nm CMOS with measured resistance to correlation power/electromagnetic analysis (CPA/CEMA) in both time and frequency domains is demonstrated. While previously reported linear low-dropout regulators (LDOs) offer improvements in minimum-time-to-disclose (MTD) of extracted key bytes in the time domain, their transformations are less effective against frequency-domain attacks. This article describes a non-linear digital LDO (NL-DLDO) with control loop randomizations that bolster SCA resistance in the frequency domain. The NL-DLDO cascaded with an AES engine augmented with arithmetic countermeasures enables > 250K× improvement in MTD, with no CPA/CEMA/DNN attacks detected after 1-B encryptions, with 8% power and 10% area overheads incurred by arithmetic techniques. The RSA-4K crypto-processor implements exponent magnitude and timing randomizations along with dynamic memory addressing to mitigate time- and frequency-domain attacks. The countermeasures enable 711× suppression in means separation in current/EM magnitudes from 3.1 mV to 4.35 μV, reducing attacker's accuracy to an ineffective random guess classification, while limiting area and performance overheads to <; 0.05% and 3.25%, respectively.
A 10nm digital Binary Neural Network (BNN) chip implements 1b activations and weights for compute density of 418TOPS/mm2 and memory density of 414KB/mm2. The chip achieves an energy efficiency of 617TOPS/W by leveraging Compute Near Memory (CNM), parallel inner product compute, and Near-Threshold Voltage (NTV) operation. The digital BNN design approaches the energy efficiency of analog in-memory techniques while also ensuring deterministic, scalable, and precise operation.
An AES engine with uniform side-channel-attack (SCA) resistance across time/frequency domains using a high-bandwidth non-linear digital low-dropout (NL-DLDO) regulator in conjunction with AES arithmetic countermeasures is fabricated in 14nm CMOS. Randomized regulator loop parameters and cascading LDO and arithmetic transformations provide >250K× increase in frequency/time-domain MTD, with no CPA attack detected on current/electromagnetic (EM) traces measured from 1 billion encryptions.
A 10-nm DNN inference accelerator compresses model size with tabulation hash-based line-grained weight sharing and increases 8bcompute density by 3.4× to 1.6 TOPS/mm 2 . The compressed model DNN implements lightweight hashing circuits to compress fully connected and recurrent neural networks. Optimized shared weight address generation reduces MUX tree area overhead by 40%. Runtime hash table generation and weight mapping circuits enable a peak energy efliciency of 9.0 TOPS/W at 450 mV, 25°C. A 128×-compressed 3-layer long shortterm memory classilies TIMIT phonemes with 85.6% accuracy for a total energy of 14 μJ/classilication, with <; 0.5% degradation in accuracy over an uncompressed network.
A.A 4900μm 2 side-channel attack (SCA) resistant AES accelerator in 14nm CMOS achieves 1200x higher minimumtime-to-disclosure (MTD) than an unprotected AES. Randomized byte-order shuffling using heterogeneous Sboxes, linear masked MixColumns and dual-rail key addition enable 9.2x lower correlation between current traces and HD/HW power models. The accelerator achieves 839Mbps throughput with 0.7% performance overhead.
A 10-nm compute-near-memory (CNM) accelerator augments SRAM with multiply accumulate (MAC) units to reduce interconnect energy and achieve 2.9 8b-TOPS/W for matrix-vector computation. The CNM provides high memory bandwidth by accessing SRAM subarrays to enable low-latency, real-time inference in fully connected and recurrent neural networks with small mini-batch sizes. For workloads with greater arithmetic intensity, such as large-batch convolutional neural networks, the CNM reconfigures into a 2-D systolic array to amortize memory access energy over a greater number of computations. Variable-precision 8b/4b/2b/1b MACs increase throughput by up to 8x for binary operations at 33.0 1b-TOPS/W.
A 10 28 challenge-response strong-PUF in 14nm CMOS, demonstrates machine learning (ML) attack resistance across 6-million training samples. The 2-stage non-linear cascaded PUF array with adversarial challenge selection limits ML attack accuracy to ~50%. The configurable cross-coupled inverter-based entropy source with stability-aware challenge pruning enables 9.8× higher array density and 0.26% peak BER across 650-850mV and 0-100°C.
Low-clock-power digital standard cell IPs in 10nm CMOS, featuring low-power shared-clock (LPSC) flip-flops (FFs), LPSC back-to-back (B2B) FFs, and pass-gate (PG) integrated clock gates (ICGs), achieve up to 14%, 45%, and 14% measured clock energy improvements, respectively, by reducing the number of clocked devices over state-of-the-art conventional transmission-gate (TG) FF and AND ICG circuits. The LPSC FF achieves a mean worst-case black-hole-time (BHT) improvement of 17ps, while the PG ICG achieves a mean enable/disable setup time improvement of 16ps/15ps, compared to conventional circuits measured at 650mV, 25°C. Power analysis of a graphics processor block with these optimized IPs results in an overall 6% clock power reduction without frequency impact.
The clock frequency of high-performance processors in the high-voltage turbo/burst operation mode [1] is governed primarily by the RC-dominated global interconnects in the communication fabrics and NoCs, since the interconnect RC delay does not improve proportionally with logic gate delay at higher voltages. Prior current-mode and pre-emphasis techniques for improved interconnect throughput and energy efficiency [2]-[5] do not address the critical turbo/burst mode latency requirements, and do not consider tunability across a wide voltage-frequency operating range in high-performance processors. Their transmit/receive circuits with complicated data recovery and differential signaling incur major latency and wiring resource overheads, especially for shorter bus distances typical in NoCs. Previous hybrid current/voltage-mode signaling used expensive 4-inversion repeaters without reconfigurable resistive termination strength, and required multi-cycle windows for current-mode operation with limited power savings at high data activities [6]. Other techniques address scalability to lower supply voltages, but do not focus on the critical interconnect RC delay bottlenecks at high voltages [7]. In this paper, we present reconfigurable current/voltage-mode driver/repeater/receiver circuits for on-die global interconnects/busses, fabricated in 10nm CMOS, with transient current-mode operation, achieving: i) up to 38% lower delay/mm vs. voltage-mode operation, ii) 5x lower transient current-mode power vs. always-on current-mode operation at 100% data activity, iii) similar noise immunity as voltage-mode operation, iv) tunable current-mode configuration for optimized performance vs. energy across PVT corners, and v) up to 8x increased repeater distance.
A 0.072mm 2 SCA-Resistant RSA-4K crypto-processor fabricated in 14nm CMOS achieves peak encryption throughput of 0.08Mbps at 750mV, 25°C. Exponent magnitude/timing randomization with dynamic RF addressing provide 711× lower means-separation in current/EM trace magnitudes, reducing attacker's accuracy to an ineffective random guess classification while limiting area/performance overhead to 3%.