The rapid proliferation of intelligent sensors has led to increased latency and privacy risks when processing data on a remote server. This work proposes a spin-transfer torque magnetic random access memory (STT-MRAM) compute-in-memory (CIM) macro to enhance inference efficiency for an artificial intelligence (AI) model on the edge device. A sparsity-adaptive design is presented, featuring coupled-1T1M (c-1T1M) computing cells to eliminate bit-level activation computing power overhead and suppress pattern-dependent deviations in the result. An activation-aware voltage comparator array (AA-VCA) reduces power consumption by up to 56 & times; and improves sensing margin by 13 & times;. At the system level, a hybrid-granularity sparsity workflow accelerates inference and aligns well with compact analog CIM architectures. In addition, a hardware-implemented variational autoencoder (VAE) is deployed to enhance processing capabilities. The test chip is fabricated in 28-nm CMOS with 60-nm MTJ, achieving energy efficiencies of 712 TOPS/W for image classification and 492.8 TOPS/W for anomaly detection, along with an accuracy of 96.3% and an F1 score of 0.97.
Spin-transfer torque magnetic random-access memory (STT-MRAM) offers inherent radiation tolerance from its magnetic tunnel junction (MTJ) structure, but remains susceptible to total ionizing dose (TID) degradation and single event effects (SEE) in CMOS circuitry. To enhance overall reliability, this work proposes a hierarchical fault mitigation architecture that dynamically distinguishes and mitigates transient and permanent faults through coordinated ECC correction and redundancy mapping. A radiation-hardened standard cell library is further incorporated to reinforce vulnerable control logic, complementing the proposed architecture for enhanced robustness. Circuit-level simulations and layout evaluation based on a 28 nm CMOS process demonstrate significant improvements in radiation robustness, maintaining stable functionality up to 70 MeV•cm2/mg and 450 krad(Si) with only 15% area overhead.
Voice devices on the edge have shown great potential in real-time speech recognition. However, their limited power resources and constant activity characteristics create a significant demand for low-power techniques. The dual-voltage method is widely adopted in low-power design as it can flexibly switch the supply voltage while maintaining performance, but its leakage current and limited driving capability pose challenges in energy efficiency and timing integrity. In this paper, we propose 1) a dual-rail gate circuit design that enables low-voltage operation to reduce cell power consumption, 2) a dynamic sensitivity-aware dual-voltage allocation strategy that effectively reduces the proportion of high-voltage cells and mitigates the additional area and power overhead caused by low-to-high voltage-driven structures. We apply this algorithm to a speech accelerator for binary weight neural network, implemented in 22nm CMOS technology, reducing power consumption by 42.7%~73.6% compared to the state-of-the-art speech recognition accelerators.
To meet the stringent requirements of ultra-reliable and low latency communications (URLLC), this paper investigates multi-stream transmission schemes for short packet communications. Specifically, we consider a multiple-input multiple-output (MIMO) eigenbeam-space division multiplexing (E-SDM) architecture operating over quasi-static fading channels. To minimize transmission latency while maintaining high reliability, we propose a packet splitting strategy that activates dominant eigenbeams. We then derive analytical expressions to evaluate the system reliability in the finite blocklength regime by leveraging the eigenvalue distributions of complex Wishart matrices. Crucially, to characterize the performance across different codeword spatial spans, we investigate two distinct transmission frameworks encompassing independent stream encoding over distinct eigenmodes and joint encoding across parallel subchannels. Furthermore, we derive closed-form asymptotic expressions for the decoding error probability in the high signal-to-noise ratio (SNR) regime to extract the achievable diversity and array gains. Extensive simulations validate the accuracy of our theoretical derivations and demonstrate that the proposed selected multi-stream transmission (SMST) strategy outperforms strategies without eigenbeam selection. Notably, the joint encoding scheme maintains a distinct reliability advantage over independent encoding, enabling the system to balance spatial multiplexing and reliability gain under strict quality of service (QoS) constraints.
Conventional FP-CIMs suffer from fixed preserved bit-width (PBW), limiting their adaptability and efficiency. This work proposes the first MXFP-CIM macro enabling wide-range adaptive PBW, featuring: (1) A serial dual-bit-sliding scheme; (2) A harmless data mapping scheme with a hierarchical hidden-bit decoder; (3) An adjustable-PBW MXFP-MAC circuit via twin-stage allocation. The 28nm MXFP-CIM macro achieves a peak energy efficiency of 127.54TFLOPS/W in the MXFP6/6 mode.
Automatic Speech Recognition (ASR) systems based on neural networks have shown outstanding performance in human-computer interaction. However, the environmental noise and growing number of speech categories result in degraded classification accuracy, larger model size and increased power consumption, making deployment on edge devices challenging. To address these issues, we propose TNNSAT, a Tree-NN based small footprint ASR processor with two key innovations: 1) An efficient Tree-NN architecture that achieves a 68.8% reduction in data memory and hardware reuse by 95%, while maintaining high accuracy. 2) This work features on-chip adaptive learning that supports real-time fine-tuning, enabling speaker registration and robust adaptation across various noise backgrounds. Fabricated in a 28-nm CMOS process, the prototype system demonstrates a high energy efficiency of 7.7 nJ/word. It can adapt to various noise backgrounds (11 noise types, SNR: 5 dB~clean) and support more keyword categories, achieving a 1.30× ~ 4.49× energy efficiency improvement over state-of-the-art (SOTA) solutions.
Large language models (LLM) have demonstrated unprecedented performance across natural language processing tasks. However, this comes at the cost of their huge model size. Three-dimensional (3D) LLM accelerators have shown potential to alleviate this challenge with highly integrated computation units and memory, yet critical challenges remain: 1) thermal wall caused by high power density, which significantly degrades chip performance; 2) imbalanced workloads between LLM operators that lead to high communication latency and limited throughput. To this end, this paper proposes 3D-TANoC, a heterogeneous 3D integrated (H3D) accelerator with hierarchical network-on-chip (NoC) for efficient LLM inference. The main contributions are: 1) a thermal-aware H3D memory-on-logic architecture with hierarchical NoC to balance the temperature distribution; 2) operator-aware dataflow mapping strategy that supports efficient communication to achieve high performance and throughput. Experiments on LLaMA and OPT inference demonstrate that the proposed accelerator achieves 13.85–14.80× higher energy efficiency and 12.90–13.71× speedup over the A800 GPU baseline.
Among new non-volatile memories, spin-transfer-torque magnetic-random-access-memory (STT-MRAM) has high advantages due to its high density and compatibility with CMOS process. In this work, a 4-Mb high reliability STT-MRAM macro is proposed with bottom-up MRAM refinement methodology, including 16 banks and the error check and correction (ECC) block. The high density MRAM bit-cell, the multi-bit voltage sense amplifier (MB-VSA), the high-speed and high-accuracy write self-termination circuit (HSHA-WT) and the endurance detection module are proposed to help realize high density and high reliability, together with the in-situ yield monitoring (ISYM) technique. Utilizing 28nm CMOS technology, this work achieves 6ns read speed, 20ns fast write time, 20 years retention time, 107 endurance cycles and 0% failure rate with ECC.
Combinatorial optimization problems, which are fundamental to fields such as finance and artificial intelligence, remain challenging for classical computers due to their NP-hard nature. Ising machines have emerged as a promising physical platform that solves such problems by searching for the ground state of the Ising model. In this work, we propose an Ising machine implemented with spin-orbit torque probabilistic bits (SOT P-Bits). A scalable network of SOT MTJ-based P-Bits was constructed and a hardware-software co-calibration strategy was introduced to ensures behavioral consistency across probabilistic units. The compressed sparse row (CSR) method was applied to compress sparse weights while simultaneously optimizing the multiply-accumulate (MAC) operations. To mitigate the issue of local minima in large-scale instances, a local optimal temperature lifting (LOTL) method was employed during simulated annealing. The proposed Ising annealer comprises 1024 fully connected nodes, representing a scalable hardware implementation for solving large-scale combinatorial optimization problems. It has been experimentally evaluated and systematically verified across the device and system levels. The final results demonstrate that the system achieves significantly higher solution accuracy.
Conventional FP-CIMs suffer from fixed preserved bit-width (PBW), limiting their adaptability and efficiency. This work proposes the first MXFP-CIM macro enabling wide-range adaptive PBW, featuring: (1) A serial dual-bit-sliding scheme; (2) A harmless data mapping scheme with a hierarchical hidden-bit decoder; (3) An adjustable-PBW MXFP-MAC circuit via twin-stage allocation. The 28nm MXFP-CIM macro achieves a peak energy efficiency of 127.54TFLOPS/W in the MXFP6/6 mode.
The escalating parameter scale of neural networks intensifies the memory wall challenge in von Neumann architectures, spurring the exploration of non-volatile memory-based computing-in-memory (NVM-CIM) architectures. This work presents an spin-transfer toque(STT)-MRAM-based computing-in-memory (CIM) macro that addresses key challenges in supporting quantized neural networks, including bit-level sparsity variation, signed-number operations, and hardware-software co-optimization. The design introduces three key innovations: a multi-domian CIM macro that eliminates dequantization overhead, a sign-bit replication strategy for efficient signed-number computation, and a dynamically adjustable accumulation-length scheme for sparse workloads. Implemented in 28nm technology, the 192Kb CIM macro achieves 24.15 TOPS/W energy efficiency. When evaluated on MNIST with LeNet-5, the system maintains 94% accuracy while consuming 2.69 nJ per inference, representing a 2.6× improvement over full-digital scheme.
Spin-torque (SOT)-MRAM offers a promising solution for computer vision (CV) applications through its superior throughput and bit-cell energy efficiency. Meanwhile, redundant features lead to dominant energy consumption in memory-access peripherals and analog-to-digital converters (ADCs). Existing delta-sigma in-memory computing (Delta Sigma IMC) only reduces logic switching power, which has limited energy benefit for non-volatile memory (NVM)-based computing cores. This paper proposes an SOT-MRAM-based Delta-bias-Sigma computing macro to address the critical energy bottlenecks in CV applications. Voltage-delta-sensitive multiplier (VDM) realizes in-SOT-MRAM multiplication of incremental values, while eliminating redundant memory-access power caused by repetitive features. Delta-sensitive analog biasing (DS-biasing) scheme reduces the required ADC dynamic range, while delta-input-adaptive (DA) variable resolution ADC eliminates unnecessary quantization energy from repetitive inputs. Simulations based on 32kb SOT-MRAM computing macro demonstrate that the VDM achieves 6.45 & times; energy reduction over conventional MRAM compute unit, while the DS-biasing combined with DA-ADC achieves 1.8 & times; A/D conversion energy reduction compared with Delta Sigma IMC. The CV computing system achieves 546TOPS/W/b on MNIST image classification task with 97.9% accuracy, while improving 34% and 5% energy efficiency in video edge detection task and MobileViT inference task, respectively.
Recently, one-time-programmable (OTP) memory has been widely used in micro controller unit (MCU) due to its storage reliability and tamper-proof. Based on the analysis of the breakdown mechanism of magnetic tunnel junction and the measurement results used for modeling, this article demonstrates a 96-Kb versatile MRAM-OTP macro integrated with a 6-Kb OTP-based physical unclonable function (PUF) in 55-nm fully depleted silicon on insulator process. The MRAM-OTP macro realize minimum 3.0 V program voltage and 100 ns for 32-bit program speed. The on-chip power supply method trims with PVT variations and completes voltage switching between two modes. The multi-bit programming design saves about 97% time and 5% energy using self-termination. The discernible dual-mode sense amplifier saves 57.4% power in OTP reading with no lacks of sensing yield. We also presented the application scheme of the proposed MRAM-OTP macro within the trimming information storage and security MCU design, with the PUF helped to complete the authentication process together with OTP and the error correcting code assisted encrypting and correcting process. The trimming information can help to enhance the reliability of STT-MRAM main area. With the reconfigurable bit-cell design, the Inter-HD and Intra-HD of the OTP-based PUF can reach 49.77% and 0.08%, respectively.
Due to ambient energy’s inherent instability, intermittent computing is essential for task completion. This work comprehensively explores the spintronic flip-flop implementation in the open-source RISC-V platform. Magnetic tunnel junction (MTJ) has great potential for non-volatile flip-flop (NV-FF) implementation because of its high density, low read and write energy consumption, and compatibility with CMOS process. To the best of the authors’ knowledge, the checkpoint preservation is firstly supported in this work. The proposed non-volatile differential sampling latch (NV-DSL) achieves 7.39 fJ/bit data transfer energy consumption. The phased write strategy reduces write energy by 24.3%. A generalized NV-FF design methodology is further established, achieving a 68.88% area reduction. The power consumption of proposed non-volatile RISC-V processor is reduced by nearly 75%. When performing atomic tasks, the energy consumption and latency are reduced by 61.4% and 43.87%, respectively, compared with the cache scheme.
Recent large language models (LLMs), driven by the scaling law, have demonstrated remarkable performance in various machine learning tasks by significantly increasing model size. However, their substantial memory footprint and computation complexity present critical challenges for the deployment of on-device intelligence. Group-wise weight quantization has emerged as an effective technique to alleviate this issue. By adopting fine-grained group-wise quantization, low-bit-width inference can be efficiently supported through mixed-precision FP-INT matrix multiplication (mpGEMM). Nevertheless, existing software-hardware co-design approaches primarily focus on supporting FP-INT computations, neglecting critical bottlenecks: (1) the exponent alignment and group-wise quantization offset the improvement of low-bit-width quantization; (2) low-bit-width dot products are inefficiently processed due to limited and highly repetitive partial result patterns; (3) full-width activation delivery is required for stall-free execution, causing considerable register and memory access overhead. To address these limitations, we present LUTIC, a software-hardware co-design framework for accelerating quantized low-bit-width LLM inference. Specifically, we propose a novel computational paradigm called compute-in-decoding, which seamlessly integrates the decoding process of our proposed symmetric lookup format with incremental computation. Based on this, an encoded-compute and deferred-decode strategy is proposed to reduce additional FP processing overheads. Meanwhile, a dual LUT-based computation efficiently processes mpGEMM. Further, we develop a dedicated accelerator leveraging our proposed optimizations, featuring a ripple-counter-based processing element, a deferred-decode unit, and a fully pipelined mixed-precision dataflow. Extensive evaluations across a comprehensive set of LLM workloads and model scales, under a commercial 28nm process, demonstrate that LUTIC achieves up to 1.93x improvement in on-chip energy efficiency and 34.44% reduction in overall system energy consumption compared to state-of-the-art accelerators.
A 51.6 mu J/token accelerator for rotation-based dual-quantized LLMs is presented. A subspace-rotation method with parallel Hadamard transposer reduces on-chip rotation power by 62.3% and area by 59.7%. A fused scale-activation unit lowers energy by 61.5% vs. the naive FP design. Rearranged bit-slice LUT computation achieves 2.28x better energy efficiency compared to a direct bit-parallel MAC implementation, while supporting flexible bit-width. The chip reduces per-token energy by 32.6% over SOTA under an equal accuracy constraint.
The rapid advancement of artificial-intelligence (Al) models has increased demand for high-precision and energy-efficient edge-Al chips. Floating-point (FP) support is essential for high-precision neural-network (NN) training and inference; yet FP incurs higher energy and area overhead due to complex FP multiplication and accumulation (MAC) operations. Digital compute-in-memory (DCIM) and floating-point CIM (FP-CIM) [1]–[10] have emerged as promising techniques to improve energy efficiency with higher accuracy. Previous FP-CIM implementations [1]–[7] achieved good performance through various alignment schemes and computing processes. However, as illustrated in Figure 14.3.1, the implementation of a digital-domain FP-CIM faces several challenges: (1) the difficulty of balancing FP-computation precision and input reusability, as alignment operations are unfriendly to CIM structure; (2) a large performance loss or area overhead due to peripheral parallel-alignment schemes; and (3) huge digital-MAC dynamic-energy consumption due to low 2's-complement (2C) negative-weight sparsity, coupled with an additional sign-bit computation overhead in digital CIM. This work presents a hierarchical broadcast-alignment non-2's-complement-MAC (B-A-N2CMAC) FP-CIM macro, featuring (1) a broadcast input and embedded lightweight convertor structure to enable BF16/LNT8 MAC operations with an improved input reusability; (2) an embedded area-efficient adaptive-alignment scheme with a dual-bit serial MAC; and (3) a format-mixed N2CMAC to reduce dynamic circuit activity and signed computation overhead. A 28nm 64kb B-A-N2CMAC FP-CIM macro is fabricated to support FP-MAC operations using BF16 and INT8 representations. This CIM macro achieved an energy efficiency of 62.84TFLOPS/W for BF16 and 90.15TOPS/W for LNT8.
Heterogeneous 3D (H3D) integration has emerged as a promising solution for high-performance computing (HPC) with growing complexity. However, the vertical-stacked integration introduces thermal challenges that severely limit reliability and performance of the system. This paper presents a fine-grained cross-layer thermal management framework with module-level Dynamic Voltage and Frequency Scaling (DVFS) and task scheduling. Our approach implements lightweight machine learning (ML)-based temperature prediction supporting Long Short-Term Memory (LSTM) and Multilayer Perceptron (MLP) models. In addition, cooperative task scheduling is introduced to balance the workload. Our framework enables thermal control for Neural Processing Unit (NPU), Digital Computing-in-Memory (DCIM), and Analog Computing-in-Memory (ACIM), while predicting temperature trends to avoid reactive delays. Experimental evaluations using a 3D-augmented HotSpot simulator demonstrate that the proposed method reduces peak temperature by over 19% and temperature variation by 62%, with less than 2% performance overhead compared to the baseline, making it feasible for predictive thermal management in 3D systems.
In 3D integrated chips with TSV-based vertical inter-connects, data transmission through Through-Silicon Vias (TSVs) can induce bit-flip errors, posing reliability challenges for neural network accelerators. We analyze the fault tolerance of mixed-precision neural networks across various data types and bit positions, revealing substantial resilience to low-significance bit (LSB) flips and high vulnerability to most-significant bit (MSB) errors. Leveraging this observation, we propose a TSV-Aware Adaptive Fault-Tolerant Coding (TSV-AFTC) scheme that applies strong error correction to critical bits, approximate coding to non-critical data, and differential transmission for locally correlated data. This approach enables accuracy preservation while reducing hardware overhead and power consumption by 25.6% and 27.4%, respectively, compared to conventional full-bit-width error correction schemes.