The development of computing-in-memory (CIM) has gone through the era of analog domain CIM, which pursued extreme energy efficiency and area efficiency, and the era of digital domain CIM, which pursued extreme computational precision. Currently, research is exploring hybrid domain CIM designs that combine the advantages of both analog and digital approaches. In order to find a better balance point between computational precision and hardware efficiency, this work proposes a bit-rotated hybrid domain CIM design which primarily employs: 1) a bit-rotated feature-in hybrid structure to achieve the optimal boundary segmentation (better computational precision) with lower hardware overhead; 2) an embedded sign-bit-processing SRAM array, enabling sign-bit extension with low hardware overhead; 3) an multi-bit-fusion scheme for low-power multi-bit quantization; and 4) a dual-granularity cooperative quantizer based on the charge domain, supporting low-power multi-cycle quantization. A 64 kb bit-rotated hybrid-domain CIM chip was fabricated with TSMC 28 nm. It supports efficient hybrid multiply-accumulate (MAC) operations and achieves 21.04–67.8 TOPS/W in energy efficiency and 0.41–1.57 TOPS/mm2 in area efficiency, which can be used to accelerate various CNN and Transformer based inference tasks.
The current source model (CSM) serves as an advanced timing model that accurately captures the current behavior of standard cells, offering superior precision over traditional non-linear delay models (NLDM) but facing significant challenges for high accuracy, large data volume, and expensive simulation cost. Most prior works focus on characterization methodology for NLDM, lacking the consideration of pronounced nonlinear and dynamic effects and inherent correlation among current-voltage waveform in CSM. To address these issues, we propose a dual-attention enhanced framework for multi-corner CSM characterization, where feature- and time-level attention mechanisms are implemented to enable accurate and robust prediction of irregular time vector and dynamic current vector respectively with Gaussian Process Regression (GPR) and Long Short-Term Memory (LSTM) networks. Experimental validation was conducted under TSMC 22nm process over 216 PVT corners for CSM waveform prediction, demonstrating excellent accuracy with 1.8% and 2.1% error for cell delay and output transition and 79.0% simulation effort reduction.
We introduce FIXME, the first end-to-end and large-scale benchmark for evaluating Large Language Models (LLMs) in hardware design functional verification (FV). Comprising 747 tasks derived from real-world hardware designs, FIXME spans five core FV sub-sets: specification comprehension, reference model generation, testbench generation, assertion design, and RTL debugging. To ensure high data quality, we developed an AI-human collaborative framework for agile data curation and annotation. This process resulted in 25,000 lines of verified RTL, 35,000 lines of enhanced testbenches, and over 1,200 SystemVerilog Assertions. Furthermore, through expert-guided optimization within the multi-agent aided flow, we achieved a remarkable 45.57% improvement in average functional coverage, underscoring the benchmark's robustness. Through evaluation of state-of-the-art LLMs like GPT-4.1, FIXME identifies key limitations and provides actionable insights, advancing the potential of LLM-driven automation in hardware design functional verification.
Conventional FP-CIMs suffer from fixed preserved bit-width (PBW), limiting their adaptability and efficiency. This work proposes the first MXFP-CIM macro enabling wide-range adaptive PBW, featuring: (1) A serial dual-bit-sliding scheme; (2) A harmless data mapping scheme with a hierarchical hidden-bit decoder; (3) An adjustable-PBW MXFP-MAC circuit via twin-stage allocation. The 28nm MXFP-CIM macro achieves a peak energy efficiency of 127.54TFLOPS/W in the MXFP6/6 mode.
Introduction.With the rapid development of trans-former-based large language models(LLMs)and deep neural networks(DNNs),the demand for both high computational throughput and massive memory capacity has grown exponen-tially[1-4].
Digital compute-in-memory (DCIM), particularly floating-point CIM (FPCIM) has emerged as a promising technique to enhance energy efficiency with higher accuracy for artificial intelligence (AI) applications. However, previous works incurred substantial hardware overhead for FP multiplication and accumulation (MAC) and failed to leverage the bit-level sparsity across data formats. This work presents SNAP-FPCIM, an embedded-serial-alignment, non-2’s-complement-MAC and adjustable computation precision FPCIM based on 6T static random-access memory (SRAM). The contributions of this work include: 1) a broadcast input and embedded light-convertor structure to enable BF16/INT8 MAC operations with improved input reusability; 2) an embedded area-efficient serial-alignment scheme with dual-bit-serial MAC; 3) a format-mixed N2CMAC flow to reduce circuit dynamic activity and signed computation overhead; 4) a mode-reconfigurable design to accommodate diverse computational precision requirements. A fabricated 28nm 64Kb SNAP-FPCIM macro supports FP-MAC operations of BF16 and INT8, achieving an energy efficiency of 62.84TFLOPS/W in BF16 mode and 90.15TOPS/W in INT8 mode.
Recent advances in generative AI and large-scale neural networks have drastically increased computational and memory bandwidth demands, exacerbating the classic von Neumann memory wall. Computing-in-memory (CIM) has emerged as a promising remedy by collocating storage and MAC operations inside memory arrays. Meanwhile, floating-point (FP) CIM is attracting growing interest for training and diffusion models, but most existing designs rely on coarse group-alignment of exponents, incurring quantization and accumulation errors that prevent deployment in high-accuracy workloads. This work presents a 28 nm word-alignment FP-CIM macro with a four-stage MAC pipeline that supports CNN training and diffusion models. The macro integrates three key techniques: (1) a word-aligned FP MAC flow with pipelined exponent comparison, mantissa shift, and partial-sum accumulation; (2) a low-bit-skip Wallace-tree BF16 multiplier that removes negligible partial products to improve power, area, and speed; and (3) a DFF-based double-function shift unit with a replica-clock scheme for low-latency and PVT-robust barrel shifting. Fabricated in 28 nm CMOS, the 16 Kb macro operates from 0.45–0.9 V and achieves 37.2 TFLOPS/W at 0.45 V and 2.56 TFLOPS/mm$^{2}$ at 0.9 V. Chip-level demonstrations obtain 72.96% top-1 on VGG16/CIFAR100, 76.14% on ResNet18/ImageNet, 57.95% mAP on YOLOv8/COCO, and 10.57 FID on a CelebA diffusion model, all within 0.34% of FP32 baselines. Overall, the proposed word-aligned FP-CIM improves the energy–area–accuracy figure of merit by 8.23–26.15× over prior FP-CIM macros.
Conventional FP-CIMs suffer from fixed preserved bit-width (PBW), limiting their adaptability and efficiency. This work proposes the first MXFP-CIM macro enabling wide-range adaptive PBW, featuring: (1) A serial dual-bit-sliding scheme; (2) A harmless data mapping scheme with a hierarchical hidden-bit decoder; (3) An adjustable-PBW MXFP-MAC circuit via twin-stage allocation. The 28nm MXFP-CIM macro achieves a peak energy efficiency of 127.54TFLOPS/W in the MXFP6/6 mode.
This paper conducts in-depth research on the process of controlled dielectric breakdown (CDB). Exploring the different mechanisms of dielectric breakdown process of silicon nitride thin films under different electrolyte solutions. By comparing the efficiency of pore formation and noise analysis of nanopore under different experimental parameters, this paper reports a method for preparing low-noise solid-state nanopore via CDB. In addition, we verify the function of the nanopores prepared by CDB through $\lambda$-DNA translocation experiments.
As an emerging type of AI computing accelerator, SRAM Computing-In-Memory (CIM) accelerators feature high energy efficiency and throughput. However, various CIM designs and under-explored mapping strategies impede the full exploration of compute and storage balancing in SRAM-CIM accelerator, potentially leading to significant performance degradation. To address this issue, we propose CIM-Tuner, an automatic tool for hardware balancing and optimal mapping strategy under area constraint via hardware-mapping co-exploration. It ensures universality across various CIM designs through a matrix abstraction of CIM macros and a generalized accelerator template. For efficient mapping with different hardware configurations, it employs fine-grained two-level strategies comprising accelerator-level scheduling and macro-level tiling. Compared to prior CIM mapping, CIM-Tuner's extended strategy space achieves 1.58$\times$ higher energy efficiency and 2.11$\times$ higher throughput. Applied to SOTA CIM accelerators with identical area budget, CIM-Tuner also delivers comparable improvements. The simulation accuracy is silicon-verified and CIM-Tuner tool is open-sourced at https://github.com/champloo2878/CIM-Tuner.git.
The growing complexity of modern hardware designs has rendered traditional functional verification increasingly time-consuming, with verification costs now dominating the design cycle. While large language models (LLMs) show promise in automating testbench generation, existing approaches struggle with real-world scalability, suffering from poor comprehension of long specifications and complex designs. To address these challenges, we propose ChatTest, a novel, end-to-end, multi-agent LLM framework for coverage-aware, agile hardware verification. Our key innovation lies in a function-mapped, divide-and-conquer architecture that integrates a Verification Description Language (VDL)—a structured, LLM-friendly DSL for precise specification encoding—with Constraint-Aware Segmental Adaptation (CASA) to enable coherent processing of long, heterogeneous design documents. By leveraging retrieval-augmented generation and supervised fine-tuning using multi-hierarchical specification-code alignment, ChatTest ensures accurate translation of functional points into targeted test stimuli. Furthermore, we introduce a coverage-driven feedback loop for automated test augmentation. Evaluated on a new benchmark of 20 complex RTL designs (up to 31K tokens of specification and 4K line-of-code), ChatTest achieves 1.46× higher toggle coverage and 2.28× higher line coverage than SOTA, with a 24.23% improvement in functional coverage, demonstrating its effectiveness in accelerating verification convergence.
As chip manufacturing processes advance to deep submicron nodes, parasitic interconnect effects increasingly dominate the performance of analog and mixed-signal (AMS) circuits and often lead to costly layout iterations. This makes early-stage estimation of parasitic capacitance and resistance important for parasitic-aware design exploration before full physical implementation. However, progress on GNN-based parasitic modeling has been hindered by the lack of public, high-fidelity RC benchmarks that support reproducible evaluation. To address this gap, we introduce ParasGB, the first open-source benchmark suite for pre-layout parasitic parameter prediction on circuit graphs. ParasGB provides large-scale, heterogeneous RC networks extracted with commercial EDA tools from tape-out-proven designs, together with a unified evaluation protocol covering node-level ground capacitance, edge-level resistance, and edge-level coupling capacitance. Within this framework, we benchmark diverse GNN architectures using a standardized training pipeline and expose challenges such as extreme label imbalance, long-tailed parasitic distributions, and strong structural heterogeneity. By establishing a physically grounded and standardized benchmark for early-stage parasitic prediction, ParasGB provides an open platform for reproducible research on circuit graph learning and parasitic-aware model development. All datasets, preprocessing scripts, and configurations are publicly available in our code repository https://github.com/ShenShan123/ParasGB.git.
The energy and thermal management systems of hybrid electric vehicles (HEVs) are inherently interdependent. With the ongoing deployment of intelligent transportation systems (ITSs) and increasing vehicle connectivity, the integration of traffic information has become crucial for improving both energy efficiency and thermal comfort in modern vehicles. To enhance fuel economy, this paper proposes a novel traffic-aware hierarchical integrated thermal and energy management (TA-ITEM) strategy for connected HEVs. In the upper layer, global reference trajectories for battery state of charge (SOC) and cabin temperature are planned using traffic flow speed information obtained from ITSs. In the lower layer, a real-time model predictive control (MPC)-based ITEM controller is developed, which incorporates a novel Transformer-based speed predictor with driving condition recognition (TF-DCR) to enable anticipatory tracking of the reference trajectories. Numerical simulations are conducted under various driving cycles and ambient temperature conditions. The results demonstrate that the proposed TA-ITEM approach outperforms conventional rule-based and MPC-SP approaches, with average fuel consumption reductions of 56.36% and 5.84%, respectively, while maintaining superior thermal regulation and cabin comfort. These findings confirm the effectiveness and strong generalization capability of TA-ITEM and underscore the advantages of incorporating traffic information.
Industrial chip development is inherently iterative, favoring localized, intent-driven updates over rewriting RTL from scratch. Yet most LLM-Aided Hardware Design (LAD) work focuses on one-shot synthesis, leaving this workflow underexplored. To bridge this gap, we for the first time formalize $Δ$Spec-to-RTL localization, a multi-positive problem mapping natural language change requests ($Δ$Spec) to the affected Register Transfer Level (RTL) syntactic blocks. We propose RTLocating, an intent-aware RTL localization framework, featuring a dynamic router that adaptively fuses complementary views from a textual semantic encoder, a local structural encoder, and a global interaction and dependency encoder (GLIDE). To enable scalable supervision, we introduce EvoRTL-Bench, the first industrial-scale benchmark for intent-code alignment derived from OpenTitan's Git history, comprising 1,905 validated requests and 13,583 $Δ$Spec-RTL block pairs. On EvoRTL-Bench, RTLocating achieves 0.568 MRR and 15.08% R@1, outperforming the strongest baseline by +22.9% and +67.0%, respectively, establishing a new state-of-the-art for intent-driven localization in evolving hardware designs.
SRAM-based compute-in-memory (CIM) offers high computational density and energy efficiency for deep neural network (DNN) accelerators, but its limited capacity causes on/off-chip data movement overhead for large DNN models. Existing CIM accelerator studies typically assume that DNN models fit entirely on-chip, leaving efficient dataflow design largely untapped. This paper introduces AccelCIM, a systematic dataflow exploration framework for SRAM CIM accelerator, which addresses two key limitations of prior work. (1) It formulates a systematic dataflow design space spanning CIM macro configurations and macro-array organizations. (2) It introduces rigorous design evaluation using cycle-accurate architectural simulation and post-layout PPA analysis. We conduct an extensive design space exploration and apply AccelCIM to representative LLM applications, providing practical insights for the principled design of CIM accelerators.
We propose an edge multioperator computing-in-memory (EMO-CIM) design that supports variable vector-wise multiply-and-accumulate (MAC) in CNN, Depthwise (DW)-Convolution, and Attention operators. It features: 1) a single EMO-CIM bank (ECB) excels in variable vector-wise MAC (V-MAC) for multioperators; 2) merging local input-shared compute units (LISCUs) with a decode-unit and adder-tree (DUAT) facilitates input/stationary-data (SD) similarity-aware computing (SAC) to improve energy and area efficiency (EF). Based on a row-wise vector separable accumulation (RVSA) strategy, the Attention/CNN task is optimized via multivector parallel crossover computation (MVPCC), while the DW task is optimized through the partial-convolution-first (PCF) computation. EMO-CIM achieves peak energy efficiencies of 93.5 and 39.3 TOPS/W for Attention/CNN and DW tasks, respectively, and area efficiencies of 4.971 and 3.728 TOPS/mm2 for the same tasks under INT8 MAC operations.
While Large Language Models (LLMs) demonstrate immense potential for automating integrated circuit (IC) development, their practical deployment is fundamentally limited by restricted context windows. Existing context-extension methods struggle to achieve effective semantic modeling and thorough multi-hop reasoning over extensive, intricate circuit specifications. To address this, we introduce ChipMind, a novel knowledge graph-augmented reasoning framework specifically designed for lengthy IC specifications. ChipMind first transforms circuit specifications into a domain-specific knowledge graph (ChipKG) through the Circuit Semantic-Aware Knowledge Graph Construction methodology. It then leverages the ChipKG-Augmented Reasoning mechanism, combining information-theoretic adaptive retrieval to dynamically trace logical dependencies with intent-aware semantic filtering to prune irrelevant noise, effectively balancing retrieval completeness and precision. Evaluated on an industrial-scale specification reasoning benchmark, ChipMind significantly outperforms state-of-the-art baselines, achieving an average improvement of 34.59% (up to 72.73%). Our framework bridges a critical gap between academic research and practical industrial deployment of LLM-aided Hardware Design (LAD).
A 51.6 mu J/token accelerator for rotation-based dual-quantized LLMs is presented. A subspace-rotation method with parallel Hadamard transposer reduces on-chip rotation power by 62.3% and area by 59.7%. A fused scale-activation unit lowers energy by 61.5% vs. the naive FP design. Rearranged bit-slice LUT computation achieves 2.28x better energy efficiency compared to a direct bit-parallel MAC implementation, while supporting flexible bit-width. The chip reduces per-token energy by 32.6% over SOTA under an equal accuracy constraint.
The rapid advancement of artificial-intelligence (Al) models has increased demand for high-precision and energy-efficient edge-Al chips. Floating-point (FP) support is essential for high-precision neural-network (NN) training and inference; yet FP incurs higher energy and area overhead due to complex FP multiplication and accumulation (MAC) operations. Digital compute-in-memory (DCIM) and floating-point CIM (FP-CIM) [1]–[10] have emerged as promising techniques to improve energy efficiency with higher accuracy. Previous FP-CIM implementations [1]–[7] achieved good performance through various alignment schemes and computing processes. However, as illustrated in Figure 14.3.1, the implementation of a digital-domain FP-CIM faces several challenges: (1) the difficulty of balancing FP-computation precision and input reusability, as alignment operations are unfriendly to CIM structure; (2) a large performance loss or area overhead due to peripheral parallel-alignment schemes; and (3) huge digital-MAC dynamic-energy consumption due to low 2's-complement (2C) negative-weight sparsity, coupled with an additional sign-bit computation overhead in digital CIM. This work presents a hierarchical broadcast-alignment non-2's-complement-MAC (B-A-N2CMAC) FP-CIM macro, featuring (1) a broadcast input and embedded lightweight convertor structure to enable BF16/LNT8 MAC operations with an improved input reusability; (2) an embedded area-efficient adaptive-alignment scheme with a dual-bit serial MAC; and (3) a format-mixed N2CMAC to reduce dynamic circuit activity and signed computation overhead. A 28nm 64kb B-A-N2CMAC FP-CIM macro is fabricated to support FP-MAC operations using BF16 and INT8 representations. This CIM macro achieved an energy efficiency of 62.84TFLOPS/W for BF16 and 90.15TOPS/W for LNT8.
Field-programmable gate array (FPGA) macro placement holds a crucial role within the FPGA physical design flow since it substantially influences the subsequent stages of cell placement and routing. With the increasing number of macros and the complex cascade shape and region constraints imposed by modern FPGAs, the routability and macro placement have become much more challenging. In this article, we propose an effective and efficient routability-driven macro placement algorithm for modern FPGAs with cascade shape and region constraints. To reserve adequate space for cell placement and guarantee routability, we first develop a routability-driven mixed-size analytical global placement that evenly distributes both macros and cells while considering cascade shape and region constraints. Particularly, the proposed global placement engine integrates a well-trained congestion prediction model, targeting benchmarks with high routing congestion to enhance overall routability. Then, we propose an integer linear programming (ILP)-based cascade shape legalization followed by matching-based macro legalization to remove macro overlaps while satisfying the region constraints. Finally, a routability-driven detailed macro placement is proposed to refine the solution. Compared with the winners of the MLCAD 2023 FPGA macro placement contest and state-of-the-art works, experimental results show that our algorithm achieves the best overall score and routability.