This letter presents a 3D-NAND-like Two-Transistor-One-Capacitor (2T1C) 3D-DRAM architecture, which features Indium-Gallium-Zinc-Oxide (IGZO) Vertical-Channel-Access-Transistors (VCAT), vertical bit-line (VBL) wiring and multi-deck stackable capability with back-end-of-line (BEOL) monolithic integration. The vertical IGZO channel, formed simultaneously by atomic layer deposition (ALD) in recessed sidewall cavity, ensures uniform performance across layers, circumventing the degradation associated with layer-by-layer processing and enabling low-bit-cost scaling. We experimentally validated a 2-layer 2T1C structure, characterizing the VCAT electrical properties, read/write operations characteristics, and a retention time of 15.7s. 3D-TCAD simulations further assessed the parasitic capacitance and sense margin of the 2T1C array, confirming the excellent potential of this 3D-DRAM architecture for ultra-high-density Gb-scale DRAM applications.
In this work, we report a demonstration of complementary field-effect transistor (CFET) and circuits with p-type carbon nanotube field-effect transistor (CNT-FET) stacked on n-type InGaZnOx field-effect transistor (IGZO-FET) at a low thermal budget of 250(degrees)C. To address the issues of hydrogen resistance and thermal stability faced by top-gate IGZO-FETs, a Y2O3 passivation layer on Indium Gallium Zinc Oxide (IGZO) was developed for the first time, which yielded > 100 x reduction in the threshold voltage shift ( triangle V-th ) during the subsequent atomic layer deposition (ALD). This enabled the fabrication of enhancement-mode top-gate IGZO-FETs with drastically enhanced hydrogen resistance. Moreover, by engineering the top-gate device structure and air annealing, the high on/off ratio was preserved. Furthermore, p-type back-gate CNT-FETs were fabricated on top to fulfill a backend-of-the-line (BEOL)-compatible CNT/IGZO CFET with symmetric P/NFET performance. CFET logic gates and ring oscillators (ROs) with sub-50 ns stage delay at micrometer-scale channel length were also demonstrated. The novel passivation method developed in this work could enable the fabrication of high-performance and high-reliability CNT/IGZO-based CFET and circuits for future 3-D integration.
Hafnium-based oxide ferroelectric memories have garnered significant attention due to their exceptional advantages, such as ultra-low operating voltages, nanosecond-level switching speeds, backend process compatibility, and potential for miniaturization. Over the past decade, research has focused on optimizing growth processes, electrode materials, and interlayers primarily to enhance the proportion of the ferroelectric O−phase, reduce coercive fields and imprint effects, and improve polarization and durability. The impact and mechanisms of antiferroelectric phase transitions have also been widely studied. This paper presents, for the first time, the variations and impacts of the antiferroelectric phase in HZO during device wake-up and fatigue processes, based on electrical characteristics and simulations using the adjusted NLS model. Phenomena such as wake-up and fatigue are influenced by the presence of the antiferroelectric phase. However, the proportion of the antiferroelectric phase does not necessarily directly reflect on the device's electrical performance, as its effects may be overshadowed by the more pronounced ferroelectric phase. That offering a new perspectives and approaches in the study of phase transitions and their effects in HZO devices.
Computational spectrometers, which rely on reconstruction algorithms to decode spectral information from raw sensor data, are of potential use in portable, in-field spectrometry. However, research on such systems primarily focuses on the front-end encoding devices, and back-end decoding hardware remains limited by severe overheads. Here we report an in situ computational spectrometer implemented on a fully integrated 576-Kb memristor chip. With systematic robustness analysis, we develop memristive regularization and filter embedding strategies to overcome the extreme sensitivity of ill-posed spectral reconstruction, achieving software-equivalent accuracy. System-level benchmarking shows that our hardware takes only 125.0 ns to reconstruct one spectrum consuming 6.7 nJ of energy, which is 26.5 times faster and 162.7 times more energy-efficient than state-of-the-art computational spectrometers. Our work illustrates the potential of memristor-chip-based computational spectrometry and provides approaches for efficiently implementing signal processing algorithms on memristor chips.
This letter presents an aging-robust 32-MHz RC frequency reference based on a frequency-locked-loop (FLL). With a temperature compensation scheme that combines BJTs and aging-robust diffusion resistors, the FLL achieves +/- 1550-ppm inaccuracy from -40 (degrees) C to 125 (degrees) C after batch calibration and a low-cost 1-point trim, which increases to +/- 2350-ppm after accelerated aging. Due to the extensive use of dynamic error-correction techniques, the FLL also achieves a state-of-the-art Allan deviation floor of 0.4 ppm.
Augmented reality (AR) contact lenses emerge as a promising immersive platform offering seamless, eyeintegrated augmented experiences. In this work, for the first time, we present M3D-FAR, a prototype Monolithic 3D integration chip featuring Flexible Augmented Reality contact lenses with three interconnected functional layers on the same flexible substrate: $1^{\text{st}}$ layer of $8 \times 8 \text{Ta}_{2} \mathrm{O}_{5}$-based resistive random-access memory (RRAM) array for digital computing-in-memory (CIM), $2^{\text{nd}}$ layer of carbon nanotube (CNT) CMOS circuits for logic and data interface, and $3^{\text{rd}}$ layer of $\text{InGaZnO}_{\mathrm{x}}$ field-effect transistors (IGZO-FET) for 8×24 photosensor array and display driver circuits. The structural integrity and proper function of the M3D-FAR chip was validated by structural analysis and electrical measurements. Furthermore, system-level benchmark in a typical image-based interactive task shows that the M3D-FAR architecture could achieve $10.45 \times$ speed-up compared to its 2D counterpart and consume $5.39 \times$ lower energy than GPU.
Memristor-based convolutional neural network (CNN) accelerators have gained considerable attention due to their low latency and high energy efficiency, making them promising candidates for edge acceleration. Alongside statically stored model weights, dynamically generated intermediate feature maps during inference occupy a significant portion of the on-chip buffer capacity and directly affect the efficiency of hardware pipeline execution. However, there is a lack of theoretical analysis and methods for efficiently allocating on-chip buffer for feature map data. To address this gap, this article implements three innovative aspects. First, a mathematical model is developed to estimate the minimal buffer size required for pipelined inference of CNNs on memristor-based accelerators, offering accurate and swift evaluations of the buffer size requirement. Second, based on this model, the article establishes mathematical conditions for buffer requirements to maintain a blocking-free pipeline during CNN inference, providing theoretical guidance for on-chip buffer allocation strategies. Third, a simulation-in-loop optimization method is proposed to further reduce latency by efficiently increasing the buffer size of critical layers. To validate our proposed model and method, evaluations were conducted on five representative models: ResNet-18, ResNet-50, YOLO-v5, U-Net, and Faster-RCNN-FPN. The results reveal a remarkably low average estimation error of only 2.6% between the mathematical model and the experimentally measured results, with the maximum error still below 10%. Moreover, our simulation-in-loop optimization strategy achieved significant latency reductions ranging from 5.3% to 57.5% across the five models.
This paper presents an embedded resistive random-access memory (eRRAM) macro fabricated on a commercial 40nm CMOS technology platform. Through a comprehensive product-level data validation, we first established a baseline HfO2-based RRAM that optimizes the trade-off between the forming and pre-cycling times during the chip probing (CP) stage. Subsequently, a novel dual-oxide resistive switching layer (RSL), enabled by an optimized atomic layer deposition (ALD) process, was developed. This dual-oxide RSL significantly enhanced the retention performance of the RRAM macro, achieving a 38.9% improvement in the post-bake memory window on the wafer level without degrading other performance metrics. The optimized 40nm RRAM macro demonstrates commercial-grade performance for the display driver integrated circuit (DDIC) application, including endurance >1K cycles, data retention >10 years at 85°C, read disturb immunity >109 cycles, and passes the solder reflow test at 260°C for 900s according to JEDEC standard [1]. Therefore, this work provides a highly competitive embedded non-volatile memory (eNVM) solution and promotes the commercialization of RRAM technology.
This letter presents an energy-efficient dynamic amplifier. It utilizes source-coupled input boosting and time-domain differential sampling techniques to boost the effective input signal by 4x compared to its floating inverter amplifier (FIA) prototype without noise or power penalties. With discharge-based dynamic biasing, the bandwidth (BW) and power of the amplifier can be scaled by 100x. Fabricated in a standard 0.18-mu m CMOS technology, the amplifier achieves a state-of-the-art power efficiency factor (PEF) of 0.19, which is 16x better than that of a standard FIA. It also achieves a scalable BW/power range from 0.5 kHz/2.3 nW to 50 kHz/206 nW.
Resistive random-access memory (RRAM) has been extensively studied as a promising candidate for high-density computing-in-memory (CIM) applications, yet its multi-level cell (MLC) precision is fundamentally compromised by conductance relaxation-an intrinsic instability driven by spontaneous oxygen vacancy $(V_{\mathrm{O}})$ diffusion. This work demonstrates a pre-cycle induced relaxation suppression (PIRS) technique on a 40 nm 16 Kb RRAM macro to thermodynamically stabilize the conductive filaments (CFs). We elucidate that the pre-cycle operation refines the CFs into enhanced “hourglass” morphology, effectively pruning unstable conduction paths. Quantitative analysis using relative deviation (RD) and relaxation velocity $(v_{\text{relax}})$ metrics reveals that pre-cycle significantly mitigates relaxation effect, securing a robust 3-bit/cell MLC with superior thermal stability $\left(25-125^{\circ} \mathrm{C}\right)$. System-level benchmark across image classification (e.g., DeiT-Tiny) and reconstruction (e.g., CAE) tasks confirms that the proposed scheme suppresses the accuracy loss by $\sim 5 \times$ while improving the peak signal-to-noise ratio (PSNR) by up to $\mathbf{8 d B}$. This approach resolves a critical reliability bottleneck for future edge AI applications.
This article introduces a circuit topology for building reconfigurable, high-accuracy CMOS current and voltage references, namely, current-locked-loop (CLL). The CLL regulates the $V_{\mathrm {GS}}$ of a current source transistor and thus its current output using a digitally assisted feedback loop, thereby enabling higher order temperature coefficient (TC) compensation and good reconfigurability simultaneously. By directing the current to an on-chip resistor, the CLL can be used as a voltage reference with a different set of digital calibration coefficients. Fabricated in a standard 180-nm CMOS technology, the prototype chip occupies an active chip area of 0.35 mm2 and consumes ~ $220~\mu $ A from a minimum 1.6-V supply. After two-point and batch calibration, it achieves ±0.027%/±0.056% voltage/current inaccuracy ( $3\sigma $ ) over the automotive temperature range from $- 40~^{\circ }$ C to $125~^{\circ }$ C, which improves the state-of-the-art by $2\times $ and $2.5\times $ , respectively. Moreover, the CLL provides a digital temperature readout with a $3\sigma $ inaccuracy of $\pm 0.06~^{\circ }$ C without further calibrations.
Edge AI urgently demand computing hardware with large on-chip storage capacity, low memory access overhead. Resistive random-access memory (RRAM) based digital computing-in-memory (DCIM) suits such scenarios by achieving a optimized performance balance in calculation precision, storage density and access overhead. However, traditional RRAM-DCIM scheme suffers from two issues: 1) RRAM multi-bit storage boosts density but introduces significant errors, while conventional ECC scheme requires massive redundant bits that offset density gains; 2) Wiring RC parasitism in large-scale arrays leads to slow and energy-intensive random access. To resolve these issues, the proposed computation-oriented DCIM (CO-DCIM) paradigm integrates two core solutions: a hierarchical multi-bit storage coding scheme and a charge-domain continuous access scheme. The designed CO-DCIM macro is fabricated in a 28nm CMOS technology. Measurement results show a single RRAM read latency of 2.5 ns, normalized read energy of 0.12 pJ/bit, and a storage density of 12.85 Mb/mm2. The macro is further evaluated on public SimpleAR-0.5B-RL and SimpleAR-1.5B-RL workloads, showing a 29.2x reduction in mapped-weight access energy relative to LPDDR5X DRAM at $256\times 256$ resolution.
Privacy-preserving data analysis is essential in health care applications to safeguard sensitive patient information while enabling medical monitoring and diagnostics. However, existing solutions generally separate security from analysis modules and memory from computation units, creating hardware and energy overheads that constrain their use in resource-limited medical devices. Here, we introduce the memristor-based colocated authentication and processing (CLAP) system, which achieves security-analysis integration through embedding physical unclonable functions within compute-in-memory architecture. To resolve the incompatibilities between these two features, we propose a differential stochastic mapping method by applying information theory principles. We demonstrate CLAP on a 130-nanometer memristor chip, validating its versatility across diverse information processing tasks. In an electrocardiogram data collection task, CLAP achieves device authentication with an area under the curve of 99.46% and efficient signal compression with a software-level percentage root mean square difference. CLAP demonstrates 146.0-fold energy efficiency gain and 17.6-fold area reduction, providing intrinsically secure hardware solutions that enhance both privacy preservation and computational efficiency for health care applications.
The rapid growth of artificial intelligence (AI) is increasingly constrained by fundamental hardware bottlenecks in computation throughput and energy efficiency. Bioinspired computing (BIC) offers a promising alternative by emulating the intrinsic advantages of biological systems, such as parallelism, adaptability, and robustness. Progress in BIC hardware demands interdisciplinary convergence that bridges materials science and device physics with neuroscience, computer science, mathematics, and information science. Therefore, the development of this cross-disciplinary field urgently requires a comprehensive roadmap that analyzes systematically and in-depth the frontier issues and the latest progress. In this roadmap, we categorize the critical challenges into three components: hardware foundations, architectures, and prototype realizations. We highlight how biological features inspire the design of BIC hardware through device physics and discuss their performance metrics and engineering challenges. We then describe how diverse signaling rules and structural organizations in BIC architectures support specific computational prototypes, including electronic and photonic BIC chips, and present a technological roadmap that outlines opportunities to expand the functional scope of BIC hardware through coordinated advances in devices, architectures, and system demonstrations. This ongoing convergence of interdisciplinary knowledge can help accelerate the shift toward high-efficiency AI hardware.
For the first time, we have developed a Drain stress-induced Reliability EnhAncement Method (DREAM) for IGZO-FETs, demonstrating significantly improved positive bias temperature instability (PBTI) reliability across a wide temperature range from 77 K to 373 K. Specifically, the ΔVth of IGZO-FETs is reduced from 210 mV to 40 mV in PBTI measurements respectively (Eox=3 MV/cm for tstress =103 s at T=300 K). By applying DREAM to an IGZO-based monolithic 3D (M3D) system design, the performance and reliability is enhanced, where the typical neural network classification accuracy can be improved by over 10%.
Deep learning has revolutionized modern society but faces growing energy and latency constraints. Deep physical neural networks (PNNs) are interconnected computing systems that directly exploit analog dynamics for energy-efficient, ultrafast AI execution. Realizing this potential, however, requires universal training methods tailored to physical intricacies. Here, we present the Physical Information Bottleneck (PIB), a general and efficient framework that integrates information theory and local learning, enabling deep PNNs to learn under arbitrary physical dynamics. By allocating matrix-based information bottlenecks to each unit, we demonstrate supervised, unsupervised, and reinforcement learning across electronic memristive chips and optical computing platforms. PIB also adapts to severe hardware faults and allows for parallel training via geographically distributed resources. Bypassing auxiliary digital models and contrastive measurements, PIB recasts PNN training as an intrinsic, scalable information-theoretic process compatible with diverse physical substrates.
In recent years, artificial intelligence (AI) has experienced rapid development, and high performance computing (HPC) has raised increasingly higher demands for hardware computational capacity. Resistive random-access memory (RRAM)-based computing-in-memory (CIM) technology is expected to overcome the bottleneck of memory wall and provide HPC solutions. However, CIM chips face critical thermal challenges, including severe hotspot formation and thermally induced performance degradation, due to increasing power density and strong data-space coupling effects. Existing thermal management solutions designed for conventional digital chips are not directly applicable to CIM architectures. In this work, we propose a comprehensive framework for thermal analysis and management tailored to CIM chips. Targeted strategies are developed across the design, pre-operation, and operation stages. During the design stage, it is essential to mitigate thermal-induced accuracy degradation by adopting optimized design strategies. During the pre-operation stage, we propose a latency-thermal co-optimization (LTCO) strategy for static thermal management. By combining LTCO with a genetic algorithm to optimize the neural network mapping scheme, we reduce the hotspot temperature by 6.5°C and the temperature standard deviation by 5.3°C, without increasing the latency. During the operation stage, we develop a dynamic thermal management (DTM) strategy tailored for RRAM-based CIM chips, considering their unique architecture and the coupling between data and space. The evaluation results show that when thermal management is triggered, the combination of LTCO and DTM achieves more than a 10
Objective With the growing demand for on-orbit information processing in satellite missions,efficient deployment of neural networks under strict power and latency constraints remains a major challenge.Resistive Random Access Memory(RRAM)-based Compute-in-Memory(CIM)architectures provide a promising solution for low power consumption and high throughput at the edge.To bridge the gap between conventional neural architectures and CIM hardware,this paper proposes NAS4CIM,a Neural Architecture Search(NAS)framework tailored for RRAM-based CIM chips.The framework proposes a decoupled distillation-enhanced training strategy and a Top-k-based operator selection method,enabling balanced optimization of task accuracy and hardware efficiency.This study presents a practical approach for algorithm-architecture co-optimization in CIM systems with potential application in satellite edge intelligence. Methods NAS4CIM is designed as a multi-stage architecture search framework that explicitly considers task performance and CIM hardware characteristics.The search process consists of three stages:task-driven operator evaluation,hardware-driven operator evaluation,and final architecture selection with retraining.In the task-driven stage,NAS4CIM employs the Decoupled Distillation-Enhanced Gradient-based Significance Coefficient Supernet Training(DDE-GSCST)method.Rather than jointly training all candidate operators in a fully coupled supernet,DDE-GSCST applies a semi-decoupled training strategy across different network stages.A high-accuracy teacher network is used to guide training.For each stage,the teacher network provides stable feature representations,whereas the remaining stages remain fixed,which reduces interference among candidate operators.Knowledge distillation is critical under CIM constraints.RRAM-based CIM systems typically rely on low-bit quantization and are affected by device-level noise,under which conventional weight-sharing NAS methods show unstable convergence.Feature distillation from a strong teacher network ensures clear optimization signals for candidate operators and supports reliable convergence.After training,each operator is assigned a task significance coefficient that quantitatively reflects its contribution to task accuracy.Following the task-driven stage,a hardware-driven search stage is performed.Candidate network structures are constructed by combining operators according to task significance rankings and are evaluated using an RRAM-based CIM hardware simulator.System-level hardware metrics,including inference latency and energy consumption,are measured.Complete network structures are evaluated directly,capturing realistic effects such as array partitioning,inter-array communication,and Analog-to-Digital Converter(ADC)overhead.From hardware-efficient networks with superior performance,the selection frequency of each operator is analyzed.Operators that appear more frequently in low-latency and low-energy designs are assigned higher hardware significance coefficients.This data-driven evaluation avoids inaccurate operator-level hardware modeling and reflects system-level behavior.In the final stage,task significance and hardware significance matrices are integrated.By adjusting weighting factors,the framework prioritizes accuracy,efficiency,or a balanced trade-off.Based on the combined evaluation,an optimal operator set is selected to construct the final network architecture,which is then retrained from scratch to refine weights and further improve accuracy while maintaining high hardware efficiency on CIM platforms. Results and Discussions NAS4CIM is evaluated on FashionMNIST,CIFAR-10,and ImageNet to demonstrate effectiveness across tasks of different scales.On FashionMNIST,the framework achieves 90.1%Top-1 accuracy in the accuracy-oriented search and an Energy-Delay Product(EDP)of 0.16 in the efficiency-oriented search(Table 4).Real-chip experiments on fabricated RRAM macros show close agreement between measured accuracy and simulation results,confirming practical feasibility.On CIFAR-10,NAS4CIM reaches 90.5%Top-1 accuracy in the accuracy-oriented mode and an EDP of 0.16 in the efficiency-oriented mode,exceeding state-of-the-art methods under the same hardware configuration.Under a balanced accuracy-efficiency setting,the framework produces a network with 89.3%accuracy and an EDP of 0.97(Table 3).On ImageNet,which represents a large-scale and more complex classification task,NAS4CIM achieves 70.0%Top-1 accuracy in the accuracy-oriented mode,whereas the efficiency-oriented search yields an EDP of 504.74(Table 5).These results indicate effective scalability from simple to complex datasets while maintaining a favorable balance between accuracy and energy efficiency across optimization settings. Conclusions This study proposes NAS4CIM,a NAS framework for RRAM-based CIM chips.Through a decoupled distillation-enhanced training method and a Top-k-based operator selection strategy,the framework addresses instability in random sampling approaches and inaccuracies in operator-level performance modeling.NAS4CIM provides a unified strategy to balance task accuracy and hardware efficiency and demonstrates generality across tasks of different complexity.Simulation and real-chip experiments confirm stable performance and consistency between algorithmic and hardware evaluations.NAS4CIM presents a practical pathway for algorithm-hardware co-optimization in CIM systems and supports energy-efficient,real-time information processing for satellite edge intelligence.
Memristor-based analogue computing in memory (CIM) offers revolutionary gains in energy efficiency and computing power for data-intensive applications such as artificial intelligence. However, it typically struggles with achieving high accuracy at the same time, owing to the noise-sensitive nature of analogue computing and the non-ideal characteristics at the device and circuit levels that inevitably result in computing errors. Although progress has been made in device engineering and hardware-algorithm co-optimization to mitigate the error and parasitic effects, many of these advances inadvertently incur a hardware or energy consumption overhead, undermining the core benefits of analogue CIM. This Review dissects the computing error sources across the CIM hierarchy from the memristor device and array to the system architecture and algorithm, and evaluates the strategies to minimize those errors. We highlight the material and device innovations, array-level techniques and algorithm-architecture co-design frameworks towards high-accuracy analogue CIM. By dissecting the trade-off between computing accuracy and implementation cost, this Review draws a roadmap for translating memristor-based analogue CIM technology from proof-of-concept prototypes to large-scale deployment for accelerating next-generation artificial intelligence.