Asymmetric light transmission at the nanoscale constitutes a fundamental challenge and a key goal in integrated photonic circuits. Polaritons in van der Waals (vdWs) materials offer a promising platform for such nanoscale light confinement and control. However, achieving complex and functional vdWs polaritons control across interfaces remains a challenge. Here, we demonstrate grating-driven on-chip asymmetric steering of phonon polaritons (PhPs) in hyperbolic vdWs crystal α-MoO3 bilayers that incorporate twisted stacks and engineered interfaces. By combining twist-induced dispersion engineering and polaritonic grating diffraction, we achieve tunable deflection and the asymmetric transmission of PhPs. PhP transmission can be switched between bidirectional and unidirectional modes by varying either the excitation frequency or the grating period. We further designed an asymmetric polaritonic lens that focuses forward-propagating PhPs while allowing normal backward transmission. These results provide a novel strategy for polaritonic steering and lay a solid foundation for designing on-chip optical isolators, diodes, and nonreciprocal routers.
This paper presents a digital Compute-in-Memory (CIM)-based macro tailored for edge Transformer acceleration. It has three features: 1) A Hierarchical Dual-Exponent Alignment scheme that separates outliers from normal values, achieving near-ideal accuracy with minimal overhead. 2) A Sign-Magnitude Input Serial Pipelined Aligner (SMSPA) reduces area and power by 41.1% and 34.1%. 3) A Shift-Concat based Sign-Inverse Adder Tree (SSAT) saves accumulation area and power by 25.9% and 18.2%. The macro achieves $24.85-80.59$ TFLOPS/W in FP8. Measurements demonstrate up to $6.66 \times$ improvements in energy efficiency, compared to state-of-theart post-alignment FP-CIM schemes, while maintaining high accuracy on different tasks.
The rapid advancement of artificial intelligence (AI) has been marked by the large language models exhibiting human-like intelligence. However, these models also present unprecedented challenges to energy consumption and environmental sustainability. One promising solution is to revisit analogue computing, a technique that predates digital computing and exploits emerging analogue electronic devices, such as resistive memory, which features in-memory computing, high scalability, and nonvolatility. However, analogue computing still faces the same challenges as before: programming nonidealities and expensive programming due to the underlying devices physics. Here, we report a universal solution, software-hardware co-design using structural plasticity-inspired edge pruning to optimize the topology of a randomly weighted analogue resistive memory neural network. Software-wise, the topology of a randomly weighted neural network is optimized by pruning connections rather than precisely tuning resistive memory weights. Hardware-wise, we reveal the physical origin of the programming stochasticity using transmission electron microscopy, which is leveraged for large-scale and low-cost implementation of an overparameterized random neural network containing high-performance sub-networks. We implemented the co-design on a 40nm 256K resistive memory macro, observing 17.3% and 19.9% accuracy improvements in image and audio classification on FashionMNIST and Spoken digits datasets, as well as 9.8% (2%) improvement in PR (ROC) in image segmentation on DRIVE datasets, respectively. This is accompanied by 82.1%, 51.2%, and 99.8% improvement in energy efficiency thanks to analogue in-memory computing. By embracing the intrinsic stochasticity and in-memory computing, this work may solve the biggest obstacle of analogue computing systems and thus unleash their immense potential for next-generation AI hardware.
Hybrid RRAM-SRAM Compute-in-Memory (CIM) architectures offer a promising solution for accelerating Transformer. However, their efficiency is severely constrained by a fundamental challenge: the prevalence of outlier values inherent to these models. These outliers create a severe quantization dilemma: hardware-friendly Block Floating-Point (BFP) suffers catastrophic accuracy loss, while Integer (INT) quantization incurs prohibitive hardware overhead. Furthermore, poor quantization directly undermines sparsity exploitation. The resulting numerical imprecision renders analog RRAM-CIM highly vulnerable to device variation noise, which is amplified during naive sparse computation. Addressing these deeply intertwined challenges requires a holistic co-design focused on robust data representation. We propose an outlier-aware hybrid CIM (OAH-CIM) accelerator built on a synergistic hardware-algorithm co-design. First, we introduce an outlier-aware BFP quantization that achieves near-FP32 accuracy at 5-bit by efficiently representing both outliers and non-outliers. Second, leveraging this high-fidelity representation, we propose balanced bit sparsity, a hardware scheduling technique that equalizes workloads to ensure reliable, low-noise sparse computation in variation-prone RRAM. Evaluations on DeiT-base show OAH-CIM improves 5-bit accuracy by 5.21% over BFP and 2.94% over INT. Furthermore, OAH-CIM achieves up to $8.6 \mathrm{TOPS} / \mathrm{W}$, yielding a $2.3-39.1 \times$ energy efficiency improvement over state-of-the-art CIM accelerators.
In-memory search has emerged as a promising solution for efficient vector discovery of the nearest neighbors in general-purpose vector databases. However, templated storage-in-array structure and VMM-based computational form of inmemory search pose challenges in supporting generic distance computations. In this work, we introduce a novel memristive in-memory similarity measure engine, MemSearch, for configurable distance calculations, including dot distance, ED, and CD. MemSearch highlights two aspects: data storage and distance computing. For data storage, we propose a Unified Similarity Element Mapping (USEM) scheme based on a pair array to accommodate various similarity calculations. For distance computing, we introduce a Reconfigurable Current Computing (RCC) circuit designed to process multiple arithmetic rules in similarity calculations, with a slightly increase of 4.4% and 9.9 % in energy consumption for ED and CD, respectively. We have tested various datasets with different modalities, including images, voice, human activity and text. Experimental results demonstrate that the MemSearch engine achieves improvements of $864 \times, 802 \times$, and $1474 \times$ in energy efficiency over CMOS-based engines for dot distance, ED, and CD calculations, respectively. The MemSearch engine highlights its potential for future highly efficient general-purpose in-memory vector databases.
Applications such as medical imaging, augmented and virtual reality, and embodied artificial intelligence (AI) depend on the ability to reconstruct complex signals from sparse observations. These applications are characterized by incomplete measurements and limited computational resources. Traditional approaches to digital hardware face the following challenges: explicit signal representations require heavy sampling and storage, data movement across the von Neumann bottleneck dominates energy and latency, and CMOS (complementary metal-oxide-semiconductor)-based circuits offer limited parallel efficiency. Here we present a software-hardware co-optimization framework for sparse-input signal reconstruction. At the software level, we use neural fields1 to implicitly represent signals using neural networks, which are further compressed by low-rank decomposition and structured pruning. At the hardware level, we design a resistive-memory-based computing-in-memory platform, featuring a Gaussian encoder and a multi-layer perceptron processing engine. The Gaussian encoder leverages the intrinsic stochasticity of resistive memory for efficient encoding, whereas the processing engine enables precise weight mapping through a hardware-aware quantization circuit. On a 40-nm 256 Kb resistive-memory macro, the system delivers 23.5×, 21.0× and 32.3× gains in projected energy efficiency, together with 10.8×, 38.8× and 6.2× gains in projected parallelism, for three-dimensional computed tomography sparse reconstruction, novel view synthesis and dynamic-scene novel view synthesis, without compromising on reconstruction quality. This work advances AI-driven signal reconstruction technology and paves the way for future efficient and robust medical AI and three-dimensional vision applications.
Threshold switching memristors are promising for neuromorphic hardware, yet conductive filament (CF) devices often suffer from poor reliability due to stochastic filament formation and uncontrolled active ion diffusion. This letter reports a vertical Ag/oxidized-ZrSe2/graphene/Au threshold switch in which an amorphous ZrOx matrix containing Ag2Se crystallites is formed from 2D ZrSe2 via in situ plasma oxidation and subsequent reaction with the Ag top electrode. The Ag2Se phase serves as a dynamic Ag-ion buffer, stabilizing filament formation and suppressing runaway diffusion. A graphene interlayer further improves the bottom-interface contact and mitigates local field concentration, enhancing switching uniformity. The device achieves a sub-0.204 mV/dec subthreshold swing, 3.7% threshold voltage coefficient of variation, stable operation from 100 pA to 10 mu A, and endurance beyond 10(7) switching cycles. These results underscore the functional role of oxidation-derived phases in two-dimensional materials and provide a materials/structure route toward reliable, low-power artificial neuronal devices.
Resistive memory (RM)-based computing-in-memory (CiM) accelerators provide a promising platform for low-power edge intelligence. However, practical edge AI requires not only efficient inference but also repeated model updates through class-incremental learning (CiL). Implementing CiL on RM substrates exposes a critical material-algorithm mismatch: the write-intensive updates required by CiL are strongly affected by the stochastic programming of filamentary RM devices. Conventional fully programming (FP) mitigates this stochasticity by repeatedly programming and verifying each cell to a precise target conductance, but this exhaustive procedure incurs substantial energy consumption, latency, and device wear across successive CiL stages. Herein, we propose a hardware-aware adaptive programming (AP) strategy that aligns CiL deployment with RM device physics. Microstructural and electrical analyzes reveal that the random spatial distribution of oxygen vacancies gives rise to unavoidable programming variability. Guided by this insight, AP does not attempt to eliminate intrinsic stochasticity through costly compensation. Instead, it updates only the most impactful weights to suppress accuracy loss caused by overall mapping errors, while leaving low-impact weights unchanged. This converts large-scale write-verify operations into targeted updates of a minimal subset of cells, reducing programming overhead without requiring device or material optimization. Validated on a hybrid analog-digital system with a 40 nm, 256 k RM-based CiM core, AP reduces programming energy by 93.0% and programming cycles by more than 90% relative to FP during five-stage CIFAR100 CiL, while achieving a final accuracy of 0.80, close to the 0.81 software baseline. For the more complex ShapeNet 3D point-cloud recognition task across eight stages, AP achieves 92.3% energy savings with only a 0.03 accuracy loss. Moreover, AP improves robustness by reducing programming-error-induced accuracy degradation by 87.7% and 88.4% for the two tasks, respectively. This work bridges algorithmic update requirements and physical programming constraints, enabling robust and energy-efficient lifelong learning on RM-based CiM platforms.
Dynamic matrix equations underpin real-time estimation, optimization and control in autonomous mobile robot (AMR) systems, where coefficient matrices must be repeatedly updated in response to changing sensory inputs and operating conditions. Memristor-based analogue computing-in-memory offers an energy-efficient approach to matrix operations, but its use in dynamic workloads is limited by analogue-computing errors and the high overhead of accurate matrix reprogramming. Here we report a memristor-based analogue iterative computing system (mAIC) for efficient and accurate dynamic matrix-equation solving. Using characterized memristor predictive models, we develop an analogue-compensated predictive one-step programming (AC-PoP) scheme that maintains accurate matrix updates without iterative write–verify operations, reducing update energy and latency by 101-fold and 345-fold, respectively. Integrated with a continuous-time analogue feedback solver and implemented using foundry-fabricated memristor macros, mAIC achieves software-comparable accuracy in representative SLAM and real-world robotic path-planning workloads. Compared with an NVIDIA Jetson platform, it provides 4.4-fold lower processing latency and 627.9-fold lower energy consumption. These results advance memristor-based computing from predominantly static matrix acceleration towards accurate and dynamically reconfigurable analogue computing.
The computational demands of modern Transformers pose a significant challenge to their energy-efficient deployment on edge devices. While RRAM-based computing in-memory (CIM) architectures alleviates the data movement bottleneck, the dynamic nature of the self-attention mechanisms limits RRAM’s efficiency and endurance. This work proposes a software-hardware co-design framework for the Fourier-based Transformer (FNet), which replaces attention with hardware-friendly Discrete Fourier Transform (DFT) operations. Our framework introduces two key optimizations: First, guided by a sensitivity analysis that reveals DFT modules are more susceptible to quantization than feed-forward networks, we employ a mixed-precision strategy augmented with hardware-aware finetuning to maintain high accuracy. Second, we propose a novel symmetric DFT mapping strategy that exploits the matrix’s inherent conjugate symmetry, reducing DFT-related RRAM storage and computation by nearly 50%. Evaluated on the GLUE benchmarks, our framework achieves accuracy comparable to 8-bit implementations while reducing power consumption by 2.96x. Critically, compared to state-of-the-art RRAM-based Transformer accelerators, our design demonstrates a 6.1x to 25.1x improvement in energy efficiency.
The memristive Hopfield neural networks (HNNs) can greatly accelerate the solving of combinatorial optimization problems. However, efficient annealing methods are required to prevent HNNs from converging to sub-optimal solutions. In this paper, we propose a novel scaling annealing (ScA) method and circuit design using the intrinsic stochasticity of the memristors to mitigate the local minima problem. Specifically, asynchronous updates are employed with the ScA method to prevent the hardware overhead from going out of control as the problem size grows. Based on the classical max-cut problems, simulation results show that on a 60-node problem, ScA finds the global minimum with 5.4× higher accuracy than the HNN without annealing. In addition, compared to the prior annealing schemes for memristive HNNs, the ScA method exhibits a better overall solution performance, and achieves 9.5× reduction in the annealing circuit area overhead for a 60-node problem.
Resistive memory compute-in-memory accelerators provide energy efficient analogue matrix vector multiplication for neural network inference, but frequent reprogramming of analogue weights remains costly because of device variability and iterative write and verify operations. This limitation hinders their use in edge model adaptation, including approximate machine unlearning and continual learning, where model parameters may need to be updated repeatedly in response to data deletion requests or newly arriving tasks. Here we present a co-design approach across hardware and software that maps frozen pretrained weights to analogue resistive memory arrays while placing trainable low rank adaptation branches in SRAM connected digital compute. By using LoRA style parameter efficient updates, the proposed scheme confines adaptation to a small set of digital parameters and avoids repeated reprogramming of the analogue backbone. To our knowledge, this work provides the first experimental demonstration of approximate machine unlearning on a fabricated resistive memory CIM accelerator. We validate the framework on a 180 nm 128x128 1T1R resistive-memory macro for face recognition, and through circuit-accurate simulations for speaker authentication and stylized image generation tasks, owing to the substantial model sizes involved. Compared with a baseline that directly updates analog weights, our hybrid mapping reduces analog training/update cost by up to 148x, on-chip deployment overhead by up to 388x, and inference energy by up to 59x, while preserving competitive task performance. These results show that hybrid analogue-digital LoRA mapping can enable efficient post-deployment adaptation on RM-CIM hardware, although formal machine-unlearning guarantees and large-scale system integration remain open challenges.
Memristive computing offers a promising route toward energy‐efficient neural inference, yet its practical efficiency is limited by intensive crossbar computation and the wide output current range required by analog‐to‐digital converters (ADCs). Here, an energy‐aware perturbation framework is presented to jointly alleviate these bottlenecks. The proposed method combines sinusoidal perturbation encoding with dual‐threshold screening to reshape layer inputs, suppress redundancy‐dominated activations, and reduce both pulse‐driven crossbar activity and ADC input range. For practical deployment, a low‐cost layer‐wise Bayesian search strategy optimizes perturbation parameters offline on a small subset, which are then fixed during inference. An 8 kb (64 × 128) 1T1R memristor‐array chip and hardware platform are developed for validation. On a hardware‐implemented MNIST convolutional neural network, the framework compresses the output current range by 2.02× while a limited accuracy drop from 95.4% to 93.8%. Additional evaluations demonstrate system energy reductions of 26.55% on ResNet‐18/CIFAR‐10 and 26.90% on VGG‐16/CIFAR‐100, with accuracy losses of 1.94% and 1.80%, respectively. On keyword spotting, selective perturbation compresses the output current range by 3.4× and reduces crossbar energy by 69.4%, while accuracy changes from 88.7% to 87.2%.
Abstract Neuronal ion channels have well-established effects on synaptic plasticity, in many cases by influencing pathways that depend on membrane excitability. Here we find that a C. elegans two-pore domain potassium (K2P) channel, TWK-40, regulates presynaptic organisation through a membrane potential-independent mechanism. Instead, this mechanism depends on TWK-40’s effects on intracellular potassium levels. Loss-of-function mutations in TWK-40 lead to excessive presynaptic protein accumulation, while gain-of-function mutations lead to depleted presynaptic components and cause synaptic transmission deficits. These abnormalities are phenocopied by transporter mutations that mimic TWK-40’s effects on intracellular potassium concentration, but not by sodium channel mutations that mimic its effects on membrane excitability. This indicates that cytoplasmic potassium promotes presynaptic assembly. This process depends on the PYK-1 pyruvate kinase, a potassium-sensitive enzyme, and three transcription factors. These findings establish a new pathway linking neuronal potassium homeostasis to the control of presynaptic organisation and synaptic function.
This work presents a hardware-algorithm co-designed framework for neuromorphic computing, enabling efficient supervised learning in spike-based neural architectures. First, synaptic updates are reformulated as low-rank outer products of forward spike vectors and backward error gradients via singular value decomposition (SVD), enabling direct parallelization on 1T1R arrays. Second, a stochastic computing scheme replaces conventional sequential updates with probabilistic pulse-driven modulation, achieving one-step full-matrix synaptic updates. Third, gradient stabilization techniques mitigate training instability in deep SNNs by addressing silent neuron and gradient explosion issues. Evaluated on the ASL-DVS dynamic gesture recognition task, the framework maintains 84.7% accuracy with hardware-realistic 1T1R characteristics, while drastically reducing hardware update steps. This demonstrates a synergistic hardware-algorithm co-design where SVD-based approximation enables parallelization, stochastic computing achieves one-step updates, and gradient stabilization ensures trainability, advancing practical neuromorphic intelligence for edge sensing systems.
Phase-change memory has attracted extensive attention due to its high storage density, low-energy consumption, and high-speed operation characteristics. Recently, 3D Vertical PCM based on the Ring switching model has demonstrated ultralow switching energy and high integration density. However, the actual switching region and the underlying microscopic switching mechanism remain insufficiently understood. In this work, a localized nanoscale switching model is proposed and experimentally supported for 3D Vertical PCM. The abnormal SET behavior with an ultralong trailing edge of up to 10 μs suggests that the crystallization process cannot be fully explained by the conventional Ring switching model. Three-dimensional electrothermal simulations indicate that the localized nanoscale switching model exhibits a more localized thermal response and better agreement with the experimentally observed transient characteristics. TEM and SAED characterizations further reveal clear phase boundaries between localized crystalline and amorphous regions inside the device structure, providing microscopic evidence supporting localized switching behavior. The results suggest that the ultralong trailing edge during the SET process originates from nucleation-limited crystallization within the nanoscale switching region. Based on this mechanism, a Two-stage SET pulse is further proposed, reducing the SET operation time by 66% and the SET energy consumption by 40.7% while maintaining stable memory window and cycling endurance characteristics. This work further clarifies the localized switching mechanism in 3D Vertical PCM and provides new insights for low-power and high-speed 3D-PCM devices.
Abstract Harvesting low-grade heat from the environment and converting it into electricity holds the potential to power devices independent of cables or batteries. However, their effectiveness is limited by weak ion selectivity and insufficient concentration gradients. Here, we introduce the use of a calix[4]pyrrole as effective anion traps to selectively capture Fe(CN)6 4– and Cl− anions, enabling simultaneous modulation of redox ion distribution and suppression of anion mobility under a temperature gradient. This strategy combines desolvation-induced entropy gain with thermodiffusion enhancement arising from the mobility asymmetry between cations and anions. This leads to a synergistic boost in thermopower to an impressive 8.1 mV K−1, and results in a 20-fold increase in output power compared to the PVA/Fe(CN)6 3–/4– system. Demonstrated through a proof-of-concept wearable device with 36 unipolar elements, our system generated nearly 3 volts under ambient conditions. This strategy offers a promising route toward thermoelectric materials with enhanced thermopower for efficient harvesting low-grade thermal energy.
Chalcogenide phase-change material (Ge2Sb2Te5) has attracted considerable research interest in recent years for tunable structural color applications, owing to its pronounced contrast in optical properties between amorphous and crystalline states. However, the strong absorption of conventional Ge2Sb2Te5 has significantly limited its applicability in the near-infrared (NIR) wavelength range. Herein, we utilize the near-zero absorption characteristic of the emerging phase-change material Sb2Se3 in the NIR band to construct a Fabry-Pérot cavity structure. This structure not only enables resonance peaks with high transmittance and large modulation range, but also allows a multi-channel filter to be achieved by increasing the thickness of Sb2Se3. This tunable multi-channel spectrally selective filter presents significant application potential in advanced fields such as multi-gas sensing, laser beam combining, and adaptive multi-spectral imaging.