
Large language models (LLMs) have demonstrated remarkable performance across diverse applications, driving an explosive growth in text data for training. This trend poses critical challenges to text data storage and transfer, highlighting the need for efficient lossless text compression systems. However, mainstream text compression schemes are generally limited by the memory-wall problem in traditional von Neumann architectures, as frequent memory access increases energy consumption and latency. In this work, we propose an in-memory text compression system based on resistive random-access memory content-addressable memory (RRAM CAM). The nonvolatile storage and high integration density of RRAM support on-chip large dictionaries for improved compression ratio, while CAM-based in-memory parallel search reduces data movement and enables energy-efficient and low-latency text compression. The proposed system is validated through a demonstration using a fabricated 1Mb 40nm 1-transistor–1-RRAM (1T1R) chip. Evaluation of a representative dictionary-matching task shows a $343.7\times $ improvement in search energy efficiency and a $13.8\times $ reduction in search latency compared with a CPU-based scheme, showing the potential of the proposed system for future in-memory compression.
This paper presents a 16-channel stimulator ASIC for prosthetic sensory feedback that incorporates a compact charge balancing method utilizing an initialization phase to preset the voltage across a series blocking capacitor. The stimulator enforces strict per-cycle charge balance between the total sourced and sunk charge without any residual current, regardless of mismatch between the current sources. To further enhance performance, a speed-boosting feedback loop is introduced to monitor the anodic phase current and temporarily increase the stimulation amplitude, enabling higher stimulation rates. The ASIC was designed and fabricated in a 130-nm BCD CMOS process. For a 1 mA stimulation current, it achieves a 99.96% reduction in unbalanced residual charge compared with a conventional, uncorrected blocking-capacitor-based approach. Incorporating the speed-boosting feedback loop yields a maximum reduction in residual voltage of 21% across pulse rates from 200 Hz to 1 kHz, and a maximum reduction of 30.7% at 200 Hz when the stimulation amplitude is increased from 1 to 4 mA. In vitro measurements with electrodes in saline yield a residual voltage of only 1.04 mV, demonstrating the effectiveness of the proposed charge balancing method.
This study proposes an energy-efficient edge-based visual odometry accelerator that achieves 130 fps for visual localization. The accelerator employs a programmable compute-near-memory architecture with a SIMD computing paradigm, offering low power consumption and versatility. The multi-precision reconfigurable adder and power-of-2 shifter are customized for vectorized multi-precision processing with low hardware overhead. To enhance the performance, we distill the essential operators that are optimally suited to the computing architecture for efficient PicoVO deployment. The operators are implemented using simple algorithms built on the limited available arithmetic elements. Critical tasks in PicoVO, namely edge detection and the Levenberg-Marquardt solver, are realized on-chip without reliance on external memory. These tasks are further optimized thanks to the programmability. Compared to ARM Cortex-M7 processors, the chip demonstrates $26\times $ speedup and $266\times $ energy saving for edge detection. Furthermore, it achieves $4.7\times $ speedup and $53\times $ energy saving for Levenberg-Marquardt solver. The fabricated chip achieves 130 fps throughput at 22.6mW power and 0.127 mJ per frame. Our experimental board shows that the accelerator enables real-time and low-power visual localization for low-end platforms based on lightweight microcontrollers.
PbS colloidal quantum dot (CQD) infrared sensors offer advantages such as low cost, high resolution, and compatibility with monolithic integration, making them promising for spectral analysis, industrial inspection, night vision, and other infrared sensing applications. Multimodal perception that combines intensity imaging and dynamic vision has therefore become an important direction for next-generation infrared sensors. However, conventional systems face two major limitations: device dark current degrades image quality, while the use of separate sensors for different modalities creates significant challenges for precise spatial alignment. To address these challenges, this paper presents a readout circuit capable of simultaneous intensity and dynamic vision detection. An adaptive in-pixel background reduction (BR) circuit is introduced to substantially suppress device dark current, and a $3\times 3$ computing-in-memory (CIM) scheme is implemented to remove isolated noise events in situ. Furthermore, a $64\times 64$ PbS CQD infrared image sensor was fabricated by monolithically integrating CQDs with the proposed $0.13~\mu $ m CMOS readout circuit. The sensor reduces the mean dark-current density to 14.1% and the dark signal nonuniformity (DSNU) to 11.6% of their original levels. Simultaneously, it effectively suppresses the majority of dynamic event noise at the front end, dropping the background event rate to 5% of its initial value. The sensor achieves peak standalone intensity and event frame rates of 500 Hz and 2140 Hz, respectively, and a 427 Hz mixed-readout rate, with each cycle comprising one intensity readout and five event readouts. By enabling rapid, pixel-level fusion of co-located data, this work paves the way for efficient and robust intelligent multimodal perception systems.
A fully synthesizable, area-optimized time-to-digital converter (TDC) with improved linearity is described. Its architecture exploits a cyclic inverter-based delay line (ring) tapped with two alternating cycle counters. A compact delay ring enables reduced hardware cost and efficient digital post-processing. The delay ring linearity is enhanced through back-annotated insertion of optional dummy inverter capacitive loads. To estimate the average inverter propagation delay, a straightforward calibration scheme is proposed. The TDC prototype fabricated in 180 nm CMOS occupies 4500 $\mu $ m2 and consumes a mean power of 2.7 mW at 1.8 V and room temperature. It achieves 52 ps resolution over a 3.2 ns full-scale range, with maximum DNL/INL of 0.38/0.7 LSB and measured single-shot precision of up to 16 ps or 0.31 LSB. Due to its low complexity, excellent PPA, and automated standard-cell-based implementation, the presented architecture is suitable for applications that require numerous digital timing channels.
This paper presents a dB-linear variable gain amplifier (VGA) based on an active-RC feedback filter topology with a ramp-based control scheme. As a core building block of a closed-loop amplifier-based VGA, a self-biased current-reuse feedforward (FF) compensated operational transconductance amplifier (OTA) is incorporated. The proposed FF-OTA achieves a unity-gain bandwidth (UGB) of 2.92GHz and a DC gain of 36.25dB while consuming only 3.1mW. To ensure smooth gain scaling, our proposed VGA incorporates a tunable ramp-based control that eliminates discrete steps typically introduced by switched-resistor transitions. This approach enables stable dB-linear operation and yields highly linear performance, achieving an OIP3 of 22.7dBm at the highest gain setting, outperforming conventional open-loop VGAs. Owing to the ramp generator’s flexibility, the VGA supports a wide 0.5V control range and achieves an effective dB-linear region spanning 25dB within a total gain range of 28dB, with a 0.4dB gain error. The prototype circuit, based on a single-stage active-RC filter, operates at a constant bandwidth of 410MHz, and the VGA core draws 6mW from a 1.2V supply. Fabricated using TSMC 65nm CMOS technology, the VGA occupies a compact core area of 0.037mm2.
This paper presents a robust, area-efficient frequency-locked loop (FLL)-based frequency reference utilizing a series–parallel double-tracking (SPDT) technique. To achieve second-order temperature compensation and enhanced process-variation tolerance, a hybrid compensation structure employing only two types of resistors is proposed, and its theoretical design considerations are thoroughly analyzed. The proposed SPDT technique induces process-tracking temperature coefficient (TC) shifts under process deviations, thereby substantially mitigating process sensitivity without complex trimming networks. Fabricated in a 180-nm standard CMOS process, the prototype occupies a compact active area of 0.024 mm2. Experimental results across 15 samples demonstrate an average TC of 7.6 ppm/°C ( $\sigma =0.7$ ppm/°C) over a wide temperature range of $- 40~^{\circ }$ C to $125~^{\circ }$ C after TC1 trimming only. Furthermore, long-term stability is validated through accelerated aging, showing a well-maintained TC of 9.6 ppm/°C. It operates at a nominal frequency of 32 MHz, consuming $44.1~\mu $ W under a 1.2 V supply, and exhibits a low line sensitivity of 0.12%/V.
We propose a low-complexity digital carrier recovery algorithm that supports multi-modulation formats and is implemented on the field-programmable gate array (FPGA) platform. Based on an iterative fast Fourier transform (IT-FFT) architecture and a polar-coordinate blind phase search mechanism, the algorithm significantly reduces the number of multiplication operations by optimizing the phase estimation process, thus effectively lowering the hardware implementation complexity. Asimplified CORDIC algorithm is proposed, which combines CORDIC pre-convergence with small-angle linear approximation to effectively reduce computational complexity and hardware resource consumption while ensuring accuracy. Meanwhile, an additional feedback loop is introduced to enhance the phase tracking performance of the system under high frequency drift scenarios and improve the robustness of the algorithm. To verify the engineering feasibility of the algorithm, a hardware implementation of 32 GBaud signal processing is completed based on the FPGA platform, and the bit error rate performance between floating-point simulation and fixed-point hardware implementation is compared. Experimental results show that the performance loss of the fixed-point hardware implementation compared with floating-point simulation is negligible, which fully demonstrates its practicability and effectiveness in high-speed optical communication systems and can provide a reliable method for real-time processing of high-speed optical communication systems.
Designing high-frequency monolithic microwave integrated circuit (MMIC) power amplifiers (PAs) typically relies heavily on designer experience and iterative manual tuning, particularly when multiple matching networks and nonlinear device interactions are involved. This work presents an artificial-intelligence-(AI)-assisted design framework for automated synthesis and optimization of multi-stage MMIC PAs. In the proposed methodology, the circuit architecture and bias conditions are specified based on domain knowledge, while the component values of the matching networks are automatically determined through a surrogate-assisted optimization workflow. A specification-driven S-parameter-assisted data-sampling-space compression (S-DSSC) strategy is introduced to iteratively refine the feasible design region using low-cost small-signal simulations, enabling efficient design-space exploration with limited training data. Kolmogorov-Arnold networks (KANs) are employed as surrogate models and differential evolution (DE) is used to perform global optimization within the refined parameter space. The framework is validated through the design and fabrication of Ka-band MMIC power amplifiers in a 150-nm Gallium-Nitride-on-Silicon-Carbide (GaN-on-SiC) process, including a three-stage power cell and a balanced amplifier. Measurement results show a peak saturated output power (Psat) of 33.4 dBm, a peak power-added efficiency (PAE) of 39.4%, and more than 30 dB small-signal gain across 17.3–21.2 GHz, demonstrating the effectiveness of the proposed AI-assisted PA design methodology.
The development of computing-in-memory (CIM) has gone through the era of analog domain CIM, which pursued extreme energy efficiency and area efficiency, and the era of digital domain CIM, which pursued extreme computational precision. Currently, research is exploring hybrid domain CIM designs that combine the advantages of both analog and digital approaches. In order to find a better balance point between computational precision and hardware efficiency, this work proposes a bit-rotated hybrid domain CIM design which primarily employs: 1) a bit-rotated feature-in hybrid structure to achieve the optimal boundary segmentation (better computational precision) with lower hardware overhead; 2) an embedded sign-bit-processing SRAM array, enabling sign-bit extension with low hardware overhead; 3) an multi-bit-fusion scheme for low-power multi-bit quantization; and 4) a dual-granularity cooperative quantizer based on the charge domain, supporting low-power multi-cycle quantization. A 64 kb bit-rotated hybrid-domain CIM chip was fabricated with TSMC 28 nm. It supports efficient hybrid multiply-accumulate (MAC) operations and achieves 21.04–67.8 TOPS/W in energy efficiency and 0.41–1.57 TOPS/mm2 in area efficiency, which can be used to accelerate various CNN and Transformer based inference tasks.
Binary and ternary weights substantially reduce model storage and weight-memory traffic in large language models. However, conventional GEMM independently reduces each output column, repeatedly generating candidate partial dot products for the same activation row and thus underusing ultra-low-bit computation. We present BAPE, a precomputation-based FPGA accelerator core for binary- and ternary-weight Transformer linear layers. BAPE schedules and retains candidates by complete activation row. Each activation group’s candidate partial dot products are generated once, indexed and reused across output columns by streamed weight patterns, and accumulated across groups on chip. The Precomputation Unit (PCU) and dual-mode Adder Tree Processing Engine (ATPE) reuse the same hierarchical adder tree for candidate generation and lookup-based accumulation. Joint exploration of precomputation granularity $\alpha $ and resource-sharing parallelism $\beta $ balances candidate-space size, parallel utilization, routing, timing, and power. Implemented on a Xilinx ZCU102 FPGA, BAPE supportsW1.58A8/W1.58A4 BitNet-b1.58-2B-4T andW1A8/W1A4 Bonsai-4B. At the 300-MHz target, it achieves 614.40–2457.60 equivalent GOPS and 9.31–37.23 GOPS/kLUT. During model-level on-board inference with these four configurations, the measured dynamic power is 5.443–6.478W, corresponding to a projected energy efficiency of 112.88–426.52GOPS/W. Relative to same-precision GPU quantized inference, the maximum ACC loss is 0.351 percentage points, and the absolute per-dataset PPL change across the 24 paired results is at most 0.11%. These results show that BAPE efficiently executes low-bit linear computation for binary- and ternary-weight Transformers while preserving model inference quality.
This paper proposes a reference-less clock and data recovery (CDR) architecture that employs two-unit-interval (2-UI) windowing for frequency acquisition. The proposed 2-UI-based frequency detector (FD) enables half-rate frequency tracking without frequency-detection errors even in the presence of duty-cycle distortion (DCD) in the received data. Conventional Extended Bang-Bang Phase Detector (XBBPD) suffers from sampling errors near data transition edges due to variations in the reference interval for frequency detection. This results in false frequency-up/down (FUP/FDN) signals and may lead to prolonged stabilization time. Moreover, conventional XBBPD-based FD requires careful tuning of the FLL gain and lock-detection tolerance to avoid false lock or lock release under DCD. To address these limitations, the proposed FD employs six sampling points over a 2-UI window, effectively avoiding DCD-affected regions while maintaining a constant reference interval, thereby reducing the probability of mis-detections during frequency tracking. The prototype, fabricated in a 28-nm FD-SOI process, tolerates up to 11% duty-cycle distortion without frequency detection errors. The proposed reference-less CDR achieves a minimum lock time of $1.5\mu $ s and an energy efficiency of 1.57 pJ/bit at 14 Gb/s.
This paper presents a low-complexity, resource-efficient streaming architecture for Channel Estimation in Zero-Padded Orthogonal Time Frequency Space (ZP-OTFS) systems. Building on an established guard-based linear interpolation algorithm, the novelty of this work lies in its streaming hardware architecture: the proposed design reconstructs the channel impulse response in the delay-time domain through a partially parallel interpolation scheme, enabling continuous processing while exploiting the sparsity of the OTFS channel matrix to reduce computational overhead. The architecture is described at the RTL level in VHDL and validated through FPGA-in-the-Loop experiments on an AMD ZCU111 board. Experimental results show a latency of $44.10~\mu s$ , a dynamic power consumption of $17~mW$ , and an energy cost of $1.45~\mu J$ per frame, confirming the feasibility of real-time and power-efficient OTFS channel estimation in high-mobility scenarios. The fixed-point implementation closely matches the floating-point reference, with only slight BER degradation for QPSK, 16-QAM, and 64-QAM. REPN and EVM are evaluated for EPA, EVA, and ETU, while REPN is also reported for UHM, where degradation is attributed to a Doppler phase-rotation validity limit. An ASIC-oriented 16nm FinFET assessment further indicates that the logic-only post-layout result reports 52,093 placed instances, 26, $384.590~\mu \mathrm {m}^{2}$ of standard-cell area, and $173.72~mW$ total power at $1~GHz$ and $0.88~V$ . These results confirm that the design is suitable for integration into fully hardware-based OTFS receivers.
The growing demand for physical AI is expanding the capabilities of autonomous systems; however, these capabilities require accurate trajectory estimation and precise 3D map reconstruction for understanding the surrounding environment. In this context, LiDAR Odometry and Mapping (LOAM) has gained attention as a widely adopted SLAM framework that offers dense geometric map reconstruction. Nevertheless, the large number of keypoints captured per frame makes real-time hardware acceleration of LOAM extremely challenging in mobile robots, and even high-performance CPUs/GPUs fail to meet real-time constraints. This paper presents a real-time and energy-efficient heterogeneous processor with four key architectural features to address the dominant memory and computational bottlenecks of LOAM acceleration: 1)a 3D spherical-bin partitioning with neighboring bin search tailored to LiDAR point distributions, reducing the number of comparisons by 99.6%; 2)a hash-page-based memory management unit that lowers the required on-chip memory capacity by 96.9% through prediction-based dynamic memory allocation exploiting spatio-temporal coherence; 3)a neighboring-bin cache that achieves a $10.5\times $ improvement in kNN energy efficiency by reducing redundant memory accesses and increasing PE utilization; and 4)a reconfigurable non-linear optimization core that shortens NLO latency by 92.3% through multi-level computational parallelization. Fabricated in 28-nm CMOS, the proposed processor achieves 36.9ms latency and 0.81mJ/frame energy consumption, demonstrating its capability for real-time LOAM acceleration in mobile systems.
This paper presents a buck-boost converter with dual-triggered adaptive-cycle (DTAC) control for thermoelectric energy harvesting. The proposed DTAC control dynamically determines the switching period according to the energy conditions at both the input and output nodes, while adaptively adjusting the on-time based on the input energy condition to maintain the energy extraction point close to the maximum power point. Based on a single-inductor multiple-output (SIMO) architecture, the converter regulates two output voltages with the assistance of an energy storage capacitor, $C_{\text {STO}}$ , which recovers the residual inductor energy during the discharge phase. In addition, a low-power hybrid zero-current detector (H-ZCD) is implemented to accurately detect the inductor zero-current point. The proposed converter was fabricated in a 180-nm CMOS process with an active area of 0.701 mm2. Measurement results show that the converter operates over an input voltage range of 0.1–0.5 V and supports a load power range of 0.01–6.6 mW. It achieves a peak conversion efficiency of 85.3%, while the use of $C_{\text {STO}}$ reduces the output voltage ripple by up to 40%. The implemented H-ZCD consumes only 800 nW.
Surface electromyographic (sEMG)-based hand gesture recognition (HGR) is a promising technique for AI glasses and VR devices control. However, its practical deployment is mainly restricted by recognition accuracy degradation caused by sEMG signal variability across individuals and time. Although a few sEMG HGR processors have proposed online learning to address the issue, they either achieve insufficient accuracy improvement or introduce significant energy overhead, making them difficult to be implemented on edge devices. In this work, a hand gesture recognition processor is proposed based on a state-of-the-art cross-domain error minimization (CDEM) learning method. The processor solves the problem with three key features: 1) a hybrid-label CDEM (HL-CDEM) online learning architecture to enable low-energy continuous CDEM learning in non-cross-day scenarios; 2) an inference-learning integrated reconfigurable (ILIR) architecture to decrease hardware resources for recognition model inference and online learning; 3) a composite-confidence-driven adaptive two-stage feature extraction (CCFE) architecture to decrease CNN inference energy consumption. Implemented and evaluated on the Artix-7 FPGA board, the processor achieves the energy consumption per recognition of $4.16\mu $ J, and SOTA inter-subject accuracy of 83.90% on the Ninapro DB1 dataset. Implemented with 40nm CMOS technology, the processor achieves a lower energy consumption of 49.75 nJ per recognition, lower than those of SOTA cross-domain recognition processors.
This paper presents a critical-timing-relaxed pre-decision feedforward equalizer (pDFFE) designed for high-speed digital signal processor (DSP)-based wireline receivers. Since conventional decision feedback equalizer (DFE) suffers from an increasing timing burden due to recursive feedback loops, which hinder parallelization and scalability, the proposed pDFFE architecture eliminates the feedback path by replacing post-cursor inter-symbol interference (ISI) cancellation with feed-forward equalizer (FFE)-based decisions. The overall system employs a hardware-in-the-loop (HIL) verification platform based on the Versal ACAP, integrating a DSP engine with sign-sign LMS (SS-LMS) adaptation. The power and area of the proposed DSP in 28nm CMOS layout are estimated to be 176.5mW and 0.201mm2, respectively. Both simulation and measurement results under −24dB channel loss at Nyquist frequency demonstrate reliable BER performance comparable to the conventional DFE, while achieving single-clock latency and improving area efficiency, from a minimum of 10 % to a maximum of 47 %, compared to conventional architectures.
Tiny machine learning (tinyML) brings machine intelligence to numerous extreme edge scenarios under battery-powered and always-on constraints, which imposes demands on tinyML processors for ultra-low power and versatile support for diverse applications and tinyML models. A heterogeneous-computing-fabric (hetero-fabric) architecture that incorporates compute-in-memory (CIM) core and multiply-accumulate (MAC) unit-based digital core (Dcore) can combine high energy efficiency and flexibility to meet such demands. However, the dataflow flexibility in the hetero-fabric architectures has not been well explored. Besides, previous CIM macros are limited by the inherently fixed compute-storage ratio, facing the dilemma of balancing compute and storage density requirements from shallow to deep layers in neural network models. This work presents a versatile tinyML SoC for diverse extreme edge ML tasks. Firstly, we propose a tightly-coupled hetero-fabric architecture for tinyML processing, where the CIM core (Ccore) and Dcore are tightly coupled through the reconfigurable data-streaming network (RDSN). Flexible inter-/intra-core dataflow orchestration for various tinyML models with reduced L1 memory access is further proposed. We also propose a layer-wise dynamic frequency scaling (DFS) scheme for hetero-fabric performance balance with adjustable inter-core clock frequency ratios. Further energy savings can be achieved when combined with dynamic voltage scaling (DVS). Besides, we propose a three-mode reconfigurable-LUT (R-LUT) CIM macro, which breaks the inherent compute-storage limitations and enables dynamically adjustable compute and storage density. A compact Twin-10T SRAM cell featuring lower read power is also proposed to construct the LUT array. Finally, we design a flexible Dcore featuring intra-core fine-grained workload partition and elaborate scheduling to coordinate with the Ccore for dataflow orchestration. The fabricated 28nm chip achieves $2.14\times $ higher SoC energy efficiency than state-of-the-art CIM-based tinyML SoC. It supports all MLPerf Tiny tasks, achieving practical performance (e.g., no less than 30 fps) with no more than 0.5mW power.
This paper proposes and validates the implementation of a Spread-Spectrum Clocking (SSC) technique in a fully-integrated 6-channel Power Management Integrated Circuit (PMIC) for power management applications of microprocessor units. Each channel of the PMIC embeds a Ripple-Based Adaptive Constant ON-Time (RB-ACOT) DC–DC Buck converter and a type-II Charge-Pump Phase-Locked Loop (CP-PLL). The converters are synchronized via the CP-PLLs to mutually phase-shifted reference clock signals derived from a centralized SSC-modulated reference oscillator. We investigate here from a theoretical perspective the impact of the CP-PLL in altering the effectiveness of the SSC technique through a design-oriented model that includes both the CP-PLL and the RB-ACOT converter. A prototype of the proposed system is fabricated in 0.18-um Bipolar-CMOS-DMOS (BCD) process. Both simulation results and experimental measurements demonstrate the effectiveness of the proposed solution in mitigating the conducted electromagnetic interference level generated by the PMIC, achieving up to 12 dB of attenuation.
This paper presents an integrated light detection and ranging (LiDAR) system alongside its corresponding 22-nm FD-SOI controller IC. The proposed architecture integrates four distinct dies—a 1550-nm DFB laser diode, a photonic IC (PIC), an $8\times 8$ silicon photomultiplier (SiPM) array, and a mixed-signal silicon controller chip—complemented by a custom resonant concave acoustic mechanical mirror specifically designed to support the multi-die assembly. A laser driver and stabilization electro-optic phase-locked loop (EO-PLL), transimpedance amplifiers (TIAs), variable gain amplifiers (VGAs), analog–to-digital converters (ADCs), a custom fast Fourier transform (FFT) engine, mirror drivers, SiPM read-out circuitry, power management and 8-Gbps low-voltage differential signaling (LVDS) driver are all integrated in planar double thick metal RF SOI process. The total final device volume target is less than 16 mm x 16 mm x 3.5 mm including the suspended mechanical concave mirror to be resonant driven from the same controller IC. A 20-mW 1550-nm distributed feedback (DFB) laser is to be flip coupled onto a 5mm x 10mm PIC along with the 4.8mm x 4.8mm SOI-CMOS die. Configured via a serial peripheral interface (SPI), the system directly transfers processed FFT data to the host processor through multiple high-speed interfaces, rendering it highly viable for cost-constrained embedded applications. The maximum system bandwidth is 4-GHz (8-GSPS 4-interleaved ADCs) with various reconfigurable chirp and ramp rates and it is designed to continuously stream out maximum throughput of 8-Gbps through source synchronous LVDS link. The digital signal processing (DSP) backend in auto-chirp mode can dynamically adjust the chirp period and chirp slope for optimal range resolution.