
Reasoning strategies such as chain-of-thought improve the accuracy of large language models (LLMs) but lengthen decoding sequences, inflating external memory access (EMA) on mobile devices. Numerous optimizations have been explored to reduce EMA, from conventional techniques such as group quantization and channel-wise pruning to sparse mixture-of-experts (MoE) and speculative decoding (SD). By holistically adopting these techniques, up to 89.7% EMA energy reduction can be achieved. However, group quantization with channel-wise pruning reduces intra-group INT MACs while increasing the proportion of inter-group FP MACs, creating distinct heterogeneous INT-FP workloads. NPU-CIM architectures, which process INT MACs on energy-efficient CIM and FP MACs on a flexible NPU, are well-suited for this workload. Nevertheless, conventional NPU-CIM fails to fully exploit MoE-based SD due to redundant expert activation, low hardware utilization from dynamic and heterogeneous workloads, and high power consumption in CIM. We present SMoLPU, a pipelined NPU-CIM architecture for energy-efficient MoE-based SD LLM inference on mobile devices with three key features. Token-adaptive expert refinement (TaER) eliminates redundant expert fetching and schedules the expert load order, achieving $2.3\times $ and $4.2\times $ energy efficiency improvements in the prefill and decode stages. An adaptive-offload NPU-CIM core (ANC) maintains high utilization under dynamically varying INT-FP ratios by bundling dense groups for CIM and offloading sparse groups to idle NPU, achieving a $3.3\times $ end-to-end speedup. Reconfigurable DRAM-based LUT-CIM enables low-power addition while supporting variable input sizes, reducing adder tree power and area by 36.8% and 26.3%. SMoLPU is fabricated in a 28-nm CMOS technology and occupies 20.25 mm2. It achieves 7.8 fJ/token/parameter, which is 43.5% lower energy per parameter than the state-of-the-art.
Environmental sound recognition (ESR) technology has gained increasing prevalence in recent years across various applications such as smart caregiving and hearing-assistive devices, where always-on operation requires both high accuracy and ultra-low power consumption. However, the existing works have two main issues. First, the existing ESR algorithms are mainly implemented on GPUs or FPGAs, which cannot meet the energy constraints of AIoT devices, especially when SOTA ESR algorithms heavily rely on deep neural networks (NNs) to improve accuracy, resulting in large computations. Second, most of the existing ESR algorithms are trained using public datasets that are recorded in different sound environmental conditions (e.g., microphone and room sound environment) from those of their actual use, leading to significant accuracy degradation. Therefore, this work presents, to the best of our knowledge, the first fabricated ASIC-based ESR processor, integrating three co-designed techniques to address the above challenges: 1) a “short-long” progressive spectrogram processing (SLPSP) architecture to reduce the energy consumption; 2) a split-and-merge processing (SMP) architecture to further reduce the energy consumption; and 3) a weight clustering-based adaptive learning (WCAL) combined with offline domain-adversarial training to improve accuracy against sound environmental condition changes with significantly reduced data-access overhead. Fabricated in 65-nm CMOS, the proposed processor achieves an ultra-low energy consumption of $86~\mu $ J per classification and achieves classification accuracy of 100%, 91.41%, and 82.4% for two-, eight-, and 50-class tasks on the re-recorded ESC-50 dataset under real environment, respectively.
This article presents an event-driven dual-mode analog front-end (AFE) IC for a capacitive touch-screen panel (TSP). To address the trade-off between energy efficiency and performance of the touch sensing system, the dual-mode AFE IC can operate in a low-power self-capacitance sensing (LP-SCS) for touch event detection during standby mode and transitions to high-frame-rate mutual-capacitance sensing (HFR-MCS) upon detecting a touch event. In LP-SCS mode, a delay-chain-based offset compensator (DCOC) digitally estimates and compensates for the large offset capacitance of TSP electrodes, enabling self-capacitance (SC) sensing over a wide input range up to 1 nF. To achieve fine-resolution and low-power capacitance conversion within this wide input range, a dual-slope-based SC detector (DS-SCD) is adopted. Fabricated in an 80-nm CMOS process, the AFE achieves power consumption of 0.62 mW in LP-SCS and 8.65 mW in HFR-MCS, realizing a 92.8% power reduction through mode switching. The proposed AFE achieves frame rates of 60 and 500 Hz in LP-SCS and HFR-MCS modes, respectively, with measured signal-to-noise ratios (SNRs) of 52.6 and 40.2 dB, demonstrating both energy efficiency and high-speed sensing capability.
Zero-knowledge proof (ZKP), one of the most popular privacy-preserving schemes, enables the prover to convince the verifier of a certain statement’s correctness without leaking any private information. Among them, the Pairing-based succinct non-interactive argument of knowledge (zkSNARK), like Groth16 and Plonk, featuring succinct proof size and constant-time verification, is deployed across a promising ecosystem. However, the heavy operations of proof generation that grow rapidly with application sizes and the frequent verification that requires costly Pairing severely hinder broader adoption. Moreover, there are many other cryptographic schemes also constructed based on Pairing, such as Boneh–Lynn–Shacham (BLS) signature, functional encryption (FE), identity-based encryption (IBE), and so on. Unfortunately, existing synthesis-based ZKP accelerators suffer from low area efficiency and incomplete operator coverage, causing poor overall speedup and expensive deployment cost. This work presents the first silicon-proven crypto-processor that supports complete proof generation/verification phases and can adapt to other Pairing-based cryptography (PBC). To achieve high flexibility while maintaining cost efficiency, we employ holistic optimizations across algorithm, architecture, and compiler levels, including the hybrid-grained instruction set, dedicated memory system, reconfigurable datapath, custom Pairing compiler, and utilization-oriented timing scheduling. Fabricated in 28-nm CMOS, it covers all related operators and achieves a $6.4\times $ area reduction compared to prior synthesis-based work LegoZK, which currently makes it the most cost-effective candidate for the practical deployment of ZKP–PBC applications. As a case study, the processor performs $2^{16}$ -gate proof generation/verification with 0.15 J/0.11 mJ, offering two orders of magnitude energy savings over a 14-core 2.4-GHz CPU. For the Pairing operator, it achieves development agility to enable fast design space exploration. It delivers $17\times $ / $1.8\times $ improvements of area-time product (ATP)/energy compared to the prior ASIC work.
This article presents 10–30-GHz reflectionless pseudo-four-path mixer-first receivers. By introducing a balanced input network at the radio frequency (RF) port, the proposed receiver requires only a two-phase local oscillator (LO) to generate differential and orthogonal baseband (BB) signals, significantly reducing the complexity and power consumption of the LO generation circuit. The crosstalk issue caused by LO overlap is topologically mitigated by the inherent isolation (ISO) between the CPL and THU ports. In addition, the balanced input network not only maintains the reflection-based out-of-band (OOB) rejection mechanism of the $N$ -path mixer-first receiver but also enables a reflectionless design. To further improve the transition-band roll-off and OOB rejection, two pairs of transmission zeros flanking the passband are introduced. The first pair of close-in transmission zeros is generated by an improved blocker-canceling technique, whereas the second pair is introduced by the $LC$ tanks. Benefiting from the balanced input network and the improved blocker-canceling technique, both the input matching and OOB rejection are decoupled from the switch size, enabling the use of small-sized switches and thereby achieving high LO-to-RF ISO. Two receiver versions (Design-1 and Design-2) are presented: Design-1 excludes $LC$ tanks, whereas Design-2 incorporates $LC$ tanks. Fabricated in a low-cost 130-nm CMOS SOI process, both designs achieve a transition-band roll-off of approximately 40 dB/decade. Across 10–31 GHz, Design-1 achieves an RF-to-BB conversion gain of 7.3–10.3 dB, a double-sideband (DSB) noise figure (NF) of 12.1–16.3 dB, a blocker 1-dB compression point (B ${}_{\mathbf {1dB}}$ ) of 2.15 dBm, an OOB input-referred third-order intercept point (IIP ${}_{\mathbf {3}}$ ) of 11.3 dBm, an OOB rejection exceeding 25 dB, and an LO-to-RF ISO of 38–56.5 dB. Operating across 10–30 GHz, Design-2 exhibits 7.7–10.7-dB RF-to-BB conversion gain, 11.9–16.2-dB DSB NF, 3.15 dBm B ${}_{\mathbf {1dB}}$ , 12.1-dBm OOB IIP ${}_{\mathbf {3}}$ , >29-dB OOB rejection, and 36.9–61.7-dB LO-to-RF ISO. The maximum power consumption of both designs is below 36.2 mW.
This article presents a reconfigurable self-clamp, continuously scalable-conversion-ratio (CSCR) switched-capacitor (SC) converter for photovoltaic (PV) solar energy harvesting (EH). By reducing flying-capacitor voltage stress without additional SC stages, the proposed architecture enables high-density capacitors and mitigates the complexity-loss tradeoff of conventional CSCR converters. A clamp-stage outphasing technique improves the power conversion efficiency (PCE), while a design methodology for selecting the number of flying and clamping capacitors is established. In addition, an insight is revealed for reducing the energy penalty of open-circuit-voltage (OCV)-based maximum power point tracking (MPPT). The reconfigurable-clamp-mode scheme extends the high-PCE voltage conversion-ratio (VCR) range and improves MPPT accuracy. Fabricated in a 180-nm CMOS process, the prototype exhibits 90% peak PCE, a high-PCE VCR ratio of 3.83 using only thirteen capacitors, and MPPT accuracy above 96.7% over a wide light-intensity (LI) range with a 334- $\mu $ s convergence time.
This article presents a 10-bit source-driver (SD) IC for high-resolution OLED-on-silicon (OLEDoS) microdisplays in augmented reality (AR)/virtual reality (VR) applications. To address the stringent constraints of aggressive area scaling, sub-microsecond settling time, and high output uniformity, three key techniques are proposed. First, a demultiplexer (DEMUX)-based channel width expansion method combined with deep-N-well (DNW)-based voltage-domain transformation enables the source driver IC (SD-IC) to fit within the 1.33- $\mu $ m sub-pixel pitch of the OLEDoS panel. This DNW-based voltage-domain transformation allows the DAC and channel logic circuits to operate in a low-voltage (LV) domain, allowing the use of compact LV devices and thereby significantly reducing the channel area. Second, a data-difference-dependent dynamic current source (D3DCS) is introduced to meet the sub-microsecond driving-time requirements of 4K $\times{\,}4$ K resolution at a 120-Hz frame rate of the OLEDoS displays. This technique selectively boosts the slew rate (SR) of the amplifier based on most significant bit (MSB) transitions without increasing static current or introducing dead-zone behavior. Third, an all-channel automatic offset calibration (ACAOC) scheme is proposed to suppress intrinsic amplifier offsets and improve output-voltage uniformity. Unlike conventional frame-to-frame polarity alternation techniques, which are less effective for OLEDoS panels due to their narrow data-voltage range, the proposed scheme directly minimizes the offset magnitude through simultaneous calibration across all channels, enabling uniform driving voltages without additional frame-level compensation. Fabricated in a 65-nm CMOS process, the proposed SD-IC occupies $2273~\mu $ m2 per channel. Measurement results show a settling time of $0.69~\mu $ s and a reduction in maximum deviation of voltage output (DVO) from 26.6 to 1.9mV. The measured differential nonlinearity (DNL) and integral nonlinearity (INL) are +0.27/−0.22 and +0.59/−0.33 LSB, respectively.
Solid-state nanopore single-molecule sensing aims to replace the commercially available biological nanopores. Array sizes have remained limited due to the area and power consumption of the transimpedance amplifiers (TIAs), thus preventing throughput and cost scaling. This article presents a 256-channel readout array for solid-state nanopore single-molecule sensing in 65 nm CMOS. Each channel contains a novel discrete-time TIA architecture, which is based on an Integrate and Reset concept and exploits source-degeneration to improve the noise performance. Each TIA is combined with a new event-detection architecture that allows disabling channels and reusing ADC resources, resulting in a $32\times $ reduction in data rate. These innovations result in a more than $8\times $ smaller channel area compared to the state-of-the-art while reaching an integrated noise performance of 193 pA ${}_{rms}$ in a 1 MHz bandwidth, thus enabling a $10\times $ channel count scaling.
A time-series analysis on multi-channel sensor signals is critical for wide-range Artificial Intelligence (AI) applications, such as in a field of mobility, robot, and infrastructure monitoring. Edge-device data collection followed by cloud-based analysis suffers from high latency and power consumption. This article presents an edge-AI-oriented sensor fusion architecture, performing analog multiply-accumulation operation by time-domain processing among eight-channel analog sensor signals to directly extract hyper-dimensional digital features, namely, direct analog sensor fusion (DASF). It completely eliminates general-purpose analog-to-digital converters (ADCs) from the system to significantly save energy and delay for on-the-spot edge-AI response. A 28-nm CMOS proof-of-concept system shows only 2.5% MAC full-scale error under $0~^{\circ } \text {C}$ – $70~^{\circ } \text {C}$ temperature and inter-chip variations and achieves 1.13- $\mu $ W retention power and a normalized energy efficiency of 591.5 TOPS/W, providing a 17.6% improvement over prior works. The proposed all-digital reservoir with sparse encoding is evaluated on three time-series analysis workloads, achieving an accuracy of 98.26%, 95.9%, and 98.11%, respectively.
Boolean satisfiability (SAT or $k$ -SAT, $k \ge 3$ ) is a binary combinatorial optimization problem (COP) that is challenging for traditional classical computing to address in polynomial time. Quantum and quantum-inspired systems have been touted to tackle these and similar nondeterministic polynomial-time (NP) problems. This work describes Medusa, a quantum-inspired $k$ -SAT solver capable of addressing problems of up to 200 variables and 1016 clauses. Implemented using analog and mixed-signal circuits, the solver benefits heavily from massive parallelism, continuous-time operation, and excellent energy efficiency, and advances the field of quantum-inspired, analog computing with best-in-class performance and energy efficiency. This performance is achieved with the introduction of three key architectural advancements: dual Make and Break feedback, a distributed logic feedback network, and multi-macro operation. The dual feedback mechanism improves the mean time to solution (TTS) for 50-variable 3-SAT problems by $30\times $ when compared with a single feedback mechanism. The distributed logic in the feedback network implements all-to-all spin connectivity with arbitrary order interactions, enabling $k$ -SAT operation and supporting the scalability of the solver, which is further enhanced by multi-macro operation. Medusa, implemented in a 28 nm process, achieves a $4.92~\mu $ s TTS for the popular SATLIB uf50-218 benchmark with a mean energy to solution (ETS) of 19.1 nJ. For the uf200-860 benchmark the solver achieves a 30.9 ms mean TTS, outperforming quantum annealing for problems of the same size.
This article presents a calibration-free fractional-N cascaded phase-locked loop (PLL) employing a sampling phase detector (SPD)-bypass-based phase-domain oversampling noise-and-spur cancellation (P-OSNSC) technique. By incorporating an oversampling cancellation path that eliminates the gain errors and noise cancellation bandwidth limitation introduced by the SPD, the proposed architecture suppresses noise from the first-stage RO-based frequency multiplier at minimal power overhead, without invoking high-frequency reference or calibration. This approach enables noise cancellation across a wider bandwidth compared to previous arts. Moreover, the proposed prototype fabricated in a 28-nm CMOS achieves 127.8-fs rms jitter (integrated from 10 kHz to 100 MHz) with 4.48-mW power consumption, corresponding to a figure of merit (FoM) of −251.4 dB.
This article presents a valuewise hybrid (VWH) compute-in-memory (CIM) macro that improves hybrid efficiency by $1.75\times $ while simultaneously reducing computation error by $2.02\times $ compared to prior bitwise hybrid schemes. The VWH approach performs accurate computations digitally forthemost significant values (MSVs) in a small proportion, while the vast majority of the values are processed by lossy analog CIM (ACIM) macros to achieve high energy efficiency. To minimize latency and enhance the energy efficiency of the ACIM macros, a heterogeneous memory fabric (HMF) is introduced, which combines the benefits of static random access memory (SRAM)-CIM and dynamic random access memory (DRAM)-CIM. In addition, an asynchronous sparsity manager (ASM) enables adaptive throughput and power consumption based on sparsity. Fabricated using a 28-nm CMOS process, the prototype VWH-CIM macro achieves energy efficiencies of 470.5 TOPS/W in four-bit mode, 109.8 TOPS/W in eight-bit mode, and 48.5TOPS/W in 12-bit mode.
This article presents a 24-channel synchronous battery monitoring IC (BMIC) with sub-1-mV accuracy and enhanced busbar fault detection for electric vehicles (EVs). Dedicated high-voltage (HV) bipolar-sampling ADCs sample all 24 cells within 1 ms, eliminating inter-channel delay. The HV bipolar sampler (HVSPL) combines active body biasing and clock doubling to extend the differential input range of −2–+5 V, using 12-V MOS devices. On-chip common-mode voltage ( $V_{\text {CM}}$ ) and input impedance ( $R_{\text {IN}}$ ) calibration engines compensate channel-dependent gain variations and offsets from finite $R_{\text {IN}}$ . Fabricated in a 0.13- $\mu $ m 120-V BCD process, the BMIC achieves ±0.73-mV total measurement error (TME) at the 3.5-V nominal cell voltage and TME within ±1-mV across −2–+ 5 V and $- 40~^{\circ } $ C– $+ 130~^{\circ } $ C (covering the Automotive Electronics Council (AEC)-Q100 Grade 1 temperature range). At room temperature (RT), a 1P24S 21700-cell EV battery test demonstrates 28.23- $\mu $ V precision with each ADC drawing 0.056 mA from an 84-V battery stack through the integrated DC–DC converter.
The pipeline-successive approximation register (SAR) analog-to-digital converter (ADC) extends the energy efficiency and the scaling properties of SAR ADCs to the medium-to-high resolutions and high bandwidths typical of classical pipeline ADCs. However, the residue amplifier (RA) is often a performance and energy efficiency bottleneck, especially for high-speed designs in scaled CMOS processes. This work proposes a $3\times $ -cascoded floating inverter amplifier (FIA) RA, powered from a low supply voltage of 1 V thanks to a virtual supply extension technique, providing a robust process, voltage, and temperature (PVT) operation without requiring analog trimming, bias, or nonlinearity calibrations, except for the linear gain. The proposed amplifier is integrated in a 500-MS/s, 12-bit pipeline-SAR ADC implemented in a 28-nm CMOS process. The prototype achieves an SNDR and an SFDR of 63.4 dB and 81.0 dBc, respectively, while dissipating 6.64 mW including the on-chip digital calibrations, resulting in a Schreier figure of merit of 169.2 dB.
This work presents a fractional- $N$ digital phase-locked loop (PLL) featuring a supply-insensitive variable-slope digital-to-time converter (SI-VS-DTC). The proposed architecture neutralizes DTC supply sensitivity—a primary bottleneck for high-performance frequency synthesis in noisy environments—without compromising inherent linearity or baseline noise. Robust background calibration algorithms are deployed to ensure stable technique operation across process, voltage, and temperature (PVT) variations. Implemented in 28-nm CMOS, the prototype synthesizes an 8.85–10.25 GHz output from a 150-MHz reference clock, occupying an active area of 0.28 mm2 while consuming 22.5 mW. Under a 10-mVpp sinusoidal ripple or 6.2-mVrms wideband noise applied directly to the DTC supply voltage, the PLL maintains an integrated rms jitter below 69 fs and spurious tones below −62 dBc at 8.85-GHz near-integer channels.
This article presents a high-bandwidth eight-channel single-ended four-level pulse amplitude modulation (PAM-4) transmitter for next-generation pseudo-open drain (POD)-terminated memory interfaces. To reduce power consumption in POD termination, a PAM-4 least-significant-bit (LSB) data bus inversion (DBI) encoding scheme is implemented and compared to previous PAM-4 encoding schemes. Based on a theoretical analysis of termination power, the PAM-4 LSB DBI encoding achieves termination power savings comparable to those of previous schemes while reducing the overall system power by simplifying the DBI flag lane. A series-resistor-shielded, source-series-terminated (SST) driver improves output bandwidth by isolating the parasitic capacitances of pull-up and pull-down segments from the output node. A four-tap ZQ-based feed-forward equalizer (FFE) provides fine coefficient control to alleviate inter-symbol interference (ISI) with a small number of driver segments. Moreover, the transmitter supports two selectable FFE modes (high-ISI mode and low-ISI mode) to optimize power consumption across different channel conditions. A ZQ calibration for the remaining tap segments compensates for impedance mismatch and improves the ratio of level mismatch (RLM). An eight-channel prototype was fabricated in a 28-nm CMOS process, occupying an area of 0.0222 mm2 per lane. The ZQ calibration ensures a minimum RLM of 0.985. Under multi-channel operation, the transmitter achieved a data rate of 42 Gb/s/channel at a worst case channel loss of −4.41 dB, with an energy efficiency of 1.2 pJ/bit, excluding the global clock path.