
A single-channel 12 GS/s 7 b time-domain analog-to-digital converter (ADC) in 28nm CMOS is presented. The proposed ADC inserts a programmable PVT-robust time-difference amplifier (TA) between pipeline SAR-TDC stages and separates coarse and fine stage delay implementations, enabling simultaneous optimization of speed and time resolution (TLSB). Furthermore, a self-adaptive TA calibration is proposed to align the TLSB of the coarse and fine stages. The ADC consumes 52.9 mW and achieves 34.9-dB SNDR and 44.9-dB SFDR for Nyquist inputs with only occupying 0.0035mm2. The small input capacitance of the VTC allows an effective resolution bandwidth (ERBW) exceeding 16.16 GHz.
This paper presents a dual-mode burst-mode receiver (BMRX) in 28-nm CMOS. A pipelined feedforward equalizer (FFE) achieves higher summer bandwidth and extended summing time margin with 9 taps. Moreover, the successive approximation register (SAR) based auto gain control (AGC) logic and data reference level (dLev) estimation technique are incorporated to provide a fast and robust adaptation process across a wide range of received optical power (ROP), achieving 6.4-ns amplitude settling time and 10.9-ns phase locking time. This work demonstrates the first 50-Gbaud non-return-to-zero (NRZ)/PAM-4 BMRX with near 1/10th the power and area compared to digital signal processing (DSP)-based solutions, achieving the highest sensitivity of −26dBm/−16dBm in mixed-signal schemes.
This work presents a 230-GHz, 128-unit, λ/2-spaced planar phased array receiver (RX) utilizing a 1-D SSPLL and a DTC phase shifter. To construct a large-scale phased array system with a precise λ/2 antenna spacing above 200 GHz, this work addresses three key challenges. First, to achieve cross-chip synchronization among multiple chips, a 1-D SSPLL is proposed to enable arbitrary local oscillator (LO) synchronization for each tile while reducing clock jitter, achieving a measured ultra-low jitter of 44 fs. Second, to reduce the loss of LO driving power while maintaining precise beam steering, a 7-bit DTC phase shifter cascaded with a subsampling multiplier is developed—with the measured integral nonlinearity (INL) of the DTC phase shifter being ≤+1.6/−1.0 LSB and the differential nonlinearity (DNL) is ≤+0.5/−0.7 LSB. Additionally, to expand IF bandwidth, a Mixer-First Zero-IF architecture for each RX unit is realized through the integration of an IQ sub-harmonic mixer. Measurement results demonstrate a 10-GHz IF bandwidth, with an IQ amplitude imbalance of ≤−0.4/+0.8 dB and an IQ phase imbalance of ≤−2.4°/+0.8°. Subsequently, 230GHz, 16-unit, and 128-unit CMOS planar phased arrays are packaged and tested. The 16-unit array achieves an equivalent isotropic conversion gain (CG) of 16–18 dB, an equivalent isotropic single-sideband (SSB) noise figure (NF) of 6–14 dB, and a scanning angle of ±45° in both horizontal (H) and vertical (V) planes, while the 128-unit array delivers an equivalent isotropic CG of 20-21 dB, an equivalent isotropic SSB NF of 4-10 dB, and an extended scanning angle of ±60° in both H and V planes.
This paper presents a zero-slope digital-to-time converter (DTC) structure that achieves high linearity and low jitter for fractional-N PLLs. The proposed zero-slope DTC sequentially switches unit capacitors in a CDAC to maintain a constant voltage on the current integration node. This significantly improves the linearity by allowing the current source to retain a nearly constant VDS during delay operation and minimizing voltage distortion due to the voltage-dependent capacitance. The prototype chip fabricated in a 22-nm FDSOI process consumes 0.843 mW and supports an 11 bit range with 140-fs resolution. The measured performance demonstrates 246 fs rms jitter and 683 fs INL, validating the effectiveness of the proposed technique.
This paper presents a ROM-based LUT tensor engine that supports an output-stationary dataflow with an area-efficient complementary ROM structure, fully exploiting unstructured sparsity for fine-tuning-free ML inference. A latch-based PE efficiently exploits irregular zero patterns while avoiding redundant switching, minimizing energy for both dense and sparse inputs. The design achieves peak energy efficiency of 16.1 TOPS/W, outperforming prior LUT-based CIM macros and delivering 1.99× improvements compared to a state-of-the-art tensor engine.
The residue amplifier (RA) often dominates the noise and power in a high-precision 2-step SAR ADC. In this work, we propose to use the continuous-time (CT) SAR as the second stage, eliminating the sampling operation after RA and thus preventing the RA noise from aliasing into the in-band. Meanwhile, combining with the tracking averaging based LSB repeating technique, the wideband RA noise is reduced by 7dB.The prototype ADC implemented in 180nm CMOS measures ~100dB SNDR with 5MS/s sampling rate. The power consumption is 3.6mW, leading to a Schreier FoM of 187.8dB with 100kHz input frequency.
This paper presents a single-mode, fine-multilevel buckboost (SMFL-BB) converter designed for flicker-free mobile OLED displays, addressing the critical trade-off between inductor volume and the sub-8mV output ripple required for high-quality visuals. Its core innovation is a Boolean logic-driven power topology that unifies fine-granularity multilevel switching with a single-duty-cycle (D) control, enabling seamless operation across the buck, boost, and buck-boost regions. This approach ideally achieves near-zero inductor current ripple in the critical buck-boost region, where mobile-OLED PMICs primarily operate, thereby minimizing output voltage ripple and allowing for a significantly smaller inductor. Fabricated in a 180nm BCD process, the converter demonstrates exceptionally smooth transitions with only a ≤1.6mV output deviation during a full buck-to-boost sweep. When tested with a compact 560nH (1.02mm3-volume) inductor, the chip maintains an output ripple of ≤7.8mV at an L×CO FoM of 5.6 p•(rad/s)–2 across all voltage and load conditions, achieves a peak efficiency of 97.1% at 200mA, and shows strong immunity to TDMA noise.
Chinese handwriting decoding in Brain-Computer Interfaces (BCIs) is challenging due to character complexity, high power consumption, and bulky implementation of current decoding systems. To address these challenges, this paper presents a low-power, high-precision application-specific integrated circuit (ASIC) for Chinese handwriting trajectory decoding. The ASIC architecture is designed based on a dynamic ensemble Bayesian filter (DyEnsemble) algorithm with three core technical approaches, including 1) a task-level pipeline that implements the DyEnsemble algorithm to adaptively update model weights based on non-stationary neural activity, 2) a sequencer-managed near-memory dataflow that enhances computational efficiency and system portability by reducing power consumption from kW to μW level, and 3) an algorithm-hardware co-design framework, which integrates state-driven power gating and shared floating-point resources to achieve an optimal trade-off between decoding accuracy and hardware efficiency. Fabricated in 65nm CMOS technology, the proposed ASIC achieves a high trajectory decoding accuracy of 94.2% while consuming only 310μW at 0.8V supply voltage. The designed ASIC can be applied to decoding 2500 commonly-used Chinese characters, and also provide an efficient and scalable hardware solution to enable the development of future portable BCI-based Chinese communication and assistive applications.
Transformer-based Large Language Models (LLMs) present severe memory and bandwidth bottlenecks on edge devices. While SRAM-based compute-in-memory (CIM) effectively minimizes data movement, existing designs face a critical trade-off between inference accuracy and system efficiency. This paper presents a fully digital SRAM-based CIM accelerator featuring algorithm-hardware co-optimization to break this trade-off. On the algorithmic front, we propose a Ternary Weight Splitting (TWS) algorithm that binarizes full-precision BERT models, achieving a 32× compression rate with negligible (0.5%) accuracy loss. On the hardware front, we propose: 1) a BF16×1-b processing element (PE) utilizing a bit-parallel SRAM macro with batch-wise mantissa alignment; 2) a Group-Vector CIM architecture capable of storing entire BERT-Tiny model columns to maximize stationary data reuse; and 3) a Data Frame Standardization scheme to eliminate tensor reshaping overhead. Fabricated in 28nm CMOS, the accelerator achieves a peak area efficiency of 3.04TOPS/mm2 at 370MHz. Measurement results demonstrate energy efficiencies of 24.46TOPS/W (BF16×1-b) and 36.45TOPS/W (8-b×1-b). Compared to state-of-the-art counterparts, this work delivers 9.5× and 2.2× improvements in area and energy efficiency, respectively, enabling efficient high-accuracy LLM inference at the edge.
This paper presents a fully integrated continuous-scalable conversion-ratio (CSCR) switched-capacitor voltage regulator (SCVR) in Intel’s 18Å Ribbon FET gate-all-around (GAA) technology. A dual-loop regulation architecture achieves fast transient response and accurate remote-load DC regulation, while a replica-based adaptive dead-time generator improves robustness to process and temperature variations. Integrated charge pumps reduce power-switch leakage through negative gate-source biasing, and an active-decay path accelerates voltage ramp-down under light load conditions. The regulator delivers 84.3% peak efficiency for a 1.25-to-0.7 V conversion, 1.42 A maximum load, and achieves an added active-area current (power) density of 9A/mm2 (6.3W/mm2), enabling compact, high-performance on-die power delivery for advanced SoCs.
This paper presents a crystal-less WiFi 802.11b backscatter tag, which generates directly-decodable WiFi packets by leveraging ambient, uncontrolled WiFi signals as carriers. This is accomplished by: 1) implementing an envelope-based demodulator that exploits the WiFi Barker-code property to extract pulse-width signatures for symbol distinction and synchronization, while recovering a 1 MHz clock to calibrate the digital controlled oscillator (DCO) for intermediate-frequency (IF) clock generation; 2) proposing an on-the-fly modulation scheme that enables arbitrary 802.11b WiFi carriers to be converted into directly-decodable WiFi packets. The proposed system achieves a measured synchronization time of less than 120 ns, ensuring reliable WiFi backscattering. Experimental results demonstrate a downlink (DL) sensitivity of −36.5 dBm with BER < 10−3, an ultra-low power consumption of 12.5 µW for downlink, and 6.4–15.7 µW for uplink (UL). Wireless over-the-air measurements show a maximum TX-to-tag range of 15 m for the downlink and a worst-case TX-to-tag-to-RX range (equidistance of TX-to-tag and tag-to-RX) of 30 m for the uplink, when excited by a 30 dBm EIRP WiFi AP.
This paper presents a fully-integrated 256-channel wireless wideband neural-recording SoC that integrates AFEs with distributed neural processors and a low-power bidirectional transceiver. A nearlossless adaptive-delta algorithm reduces data rate by 60% while preserving features. A power-efficient tri-state bus enables streamlined transfer. The 3.2 GHz/433 MHz OOK links with hybrid error-resilient coding boost reliability. Fabricated in 65 nm CMOS, this SoC consumes 3.49 mW at 1.2 V and supports 30 kHz/channel.
In closed-loop neural interfaces, simultaneous recording and stimulation impose significant challenges on the recording path, demanding a frontend capable of accommodating large amplitudes and a wide dynamic range. This paper presents a Noise-Shaping SAR ADC with large input amplitude and high dynamic range for closed-loop interfaces, featuring in-loop Serial-to-Parallel Amplification (SPA) and a fully dynamic Capacitor-Stacking-and-Buffering (CSB) integrator with a Fully Dynamic Voltage Buffer (FDVB). Fabricated in 65-nm CMOS, the proposed design achieves 2-Vpp input range, 95-dB SNDR, and 178.5-dB FOMs with only 2.26 µW within a compact 0.08 mm2 area. In-vitro experiments using pre-recorded neural data with stimulation artifacts demonstrate robust, artifact-resilient neural recording.
This paper proposes a four-channel (4-ch) time-interleaved (TI) high-speed current-steering Digital-to-Analog Converter (DAC) with hybrid segmentation that employs a single-stage analog multiplexer followed by a current summation. Compared with conventional 4-ch structures, the proposed architecture simplifies the clocking scheme by reducing the number of required clocks and eliminating the need for high-speed clock signals. In addition, the proposed design does not have the dead-zone in the middle of the bandwidth and reduces the sensitivity to duty-cycle mismatches. The DAC was fabricated in 28 nm CMOS, incorporating a simple differential clock timing calibration to ensure precise synchronization of sub-DACs. Experimental results demonstrate that the proposed 10-bit, 20-GS/s DAC achieves >40 dBc spurious-free dynamic range (SFDR) for a signal frequency up to 6.3 GHz, representing an improvement over state-of-the-art prior 4-ch time-interleaved DACs
Among new non-volatile memories, spin-transfer-torque magnetic-random-access-memory (STT-MRAM) has high advantages due to its high density and compatibility with CMOS process. In this work, a 4-Mb high reliability STT-MRAM macro is proposed with bottom-up MRAM refinement methodology, including 16 banks and the error check and correction (ECC) block. The high density MRAM bit-cell, the multi-bit voltage sense amplifier (MB-VSA), the high-speed and high-accuracy write self-termination circuit (HSHA-WT) and the endurance detection module are proposed to help realize high density and high reliability, together with the in-situ yield monitoring (ISYM) technique. Utilizing 28nm CMOS technology, this work achieves 6ns read speed, 20ns fast write time, 20 years retention time, 107 endurance cycles and 0% failure rate with ECC.
A discrete-time amplifier with multi-path kT/C noise cancellation is presented for ultra-low-power IoT applications. By pre-storing kT/C noise across N time-multiplexed transconductance paths and cancelling it at the series output, the design decouples noise performance from switching frequency and capacitor size. Fabricated in 180-nm CMOS, it achieves an input-referred noise of 1.98 μVrms within a 2-kHz bandwidth, corresponding to an NEF of 0.32 and a PEF of 0.08 under 36 nA from a 0.8-V supply. This work achieves the best trade-off between current and NEF, as well as between power consumption and PEF, while reducing power by ~8.3× compared to a 22-nm design and occupying a similar area.
This paper presents a fully integrated continuous scalable conversion ratio (CSCR)-based single-stage piezoelectric energy harvesting (PEH) interface suitable for arbitrary periodic vibrations. The proposed CSCR switched-capacitor (SC) converter with bi-directional ring-connected register sequence control can provide seamless buck/boost switching under continuous voltage conversion ratios (VCRs) while eliminating the reverse conduction loss. This single-stage PEH interface can also achieve simultaneous fine-grained piezoelectric transducer (PT) biasing, energy extraction and load regulation from distinct inputs/outputs. The capability of independent PT biasing for both the positive and negative half cycles under asymmetric excitations, together with the proposed PIN-based dual-sided MPPT scheme and energy investment (EI)-assisted bias flip technique, can further enhance the system efficiency. Fabricated in 180nm BCD process, this work successfully demonstrates a fully integrated single-stage implementation with a total on-chip capacitance of only 468pF, while achieving a measured voltage flipping efficiency (VFE) of 84.8% and power conversion efficiency (PCE) of 85.3%. It also demonstrates a high MOPIR of up to 12.67×, corresponding to a 36% improvement over similar prior arts.