
This paper presents a 24.25-29.5-GHz transmitter designed for fifth-generation (5G) millimeter-wave (mm-wave) applications, featuring a fully integrated self-calibration system for local oscillator feedthrough (LOFT) and in-phase/quadrature (I/Q) imbalance. The system incorporates an on-chip LOFT and image detection chain, together with a self-convergent digital control circuit. The existing on-chip calibration methods rely on Analog-to-Digital Converters (ADCs) to analyze the LOFT and image signals in the digital domain, or employ complex frequency-division circuits to down-convert the two impairments individually. However, these methods are confined to zero-IF or narrow-band low-IF systems, due to the limited bandwidth of the ADCs and frequency-division circuits. Furthermore, many existing approaches rely on external processing devices to implement the calibration algorithm, thereby increasing overall system complexity and operational burden. To overcome these limitations, the proposed architecture integrates a broadband detection chain that converts both LOFT and image signals into a unified DC output for high-sensitivity monitoring, which significantly extends the detection range. Subsequently, a fully built-in cyclic coordinate descent (CCD) algorithm jointly suppresses LOFT and image components, by autonomously optimizing the bias voltages of the mixer's transconductance stage. Fabricated in 65-nm CMOS technology, the transmitter achieves a conversion gain of 7.5 +/- 1.5 dB across a wide IF bandwidth from 2.4 to 7.5 GHz, and delivers a maximum output 1-dB compression point (OP1dB) of 12.8 dBm. Post-calibration measurements reveal over 40-dB LOFT suppression and 47-dB image rejection ratio (IRR) across the operating bands. To the best of the authors' knowledge, this work demonstrates the first reported fully integrated calibration system for joint LOFT and I/Q imbalance mitigation for 5G n257, n258, and n261 bands.
This paper proposes a sub-THz voltage-controlled oscillator (VCO) employing an active negative-conductance boosted architecture. The proposed architecture contributes to relaxing start-up conditions, improving tuning linearity, and increasing swing. A triple-coil coupled transformer is introduced as the tail resonator, simultaneously providing the broadband frequency selectivity of the feedback loop and further achieving conductance enhancement. The head resonator further improves phase noise (PN) performance while efficiently extracting second-harmonic components without requiring additional structures. And filtered output is achieved by a compact coupled-line-based buffer. Fabricated in 28nm CMOS process, the proposed VCO achieves a frequency tuning range (FTR) of 7.1% while operating at a center frequency of 109.6 GHz. At a 10 MHz frequency offset, the measured PN is -112.8 dBc/Hz, corresponding to a figure of merit (FoM) of -180.4 dBc/Hz. The VCO operates under a 0.5-1 V supply and occupies a core chip area of 0.035 mm2.
This paper presents a rail-to-rail buffer architecture for TFT-LCD driver ICs that addresses two practical challenges: bandwidth degradation near the supply rails and potential instability during on-wafer testing under heavy capacitive loading. Four auxiliary current buffers are introduced into the floating class-AB core to shorten internal control paths and maintain wide bandwidth across the full input range. To ensure stable operation with an 80-pF on-wafer load, the second-stage amplifier is partitioned and an impedance RZ is inserted to generate a left-half-plane zero for effective pole-zero compensation. In addition, a compact MUX structure with integrated ESD protection is co-designed to reduce the effective output-path resistance without enlarging switch devices. By redistributing the ESD resistance across parallel branches, the proposed design shortens the RC time constant and improves area efficiency. Fabricated in a 0.18-& micro;m CMOS process, the 20-channel prototype demonstrates stable, oscillation-free operation under 80-pF on-wafer loading and achieves 90% settling times of 346 ns (0.2-2.3 V) and 248 ns (2.7-4.8 V), with corresponding 99% settling times of 1.65 & micro;s and 1.51 & micro;s, respectively, under a 5-k Omega/400-pF five-stage distributed R-C load emulating a TFT-LCD data line.
An SSHI-based interface circuit for piezoelectric vibration energy harvesters is presented and experimentally validated. The developed solution relies on current-driven regulation of the rectifier output voltage downstream of a parallel SSHI stage. By measuring only the rectifier output current, the proposed Current-Driven SSHI (CD-SSHI) interface directly and continuously sets the rectified DC voltage to its optimal value under time-varying vibration conditions, without resorting to the perturbative approaches of classical Maximum Power Point Tracking (MPPT) techniques. The CD-SSHI reaches and tracks the Maximum Power Point (MPP) without the steady-state oscillations typical of perturb and observe MPPT, and without the power interruptions characteristic of the fractional open circuit voltage MPPT. Moreover, the proposed CD-SSHI interface significantly increases dynamic performance compared to an SSHI coupled with a conventional implementation of the P&O MPPT control. A theoretical analysis based on a general equivalent circuit model is developed to support the proposed circuit. Experimental results obtained from a prototype implementation demonstrate the effectiveness of the CD-SSHI interface in terms of both steady-state efficiency and dynamic response.
ReRAM-based in-memory computing (IMC) architectures enable efficient neural-network inference on edge devices, due to their high density and non-volatility. However, ReRAM is prone to the stuck-at faults (SAFs) which would distort weight mappings and reduce model accuracy substantially. To tackle this problem, a high-efficiency fault-resilient (HEFR) framework is introduced in the paper, which couples layered-precision quantization based on the generalized gaussian distribution cumulative distribution function (GGD-CDF), preserving weight distributions with the minimal information loss. Further, a fault-aware weight re-decomposition method is proposed. Specifically, a Q-agent-based method is proposed for sparse SAFs, which employs offline reinforcement learning to construct a globally optimized hash table with reduced compilation complexity, while incorporating both weight matching and cell-state stability into the reward function to suppress electro-stress. For dense SAFs, we exploit a greedy search over the remaining cells to provide rapid and accurate mappings. Experimental results show that GGD-CDF quantization improves accuracy by 2.18% over Float32. Under fabrication faults, HE-FR surpasses a fault-free method by 12.97% and accelerates compilation by 197 & times; with only 0.024% picojoule-level energy overhead. In electro-stress evaluations, HE-FR reduces performance degradation by 46.27%. It demonstrates that the proposed framework has advantages in robustness, efficiency, and reliability.
This paper proposes an implementation-friendly algorithm for multi-coset sub-Nyquist sampling-based wideband spectrum sensing (WSS). To the best of authors' knowledge, this algorithm has been designed for the first time using the iterative power method with deflation technique and the exponential fitting test to reduce its computational complexity. In addition, it has been incorporated with a memory-efficient and hardware-friendly interpolation algorithm. In response to the suggested WSS algorithm, a hardware-efficient architecture for wideband spectrum sensor (WBSS) is presented in this work from system to micro-architectural level. Furthermore, a dual-clock architectural technique has been suggested for signal acquisition in the proposed WBSS architecture to enhance the sensing bandwidth. In this paper, computational complexity and performance analyses of the suggested WSS algorithms are presented in a comprehensive way. Finally, the WBSS architecture is implemented on the AMD Zynq UltraScale+ MPSoC-ZCU104 FPGA-board that is used in a real-world test setup for functional validation of the measured results. Compared to state-of-the-art implementations, the proposed WBSS has achieved 33.75% better hardware efficiency, $3.6{ imes }$ lower memory requirement, and $6.34{ imes }$ wider sensing bandwidth.
The rapid scaling of large language models (LLMs) has made memory capacity, bandwidth, and energy efficiency critical bottlenecks for practical deployment. Quantization is an effective approach for mitigating these costs, but conventional tensor- and channel-level quantization suffers from severe accuracy loss due to outliers in LLM weights and activations. Although block quantization improves robustness by localizing outlier effects, straightforward block-size reduction introduces substantial metadata overhead and significantly increases the hardware cost of post-processing units (PPUs) and accumulators. Furthermore, existing outlier-aware block quantization methods still exhibit limited block-size scalability and depend on simplified online outlier detection criteria that fall short of mean square error (MSE)-based optimality. This paper presents BGM, a GEMM accelerator for LLM inference that co-designs outlier-aware mixed-block-size quantization and dedicated hardware support. The proposed merge-and-split block (M&SB) quantization employs a 1:1:2:4 mixed-block-size structure that selectively splits only outlier-containing sub-blocks while merging the remaining sub-blocks, thereby reducing quantization error and compensating for metadata overhead. To support this mixed-block-size execution efficiently, we design a block processing unit (BPU) with a Sub-PPU and a 2-depth FIFO-based dynamic scheduling mechanism, which avoids the linear hardware overhead growth of naive designs. We further propose zero-counting and variance-based criterion (ZVC), a lightweight online criterion that enables split/merge decisions without the expensive quantization-dequantization-error accumulation required by MSE-based selection. RTL synthesis and accelerator-level evaluations demonstrate that BGM achieves 2.95 & times; higher core energy efficiency than the strongest baseline accelerator, together with up to 3.04 & times; speedup and 2.90 & times; energy reduction in end-to-end simulations. These results establish BGM as an efficient and scalable hardware architecture for practical quantized LLM inference.
This paper proposes a low-power voltage reference that generates 736 mV over a temperature range of 0 degrees C to 170 degrees C, targeting low-power, high-temperature miniature sensing systems. Operating in the subthreshold region, a bipolar junction transistor (BJT) diode produces a process-insensitive complementary-to-absolute-temperature (CTAT) voltage, while stacked complementary metal-oxide-semiconductor (CMOS) transistors provide temperature compensation by adding a proportional-to-absolute-temperature (PTAT) voltage. Measurements from 76 samples across three wafers fabricated in a 180-nm CMOS process demonstrate a +/- 3 sigma inaccuracy of 3.64% from 0 degrees C to 170 degrees C without trimming. The reference consumes 31 pW at 27 degrees C and 113 nW at 170 degrees C from a 0.9 V supply. Compared with the state-of-the-art voltage reference covering 170 degrees C, it consumes 4.36x less power, achieves a 50% lower minimum supply voltage (0.9 V), and reduces the average temperature coefficient (TC) by 2.37 x without per-chip trimming (27 ppm/degrees C).
Parity-time (PT) symmetry enables operating characteristics insensitive to parameter variations in resonant networks. When applied to inductive power transfer (IPT) realizations, PT-symmetric operation leads to coupling-independent transfer characteristics. However, it requires the coupling coefficient to exceed a critical value, which limits the achievable transfer distance. Relay coils are commonly introduced to enhance the effective coupling in multi-coil resonant networks. Existing analyses of relay-coil-coupled resonant networks are mainly based on coupled mode theory, where higher-order coupling terms are neglected and the approximation error becomes non-negligible as the coupling increases. To address this issue, this paper analyzes relay-coil-coupled networks using electric circuit theory, taking two three-coil networks with series-series-series (S-S-S) and parallel-parallel-parallel (P-P-P) compensations as examples. The exact PT-symmetric characteristics, including the PT-symmetric frequencies, the critical coupling coefficient, and the load range, are derived together with the corresponding coupling-independent transfer characteristics. Furthermore, an equivalent cascaded two-coil model is established to reveal the circuit-level relation between relay-coil and basic two-coil networks. The three-coil network can be exactly modeled as cascaded two-coil subnetworks operating under the same PT-symmetric condition. This cascaded model explains the extension of the PT-symmetric operating region in terms of equivalent coupling enhancement. Finally, a three-coil IPT prototype is built to verify the above analysis.
While high-resolution inputs improve recognition performance, they impose significant hardware overhead due to the computational intensity and memory bandwidth demands of conventional image signal processing (ISP). This inefficiency primarily stems from the structural mismatch between pixel-level ISP operations and the semantic information required by back-end neural networks. To address this, we propose an end-to-end RAW-based frequency-domain vision pipeline designed for efficient high-resolution detection. Under normal imaging conditions, the proposed pipeline bypasses the full human-perception-oriented ISP and directly processes RAW sensor data through lightweight channel mapping and a hardware-friendly approximate discrete cosine transform (DCT). To minimize front-end data movement, we reformulate the color-space transformation within the frequency domain and retain only a compact set of task-relevant low-frequency coefficients. Furthermore, we employ a multiplier-less, shift-and-add-based arithmetic architecture to simplify control logic and enhance energy efficiency. The pipeline is fully compatible with standard convolutional neural networks, requiring no architectural changes to the back-end detector. We validated the design using a Field-Programmable Gate Array (FPGA)-based prototype on the Xilinx ZCU102 platform integrated with a YOLOv3-Tiny detector. The system achieves real-time processing of 6144 & times;4096 RAW inputs at 33.18 frames per second (FPS). Experimental results demonstrate that the proposed frequency-domain approach maintains detection accuracy comparable to RGB-based pipelines while significantly reducing hardware resource utilization and power consumption.
This paper presents a multi-disturbance resilient CMOS bandgap reference (BGR) that achieves high precision and robust stability across a wide range of capacitive loads. To overcome the fundamental accuracy bottlenecks of high-precision bandgap references, the proposed architecture systematically addresses four major error sources. temperature drift, power supply fluctuations, internal circuit noise, and output load disturbances. A high-order curvature compensation network is employed to substantially eliminate nonlinear base-emitter voltage variations, achieving an ultra-low temperature coefficient. A three-stage chopper-stabilized operational amplifier is utilized to effectively attenuate low-frequency flicker noise and minimize offset drift. To suppress external supply noise, a pre-regulated low-dropout (LDO) regulator is integrated to isolate the sensitive bandgap core. At the output stage, a nested source-feedback (NSF) buffer topology is proposed to actively reduce the open-loop output impedance, ensuring absolute stability across a wide capacitive load range of 0 to 10 mu F. Fabricated in a 0.18 mu m CMOS process, the BGR occupies a total area of 0.8 mm(2). Measurement results confirm that, operating with a quiescent current of 150 mu A, the proposed design achieves a maximum temperature drift of only 3 ppm/degrees C over a temperature range of 40 degrees C to 125 degrees C after single-point calibration. Additionally, the BGR delivers a high power-supply rejection ratio (PSRR) of 120 dB at 10 Hz, and limits the 0.1 to 10 Hz integrated noise to 1.5 mu Vrms.
With the growing reliance on third-party intellectual property (3PIP) in integrated circuit design, the threat of functional failures induced by hardware Trojans embedded within these components has become increasingly critical. The high stealthiness and sophistication of hardware Trojans often render existing detection techniques insufficient for comprehensive coverage. As a result, runtime detection and recovery techniques have emerged and been proposed as a crucial last line of defense. However, these solutions typically introduce substantial hardware overhead, limiting their practicality in resource-constrained applications such as edge computing. To address this critical challenge, this work leverages approximate computing to develop low-cost runtime recovery strategies against hardware Trojans. Specifically, it begins by analyzing the challenges and potential opportunities introduced into existing security schemes when approximate computing is applied. Based on this analysis, targeted solutions and optimization techniques are proposed. These are then integrated into a unified framework that explores the trade-off between hardware resource usage and computational accuracy while maintaining circuit-level security. Experimental results across several commonly used applications demonstrate that the proposed framework can achieve over 20% of hardware resource savings with only a 5% reduction in accuracy. To the best of our knowledge, this work is the first to systematically integrate approximate computing with runtime hardware Trojan recovery, providing a new cost-effective direction for circuit-level security in resource-constrained systems.
This paper presents a high-precision successive approximation register (SAR) analog-to-digital converter (ADC) with a g(m)-constant voltage-controlled oscillator (VCO)-based comparator for enhanced energy efficiency. A cascode voltage-controlled current source and an isolation transistor are integrated in the comparator to stabilize the drain voltage of the current limiting transistor, reduce noise induced by operating region variations, and suppress kickback noise of the comparator to improve the precision of the ADC. The prototype SAR ADC fabricated in 65 nm CMOS process achieves a SNDR of 84.2dB at a 1 MS/s sampling rate, while consuming only 100.5 mu W, corresponding to a Schreier FoM of 181.2 dB.
This paper proposes a comprehensive frequency-domain aliasing-tones analysis for filtering-by-aliasing (FA) behaviour based on a linear aperiodic discrete-spectral (LADS) representation, which provides a unified framework for analysing time-varying devices and aliasing effects. In the analysis of aliasing tones, the FA behavior is interpreted in terms of spectral spreading, spectral weighting, and spectral summation, thus intuitively explaining the frequency-selective cancellation caused by spectral aliasing. Furthermore, the corresponding frequency-domain transfer-function expressions are derived to enable the filtering capability to be precisely analysed and quantitatively evaluated with regard to blocker frequencies and adjacent-channel rejection, while explicitly revealing the trade-offs among key performance metrics. To validate the proposed analysis, simulation results are presented to show that the frequency responses obtained using the frequency-domain aliasing tones are consistent with those derived from conventional linear periodically time-varying (LPTV) analysis. Finally, the frequency-domain aliasing tones are applied to analyse practical non-ideal effects, including non-ideal series impedance, LPTV impedance, and non-ideal integrator response, thereby revealing their underlying physical origins and providing insight for further optimisation of the FA technique.
This paper presents a phase-interval balanced ring voltage-controlled oscillator (RVCO) with ultra-low power and ultra-wide frequency tuning range (FTR), for wearable or implanted brain-computer sensors. An approximate phase-interval error analysis model is developed to quantify the phase imbalance introduced by switched capacitors in phase-combined RVCOs. Based on this model, digitally controlled switched capacitors are incorporated into the low-power RVCO core to significantly expand the FTR while reducing phase-interval imbalance. An enhanced time-window phase combiner (E-TWPC) is further proposed to improve tolerance to residual phase-interval errors and mitigate phase noise degradation over a wide operating range. Prototyped in 65-nm CMOS, the proposed oscillator consumes only 16-144 mu W and achieves an FTR of 440 MHz-2.6 GHz (142%) under 4-bit digital control. This RVCO maintains comparable phase noise performance to existing RVCOs, while occupying a compact area of 0.00057 mm(2), achieving a FoM(TA) of 218.2 dBc/Hz at 1 MHz offset.
This paper proposes a chip-PCB hybrid switched-capacitor (SC) physical unclonable function (PUF) to detect PCB-level attacks using only two, or even a single, sense IO. An on-chip capacitor array is employed to compensate for and balance the capacitances between the two sense IOs (or between one sense IO and one on-chip capacitor). The sense IOs then form a sense SC circuit, while two on-chip capacitors constitute a reference SC circuit to which a small threshold capacitor is further introduced. An unchanged relationship between the output voltages of the sense and reference SC circuits under variation of the threshold capacitor in the reference SC circuit indicates that the capacitance mismatch between the two sense IOs (or between a sense IO and an on-chip capacitor) exceeds a predefined threshold. This reflects a significant change in the capacitance of the sense IO, potentially caused by PCB-level desoldering, resoldering, or probing attacks. Furthermore, the two capacitors in the reference SC circuit are partitioned into multiple sub-capacitors to form multiple sub-reference SC circuits. Together with the sense SC circuit, these structures constitute multiple SC PUF units that generate PUF keys strongly correlated with the parasitic capacitance of the sense IOs. The proposed anti-PCB-level attack scheme is fabricated in a 180nm CMOS process for silicon verification. Measurement results demonstrate that the SC PUF output keys effectively reflect IO capacitance variations caused by PCB-level attacks, thereby providing reliable protection for the chip against such attacks.
As wireline data rates continue to scale, pre-cursor inter-symbol interferences (ISIs) become increasingly significant due to channel dispersion and bandwidth limitations. Conventional decision feedback equalizers (DFEs) are inherently limited to cancel post-cursor ISIs only, while feedforward equalizers (FFEs) suffer from fundamental trade-offs, either reducing signal swing at the transmitter or amplifying noise at the receiver. To address these problems, a bi-directional DFE (BiDi-DFE) architecture is introduced to equalize both pre-cursor and post-cursor ISIs without noise amplification. Leveraging the opposite decision-propagation directions of the two DFE paths, a maximum likelihood sequence cross-detection (MLSCD) scheme is further proposed to detect error events and selectively enable maximum likelihood sequence detection (MLSD), improving energy efficiency. A 112-Gbps DSP-based wireline transceiver deployed over a 45 dB insertion-loss channel achieves a bit error rate (BER) reduction from 1.5 & times;10(-6) to 1.6 & times;10(-7) using BiDi-DFE, and further to 7.5 & times;10(-9) with MLSCD, incurring only an estimated 20-mW power increase in post-layout compared with a 1-tap DFE while achieving an estimated 40-mW power saving in post-layout relative to a full-state MLSD (FS-MLSD). These results indicate that the proposed BiDi-DFE and MLSCD framework offers a scalable solution for future high-loss, high-speed wireline transceivers.
We present a self-referenced, calibration-free control circuit for thermally stabilizing silicon photonic microring modulators. The technique monitors the through and drop optical output powers of the ring and locks operation at the point of equal intensity, a bias close to the maximum optical modulation amplitude (OMA). Unlike prior schemes that require reference calibration or external dithering, this method derives its error signal directly from the device outputs, enabling a fully automatic tuning. The all-analog feedback circuit, implemented on a printed circuit board (PCB), achieves stable operation for modulation up to 14 Gb/s with a response time below 1 ms, and a wavelength tuning range of 5.88 nm. The system maintains lock under +/- 6 dB input-power fluctuations, 1 nm laser drift, and 10 degrees C ambient variation, demonstrating strong immunity to environmental and process variations. The approach offers a compact, low-power solution readily adaptable to monolithic electronic-photonic integration, advancing practical stabilization of high-speed silicon photonic transmitters.
This paper presents an 8-module, fully integrated voltage regulator (FIVR) designed for XPU power delivery, operating from 1.8V to 1V with a switching frequency of 50 MHz and supporting up to 90 A total output current. To address the current imbalance issue in multi-module parallel systems, a novel matrix-based current-sharing scheme is introduced. In this architecture, each power module exchanges sensed inductor current information with its two adjacent modules, forming overlapping horizontal and vertical control loops that ensure uniform current distribution across the entire module matrix. The proposed method eliminates the need for long-distance signal transmission, enhances signal integrity, and reduces the risk of single-point failure while maintaining high scalability and design consistency. A detailed analysis is demonstrated to clarify the stability of this complicated multi-loop system and provides a practical method to stabilize the current-sharing loop without introducing complex compensators. The prototype, fabricated in a 28 nm CMOS process, achieves a peak efficiency of 85.6% with 8 modules in parallel and exhibits an estimated current-sharing accuracy within 10.6%. Under the full-load condition, the proposed current-sharing method in such an 8-module FIVR holds the temperature variation among modules less than 10.5 degrees C.
Semiconductor-based cryogenic quantum processors require the accurate biasing of a large number of gate electrodes, which are typically individually wired to room-temperature DACs. To prevent the wiring bottleneck when scaling to future very large processors, this work proposes a scalable cryo-CMOS DAC that can operate with a S&H demultiplexer close to the quantum processor. By adopting an integrator-based switched-capacitor DAC with a dynamically-biased high-voltage output stage, both the required large output range and high resolution can be achieved while sharing the DAC over a large number of electrodes, thus improving the power efficiency. Fabricated in a 22-nm FinFET technology, the DAC occupies 0.076 mm(2). At 4.2 K (RT), it achieves an LSB of 57.1 mu V (68.3 mu V) over a 3 V range with 188 mu V-rms (192 mu V-rms) noise and 36.5-LSB (12.6-LSB) INL while dissipating 157 mu W (138 mu W). The low power dissipation and the potential to drive more than 30,000 electrodes paves the way for scalable biasing of quantum processors operating down to mK temperatures.