
In this paper, we introduce a new method for shifting-in-memory exponent normalization and shifting-in-memory mantissa alignment. This approach allows us to add exponents and find the maximum sum simultaneously. We also replace the traditional subtraction and shifting for mantissa alignment with CA and shifting-in-memory operation, enabling FP-CIM to be achieved with lower delay, area, and power overheads. The macro is designed in TSMC 28nm process, with a memory size of 32Kb. Simulation results indicate that the computing core operates at a frequency of 280MHz at a 0.55V supply voltage, and achieves an energy efficiency of 13.5 TFLOPS/W.
In real-time video conferencing systems, webcams often apply image resizing methods, such as nearest-neighbor, bilinear, bicubic, and Lanczos interpolation, to highlight facial regions and enhance user experience. However, interpolations introduce significant high-frequency artifacts that distort perceived image quality. Our work discovered that existing state-of-theart image quality assessments (IQA) greatly overestimate the sharpness in images resized by nearest-neighbor interpolation, failing to extract useful information in the image's high-frequency components. To address this, we propose FRIEREN (Face Resizing Image Detail Quality Evaluation via Robust Estimation of Image Naturalness), a novel IQA for detail quality evaluation that integrates measures of image naturalness, including motion noise, spatial noise, and HVS-based sharpness. Designed features are fed to Kolmogorov-Arnold Networks (KANs) for quality prediction. Experimental results show that FRIEREN effectively and accurately evaluates the face image detail quality scaled by different interpolations with low computational complexity, making it suitable for quality-aware vision systems.
This work presents an energy-efficient multi-bit (8-b) current-based analog SRAM compute-in-memory (CIM) architecture for MAC operations. Our key idea is to group 8-bit inputs and weights into 2-bit segments for performing MAC operations. This improves signal margin without requiring external circuitry. We leverage a 1C-2C capacitive network for efficient weight encoding and also use an analog shift-and add technique for accumulating partial sums with only 1C and-4C capacitors, thus avoiding the need for higher capacitor values. This approach enhances charge redistribution efficiency and minimizes ADC overhead. The proposed design features a 64x128 6T SRAM array, which can perform 2048 8-bit MAC operations in four cycles with a latency of 15ns. The architecture is implemented in TSMC 65nm technology at 1.2V. The architecture maintains linear MAC operations across different input and weight cases and process corners. The proposed CIM macro delivers a throughput of 136.5 GOPS and MAC energy efficiency of 6.18 TOPS/W while maintaining a signal margin of 32 mV. Our approach achieves an inference accuracy of 98.62%, 90.38% and 69.46% on MNIST, CIFAR-10 and CIFAR-100, respectively.
This paper proposes an edge computing system for real-time arrhythmia classification in the elderly, leveraging a novel one-dimensional SqueezeNet optimized for ECG signal. The special design of Fire Module leverages point-wise (1x1) convolutions to compress channel dimensions and combines them with wider (3x1) convolutions to effectively balance parameter efficiency and temporal feature extraction. By directly receiving the raw input, unnecessary signal conversion operations are effectively avoided. The compact architecture (0.36 million parameters) attains 91.41% accuracy, 99.56% recall, 98.56% precision, and 99.06% F1 score on the MIT-BIH Arrhythmia Database. These results show better performance for classifying previously unseen data of the proposed model than the state-of-the-art works. Evaluated only on elderly subjects, accuracy rises to 96.25% without any decline in other performance metrics, emphasizing the necessity of conducting separate training for the ECG of the elderly. The trained model is quantized and deployed to an Arm Cortex-M7 processor, yielding an inference latency of 146 ms with only 121 kB RAM usage to enable on-device processing. This demonstrates that high-performance elderly-specific arrhythmia classification can be realized on resource-constrained edge computing system without cloud dependency.
The multi-arm series capacitor buck converter (SCBC) improves output current and reduces ripple through interleaving control, making it ideal for power conversion applications. However, conventional control methods cause current imbalance between arms when the duty cycle exceeds 50 %. This paper proposes a simple interleaved control strategy that fixes one arm's duty cycle while adjusting the others to maintain balance. A three-arm SCBC with an interleaved controller is designed and implemented, achieving balanced inductor currents without additional current sampling. SIMPLIS-based simulations verify the effectiveness of the proposed strategy, demonstrating current balance and robust performance under dynamic conditions.
This paper proposes a TC (Temperature Coefficient) improved dual-output Bandgap Reference (BGR) circuit using the Bowl-shaped Curvature Compensator (BCC) and digital PVT Detector. Dual BCCs are used to generate two bowl-shaped currents for the curvature compensation of the two outputs of the BGR. The digital PVT detector generates 13 digital control signals to calibrate the current mirror transistors of the BGR and improve the TC against the PVT variations. The proposed design is implemented with TSMC 0.18 mu m BCD process. The core area is 302.56 x 269.29 mu m(2). The simulated output reference voltage is 0.7 V and 1.41 V with averaged TC of 3.35 ppm/degrees C and 4.18 ppm/ degrees C, respectively, at the range from -40 degrees C to 125 degrees C. The improvement of averaged TC are 832.2% and 828.7% after the BCC and PVT compensation.
As on-device inference gains traction for privacy protection and low-latency AI services in edge environments, the demand for flexible and lightweight SoC architecture is increasing. This paper proposes a novel SoC design that directly executes IREE (Intermediate Representation Execution Environment) VM Bytecode through an interpreter-based approach implemented on a RISC-V core. The system is validated on a Zynq ZC706 FPGA board and consists of a compiler-hosted PS (Processing System) and an interpreterexecuting PL (Programmable Logic). To improve performance, DMA and cache modules are integrated to optimize data movement and memory access during inference. The proposed system successfully executes various machine learning models, including MUL, MMT, and MNIST, and achieves up to 29 % reduction in inference time depending on the model characteristics. This work demonstrates the feasibility of compiler-driven, reconfigurable inference architectures for embedded AI applications without the need for model-specific hardware redesign.
Minimum Mean Square Error (MMSE) detection is widely used in Multiple-Input Multiple-Output (MIMO) systems due to its ability to mitigate noise, reduce interference, and enhance system performance. However, its practical hardware implementation is limited by the high computational cost of matrix inversion. To address this challenge, we propose a hardware-efficient detection framework based on the limited-memory Broyden-Fletcher-Goldfarb-Shanno (L-BFGS) algorithm. The framework features a Dual-Phase L-BFGS (DP-LBFGS) structure that enables efficient parallel processing. Additionally, a hybrid optimization strategy combining input bit-width expansion and adaptive iteration termination is introduced to reduce quantization errors and computational overhead. Simulations on an 8x16 Quadrature Phase Shift Keying (QPSK) MIMO-OFDM system demonstrate that the proposed detector achieves Bit Error Rates (BERs) as low as 10(-7), maintaining over 95% of full-precision performance with significantly reduced complexity. Implementation on a Xilinx Virtex-7 FPGA shows a throughput of 848 Mbps and a hardware efficiency of 4.79 Mbps/K slices, outperforming existing designs by a factor of 2.25x to 9.04x.
This work presents a fully differential, offset-stabilized amplifier achieving sub-mu V offset and a 55 MHz gain-bandwidth product (GBW). A two-step offset calibration strategy is employed, combining digitally-assisted auto-calibration with dynamic offset compensation (DOC). Slow-settling auto-zeroing (AZ) is integrated with chopping through a customized clocking scheme, enabling bandwidth extension without compromising offset and noise performance. Implemented in a 180 nm BCD process, the prototype occupies an area of 1.32 mm(2). Simulation results show a 3s input-referred offset reduction from 1.6 mV to 615 nV and a nearly flat low-frequency noise spectrum without tones at the switching frequency. A noise floor of 7 nV/v Hz has been achieved, corresponding to a noise efficiency factor (NEF) of 13. These results demonstrate its potential to be widely adopted in high-accuracy and high-bandwidth applications.
The ultra-low-power operational transconductance amplifier (OTA) plays a crucial role in amplifiers and filters, which are implemented in many ultra-low-power Internet-of-Thing devices. Depending on the different specifications of the OTAs, the performance of those amplifiers and filters could be greatly changed. In addition, the effort to reimplement one OTA topology to target a different optimization is high. Therefore, this paper introduces an inverter-based OTA topology with common-mode feedback based on the NAND2-and-NOR2 gates to target different optimizations. By changing the different drive strengths of digital standard cells with fully automatic place-and-route flow, one can choose which optimizations are highlighted. This method significantly reduces the implementation effort. In this paper, we compare two different-strength designs from the same topology to target two different optimizations. The first OTA aims to have a state-of-the-art gain-bandwidth product, which is about 87.46kHz@C-L =150pF. The other is for a state-of-the-art area cost, which is about 6.8 mu mx3.6 mu m. Both designs are implemented using a TSMC 65nm CMOS technology node, which can operate under aggressively scaled supply voltages to 0.3V.
This paper proposed a Combined-Head Low-Rank Projection (CLRP) method and a Unified-PE Candidate Selection (UCS) method to mitigate storage and computational overheads caused by attention mechanisms in transformer models. The CLRP method exploits the cross-head low-rank properties of the projection weights and reduces their hidden dimensions by applying singular value decomposition (SVD) to their products. It achieves a 1.27x compression ratio and maintains high inference speed through cross-head interaction analysis. The UCS method rearranges the computational stages of Unified Attention PEs for efficient candidate selection and reduces the computational workload by pruning unimportant candidates from long-sequence inputs. An efficient accelerator integrated with the above methods is designed. Compared to four state-of-the-art designs of attention-only accelerators, the proposed accelerator achieves an average 1.99x energy efficiency and 2.15x area efficiency.
This paper achieves energy savings and memory reduction in an FPGA implementation of speech recognition using a CNN by applying stream processing and model miniaturization. Typically, the processing of the next convolutional layer begins only after the completion of the previous layer. Therefore, a large-capacity interlayer memory is required to hold all the outputs of the previous layer. In this study, we leverage the fact that speech data is one-dimensional and introduce stream processing, where subsequent convolution processes are initiated midway through the previous convolution process to consume the output of the previous layer. This presents a reduction of 75% in the interlayer memory usage. We also introduce parallel processing to speed up the first convolutional layer without requiring a larger memory bandwidth. A reduction of 49% in the run time and 41% in the energy consumption could be achieved. To further reduce the memory size, we reduced the sampling rate of the input audio to 8 kHz, designed a compact convolutional neural network (CNN) model, and quantized the weight parameters of the layer with the largest model size to 4 bits. Consequently, the model size is reduced to 21-27 kB and recognition accuracies of 74.1-77.6% are achieved.
This paper presents the development of a knee rehabilitation device designed in strict compliance with ISO 13485, the international standard for Quality Management Systems (QMS) specific to medical devices. The aim is to create a safe, effective, and user-friendly rehabilitation solution that supports patients recovering from knee injuries or surgeries, with a particular focus on elderly patients who often experience reduced mobility and longer recovery periods. The development process integrates interdisciplinary engineering approaches, including biomechanical analysis, embedded systems, and ergonomic design. Key features of the device include adjustable motion control, data logging for physiotherapist evaluation, and user interface optimization to enhance patient engagement and safety. Risk management, traceability, and documentation were rigorously implemented throughout the design phases in accordance with ISO 14971 and ISO 13485 requirements. Testing was conducted to verify device performance, safety, and compliance with regulatory standards. The resulting system demonstrates both functional reliability and the ability to improve rehabilitation outcomes. This work highlights the importance of aligning product development with international medical device standards to ensure market readiness and patient safety.
Deep Neural Networks (DNNs) excel in various artificial intelligence applications due to their ability to process large amounts of data and their computational power. However, high power consumption poses a significant challenge in DNN hardware design. Spiking Neural Networks (SNNs) offer a promising alternative, mimicking biological neural systems and processing information through the firing of neurons, which makes them more energy-efficient for time-series data and complex pattern recognition tasks. Despite their advantages, existing SNN learning methods often struggle with low computing accuracy due to the improper spike inhibition. To enhance learning efficiency, we consider the conventional convolution operation in contemporary DNN models and present a novel Block Spiking Neural Network (BSNN) in this work. Similar to the localized feature extraction in the convolution layer of DNNs, the proposed BSNN only allows excited neurons to inhibit the other neurons in a restricted region, called a block. Because the proposed BSNN retains spikes according to the individual block, the features of the input data are easier to consider, thereby improving the computing accuracy. Experimental results demonstrate our proposed BSNN significantly improves performance, increasing accuracy by 3.57% to 5.57% over a conventional SNN approach while simultaneously improving hardware efficiency by 95.62% to 96.07%.
Diffusion models have achieved strong performance in generative vision tasks but suffer from high compute and memory demands due to their iterative U-Net structure. This paper presents a hardware-efficient diffusion accelerator with three key architectural optimizations. First, a Winograd-Enhanced Dual-Mode Dispatcher improves convolution throughput by eliminating im2col overhead and maintaining high PE utilization. Second, a Dynamic Sparse Attention Engine predicts and prunes lowimportance attention scores at runtime, reducing unnecessary multiplications. Third, a Reconfigurable Mixed-Precision Processor adapts different bit-width workload demands to avoid resource waste. The proposed accelerator is implemented in TSMC 28nm CMOS technology and achieves a peak energy efficiency of 120.59 TOPS/W under 90% sparsity during attention computation, demonstrating its suitability for real-time and low-power diffusion inference.
Single image dehazing is a critical computer vision task that aims to recover a clear image from a hazy input. In recent years, both learning-based and prior-based dehazing methods have been rigorously developed, yet each has limitations in haze removal performance. In this work, we propose the Adaptive Patch Size Dark Channel Prior Network (APSDCP-Net), a hybrid dehazing network that combines the dark channel prior with a learnable architecture. By reframing the dehazing problem as a patch size selection task, the proposed model combines multiple patch-sizes dark channels using learned soft weights, resulting in more accurate transmission estimation and enhanced dehazing performance. Experiments show that the proposed APSDCP-Net achieves superior performance compared to state-of- the-art dehazing methods.
The paper presents a 4th-order active-RC low-pass filter enhanced with negative-resistance (negative-R) circuit to mitigate the gain and bandwidth limitations of Operational Transconductance Amplifiers (OTAs). The imperfect virtual ground at the input of an OTA induces distortion and input current losses. With the assistance of negative-R, these imperfections are eliminated by gain enhancement. Implemented in open-source 130nm PDK, this concept is extended to a 4th-order active-RC filter, which is realized as two cascaded Tow-Thomas biquads. The negative-R assisted filter achieves a 10MHz bandwidth extension, i.e. 71% improvement over a conventional active-RC implementation. The filter achieves a spurious-free dynamic range (SFDR) of 57.15 dB by consuming 3.15mW from a 1.8-V power supply.
Statistical performance analysis of low-noise amplifiers (LNA) under process, voltage, and temperature (PVT) variations and device mismatches is crucial for ensuring design robustness. Unfortunately, traditional Monte Carlo (MC) analysis is often computationally prohibitive for complex RF circuits. In this work, a generalized polynomial chaos expansion (gPCE) model is constructed as a surrogate model from a limited set of circuit simulations to approximate a comprehensive figure-ofmerit (FoM). The results demonstrate that the third-order gPCE model achieves high predictive accuracy, confirmed by a high coefficient of determination (R-2 = 0.925) and a low root mean squared error (RMSE = 0.007). In addition, the proposed model enables the extraction of these statistical moments and tail-risk metrics over 26 times faster than the benchmark MC simulation.
This work proposes a capacitor-less low-dropout regulator (LDO) based on super source follower (SSF) architecture in a 180nm process achieving improved wide band power supply rejection (PSR) performance and fast transient response. With a novel topology that couples the LDO and a symmetric-load voltage-controlled oscillator (VCO) with bias, part of the bias circuit proportional to the oscillation intensity flows into the LDO, optimizing the LDO's PSR performance in the mid-to-low frequency. This work is implemented in a 180nm CMOS process. Simulation results demonstrate that incorporating the additional LDO current improves mid-to-low frequency PSR by at least 10dB compared to the condition without the current. Furthermore, the PSR performance remains -28dB at frequencies up to 10GHz, indicating robust PSR performance within wide band. The system takes 37ns to jump from zero load to 20mA full load, with a 210mV overshoot.