
This paper proposes a HW/SW Co-Optimized OnLine 3D Gaussian Splatting modeling accelerator, COOL-splat, for real-time AR/VR applications on mobile devices. To address high external memory access energy and redundancies in computation, COOL-splat introduces novel contributions across SW, core, and PE levels. Projected gradient interpolation-based masked training (SW), bi-directional rendering core (Core), and reconfigurable gaussian computation units supporting mixed-precision (PE). These enhancements achieve an average modeling efficiency of 10.1 mJ/iter, outperforming state-of-the-art processors by 1.49×.
This paper presents a multicore FPGA-based hardware accelerator for RLWE encryption and decryption using the TFHE scheme. The proposed modular architecture consists of two main components, namely the encryption and decryption units, each implemented with up to four parallel RTL cores. The implementation on an Xilinx ZCU104 FPGA operates at maximum frequencies of 214 MHz for encryption and 300 MHz for decryption, supporting polynomial degrees up to 4096. The proposed design achieves throughputs of 9.07 Gbps and 12.73 Gbps for encryption and decryption, respectively. (Keywords: TFHE, RLWE, Multicore, Torus)
Audio-visual classification on edge and neuromorphic platforms requires high recognition accuracy under tight power constraints. Existing attention-based spiking Transformers still suffer from quadratic attention complexity with respect to audio-visual token length, making deployment on resource-constrained systems difficult. This paper proposes Spikceiver, an audio-visual spiking Transformer that targets efficient multimodal fusion with a compact model size. Spikceiver integrates spiking bottleneck fusion using a small set of learnable Spiking Bottleneck Tokens that attend to each modality to produce compact latent representations for fusion, and an efficiency-oriented Spiking Depthwise-Separable Splitting front-end that replaces vanilla convolution for tokenization. On UrbanSound8K-AV and enhanced CIFAR10-AV (CIFAR10-AV(e)), Spikceiver achieves 95.05% and 89.78% top-1 accuracy with 4.84M MACs and 1.271M and 1.273M parameters, resulting in Efficiency-Accuracy Scores (EAS) of 0.963 and 0.895. Compared with the Multimodal Bottleneck Transformer (MBT-4-256), Spikceiver improves accuracy (95.05% compared with 84.01%, and 89.78% compared with 83.46%) while reducing MACs (4.84M compared with 7.26M) and parameters (about 1.27M compared with 6.40M). Compared with the Spiking Multimodal Transformer (SMMT-1-256), Spikceiver reduces MACs by more than eight hundred times with fewer parameters, demonstrating a favorable efficiency-accuracy trade-off for audio-visual spiking inference on resource-constrained platforms.
Nowadays, Fast Fourier Transform (FFT) and inverse FFT (IFFT) serve as foundational computational kernels in modern signal processing, and notably, as a critical building block in Falcon, a prominent post-quantum cryptography (PQC) digital signature scheme. Unfortunately, most existing memory-based FFT accelerators remain constrained by bank conflicts and arbitration overhead, which limit sustained throughput and energy efficiency in edge and SoC deployments. To address these challenges, this paper proposes Conflict-Free and eXponent-scaled FFT/IFFT (CFX), a co-designed architecture that targets high performance with minimal resource footprints. Specifically, CFX incorporates three key innovations: conflict-free logical bank mapping, stage-wise exponent normalization, and schedule-aware twiddle compaction. The CFX architecture has been implemented and verified at the system-on-chip level on the Xilinx ZCU104 FPGA operating at a maximum frequency of 220 MHz. The SoC design utilizes 13,144 LUTs, 14,604 FFs, 6 BRAMs, and 24 DSPs, with a total power consumption of 3.788 W. Real-time comparisons with powerful CPUs demonstrate that CFX achieves up to 81.4 times higher throughput and up to 524 times better energy efficiency (PDP). Furthermore, compared to state-of-the-art FPGA-based accelerators, CFX delivers throughput improvements of up to 20.88 times and area efficiency improvements of up to 41.8 times.
In heterogeneous multi-core NUMA systems, thread-to-core/node mapping simultaneously affects load balancing, NUMA locality, and synchronization waiting; therefore, no single mapping strategy is optimal for all applications. This paper proposes a heuristic key-metric-driven random-forest mapping selector (HML-RF): with a lightweight Scatter profiling run on a target application to extract a small set of features, HML-RF predicts an energy–delay product (EDP)-optimal strategy among five mapping strategies (Packed/Scatter/HPO/MPO/SPO); in particular, the Synchronization-aware Priority Option (SPO) constructs synchronization priorities using workload and wait-intensity features from profiling data to reduce waiting times. We evaluate on 10 PARSEC benchmark applications (large) in the Sniper simulator, and the results show that HML-RF achieves an average EDP regret of 0.299%, a geometric-mean EDP/Oracle of 1.0030, and significantly outperforms POSM and the best single mapping strategy in statistical tests.
This paper presents the RX VPU, a newly developed Vector Processing Unit (VPU) designed as a coprocessor for the RX series of embedded microcontrollers. The RX VPU adopts a variable-length vector architecture and features a compact, low-power implementation suitable for embedded systems. It is specifically designed to accelerate CNN inference on embedded devices while maintaining low power consumption. The RX VPU provides instructions optimized for int8-quantized CNNs, which are widely used in embedded inference workloads, enabling efficient execution of compact CNN models. The architecture is particularly optimized for anomaly detection workloads using int8-quantized CNN models, where inner-product accumulation operations account for approximately 40% of all executed instructions. This work introduces a 4× widening reduction MAC instruction that accelerates these operations within a single instruction. Using this instruction, inference throughput on the Arm ML-Zoo anomaly detection model improves by 1.85×, achieving up to 30.0 inference/s. The coprocessor also achieves 1025.6 GFLOPS/W at 600 MHz, demonstrating that highly efficient ML inference can be achieved even under stringent power and area constraints.
A cyclic pipeline analog-to-digital converter (ADC) for light and rainfall detection sensors is presented in this paper. This architecture meets the requirements of real-time and accurate acquisition of key parameters. The proposed ADC employs a second-stage cyclic ADC which consists of a second-stage pipeline ADC, operating at a frequency of 4 MHz to achieve high energy efficiency. To minimize the chip area, the sub-ADC is shared by the second pipeline stage. The amplifier incorporates a folded cascode topology within a Class AB output stage to achieve high speed, resolution, and energy efficiency. Fabricated in a 180-nm standard CMOS process, the prototype achieves a peak signal-to-noise ratio (SNR) of 72.2 dB and spurious-free dynamic range (SFDR) of 73.5 dB, resulting in a Schreier figure of merit (FoM) of 0.3 pJ/conv-step, while consuming only 2.8 mA from a 1.8 V supply. The ADC provides an energy and area efficient solution for sensors to monitor autonomously.
This paper presents a real-time and energy-efficient 3D Gaussian Splatting (3DGS)-based geometric and semantic mapping processor for embodied agents. The proposed processor exploits two levels of redundancy: 1) Embedding Redundancy Handling Unit (ERHU), which skips occluded-object embeddings and prunes low-importance channels to reduce 65.6% computation and 77.7% memory access, and 2) Temporal Redundancy Handling Unit (TRHU), which restricts mapping regions to moved objects using bounding-boxes, reducing 93.4% mapping latency. Fabricated in 28nm FDSOI, the processor achieves 70.7 FPS at 0.02 μJ/point for 640×480 dense semantic mapping.
In modern processor design, vector instructions are widely used in general-purpose CPUs. One challenge in introducing vector instructions into general-purpose CPUs is that naively applying out-of-order execution to them incurs a prohibitively large area overhead. We found that many vector instructions do not benefit from instruction reordering, whereas only a subset does, namely cache-miss-prone vector loads and their producer instructions. Based on this observation, we propose a novel method that significantly reduces hardware resources, such as vector physical register files and load-store queues, by allowing reordering only for vector instructions that benefit from it. We evaluated the proposed method using a cycle-accurate simulator. The results show that our method reduces energy consumption by 22.34% with only a 0.41% performance degradation.
AI scaling is severely bottlenecked by the latency and energy costs of von Neumann memory-processor data shuttling. This work investigates neuromorphic analog computing to execute massive Matrix-Vector Multiplications (MVMs) efficiently in-memory. We comprehensively analyze a 16 nm FinFET analog multiplier leveraging a modified Leaky Integrate-and-Fire (LIF) neuron topology. SPICE simulations validate its function as a hardware-efficient multiplier (fout ∝ I1 × I2), achieving 4.84 GHz peak throughput, high linearity (R2 ≈ 0.99), and ≈ 2.2 fJ/spike efficiency. Furthermore, parametric sweeps demonstrate robust mode-locking domains against ≈ 0.2 μA input tolerances. This stability allows the continuous transfer function to be actively discretized, providing intrinsic noise immunity for reliable analog arithmetic.
Emerging nanoscale X-ray imaging techniques are expected to eventually require Bragg peak localization at higher rates (e.g., 1MHz or higher) on streaming detector data. However, conventional offline analysis and even deep-learning-based approaches on host machines require substantial computational costs and hinder real-time feedback for in-situ experiments. BraggNN, a CNN-based approach to Bragg peak localization, is expected to be an enabler of real-time peak localization integrated with an AI inference hardware accelerator running at the edge. In this paper, we evaluate BraggNN inference on the Gemmini accelerator within the Chipyard RISC-V SoC framework, comparing the ONNX Runtime and hand-tuned C API approaches to analyze bottlenecks and explore optimization strategies through cycle-accurate simulation and hardware profiling. The optimized C implementation achieves 38 microseconds per single-patch inference with an 892 times speedup over the CPU baseline without accuracy degradation.
Manufacturing process variations increasingly cause nominally identical chips to differ significantly in power consumption. Rather than treating this variability as a defect to mitigate, we show it creates a natural efficiency hierarchy that system schedulers can exploit. We propose strategic job deferral: selectively delaying workloads whose deadlines permit deferral to run on the most power-efficient nodes in a cluster. This requires sorted assembly, which pairs CPUs and DRAM modules by measured quality to form predictable efficiency tiers. On a simulated HPC cluster with synthetic workloads under realistic submission patterns, evaluations demonstrate up to 16.2% energy savings with minimal impact on job completion times and no hardware modifications. These findings reframe chip-level manufacturing variability from a packaging problem into a system-level scheduling opportunity.
This paper presents MicroXPU, a 28 nm energy-efficient large language model (LLM) accelerator featuring a microscaling (MX) format-tailored digital CIM macro that addresses the key challenges of combining DCIM with MX quantization. It features exponent-based cyclic mantissa alignment (ECMA & CPCA), minimum exponent post-detection (MEDU), bit-wise dynamic approximation with reference voltage selection (BARS), and shared scale-based block sorting (SSBS). Fabricated in Samsung 28 nm CMOS, MicroXPU achieves 96.3 TFLOPS/W macro energy efficiency and 28.2 TFLOPS/W system energy efficiency under MXFP4, demonstrating 1.58× and 3.32× system and macro energy efficiency improvements over state-of-the-art DCIMs.
Methods for recording brain activity, particularly electroencephalography (EEG), which is a noninvasive and low-cost technique, have been widely investigated. In particular, non-contact electrodes have attracted considerable interest because of their ability to perform long-term and unobtrusive EEG measurements. In these methods, the buffer circuit is a crucial component that preserves the accuracy of the measured signals. However, for EEG signal acquisition using noncontact electrodes, the source impedance is considerably high and requires a buffer circuit with an even higher input impedance, in addition to low-power operation. We fabricated and measured/simulated a low-power and high-input-impedance buffer circuit that uses a self-cascode structure and is intended mainly for low-frequency signals, including EEG signals. In simulation, the buffer circuit achieved a power consumption of 5.68 μW and an input impedance of ≥ 45.8TΩ in the frequency range of up to 100 Hz at a supply voltage of 3.3V.
Cryptographic operations are critical for securing IoT, edge computing, and autonomous systems. However, current RISC-V platforms lack efficient hardware support for comprehensive cryptographic algorithm families and post-quantum cryptography. This paper presents Crypto-RV, a RISC-V co-processor architecture that unifies support for SHA-256, SHA-512, SM3, SHA3-256, SHAKE-128, SHAKE-256 AES-128, HARAKA-256, and HARAKA-512 within a single 64-bit datapath. Crypto-RV introduces three key architectural innovations: a high-bandwidth internal buffer (128x64-bit), cryptography-specialized execution units with four-stage pipelined datapaths, and a double-buffering mechanism with adaptive scheduling optimized for large-hash. Implemented on Xilinx ZCU102 FPGA at 160 MHz with 0.851 W dynamic power, Crypto-RV achieves 165 times to 1,061 times speedup over baseline RISC-V cores, 5.8 times to 17.4 times better energy efficiency compared to powerful CPUs. The design occupies only 34,704 LUTs, 37,329 FFs, and 22 BRAMs demonstrating viability for high-performance, energy-efficient cryptographic processing in resource-constrained IoT environments.
Pauli strings are a fundamental computational primitive in hybrid quantum-classical algorithms. However, classical computation of Pauli strings suffers from exponential complexity and quickly becomes a performance bottleneck as the number of qubits increases. To address this challenge, this paper proposes the Pauli Composer Accelerator (PACOX), the first dedicated FPGA-based accelerator for Pauli string computation. PACOX employs a compact binary encoding with XOR-based index permutation and phase accumulation. Based on this formulation, we design a parallel and pipelined processing element (PE) cluster architecture that efficiently exploits data-level parallelism on FPGA. Experimental results on a Xilinx ZCU102 FPGA show that PACOX operates at 250 MHz with a dynamic power consumption of 0.33 W, using 8,052 LUTs, 10,934 FFs, and 324 BRAMs. For Pauli strings of up to 19 qubits, PACOX achieves speedups of up to 100 times compared with state-of-the-art CPU-based methods, while requiring significantly less memory and achieving a much lower power-delay product. These results demonstrate that PACOX delivers high computational speed with superior energy efficiency for Pauli-based workloads in hybrid quantum-classical systems.
This panel will discuss sustainable AI and quantum computing systems. Key topics will include energy efficiency, scalability, and reliability of each computing technology from each panelist's perspective. The discussion will continue with comments and questions from the audience to help us understand future directions.
In the CRYSTALS-Kyber encryption algorithm, polynomial multiplication is one of the most time-consuming tasks, making the optimization of its hardware implementations crucial. In this paper, we propose a hardware architecture for polynomial multiplication in the CRYSTALS-Kyber encryption algorithm, including butterfly units for implementing numbertheoretic transform (NTT), inverse NTT (INTT), and point-wise multiplication (PWM) operations, as well as a conflict-free implementation for constant-geometry memory access on dualport memory. Unlike previous works that focus solely on optimizing NTT and INTT circuits, our approach also optimizes the PWM circuit. Experimental results show that our approach achieves excellent performance in terms of both clock frequency and chip area. Compared to hardware designs that use the same design parameters, our approach can achieve a smaller areatime product.
Convolutional neural networks (CNNs) face significant deployment challenges on edge devices due to their high memory and computational demands. Depthwise separable convolution (DSC), with its reduced complexity and comparable accuracy, offers a promising solution. In this work, we propose an algorithm-hardware co-design that integrates a group-wise uniform pruning (GUP) strategy specifically tailored for DSC. This approach optimizes depthwise convolution (DWC) by channel pruning and pointwise convolution (PWC) by kernel pruning, with pruned channels in DWC propagating to PWC. Driven by the GUP strategy, we design a dual-engine DSC accelerator that fetches only unpruned weights and activations. Additionally, we optimized the dataflow and timing to minimize latency. The proposed GUP DSC accelerator is implemented in a 22nm FDSOI technology, operating at a frequency of 1 GHz after signoff, and occupies an area of 0.4 mm(2). With 50% pruning applied to the MobileNetV1 model, it reduces MAC operations by 74.22%. The accelerator achieves an average energy efficiency of 24 TOPS/W and a throughput of 8365 GOPS.
Nowadays, traditional hash functions and stream ciphers, including SHA256, SM3, BLAKE256, BLAKE2s, ChaCha20, and Salsa20, are widely used for data security and protection. Unfortunately, most existing cryptographic platforms supporting these algorithms suffer from limitations in performance, flexibility, and hardware efficiency. To address these challenges, this paper proposes the Universal Efficient Crypto Engine (UniCrypt), a highly flexible architecture that efficiently supports six traditional cryptographic algorithms with low power consumption, minimal resource usage, and high performance. Specifically, UniCrypt incorporates three key innovations: a dynamically configurable memory organization, a unified configurable ALU, and system pipeline scheduling. The UniCrypt architecture has been implemented and verified at the system-onchip level on the Xilinx ZCU102 FPGA. Real-time comparisons with powerful CPUs demonstrate that UniCrypt achieves 1.2 times to 9.3 times higher energy efficiency. Moreover, FPGA synthesis results further show that UniCrypt delivers throughput improvements ranging from 1.35 times to 92.3 times and area efficiency improvements from 1.14 times to 280.5 times compared to related FPGA-based accelerators. Finally, ASIC synthesis results demonstrate that UniCrypt achieves a 78.7% unified efficiency while attaining a computational throughput improvement of 3.34 times to 54.44 times compared to related works.