
The rapid proliferation of edge devices, cyber-physical systems, autonomous platforms, and large-scale IoT infrastructures has fundamentally transformed how intelligence is computed and deployed. Traditional centralized cloud-based AI architectures are increasingly limited by latency, bandwidth, privacy, energy efficiency, and reliability constraints. As a result, distributed intelligence—where sensing, learning, inference, and decision-making are performed collaboratively across networked nodes—has emerged as a critical paradigm shift. While significant progress is made in creating smart AI algorithms and systems that can learn from each other, but there's still a big hole when it comes to the underlying technology that makes it all work. Basically, the circuits and systems that are developed aren't really built for this kind of distributed intelligence, where lots of devices are working together and sharing information in real-time. Most of the current hardware is just adapted from old centralized computing systems, which aren't ideal for collaborative, on-device, and federated intelligence. Hence there is a need for major innovations at the circuit and system level to make distributed intelligence scalable, secure, energy-efficient, and fast.
B5G/6G extended reality (XR) applications motivate encoder-side channel-coding hardware that contributes to terabit-class throughput, low processing latency, and high energy efficiency for wearable deployment. While quasi-cyclic low-density parity-check (QC-LDPC) encoding provides a standard-compatible coding structure for these workloads, its vast configuration space of two base graphs (BGs) and 51 lifting sizes makes efficient hardware implementation a formidable multi-objective optimization problem. Such a large parameter set can only be accommodated by universal architectures, often at the cost of prohibitive hardware overhead and efficiency loss in VLSI. Conversely, manually developing dedicated ASICs for distinct XR scenarios is infeasible due to the significant design cost. To tackle this trade-off, we propose the Standard-Compatible Automated QC-LDPC Encoding accelerator (SCALE). SCALE adopts a unified formula-driven microarchitecture and a multi-dimensional parametric model that exposes fine-grained control over datapath parallelism. We establish explicit analytical models for throughput, latency, power, and area to systematically explore the design space and synthesize hardware optimized for diverse XR service constraints. Experimental results demonstrate that the generated RTL achieves 542 Gb/s throughput for immersive XR streaming, 6.04 ns latency for motion-critical interaction, and 2.17 Tbit/J energy efficiency for wearable terminals. A 65 nm CMOS implementation confirms silicon feasibility, achieving 145 Gb/s throughput and 1.4 Tbit/J energy efficiency at 103 mW for encoder-side LDPC processing in B5G/6G XR systems.
The computational complexity of deep learning algorithms has given rise to significant speed and memory challenges for the execution hardware. In energy-limited portable devices, highly efficient processing platforms are indispensable for reproducing the efficiency achieved by much bulkier platforms. In this work, we present a low-power Leaky Integrate-and-Fire (LIF) neuron design fabricated in TSMC’s 28 nm CMOS technology as proof of concept to build an energy-efficient mixed-signal Neuromorphic System-on-Chip (NeuroSoC). The fabricated neuron consumes 1.61 fJ/spike and occupies an active area of 34 μm2, leading to a maximum spiking frequency of 300 kHz at 250 mV power supply. These performances are used in a software model to emulate the dynamics of a Spiking Neural Network (SNN). Employing supervised backpropagation and a surrogate gradient technique, the resulting accuracy on the MNIST dataset, using 4-bit post-training quantization stands at 82.5%. The approach underscores the potential of such ASIC implementation of quantized SNNs to deliver high-performance, energy-efficient solutions to various embedded machine-learning applications.
This work reports a first-principles study of a hybrid-barrier magnetic tunnel junction (MTJ) with the stack Co₂FeAl/MgO/WSe₂/MgO/Co₂FeAl, modeled using density-functional theory (DFT) combined with a non-equilibrium Green’s-function (NEGF) transport framework in QuantumATK. The symmetric device yields a tunneling magnetoresistance (TMR) of about 170% at 300 K, indicating strong spin-dependent tunneling suitable for room-temperature spintronic operation. To assess the impact of structural asymmetry, two modified hybrid-barrier junctions were examined in which the left-side (skip-L) or right-side (skip-R) MgO sublayer was removed, giving WSe₂/MgO and MgO/WSe₂ hybrid-barrier. In these asymmetric configurations, the TMR drops to ~ 19% at 300 K, demonstrating that the loss of symmetric MgO confinement substantially weakens Δ₁-like coherent tunneling and spin-selective transmission. Overall, the results show that a symmetric MgO/WSe₂/MgO barrier can effectively merge the coherent tunneling characteristics of MgO with the spin-filtering behavior of WSe₂, leading to a thermally robust, high-TMR Co₂FeAl-based MTJ that is well suited for spintronic memory applications.
Scalable neuromorphic systems with online learning capability are essential for processing time-varying and data-intensive signals under strict power and hardware constraints. Spiking neural networks (SNNs) provide an event-driven and temporally rich computing paradigm; however, achieving flexible network scalability together with efficient on-chip learning remains a significant challenge. In this work, we propose a scalable spiking neural network (S-SNN) architecture that integrates multiplexing ISI–phase temporal encoding, reconfigurable triplet spike-timing-dependent plasticity (STDP), and reward-based supervised learning to enable high-dimensional spatial–temporal processing with low hardware overhead. The proposed architecture supports dynamic adjustment of input and output dimensions while maintaining stable online learning behavior. A prototype S-SNN chip fabricated in GlobalFoundries 22FDX CMOS technology occupies an area of 670 × 1620 μm2 and consumes 13.0 mW under typical operating conditions, demonstrating compact and energy-efficient neuromorphic implementation. System-level simulations show that the S-SNN achieves approximately 98% classification accuracy on MNIST and around 81% on CIFAR-10. Moreover, evaluations on size-varying MIMO wireless channel prediction datasets demonstrate that the scalable architecture naturally enables transfer learning across changing dimensions, reducing training time by up to 3× while achieving comparable or improved NMSE performance compared to training from scratch. These results indicate that the proposed S-SNN provides a practical and scalable neuro-morphic computing framework for adaptive edge computing and communication-oriented applications.
Memristor crossbar arrays have emerged as a promising technology for future memory and computing systems due to their high density, low power consumption, and non-volatile behavior. As the demand for more powerful and energy-efficient electronic devices continues to grow, large-scale memristor crossbar arrays have gained significant attention in both academic research and industrial applications. However, the practical deployment of these arrays faces various challenges related to device variability, which can significantly impact their performance and reliability. This research proposes a novel framework for estimating variability in large memristor crossbar arrays using surrogate-assisted metaheuristic approach. The use of Gaussian Process Regressor (GPR) as the surrogate model is explored in this work. The study aims to understand the effect of variability, quantify its impact on array performance, and demonstrate the efficacy of the proposed method for fast estimation of variability. Each individual cell in the crossbar array is made up of one transistor one memristor (1T1M) unit. The variability obtained from the surrogate-assisted method is compared with conventional metaheuristic approaches, polynomial chaos expansion (PCE) and Monte Carlo (MC) simulations. The surrogate model is further validated at the distribution level against MC. Here, when training database reaches the order of the parameter count, the GPR reproduces the MC statistics (i.e. the probability density, standard deviation, skewness) with a point-wise correlation of r = 0.997 on unseen samples.The results are cross-verified on an independent, SPICE testbench (ngspice), and we observe that the two testbenches agree with a correlation of r = 1.000 on identical parameter sets, with the the Monte Carlo statistics match within 1%.
The rapid evolution of Extended Reality (XR) drives a critical need for ultra-high-resolution, low-latency rendering, creating a significant burden on the processing capabilities of power-constrained edge devices. Although 3D Gaussian Splatting (3DGS) is a pivotal reconstruction and rendering technique, its real-time execution is bottlenecked by massive throughput demands for Gaussian exponential evaluations, causing severe structural hazards and pipeline stalls on conventional Special Function Units (SFUs). To address these bottlenecks, this paper proposes BitSplat, a bit-level algorithm-hardware co-optimization paradigm. At the algorithm level, we exploit the bit-level representation of IEEE 754 floating-point numbers to reformulate latency-intensive exponential calculations into a compact Bitcast-FMA datapath. The proposed kernel ensures rigorous rendering fidelity by leveraging a hardware-native boundary absorption mechanism, where the bit-level reformulation allows the kernel to naturally taper to zero within its compact support. At the hardware level, BitSplat leverages direct physical wire routing and ubiquitous integer FMA units to achieve a single-cycle processing pipeline, effectively minimizing silicon area and dynamic power while maximizing rendering throughput.We evaluate BitSplat in 28nm technology. BitSplat outperforms the state-of-the-art GCC by reducing silicon area and power consumption by 16.3% and 25.5%, respectively. Furthermore, in-situ GPU profiling confirms that BitSplat significantly alleviates instruction-level bottlenecks, delivering a 20% system-level speedup in end-to-end runtime.
The emergence of novel non-volatile memories, such as resistive random access memory (ReRAM), enables deployment of neuromorphic systems and algorithms on specialized neuromorphic hardware. Spiking neural networks (SNNs), a type of bio-inspired Neural Networks, is the one example to benefit the most from dedicated neuromorphic hardware. This work explores Current Controlled Oscillator (CCO) based realization of leaky integrate and fire (LIF) neurons for neuromorphic hardware-based SNNs. The presented circuit uses CCO for analog integration of weighted input spikes instead of traditional neurons that rely on large capacitors and/or operational amplifiers, thus allowing significantly smaller on-chip area. We demonstrate a compact CCO-based LIF neuron in Skywater 130nm CMOS process, compatible for integration with ReRAM crossbar synapses. One of the main advantages of the proposed design is sharing of the reference CCO for an array of neurons by introducing techniques for rapid phase reset of the CCO and phase drift compensation, allowing to significantly increase the density of readout circuits. Using mostly digital circuit elements allows the proposed design to be scaled down to low-supply voltage and advanced CMOS technologies.
The transistor-memristor (1T-1M) integration has been recognized as an effective strategy to suppress crosstalk between neighboring cells for the development of a practical crossbar architecture, though its adoption for in-memory Boolean logic gate implementation is still missing. In this regard, we systematically investigate HfOX-based 1T-1M cell for in-memory logic gates implementation employing the Memristor-Aided loGIC (MAGIC) design style. This performance analysis is carried out using a fully calibrated and physics-based Stanford-PKU model, which captures full I-V switching dynamics, including temperature and stochastic effects. Our results show that the 1T-1M cell with the Gate and Source-controlled Off-State RESET (GSOSR) scheme substantially increases the high resistance state (HRS), which enables nearly $2.2 imes $ higher on-off resistance ratio compared to the 1M cell and achieves significant energy saving in execution operation for AND, NAND, XOR, and XNOR gates. Despite the 1T-1M cell exhibiting a slight increase in delay due to additional capacitances, the associated energy reduction and improved reliability outweigh this penalty. Variability analysis further reveals that the XOR gate with 1T-1M cell suppresses resistance fluctuation in both low resistance state (LRS) and HRS by up to $3 imes $ under worst-case corners, which ensures more robust operation over 1M cell. Overall, our findings demonstrate that the 1T-1M cell offers a viable path toward a scalable, energy-efficient, robust, and crossbar-compatible in-memory computing system, surpassing the widely employed 1M cell architecture.
Local Binary Pattern Networks (LBPNet) eliminate multiply-accumulate operations by replacing learned convolutions with binary pixel comparisons, offering a promising paradigm for energy-constrained edge vision. However, baseline LBPNet processes all spatial locations uniformly, allocating identical computational budgets to discriminative regions and uninformative backgrounds alike. This paper presents a heatmap-guided framework that introduces spatial awareness into LBPNet through three complementary mechanisms. First, we construct dataset-level importance maps via structured occlusion analysis, aggregating per-class saliency into global binary masks that identify discriminative spatial regions. Second, we propose adaptive neighbor selection that varies the LBP comparison budget P based on local pixel statistics derived from these heatmaps, concentrating sampling density on informative areas while minimizing computation elsewhere. Third, we introduce progressive P-decay across network depth, exploiting the observation that deeper layers represent lower-frequency abstractions requiring fewer fine-grained comparisons. To fully realize these algorithmic gains, we develop a Processing-in-Sensor (PIS) hardware architecture that promotes learned spatial priors into the physical domain through mask-aware selective readout, analog-domain LBP computation via Sample-and-Hold comparators, and programmable feature selectors. Circuit-level simulations in 65nm CMOS demonstrate over 200× improvement in power-delay product compared to conventional sensor-plus-ADC pipelines. Comprehensive evaluation across six datasets of varying complexity shows that the proposed framework achieves up to 95.96% accuracy on CK+ (exceeding even CNN baselines) while reducing operations by over 10× and model size by over 100× relative to conventional CNNs. The framework is particularly effective for structured visual domains with spatially concentrated features, positioning heatmap-guided LBPNet as a practical solution for ultra–low-power in-sensor visual intelligence.
Extended Reality (XR) wearables require always-on perception within tight power envelopes of a few watts and motion-to-photon latency budgets below 20 ms, leaving only a few milliseconds for neural-network inference. Bit-serial computing is attractive for such energy-efficient neural network acceleration, but many existing architectures still process all bits even when ReLU sets the final output to zero. This paper presents BitFair, a software-hardware co-designed bit-serial CNN accelerator with learnable bit-level early termination and adaptive bit ordering, working under the ultra-low-power and strict latency requirements of XR applications. BitFair exploits dynamic bit-level sparsity by learning per-layer thresholds that trigger early termination when partial sums reliably predict that the final ReLU output will be zero. Furthermore, it searches for layer-wise bit orders that prioritize informative bits, maximizing early termination without sacrificing accuracy. A GlobalFoundries 12-nm FinFET implementation with a core area of 0.34 mm^2, 104 KB on-chip memory, and voltage scaling from 0.55 to 0.70 V achieves sub-millisecond latency, up to 117.0 BTOPS/W, and 0.07 pJ/SOP. On IBM DVS128 Gesture and N-MNIST, BitFair achieves 96.5
With the proliferation of mobile Extended Reality (XR) devices, achieving photorealistic mobile immersion has become critical functionality. While 3D Gaussian Splatting (3DGS) offers superior quality, real-time stereoscopic rendering suffers from heavy sorting loads, redundant binocular computations, and excessive memory traffic in conventional GPUs. To address these bottlenecks, we propose a dedicated stereoscopic 3DGS rendering processor combining (i) a Gaussian streaming architecture and (ii) a pre-blended fragment warping. The Gaussian streaming architecture replaces inefficient tile-binning with a stream-generating rasterizer and pixel blender, enabling fully on-chip processing of sorted Gaussian streams. Furthermore, the pre-blended fragment warping pipeline achieves high energy efficiency by sharing redundant workloads between binocular views at the fragment level. It computes splat fragments, which consist of color and opacity, only once for the reference view and reuses them for the target view via disparity-based shifting, effectively halving the complexity of exponential and Spherical Harmonics (SH) evaluations. Experimental results show that the proposed processor achieves 112.2 FPS at 1552×1040 per eye, while reducing off-chip memory bandwidth by 61.8% and improving energy efficiency by 15.8× over a state-of-the-art mobile GPU implementation.
Matrix computations have become increasingly significant in many data-driven applications. However, Moore’s law for digital computers has been gradually approaching its limit in recent years. Moreover, digital computers encounter substantial complexity when performing matrix computations, requiring extensive computation time. Existing analog matrix computation schemes, on the other hand, require a large chip area and high power consumption. This paper proposes a linear algebra system of equations based on integrators, which features low power consumption, compact area, and fast computation time. The demonstrated scheme is capable of performing first-order ordinary differential equation (ODE) computations to realize a 8 × 8 computation matrix and can be expanded further depending on the available chip area. Due to its simple structure, the ring oscillator-based integrator exhibits a compact area and low power consumption. Therefore, ring oscillator-based integrators are introduced into the linear algebra system of equations. This system can be used to compute the linear algebra equations of the matrix with either positive or negative values. This paper provides a detailed analysis and verification of the proposed circuit structure. Compared to similar circuits, this work has significant advantages in terms of area, power consumption, and computation speed.
Wearable devices and Internet of Things (IoT) sensors require on-sensor processing of biosignals and environmental data, including computationally demanding operations such as nonlinear activation functions for neural network inference, sensor calibration curves to map raw readings to physical units, and signal preprocessing functions like logarithmic compression and power operations for feature extraction. These functions exhibit significant complexity, often involving transcendental operations and multivariate dependencies that are costly to implement digitally. Analog function approximation provides a power-efficient alternative by performing these computations in the analog domain, thereby reducing the energy overhead associated with analog-to-digital conversion and subsequent digital processing. Flexible Electronics (FE) present a particularly attractive platform for wearable applications due to mechanical flexibility and low-cost fabrication, but impose strict constraints on circuit density and power consumption, making efficient analog implementations critical but challenging. This work introduces Analog Kolmogorov-Arnold Networks (AKANs), developed via hardware-software co-optimization, to approximate these complex multivariate functions accurately under hardware imperfections. Our method incorporates circuit-level error modeling during training and applies pruning at both software and hardware levels to reduce area and power. Validation across multiple benchmarks demonstrates that our proposed pruning methodology not only reduces hardware cost but can also improve approximation accuracy by regularizing spline parameters. Results show up to 55
Analog in-memory computing (AIMC) has gained popularity as an alternative to conventional von Neumann architectures for deep learning inference, offering significant energy and latency advantages. By exploiting the analog properties of memory devices, matrix-vector multiplications can occur within memory arrays, thereby eliminating the movement of weights and associated costs. However, operating in the analog domain introduces a range of non-idealities that hamper the computational precision. These non-idealities can be stochastic or deterministic and are often difficult to characterize analytically due to their complexity and interplay with one another. In this work, we present a data-driven modeling framework that captures the behavior of AIMC hardware using experimental data from the hardware itself. Our approach is grounded in a first-order Taylor approximation of the matrix-vector multiply operation, which leads to the decomposition of the overall error into two components: linear perturbations, mainly stemming from device variability, and non-linear or recurrent residual error, typically induced by circuit-level effects and recurrent noise sources. We demonstrate that this two-model approach can faithfully reproduce the behavior of a real phase-change memory-based AIMC chip. We also show generalization to downstream inference tasks on four different networks, two ResNets for image classification and two LSTMs for character prediction and image captioning, demonstrating that the model predicts the hardware error and task performance in a diverse set of tasks.
Oscillatory Neural Networks (ONNs) exploit the rich phase dynamics of coupled oscillators to perform parallel, in-memory, and inherently analog computation. By encoding information in relative phase differences, ONNs can efficiently solve complex optimization and pattern-recognition tasks while mitigating data-movement bottlenecks inherent to conventional von Neumann architectures. To date, most ONNs have relied on CMOS-based ring oscillators, whose scalability is limited by field-effect transistor constraints and complex coupling circuitry. This work explores alternative device technologies compatible with Back-End-of-Line (BEOL) integration, combining Vanadium Dioxide (VO2) relaxation oscillators with analog Resistive Random Access Memory (RRAM) coupling elements. Compact Verilog-A models for both VO2 and RRAM devices are developed and validated against experimental data, enabling circuit-level exploration of coupled dynamics, biasing conditions, and device compatibility. These models and simulation tools establish a foundation for co-designing scalable VO2–RRAM ONN architectures, supporting the development of energy-efficient, BEOL-integrated in-memory computing systems.
This work demonstrates a compute-in-memory (CIM) capacitive crossbar array that solves the core least-squares problems of broadband Slepian beamforming non-iteratively for batch processing and in real-time for streaming data, eliminating the need for iterative algorithms, auxiliary crossbars, or complex digital peripherals. Our analog CIM architecture directly computes the solution to Aα = b through charge-domain matrix inversion. The design exploits non-volatile ferroelectric capacitors (nvCAPs) to store matrix weights, enabling in-situ computation with zero static power and 1000× higher energy efficiency than the resistive CIM counterpart. SPICE simulations show 0.1% normalized mean squared error (MSE) and < 0.5 dB beam-forming gain loss. By circumventing iterative algorithms and machine learning (ML)’s data dependency, this work provides a physics-based, hardware-efficient path for real-time broadband beamforming.
Implantable multi-channel neural interfaces are essential for high-resolution, long-term brain-machine interface applications, yet conventional designs are constrained by power consumption, data bandwidth, and silicon area. We present NeuroSEED, a 32-channel neuromorphic scalable neural interface inspired by biological neuron signaling. This system introduces four key innovations to address the bottlenecks of scalability and efficiency: (1) an area-efficient time-division multiplexed analog front-end (0.0032 mm2/channel) with a dedicated digital-servo electrode DC offset cancellation loop; (2) an analog-to-spike converter (ASC) that employs adaptive event-driven sampling; (3) a clockless spike detector utilizing valley-to-peak detection for robust spike identification; and (4) an event-triggered impulse-radio ultra-wideband (IR-UWB) transmitter for energy-efficient wireless spike telemetry. Fabricated in 40-nm CMOS, NeuroSEED achieves 1.38 μW per recording channel and 4.6 μW transmitter power, with more than 500× data reduction compared to Nyquistrate sampling. Experimental results, validated with in-vivo neural datasets, confirm low noise, compact area, and robust system-level operation, positioning NeuroSEED as a highly efficient SoC for high-density neural recording.
Stochastic computing has emerged as a promising alternative to traditional floating-point probability computations, offering simple and area-efficient implementations specifically for classification tasks with limited training data where even the resource-intensive deep neural networks yield poor accuracies. However, a key prerequisite for stochastic computing is the probability conversion circuit (PCC) which relies on entropy sources to encode floating-point probabilities into low-discrepancy stochastic bitstreams. Most of the PCCs typically rely on bulky CMOS entropy sources exhibiting pseudo-randomness, which not only consume large energy but also introduce non-uniformity, autocorrelation and cross-correlation in the generated stochastic bitstream limiting the overall performance and efficiency. Although recent studies have explored the efficacy of software-generated Sobol and Halton sequences for behavioral modeling of stochastic computing, hardware implementation of these sequences is complex. To realize the true potential of stochastic computing, in this work, for the first time, we propose a probability conversion circuit designed using a ferroelectric (Fe) FinFET-based entropy source which exhibits true randomness and successfully passes all the statistical tests of the NIST SP800-22 test suite. We perform an extensive analysis of the efficacy of the Fe-FinFET-based PCC by measuring the information loss during the conversion from floating-point to binary bitstream representation, using the statistical measure Kullback-Leibler (KL) divergence. Our extensive investigation reveals that the proposed PCC implementation is not only compact, accurate and ultra energy-efficient but also outperforms the software generated Sobol sequences for practical applications such as Bayesian inference-based spam filtering.
In-memory-computing promises to reduce power consumption by tightly integrating analog compute logic and memory, thus avoiding costly data transfers. This requires analog storage with high density, sufficient precision and long retention. However, increasing retention at the device level often comes at the expense of excessive area, write energy and/or precision. We instead propose an algorithmic solution to realize dynamic semi-volatile memory, rooted in large-margin statistical learning theory, that dynamically refreshes the analog vector contents of an in-memory computing accelerator towards one among a large discrete set of stable states, and can thus extend its retention time indefinitely. We present three different variants of this algorithm, analyze its capabilities and limitations, and discuss potential applications.