
The paper describes a universal approach to formal verification of modular multipliers and other modular arithmetic circuits. It is based on a hardware system that reverses modular reduction of $X=A \cdot B$, subtracts the remainder $R$, and divides $X-R$ by modulus $m$. In a correct modular reduction circuit, this division should yield a zero remainder $R_{m}$. Verification reduces to checking if every bit of $R_{m}$ is zero. The verification system is constructed from pre-verified arithmetic components (multiplier and divider) and synthesized using standard optimization tools such as Yosys and ABC. The method applies to any modular reduction circuit, regardless of its structure, operand width, or modulus value. We tested the method on various modular multiplier types, including large modulo $\left(2^{n}+1\right)$ multipliers and Barrett multipliers with different operand and modulus sizes, showing promising results. To the best of our knowledge, such a method has not been proposed before. Except for structured modulo $\left(2^{n}+1\right)$ multipliers, no suitable reference was found for comparison.
With phenomenal progress in SAT solvers, Bounded Model Checking (BMC) tools have grown in prominence today. A popular and important line of research in the domain of SAT Solvers is efficient Learnt Clause Management (LCM). Clauses are learnt through conflicts during SAT solving which are crucial for pruning the search space, hence reducing runtime. Most solvers usually keep an upper bound on the number of learnt clauses to be retained as too many redundant clauses can slow it down. Hence, it becomes imperative to determine the comparative importance of the learnt clauses. One of the popular metrics to evaluate this relative importance in the SAT solver parlance is the Literal Block Distance (LBD). Our paper extends this metric of clause importance to the BMC context. In this paper, we present a novel mechanism based on the incremental nature of BMC to predict the importance of learnt clauses across future BMC iterations. We use it to propose a new LCM policy and integrate it with the LBD based technique. The motivation for this work stems from the fact that the expense of generating (not solving) the SAT formula of the future BMC iterations is minimal but it can give a lot of insight about which learnt clauses to retain. Experiments on HWMCC benchmarks show comparable or improved results for 89 % of the benchmarks on average.
This work proposes a vertical fixed pattern noise (VFPN) reduction technique in CMOS image sensors. The technique incorporates the use of double sampling circuit wherein the offset voltage of the amplifier is stored during the sampling phase and canceled in the amplification phase without any increment in chip area and clock cycles. The proposed technique is implemented as a column-parallel readout with the pitch of 9 µm and verified over silicon for a 128 × 96 pixel array. The measurement results show the reduction in vertical FPN to 0.069%. The prototype image sensor is designed and fabricated in AMS 350 nm CMOS OPTO process using 3.3 V power supply. The designed image sensor occupies an area of 1.69 mm × 5.75 mm and consumes 55 mW of power.
This paper introduces an intelligent digital microfluidic biochip (DMFB) system that integrates deep learning for automated droplet manipulation with an intuitive interface for manual control. Our approach combines a custom, low-cost $8 \times 8$ electrode array fabricated on a printed circuit board (PCB) with a real-time vision system. The system utilizes copper-coated electrodes with cost-effective dielectric and hydrophobic layers, driven by a microcontroller-based high-voltage actuation circuit. Real-time, autonomous droplet manipulation is enabled by a deep learning-based pipeline that features a fine-tuned You Only Look Once (YOLOv11n) model for robust droplet detection and an intelligent routing algorithm for generating optimal, collision-free routing paths. Experimental validation demonstrates high system performance, with the automation pipeline achieving a 90% routing success rate and the YOLOv11n detector attaining a mean Average Precision (mAP@0.5) of 0.993 on a custom-created dataset. Furthermore, the system incorporates a graphical user interface (GUI) for manual control, which exhibited 98% task accuracy. By successfully integrating affordable hardware with advanced computer vision, this work presents a scalable and accessible platform for automated biochemical assays, demonstrating significant potential to advance applications in point-of-care diagnostics, drug discovery, and personalized medicine.
The memory wall problem significantly hampers the performance of modern data-intensive applications (such as AI/ML) in conventional computing architectures. Near-Memory Computing (NMC) offers a promising alternative by embedding computation units within main memory; however, it necessitates an effective offloading approach. Current offloading approaches primarily concentrate on performance metrics, often neglecting energy considerations during decision-making. Furthermore, many existing approaches require conducting preliminary trials on both the host and NMC sides, which can lead to considerable overhead. To address these challenges, we propose EP-Off, an estimation-based dynamic offloading approach. Initially, EP-Off executes a few iterations of the offloadable region exclusively on the host CPU, allowing for the estimation of the Energy-Delay Product (EDP) for both host and NMC sides. The remaining portion is then offloaded to the side with the lower estimated EDP. This approach ensures a well-balanced decision-making process that strikes a balance between performance and energy efficiency, while reducing overhead by eliminating unnecessary preliminary trials on the NMC. Through extensive experiments with diverse data-intensive applications, EP-Off exhibits substantial performance improvements ($2.1 x$ & $1.5 x$), energy saving (54% & 35%), and reductions in off-chip data transfers (65% & 27%) compared to both conventional and state-of-the-art approaches. Furthermore, EP-Off reduces EDP by 79% and 52% compared to conventional and state-of-the-art approaches, respectively.
Modern manufacturing demands physically feasible, optimized schedules, but traditional methods often neglect machine dynamics, yielding theoretically optimal yet not robustly feasible plans in practice. This paper presents a comprehensive workflow bridging this gap for cyber-physical production systems. Our first contribution enhances production line schedule synthesis with variable discretization steps for mode sequences, balancing accuracy and efficiency. Second, we propose a novel validation framework using Frost, a deterministic digital twin built on Lingua Franca. This framework models physical dynamics using adaptive ODE solvers and the full software/communication stack. Crucially, significant deviations in Frost simulations trigger an iterative re-synthesis loop. This closed-loop approach, underpinned by Lingua Franca's determinism, ensures schedules are rigorously validated and iteratively improved for practical feasibility in complex cyber-physical settings.
This study explores the effectiveness of an aluminum nitride (AlN) cap layer that can be suppress self-heating effects (SHE) and improve the performance of AlGaN/GaN high electron mobility transistors (HEMTs). Using calibrated Sentaurus TCAD simulations, three device configurations are compared such as without a cap, with a GaN cap, and with an AlN cap. Results demonstrated that the hotspot lattice temperature decreases from 321 K (no cap) to 314 K (GaN cap) and further to 309 K (AlN cap) owing to AlN's superior thermal conductivity. The AlN cap also improves electron confinement which raises the 2DEG density to $1.32 \times 10^{13} ~\text{cm}^{-2}$ compared to $1.11 \times 10^{13} ~\text{cm}^{-2}$ in the conventional device. Performance improvements include a 110 % increase in drain current ($\mathbf{I}_{d s}$), 73% higher transconductance $g_{m}$, and a 22 % reduction in peak electric field (EFP) at the gate edge. These results emphasize that the AlN cap layer is an effective technique for mitigating SHE and enhancing reliability in high-power, high-frequency GaN-based electronics.
Photonic accelerators addresses the limitations of bandwidth, latency, and energy found in traditional electronic computing, especially for data-heavy applications like AI inference. Photonic compute-in-memory (PCIM) architectures bring together storage and computation within the optical realm, avoiding the expensive process of optical-electrical conversions. This study introduces a fully passive optical memory cell that is semiconductor optical amplifier (SOA) free and its expansion into a PCIM bitcell capable of performing AND operation. The design comprises directional couplers, a phase shifter, nonlinear fiber, and a delay loop, achieving bistable performance at $50 ~\text{Gb}/\mathrm{s}$ with switching times under one nanosecond, balanced rise and fall times, and a bit error rate (BER) of less than 10−9, surpassing earlier memory configurations. In its PCIM mode, the bitcell can perform logical AND operation, creating a memory-logic system. A scalable 16 Kb array of photonic memory is proposed, providing row and column addressability along with adaptability for integration into advanced computing systems. Future work will focus on deploying convolutional neural networks, like LeNet-5, on the proposed PCIM array for real-time inference, aiming for high accuracy while minimizing latency and power use which highlights the potential of completely passive PCIM architectures for next generation high-performance computing.
Test Chips (TC) are developed to perform the silicon qualification of Analog to Digital Converters (ADCs). The IO Ring serves as a critical interface between the ADC circuitry and external laboratory equipment. Its design is pivotal in ensuring that ADC's performance is not limited by the TC interface. This paper details the key design considerations and proposes an optimized IO-Ring architecture tailored for high resolution ADCs (9 to 12 bits) operating at high sampling rates (40 to 160 MSPS). We analyze the limitations of conventional low-speed IORings when applied to high-speed, high-resolution ADC testing. Furthermore, simulation results comparing a 9bit 320 MSPS ADC interfaced with both low-speed and high-speed IO-Rings are presented, demonstrating the impact of IO-Ring design on overall ADC performance.
In-Memory computing presents a viable approach to overcome the data transfer limitations imposed by the von Neumann bottleneck. This work presents a novel current source based sense amplifier (SA) which is capable of performing bitwise Boolean logic operations including $A N D/N A N D/O R/N O R$. The proposed SA integrates two additional current sources to introduce asymmetricity in the balanced current latch sense amplifier. The designed circuit can be reconfigured to both normal read mode and compute mode without sacrificing the read delay. The proposed design is reconfigurable and adaptable in nature. The post-layout simulations show the worst case sensing delay of 47 ps and energy consumption of 15.6 fJ/bit. The worst case yield of 99.6% is observed by performing Monte Carlo simulations for 1000 runs. The designed circuit achieves notable improvement in energy efficiency and performance, establishing its suitability for logic-in-memory applications.
This paper presents a power-efficient IoT sensor network architecture comprising sensor nodes realized using ESP32 microcontrollers with ultra-low power coprocessors being used for always-on detection of threshold-level crossing events of the sensed parameters. By suitably leveraging the available deepsleep modes, event-driven wake-ups, and ESP-NOW protocol for data transmission, the system minimizes active time and power requirement in the sensor as well as in back-end processing. The proposed system is suitable for a multitude of applications that involve environmental monitoring and actuation. The system is tested in a smart agriculture demonstration whereby multiple sensor nodes monitor soil-moisture level at different sites across the field, and send the salient data to a centralized access-point (receiver). Cloud integration is handled via the ESP32-based receiver node that can upload the data to cloud platforms like Google Sheets. The sensor nodes also have capability of actuating irrigation at the corresponding location for ensuring optimum plant growth and water-usage throughout the cultivated field. The deployed network achieves an average consumption of just around 0.74 mAh/day in each of the sensor nodes. Compared to similar systems relying on periodic polling and higher baseline currents, the proposed design achieves a state-of-the-art operational lifetime of over 5-8 × when powered from battery - e.g., almost 5 years of usage on a small 1500 mAh battery under typical conditions.
This work explores the design methodology and FPGA-based realization of asynchronous circuits composed of static data-flow handshake components governed by the two-phase bundled-data (BD) protocol. The methodology integrates tutorial-driven exposition with formal architectural analysis, emphasizing a rigorous approach to initialization strategies and the construction of coupled ring structures with arbitrary token distributions. Central to the design flow is the use of Mousetrap-style control templates, enabling localized, transition-driven handshake logic. The paper substantiates its methodology through a comprehensive power-performance-area evaluation, including gate-level synthesis of core handshaking primitives in a 32 nm CMOS technology node. Furthermore, it incorporates peephole-level optimizations, wherein composite components such as merged Join-register or MUX-register structures are synthesized for improved design efficiency. The practicality of the design flow is ultimately validated through the implementation of a Fibonacci sequence generator, showcasing correct functionality and handshake compliance across all components. All modules and the target circuits are verified through full-system deployment on Xilinx Basys 3 FPGA board, confirming robustness and functional correctness of the protocol under real hardware conditions. The average PPA metric of the Mousetrap controller has improved to 38%.
The comparators are crucial components in data converters, e.g., analog-to-digital converters (ADCs), sensors, and control systems to enable rapid binary decisions derived from specific thresholds in analog signals. This paper proposes a novel latch reset technique for a double-tail comparator to attain lowpower operations. To reset the latch, the additional pull-up (PMOS) transistors are included in the pre-amplifier stage, which are switched by the derived clock signal (MCLKB) from the baseline clock (CLK). The delayed version, i.e., MCLKB, resets the latch to $V_{D d}$. Therefore, the proposed method enhances the energy efficiency of the comparator by minimizing dynamic power consumption. Using optimized timing and control of the latch stage, the simulation has been performed for the $\mathbf{4 5 ~ n m}$ CMOS process, demonstrating that the comparator achieves a power consumption of $5.85 \mu ~\mathrm{W}$ at a 1 V supply. Monte-Carlo simulation has been done for 200 samples, showing the average power distribution with a standard deviation of $0.8 \mu ~\mathrm{W}$. The effective layout area of the proposed design is $60 \mu ~\mathrm{m}^{2}$. Thus, the proposed technique is needed to design the low-power analog front ends.
Large Language Models (LLMs) have recently emerged as powerful assistants for hardware design, translating naturallanguage specifications into Hardware Description Languages (HDLs), yet fine-tuning these models on domain-specific corpora routinely exceeds the memory capacity of commodity GPUs and triggers Out-ofMemory (OoM) errors. We present AAMLA, an autonomous agentic framework that converts this pain point into a push-button experience. AAMLA incorporates memory awareness by coupling (i) a predictive memory profiler-LLMem++-that significantly extends the original LLMem framework to support a diverse set of memory-efficient finetuning strategies, including adapter-based (LoRA, DoRA), gradientfree (MeZO), token-sparse (TokenTune), and optimizer-modified (APOLLO) methods, with (ii) a portfolio of complementary memoryefficient adaptation techniques-leading to a complete synthesis flow. The system allows for a two-step early design space exploration workflow: it first prunes any method whose predicted footprint would violate the user's GPU budget, then consults an offline accuracy-\mono{latency} atlas to recommend the Pareto-optimal strategy that aligns with the designer's stated priority (accuracy or turnaround time). Guided by this workflow, an agentic controller configures and applies the chosen technique, guaranteeing OoM-free execution without manual trial-and-error. The framework is model- and dataset-agnostic, easily extensible with new tuning primitives, and exposes a simple interface that accepts natural-language prompts and emits synthesizable Verilog, thereby lowering the barrier to LLM-assisted hardware design for researchers and organizations with limited computational resources. We open-source AAMLA at https://github.com/rajatbh21/AAMLA.
Efficient implementation of Depthwise Separable Convolution (DSC) inference on FPGAs remains a significant challenge. Conventional accelerators either suffer from suboptimal resource utilization with unified engines or incur excessive hardware footprints with fully dedicated designs. This work presents a new FPGA accelerator that resolves this trade-off through three specialized Processing Elements (PEs) dedicated to expansion, depthwise, and projection convolutions. These PEs are iteratively reused across channel slices, maximizing hardware sharing while retaining the benefits of highly task-specific computation. The suggested architecture adopts a multi-stage, streaming dataflow interconnected by tightly synchronized on-chip buffers, fully eliminating costly off-chip memory transactions. A fused quantization unit integrates bias addition, ReLU6 activation, and shortcut connections, enabling efficient integer-only arithmetic and significantly reducing overall computation overhead. Implemented on a Xilinx ZCU102, the proposed design operates at 137 MHz clock frequency with a mixed 4/8-bit quantized data path, achieving an inference rate of 20 frames per second. Compared to state-of-the-art work, our design demonstrates superior efficiency, reducing Look-Up Table (LUT) usage by 12.3%, Flip-Flop (FF) usage by 16.5%, and DSP usage by 77.8%, while achieving a substantial power reduction of 64.3% at 1.78 W. These combined features deliver a compact and highly energyefficient solution, making it ideal for power and area-constrained edge Artificial Intelligence (AI) deployments.
This work aims to create machine learning (ML) models for register transfer level (RTL) designs. We train an ML model on a small dataset generated from a set of test cases of RTL modules which can predict output of the RTL designs for a given input. One pitfall of traditional ML-based solutions is false-positive predictions. To address this issue, we propose probabilistic neural network (PNN) models for RTL designs based on confidence scores which help us to segregate the inputs for which the model fails to provide confident predictions. One of the use cases will be to expedite the time-consuming RTL simulation. If we use PNN model prediction during RTL simulation, test cases to be run on actual RTL simulation can be reduced by up to 98% and 89% on average.
In this work, we investigate the key challenges that arise in nanoscale GaN HEMTs with gate lengths down to 20 nm and source/drain access spacings as low as 10 nm. As GaN HEMTs scale toward advanced technology nodes for high-speed and high-frequency applications, understanding their reliability and physical limits becomes critical. Moving beyond conventional RF performance benchmarks, this work focusses on reliability limiting factors such as short-channel effects, trapinduced degradation, and RF performance degradation. Using numerical simulations, we assess device performance and transconductance degradation, under scaled geometries for DC and RF operation. The findings highlight key reliability considerations and scaling challenges that designers should account for when evaluating GaN HEMT miniaturization, offering guidance for balancing performance and long-term stability in future DC and RF applications.
This paper presents design of low-power, low-area, PVT-invariant 3-stage differential ring oscillator (RO) targeting a frequency of 1 GHz. The design uses a process-trimmed, temperature-compensated robust current reference, designed in TSMC 65 nm CMOS technology, achieving a temperature coefficient as low as 21.67 ppm/° C across temperatures from −40° C to 125° C. A detailed, quantitative simulation focussed design methodology and implemenation is proposed for the design of the current reference and mirroring circuit. The resulting oscillator shows less than 0.5% frequency variation across temperature and a worst-case variation of only 4.78% across PVT. To ensure sub-1V supply operation, only lowthreshold (LVT) MOSFETs are used throughout the design. The proposed design is a low power design and the oscillator consumes approximately 0.36 mW of power at 800 mV supply. This power efficient architecture is well-suited for PVT-robust on-chip clocking in modern low-voltage systems.
SNNs are a bio-inspired and energy-efficient framework for processing temporal information. However the design of optimal SNNs has been a complex, multi-faceted problem to this date. This paper proposes EvoSNN-MO a novel evolutionary framework for multi-objective optimization of SNNs. EvoSNN-MO performs the co-evolution of network architecture, synaptic dynamics per-neuron parameters and intrinsic plasticity mechanisms using the NSGA-II algorithm. The five primary objectives focused on maximizing accuracy while minimizing energy consumption and structural complexity optimizing the average firing rates of hidden neurons, and improving sparsity. As applied to a temporal XOR pattern recognition task EvoSNN-MO found SNNs which achieved 100% accuracy. One representative solution from the Pareto front produced an exceptionally energy-efficient solution of 6.253 nJ inference with a compact structure composed of 169 active synapses while yielding robust performances. The resulting Pareto front consisted of 5,890 non-dominated solutions after 300 generations and illustrated a diverse array of high-performing and resource-efficient SNN configurations. This work demonstrates the effectiveness of multi-objective neuro evolution in automatically generating SNNs tailored for complex temporal tasks.
This work provides theoretical evidence for longer battery life for cardiac pacemakers powered by a hybrid batterysupercapacitor system with a single supercapacitor ($B$-SC). The C-rate expressions of the battery are derived for the cardiac pacemaker powered by the battery $(B)$ and $B$-SC. A detailed analysis is performed to prove that the peak C-rate of $\boldsymbol{B}$-SC is less than that of $B$. Expressions for the minimum charge storage capacity of the supercapacitor ($S C$), the reduction in the peak C -rate, and the improvement in battery life of the cardiac pacemaker powered by $\boldsymbol{B}$-SC are obtained. A cardiac pacemaker powered by a battery $B$ with a $V$ volt - $A_{B}$ ampere hour with a lifespan of $T$ seconds requires an $S C$ of $\left[1-e^{-\frac{t_{p v}}{R C}}\right]^{-1} \times \frac{Q_{p v} \times q}{V}$ Farads. The peak C-rate is reduced by $\frac{t_{s v}}{t_{p v}}\left[\frac{Q_{p v}}{Q_{s v}+Q_{p v}-Q_{p a}}\right] \times$. The lifespan of $B$ is enhanced by $\frac{2 Q_{p a}-Q_{p v}}{Q_{s a}+Q_{p a}+Q_{s v}+Q_{p v}} \times T$ seconds in the worst-case scenario, where $t_{s v}$ and $t_{p v}$ are the duration to sense and pace the ventricle; $Q_{s a}, Q_{p a}, Q_{s v}$ and $Q_{p v}$ are the charge units required to sense atrium, pace atrium, sense ventricle and pace ventricle, respectively; $\boldsymbol{q}$ is Coulombs per charge unit, $\boldsymbol{R} \boldsymbol{C}$ is the time constant of $\boldsymbol{S} \boldsymbol{C}$, and $\frac{t_{p v}}{R C} \geq 1$. The simulation result shows that the 7 -year lifespan of a $2.5 ~\mathrm{V}-2 \text{AH}$ battery with $S C$ of $6.38 \mu ~\mathrm{F}$ is extended by 11 days in the worstcase scenario. The peak battery $\mathbf{C}$-rate of the pacemaker powered by $\boldsymbol{B}$-SC is $205 \times$ less.