
This study explores the versatility of 1T-nC ferroelectric RAM (FeRAM) memory capable of operating in both secure storage and highly energy-efficient Compute-in-Memory (CiM) modes without incurring any extra design overhead. In contrast to conventional, volatile 1T-1C Dynamic Random Access Memory (DRAM), having significant energy cost due to the continuous refresh cycles, FeRAM presents a robust non-volatile memory solution, enabling efficient data retention with minimal power consumption. Bitwise logic operations (AND, OR) are efficiently implemented using a single, unmodified 1T-nC FeRAM cell and single-row activation (SRA), while NOT operation is realized by a modified sense amplifier. These facilitate bulk bitwise operations with significant energy benefits compared to multi-row based DRAM solutions. Furthermore, we demonstrate key-based (K) cryptographic operations which help in storing Plain Text (PT) in Cipher Text (CT) format tightly coupled with the memory write/read cycles in secure-mode operation with minimal overload while maintaining robust security. In addition, the advantages of vertical 3D integration of 1T-nC FeRAM are evaluated, and thermal profiling of the remnant polarization (PR) is done. Furthermore, we demonstrate that 1T-nC FeRAM memory reduces energy consumption by 31% compared to DRAM by assessing 8 real-world data-intensive workloads utilizing bulk bitwise operations.
Non-volatile in-memory computing (iMC) has emerged as an energy-efficient paradigm well suited to AI work-loads. Its implementation using 1T1C FeMFETs (Ferroelectric Memory Field Effect Transistors), a best-in-class emerging nonvolatile memory technology that integrates BEOL ferroelectric devices with FEOL transistors, is of particular interest. This interest stems from their potential to enable large-scale multiply-accumulate (MAC) operations in both digital and analog domains. However, realizing tangible performance benefits requires comprehensive cross-layer exploration of both design and technology parameters, extending up to accelerator level. In this work, we propose a bitcell-level multi-objective optimization methodology to identify and extract optimal sizing solutions that provide tractable trade-offs between key performance indicators (KPI). We further demonstrate how this approach facilitates cross-stack exploration of accelerator architectures. Results are presented as Pareto fronts spanning 2-4 KPIs: a 2-KPI problem illustrates the methodology, while a 4-KPI problem represents a realistic design scenario. Comparison is made between 130nm and 28nm technologies demonstrating a decrease in the average of write energy and area up to 24X and 30X respectively.
Routing-dominated multi-path architectures play a critical role in enhancing data throughput and resource efficiency in computing, communication, and RF systems. Hence, ensuring matched point-to-point signal arrival is critical to minimizing distortion, latency, and data corruption while managing increased routing density and interconnect complexity. In this paper, delay mismatch in routing-dominated systems is investigated while assuming negligible mismatch within circuit blocks. A variance-based framework is proposed to quantify and optimize the delay mismatch components across multiple levels of abstraction: the system, block, and layout. To validate this model, systematic routing optimization techniques are applied to an impulse-radio ultra-wideband (IR-UWB) edge-combiner transmitter front end (TFE) operating at 6.5-8.0 GHz center frequency. The TFE requires precise pulse widths in the picosecond range, making it highly susceptible to interconnect-induced delay variability. After optimization, the total variance of propagation delay was decreased by 94.37% based on post-layout simulations.
This work presents a time register designed for time-domain signal processing, employing a delay line in a ring configuration topology combined with a digital counter to achieve a high dynamic range. The proposed architecture supports both addition and subtraction operations, offering a practical framework for understanding the implementation of such functionalities in the time domain. The register is capable of efficiently processing signals with durations ranging from picoseconds to milliseconds and demonstrating feasibility for physical implementation.
This work presents the evolution of a Three-Level Flying Capacitor Converter (3L-FCC) ASIC design from earlystage modeling and simulation towards full integration in a mixed-signal ASIC. As an intermediate step, the power converter core was first implemented in an analog ASIC, complemented by a System-on-Chip (SoC) integrating an embedded CPU for the digital control loop and FPGA fabric for the modulation scheme. This hybrid platform enabled functional validation and post-silicon testing of the analog stage before migrating the remaining digital control and phase-shifted pulse-width modulation (PS PWM) logic into the ASIC. In the current stage of the project, the control architecture is being migrated to a hardware description language (HDL) implementation, with the goal of enabling full integration into a standard digital design flow for mixed-signal ASICs. An FPGA-based prototype is being used to verify functionality and assess timing performance prior to full integration, providing a validation path towards future monolithic implementation.
This paper introduces a novel time register for time-domain signal processing, utilizing a delay line within a ring configuration topology and a digital counter to achieve high dynamic range. The proposed design was implemented and simulated using the AMS 350 nm CMOS process and tested under nominal supply voltage of 3.3 V. Corner simulation, considering supply voltage, temperature, and process variations, demonstrated an error rate below 2% for inputs exceeding 50 ns. A key advantage of the proposed solution is that its range can be easily scaled by increasing the number of counter bits, while the layout area overhead introduced by the counter remains minimal. This is the first CMOS time register capable of efficiently processing signals from picoseconds to milliseconds, providing an output pulse equal to the input pulse, and being feasible for physical implementation.
Deploying modern applications with significant resource demands is often challenging, but approximate designs offer a promising alternative by delivering high performance with minimal compromises in output quality. Traditionally, approximate hardware accelerators have been developed through search-based iterative frameworks, which suffer from long runtimes due to the exponential growth of the design space. A significant portion of the runtime is consumed by either invalid nodes or valid nodes that offer minimal improvements in performance metrics, such as runtime or power consumption. This severely limits the thorough exploration of the design space. In this paper, we introduce a novel approach for synthesizing approximate accelerators that leverages sparsity to reduce the complexity of the design space exploration problem. Our method employs a generative adversarial network (GAN) to rapidly generate a diverse set of high-quality design nodes, eliminating the need for costly node evaluations. This enables the swift creation of approximate accelerators generated for any given error threshold in a fraction of time as compared to a simulation-based framework. We conducted experiments on a suite of benchmarks from real-world domains, demonstrating that our methodology can generate approximate hardware designs with significant area and power savings, comparable to state-of-the-art search-based approaches. In a comparative evaluation against two leading methods, our approach achieved equal or better quality results for two out of four benchmarks while reaching up to 55% area savings, thus effectively demonstrating a new avenue for automated generation of approximate accelerators.
The fact that we can stream video on multiple devices in our homes, on the go using mobile devices, or even while video chatting across the globe, even with low bandwidth, is owed to video coding. Versatile Video Coding (VVC) is the latest video coding standard, which introduces a range of innovative tools. One example is adopting an alternative fractional motion estimation (FME) filter that is part of the Advanced Motion Vector Resolution (AMVR) extension. This work introduces a lowpower hardware architecture accelerator specifically designed for Fractional Motion Estimation (FME) with support for the AMVR extension of VVC.
This article presents a structured methodology for migrating digital control algorithms developed in high-level languages into synthesizable HDL suitable for ASIC implementation. A proportional-integral controller for a three-level flying capacitor converter is used as a representative case study to demonstrate the design flow. The methodology encompasses initial modeling using floating-point arithmetic, conversion to fixed-point representation, generation of reference models for validation, translation to C/C++, and high-level synthesis for HDL generation. Preliminary results consider co-simulation to ensure behavioral consistency between the high-level description and the generated HDL. The HDL was successfully synthesized using the open-source LibreLane tool-chain targeting the SkyWater SKY130 PDK. Ongoing work includes FPGA-based prototyping and detailed characterization of the ASIC. The proposed approach establishes a verified and systematic pathway for transitioning high-level control logic to ASIC-ready hardware implementations, preserving the original algorithmic behavior across abstraction levels.
This paper presents a comparative study between the measurement data provided in the IHP-Open-PDK and the simulation results obtained using open-source tools. Our primary objective is to evaluate how accurately open source simulators, specifically Ngspice and Xyce, can replicate real-world performance when used in conjunction with open-source device models developed by IHP. To ensure a meaningful comparison, we recreated the measurement conditions within our simulation environment, carefully aligning the test scenarios and circuit setup. The study focuses on key discrepancies observed between our open-source simulation results and reference simulations performed with proprietary EDA tools, as documented in the IHP-Open-PDK repository. We analyze these differences in detail and discuss possible causes, including modeling limitations, simulatorspecific behaviors, and configuration variations. Our findings highlight areas where open source tools currently diverge from industry-standard results and suggest targeted improvements for future tool and model development. By systematically revealing these differences, our work contributes valuable feedback to the open-source hardware design community. The insights gained from this study will help developers and users of open-source PDKs and simulation tools to improve model accuracy, improve tool compatibility, and ultimately promote broader adoption of open-source methodologies in integrated circuit design workflows.
The increasing deployment of artificial intelligence (AI) at the edge, particularly convolutional neural networks (CNNs) in resource-constrained devices, has created new challenges for ensuring system reliability and safety. Market analysts project a 21% annual growth rate in the edge AI market size over the next five years. These devices are being used in safety-critical applications such as autonomous vehicles, industrial control systems, and medical devices, where malfunctions due to radiationinduced soft errors can have severe consequences, ranging from degraded performance to life-threatening situations. Soft errors, caused by energetic particles, can corrupt data and instructions, resulting in unpredictable system behaviour. To meet safety standards in these domains, reliability engineers must proactively explore and implement efficient mitigation solutions during the initial design cycle.
Drones are nowadays essential in construction and structural maintenance, providing high-resolution images for structural assessments and precise interventions. However, complex maneuvering in confined spaces and reduced communication with external positioning systems, such as GPS, pose challenges for conventional drones. Micro-drones, with their small size that minimizes failure impact and enhances adaptability in restricted or hazardous environments, have gained increasing research interest for expanding drone applications across various fields. This study presents an energy-efficient design that incorporates an optimized combination of basic color separation and a custom downsized version of the Sobel operator to detect markers placed along the trajectory of the drone in real time. The flight instructions stem from the marker identification to fine-tune altitude, horizontal position, and orientation of the micro-drone before moving forward to the next marker. The proposed strategy demonstrated accurate marker detection under favorable lighting conditions 97% of the time, while significantly reducing BlockRAM usage.
The measurement of breath rate (BR) is essential for comprehensive human health monitoring across a wide range of scenarios. Several studies in the literature have explored the estimation of BR using millimeter wave (mmWave) technology. However, these approaches typically focus on a single subject at a time. To enable multi-person estimation, researchers have often relied on data fusion with camera systems or employed specialized hardware configurations. On the contrary, this paper proposes a methodology that employs only one Frequency Modulated Continuous Wave (FMCW) radar to estimate the BR of multiple subjects stationary in the environment. The proposed methodology includes a pre-processing pipeline to refine the radar-captured signals, followed by frequency-domain analysis to distinguish between subjects. Finally, phase variations in the reflected signals caused by chest movements are analyzed to estimate the BR. Advantages and limitations of the approach are discussed on the basis of an experimental campaign.
The globalization of digital circuit manufacturing has reduced costs and enabled large-scale production. However, reduced visibility into the supply chain can lead to supply issues and the insertion of malicious modifications into the original design. Hardware Trojans are alterations introduced into circuits to degrade performance, change functionality, or leak information. This paper proposes the application of the detection method for Null Convention Logic (NCL) combinational circuits, targeting the valid data states (’01’ and ‘10’). Furthermore, it reviews existing methods for detecting Hardware Trojans in NCL architectures. The propose method investigated is based on probabilistic transition calculations. Within this framework, the output probabilities of each logic element-specifically, the likelihood of transitioning to a ‘ 10 ‘ or ‘ 01 ‘ state-are computed based on the corresponding input probabilities. Consequently, should unauthorized modifications be present, the algorithm identifies them by detecting statistical divergences between the probabilistic values calculated for the golden (original) circuit and the circuit under test.
As technology nodes advance, design rules restrictions are becoming increasingly complex due to physical limitations. This has made automatic cell layout generation a crucial and rapidly evolving area of study. A core stage of the layout design flow is the placement of transistors, which directly impacts the intracell routing and overall layout characteristics. This work presents a concise review of the different approaches that tackle the transistor placement problem to meet cell design goals.
This work addresses the acceleration of convolutional neural network (CNN) inference in manycore architectures using coarse and fine-grain parallelism. Prior approaches focus on dedicated accelerators or modified NoCs, limiting flexibility. This work proposes integrating a RISC-V processor extended with the vector extensions (RVV) as general-purpose processing elements in a NoC-based manycore. The implementation applies depthwise convolution mapped across PEs and uses auto-vectorization provided by the compiler. Experiments on a 4x4 manycore running the first AlexNet layer achieved up to 5.70x speedup and reduced execution cycles by 82.45% compared to a scalar single-core baseline.
Several new shorter floating-point formats have been proposed to match requirements of emerging application workloads. To simplify hardware development in the presence of an increasing number of formats, one practical design option is to use as much as possible preexisting hardware, such as standard 32-bit IEEE-754 (FP32) floating-point units, to handle emerging, less complex formats. We evaluate the case where we use an FP32 multiplier to run Nvidia TensorFloat32 data. While the FP32 multiplier area is not as small as a dedicated TensorFloat32 multiplier, we show that energy per operation scales well with the mantissa width reduction and that smart pin assignment can leverage uneven input vector switching activities to significantly decrease energy for reduced precisions.
The objective of this work is to unify, with mathematical rigor, the analyses to estimate the effects of process variability and the effects caused by radiation in digital circuits using a single analysis. First, we proved that to estimate the effects of process variability on delay and power, it is enough to estimate at least one of the two. Later, a relationship between the effects of radiation and the effects of variability was obtained through a common element between the two: the entropy of the circuit. We applied the presented methods to compare circuits with the operations: matrix multiplication, graph searches, encryption algorithms, and neural networks. All of the circuits are designed using a 180 nm technology. We were able to verify which circuits are more sensitive to the effects of variability and radiation simultaneously through the proposed method.
This work presents the design, simulation, and comparative analysis of a 6-transistor (6T) static random-access memory (SRAM) cell optimized for ultra-low-power operation. Emerging device technologies-Fin Field-Effect Transistor (FinFET), Tunnel Field-Effect Transistor (TFET), and Carbon Nanotube Field-Effect Transistor (CNFET)-are evaluated individually and in hybrid configurations to overcome the limitations of conventional CMOS SRAM cells. Device-level models are calibrated to ensure fair comparison, matching leakage currents and parasitic capacitances. Using NGSPICE and HSPICE simulations at subthreshold supply voltages (0.6-0.8 V), we measure static, dynamic, and total power consumption for each configuration. Results show that the FinFET-TFET hybrid achieves the lowest total power dissipation, combining the low leakage of TFETs with the high speed and channel control of FinFETs, making it suitable for energy-constrained applications such as IoT and implantable devices.
Emerging Post-Quantum Cryptographic (PQC) schemes such as FALCON demand highly optimized hardware implementations to meet strict area and execution time constraints on embedded devices. Traditional hardware designs rely heavily on expert-crafted Register Transfer Level (RTL) or High Level Synthesis (HLS) code, which is time-consuming and error-prone. In this work, we explore the use of large language models (LLMs) for accelerating the development of cryptographic hardware, focusing on FALCON’s performance-critical Samplerz subroutine. We propose a design flow that iteratively leverages LLMs to generate, refine, and evaluate synthesizable C code using HLS tools. We analyze generated designs across a range of models (e.g., GPT-4, Claude, Gemini, Grok), compare them with prior hand-crafted RTL designs, and report implementation metrics including Area-Delay Product (ADP) and synthesis convergence. Alongside achieving implementations within 4% execution time and 30% area of expert-tuned code, our results demonstrate that LLMs can discover novel hardware optimizations. We finally identify key challenges in prompt engineering, numerical stability, and testbench overfitting, and provide actionable recommendations for future AI-assisted hardware design frameworks.