
Due to the desire to further improve the energy efficiency and performance of ultra-low power (ULP) radio frequency (RF) circuits. In recent years, the growing demand for Internet of Things (IoT) devices has required efficient design methodologies. However, confronting the trade-off becomes difficult. Therefore, this paper proposes a systematic sizing methodology for low-noise transconductance amplifiers (LNTAs) based on normalized biasing metrics and pre-calculated lookup tables. This approach aims to optimize the inherent tradeoffs between gain, noise, linearity, and power consumption. The methodology was applied to design and compare six variations of a 5.8 GHz LNTA implemented in 65 nm CMOS technology. The comparative analysis demonstrates that the normalized biasing metric offers the best compromise, achieving a Figure of Merit (FoM) of 180.8 GHz with a power consumption of only 147.5 µW. This approach provides a structured design flow that allows designers to quickly identify an optimal operating point and thus accelerate the development of efficient RF front-ends for power-constrained IoT applications.
The continuous scaling of semiconductor technology drives the demand for integrated circuits with higher performance and smaller area. In this context, standard cell libraries play a key role in optimizing digital designs. This paper presents the ASIC implementation of the PicoRV32 RISC-V CPU using a 4.5-Tracks multi-height standard cell library tailored for 7nm FinFET technology. While multi-height libraries are established, the novelty of this work lies in the practical demonstration of a complete back-end implementation, including logic synthesis, placement, routing, and area optimization. The utilization of the proposed 4.5-Tracks library achieves a 29% reduction in overall core area compared to a conventional 6-Tracks library, with DFFR and XNOR2 cells contributing most to this reduction. These results highlight the effectiveness of multi-height cells in real-world designs and provide practical guidance for future area optimized RISC-V ASIC developments.
The symmetrically built MOS transistors of integrated circuits exhibit symmetric electrical behavior if the source and drain terminals are interchanged. Additionally, a series association of transistors is electrically “equivalent” to a single transistor. However, some of the compact MOSFET models do not comply with the requirements of symmetry and transistor equivalence. This paper reports tests of symmetry and of series association of transistors of some compact models available in circuit simulators. We show that the ACM2, a chargebased model in which the terminal voltages are referred to the substrate, is fully compliant with the transistor symmetry, but that some popular models are not. To test the symmetry property, we show examples of transistor current-voltage characteristics and derivatives up to the fifth order, and capacitance-voltage characteristics, all tests around $V_{D S}=0$. A MOSFET binary current divider is employed to test the consistency of the model applied to a series association of transistors.
This paper presents the design of an open-source wireless power and data transfer (WPDT) unit for implantable and wearable medical devices (IWMD) using a parameterized-cells (P-Cells) library of commonly used building blocks. The P-Cells are designed with open-source design tools Ngspice and Xschem, and use Python as modeling and scripting language. Based on input specifications, four blocks including a capacitor bank, load-shift keying modulator, full-wave rectifier, and low-dropout regulator have been automatically designed using the P-Cells. To highlight the versatility of our approach, two units with different load specifications have been implemented in the open-source IHP SG13G2 130 nm process. Simulation results demonstrate agreement with the specifications and the viability of our approach.
This paper presents an ultra-low-power hardware architecture for muscle fatigue detection based on the discrete Haar Wavelet transform (DHWT). Surface electromyography (sEMG) signals are processed using a 4-level (D4) discrete Wavelet transform (DWT) to extract mean (MNF) and median (MDF) frequency features, key indicators of muscle fatigue. Our DHWT explores the techniques of pruning and approximation. The pruning eliminates redundant arithmetic modules that are unimportant for muscle fatigue detection. The approximation DHWT (ADHWT) replaces conventional adders with approximate adders (AxAs), such as COPY, ETA-I, LOA, Truncation (TRUNC), and AxPPA. The ADHWT maintains the effectiveness of the muscle fatigue detection, while significantly saving 96.01% energy and 94.16% area. The ADHWT with the AxPPA represents the best trade-off between the energy- and area-savings versus maintaining the ability to detect muscle fatigue.
Neurostimulators are key devices for treating neurological disorders, where closed-loop operation remains a major challenge. Noise in stimulation current drivers, though rarely analyzed, can critically affect closed-loop performance when stimulation and sensing bandwidths overlap. This work studies noise in current sources for neurostimulation, focusing on scenarios where stimulation artifacts (SA) and evoked compound action potentials (ECAP) share the same frequency band. A cascoded current sink was designed using the $gm/ID$ methodology to meet a set of requirements, while exploring trade-offs among noise, power consumption, compliance voltage, and silicon area. Simulations benchmark the design against state-of-the-art references, showing that reaching microvolt-level noise demands significant overheads in area and power. The results suggest that alternative architectures may be needed to suppress dominant noise sources.
Wireless networks are widely used in industrial environments due to their practicality and mobility, however, they face challenges such as interference, bandwidth limitations, and security issues, which impact their reliability and energy consumption. This work employs the concept of a “Digital Twin” with the purpose of virtualization the network to perform prior modifications on the physical network, focusing on Industrial Wireless Sensor Networks (IWSN), in order to analyze their performance, diagnose problems, predict future outcomes, and implement improvements more safely and efficiently. The project collects data from a real network, such as signal levels, packet loss rates, and overall performance. These data are used to build a virtual model, which is subjected to analyzes based on social network theory concepts, such as betweenness, degree, and closeness centralities, to identify bottlenecks and critical nodes. From this analysis, interventions such as adding repeaters, creating new links, and changing the routing algorithm are simulated to evaluate their implications for network scalability.
Neural video codecs, such as DCVC-RT, represent the state-of-the-art in real-time neural video compression, achieving high coding efficiency. However, their deployment on resource-constrained hardware is limited by their high computational complexity, particularly in activation functions like the Weighted Sigmoid Linear Unit (WSiLU). To address this challenge, a hardware-friendly polynomial approximation of the WSiLU function is proposed, for which a dedicated high-throughput hardware architecture was also designed. This approximation reduces by almost 10 times the number of arithmetic operations with a minor impact on the DCVC-RT coding efficiency of 0.49 % in BD-Rate. Synthesis results targeting a 40 nm technology demonstrate that the designed hardware can process Full High Definition (FHD - $\mathbf{1 9 2 0} \boldsymbol{\times} \mathbf{1 0 8 0}$ pixels) at $\mathbf{3 0}$ frames per second (30 fps) video in real time, with an area of 209.81 kG ates and a power dissipation of 29.97 mW. To the best of the authors' knowledge, this is the first work in the literature proposing a hardware design for the WSiLU activation function.
This paper presents a comparative study of inverter cell topologies for use in a single capacitively coupled (SCC) ring VCO targeting multi-phase RF applications. Three inverter cells are investigated: nMOS-only (N), pMOS-only (P), and complementary nMOS-pMOS (NP), each implemented with current-starved control for frequency tuning. The cells are designed in $28-\text{nm}$ FD-SOI technology under a common parameterization, ensuring that performance differences arise from intrinsic architectural properties rather than transistor sizing. Simulation results compare frequency tuning range, power consumption, output swing, signal symmetry and phase noise. The N cell achieves the widest tuning range, the NP cell offers better symmetry and the $\mathrm{P}$ cell shows limited swing and higher asymmetry but with lower power consumption. These results reveal intrinsic trade-offs between efficiency, signal integrity, and spectral purity, providing practical guidelines for inverter selection in SCC VCO designs.
The evolution of deep learning architectures has culminated in Transformers emerging as the state of the art for applications such as natural language processing (NLP) and computer vision. Still, the cost of time and energy for processing highlights the need for the use of accelerators. The present project aims to implement a hardware accelerator for one of the elementary units of a Transformer, the Scaled-dot product attention, in order to mitigate this issue. The designed models are integrated with an AXI-Stream communication interface. The solutions were employed on a RTL-GDSII flow, which included functional verification, logical and physical synthesis. In order to evaluate the implementation, the synthesis was carried out with different values of operand bit width (4,8 and 12), tokens and features (4, 8 and 16), and attention heads (1, 2 and 4). By doing so, it was found that the presented design is sufficiently robust, outperforming the state-of-the-art accelerators in throughput, power and area efficiency, also reaching higher operation frequencies.
The widespread adoption of Deep Neural Networks (DNNs) across many domains, including safety-critical applications, has been driven by the availability of sophisticated deployment algorithms that enable the development of complex and effective solutions. Their success is largely due to advanced compression and pruning techniques that preserve accuracy and performance, as well as specialized hardware accelerators, such as GPU Tensor Cores (TCs) with support for structural sparsity. However, the overall impact of soft errors on sparse CNN applications due to corrupted sparsity mechanisms in GPUs has not yet been fully investigated. This work evaluates the impact of soft errors impacting the structured sparsity mechanism in GPU TCs on the overall operation of CNN models. For the experiments, we use real GPUs that implement sparse versions of multiple Convolutional Neural Network (CNN) models. Then, several fault injection campaigns targeting the sparsity mechanism (errors in sparse indices and compressed weights) enable the characterization of errors in sparse workloads. According to our experiments, the sparsity mechanism increases the fault resilience of CNN deployment. Moreover, it is shown that sparse indices are less susceptible to soft errors than the weights. These results indicate that control corruptions are mostly masked by the inherent resilience of CNN workloads.
Hyperdimensional Computing (HDC) is a hardware-friendly machine learning paradigm that encodes data into high-dimensional binary hypervectors, enabling efficient computation through lightweight bitwise operations. Existing HDC accelerators, often implemented on FPGAs, typically rely on custom instruction sets and co-processor designs to optimize performance. However, these architectures face programmability challenges and incur significant control overhead when interacting with the main processor. To overcome these limitations, we present a methodology for designing and optimizing tightly coupled HDC accelerators. The proposed method produces accelerators enhanced with a controller capable of decoding instructions corresponding to key HDC kernels, thereby reducing communication overhead and improving throughput. To validate the approach, we implemented the accelerator at the RTL level and modeled the entire system in SystemVerilog. Experimental results demonstrate that, compared to state-of-the-art architectures, the proposed method reduces software code size by up to 98% while achieving a $(3 \times$) speedup.
A Swarm Intelligence-based metaheuristic, the Cheetah Optimizer (CO), is applied to hyperparameter tuning of Multilayer Extreme Learning Machines (MELM) for cardiac healthcare classification tasks. MELM extends standard ELM by stacking hidden layers, which increases representational capacity but also the number of hyperparameters to configure, making optimization costly. The proposed CO-based method efficiently searches for suitable configurations, evaluated on two benchmark datasets: the South Africa Heart Disease Data (SAHD) and the Mendeley Heart Attack Dataset (MHAD). Performance was measured using F1-Score and G-Mean. Results are compared with Exhaustive Search, showing that CO achieves performance close to the optimum while significantly reducing computational cost. CO was configured to explore only $\approx 1 \%$ of the fine Exhaustive Search space (step size ${=}1$), which corresponds to $\approx \text{2 5 \%}$ of the coarse Exhaustive Search actually implemented (step size $=5$). These findings highlight CO as an effective alternative for building efficient predictive models in healthcare decision-support systems, where lightweight models are especially valuable for potential deployment in embedded or low-power environments.
This paper proposes a CNN-driven fast intra-mode decision solution for the Angular modes in the VVC standard. The Angular modes of VVC are categorized into four classes according to their similar direction, while a CNN is trained to predict the most suitable class for each block, supporting all VVC block sizes. With the CNN predictions, the solution evaluates only the Angular modes within the predicted class, achieving an average time saving of 3.48 % with only a 0.22 % loss in coding efficiency. Compared with related works, the proposed solution achieves competitive results and can be integrated with existing methods to improve time savings.
Emerging edge technologies such as Advanced Driver Assistance Systems and autonomous vehicles demand exceptional dependability to mitigate risks and ensure human safety. To address these challenges, industry standards like ISO 26262 enforce rigorous safety requirements for automotive systems. While EDA tools are advancing to support safer SoC designs, the potential of Virtual Prototyping for developing and validating safety mechanisms remains underexploited. This paper introduces a fully SystemC/TLM2.0-based virtual prototype of an automotive-grade SoC (AutoSoC), designed to support early-stage design. The prototype integrates an instruction set simulator, memory, a UART peripheral connected to a terminal interface, and a custom CAN controller. It enables early-stage validation of the AutoSoC software stack and provides debugging capabilities for monitoring internal states and peripheral interactions. Simulation of the Software Test Library (STL) on the OpenRISC 1000 CPU revealed behavioral discrepancies in register operations compared to RT-level models, highlighting the VP's effectiveness for architectural exploration. Furthermore, simulating STL protection of automotive peripherals such as the CAN controller supports iterative refinement of the highlevel model to ensure functional and register-level accuracy. By reducing reliance on RT-level models during early development phases, our approach accelerates design cycles and facilitates the early detection of architectural vulnerabilities.
In the field of analog mixed-signal circuits, Sample-and-Hold (S/H) circuits are frequently used as a key component in other circuits to preserve an analog voltage for downstream processing, as done in Analog-to-Digital Converters (ADCs). When developing and validating these, simulative methods are often used on the basis of behavioral models. One of the faults to be examined is signal noise and how it affects the circuits' behavior. Unfortunately, simulations for noise-injected input voltages involve a very high computational effort. This paper presents an efficient noise injection method in which noise is only injected at crucial time points during the simulation. Thus, the number of samples instead of the pure simulated time is decisive for the computational effort. As a result, our method achieves a significantly lower required computing time equal to transient analysis without noise. An implementation in VerilogAMS is introduced and the simulation results are presented.
This paper proposes synthesizing Electrocardiogram (ECG) signals from Photoplethysmogram (PPG) recordings using a Generative Adversarial Network (GAN). This proposal integrates synchronized acquisition through a custom hardware platform, an explicit preprocessing pipeline for narrowband and broadband noise suppression, and a GAN with dual-domain discriminators operating in both time and frequency domains. The model was trained on a combined dataset comprising public sources and real-world data collected with the proposed acquisition system. Under realistic experimental conditions (varying heart rates, activity levels, and noise), the proposed method achieved a Mean Squared Error (MSE) of $0.2780 \pm 0.05$ and a Pearson correlation coefficient $\rho=0.68 \pm 0.03$ between real and synthesized ECGs, corresponding to a $\mathbf{7 2. 2 \%}$ similarity based on normalized MSE. These results demonstrate the feasibility of reconstructing ECG morphology from PPG signals in wearablecompatible devices.
This paper presents the design, logic synthesis, and characterization of a reconfigurable CIC decimation filter implemented as an ASIC for data-conversion systems. The proposed architecture provides high flexibility by enabling dynamic adjustment of both the filter order and the decimation factor. To control internal word-length growth, a bit-pruning technique is applied to reduce resource usage. The circuit was implemented in 22 nm CMOS technology and validated through post-synthesis simulations, demonstrating a compact area of $1455.86 \mu ~\mathrm{m}^{2}$, equivalent to 4952 normalized logic gates, a maximum operating frequency of 7.2 MHz, and a peak ENOB of 16.11 bits. These results confirm the feasibility and efficiency of the proposed architecture as an adaptable decimation solution for high-precision data converters.
Charge-based MOSFET compact models provide a physically consistent framework to describe transistor charges and capacitances across operating regimes. Unlike current-based approaches, they enforce charge conservation and yield reliable predictions of dynamic and RF behavior. This paper reviews the main charge-based formulations, ranging from industrial standards (BSIM, PSP, HiSIM) to academic compact models such as EKV and the recent ACM-2 five-parameter approach. We contrast their philosophies, complexity, and accuracy, highlighting the trade-offs between highly parameterized industrial models and compact analytical formulations oriented to design and education. Representative applications in analog/RF design, digital timing and power estimation are discussed. Particular attention is given to the lightweight ACM-2 model as a paradigmatic example of simplicity and analytical clarity. We conclude by outlining current challenges-advanced device architectures, quantum effects, and automated parameter extraction-and perspectives for future compact modeling in deeply scaled technologies.
This paper presents a comprehensive evaluation of the prediction tools implemented in the open-source VVenC encoder, focusing on their computational cost, prevalence, and overall efficiency. The experiments target both intra- and inter-frame prediction modules across the five VVenC coding profiles (faster to slower), aiming to capture the encoder's behavior across different scenarios. An efficiency metric was defined relating the prevalence to the measured computational complexity of each prediction tool within the video encoding process. Our results show that inter-frame prediction tools achieve significantly higher efficiency (up to 4.47) compared to intra-frame tools (up to 0.92), primarily due to their higher prevalence and similar computational costs. Thus, VVC novel intra-frame prediction tools present promising opportunities for optimization schemes. From another perspective, the inter-frame prediction's efficiency drops in more complex profiles and at higher resolutions, as heavier implementations of internal prediction tools are enabled, indicating a trade-off between complexity and performance. Looking at inter-frame prediction steps, Integer Motion Estimation (ME) plays an important role and benefits from hardware acceleration, while Fractional ME presents opportunities for adaptive optimizations. Affine ME, despite its computational cost, shows limited prevalence and emerges as the most promising candidate for complexity reduction and hardware optimizations.