This paper proposes a novel low-power approximate floating-point adder with runtime configurable precision. The design combines skipping addition, triggered when the exponent difference exceeds a programmable threshold, with tunable mantissa rounding to reduce switching activity. We present a comprehensive analysis aimed at identifying the most appropriate combination of skipping and rounding thresholds that enables the best balance between precision and power consumption. The thresholds are controlled by an accuracy control signal, which can be modified at runtime to achieve the desired power-precision trade-off. A detailed comparison with state-of-the-art designs was carried out in FP32 and FP16 formats, considering implementations in 28nm technology. The proposed approximate floating-point adder outperforms existing approaches in the power-precision trade-off, while also offering the unique advantage of runtime configurability. This flexibility comes at the cost of a moderate area increase, mainly due to the additional control logic for configurability. Power savings over the exact adder range from 45% to 60% in FP32 and from 20% to 53% in FP16, with the mean relative error distance tunable between 1.34x10(-8 ) and 6.34x10(-6) in the FP32 and from 5.07x10(-4 )to 1.78x10(-2) in FP16. The effectiveness of our configurable adder is validated in three practical applications: JPEG image compression, convolution-based image processing (filtering and edge detection), and time sequence prediction using a Long Short-Term Memory recurrent neural network. In all the tests the proposed adder consistently delivered string performance, confirming the practical relevance of the proposed technique.
Design of application-specific Digital Signal Processing elements is related to the implementation of an algorithm, initially implemented in software. The time-to-design can be extremely long, especially when design space exploration must be done and hardware description must be written from scratch. Different tools are nowadays available to generate a Register Transfer Level (RTL) circuit and corresponding testbenches from software, even with routines for design space exploration. These tools are based on High-Level Synthesis (HLS) and High-Language Hardware description (HLH) paradigms. This work compares the RTLs for an integrated digital Amplitude Shift Keying demodulator, employable either for the recognition of messages detected by transducers or in sensors networks. SystemVerilog, Catapult HLS from Siemens EDA and SpinalHDL were considered for canonical hardware description, HLS and HLH respectively. Synthesis results with three STMicroelectronics technologies prove that the RTLs automatically generated by SpinalHDL and Catapult provide area results comparable with those obtained with from-scratch design in SystemVerilog and can also achieve better performance in terms of dynamic power.
In this paper, an energy efficient HW accelerator for AI edge-computing in Human Activity Recognition is proposed. The system processes samples from a tri-axial accelerometer and classifies the human activities by using a novel Hybrid Neural Network (HNN) topology, which has been designed to reduce the computational complexity of the system while preserving its accuracy. The HW design improves the characteristics of the HNN by means of an architecture that is aimed to reduce the allocated physical resources and the memory accesses. While accuracy measured on ad-hoc dataset is 97.5 %, measurements from synthesis with CMOS 65 nm standard cells report power consumption of 6.3 μW when the sensor output data rate is 25 Hz, normally used for HAR.
NAND-based digitally controlled delay-lines (DCDLs) are employed in several applications owing to their excellent linearity, good resolution and easy standard cell design. A glitch-free DCDL behavior is often a strict requirement [e.g. spread-spectrum clock generators (SSCG) and digitally controlled oscillators]. Existing glitch-free NAND-based DCDL topologies either require two flip-flops for each DCDL delay-element (DE) or present a very long settling time which limits the maximum working frequency. This paper proposes a novel glitch-free NAND-based DCDL that joins the advantages of previously proposed topologies: uses only a single flip-flop for each DE (reducing area and power) and has relaxed timing requirements (allowing easy integration in applications like SSCG). In the paper, the glitch-free operation of the proposed circuit is firstly demonstrated theoretically and then verified experimentally, with the help of an SSCG built using proposed DCDL and implemented in 28 nm CMOS. Simulation results show that proposed DCDL results in a more that 30 % reduction of the power dissipation and a >20 % reduction in area occupation with respect to double flip-flop DCDL, without any timing constraints penalty.
Spread-spectrum clocking is an established approach to mitigate electromagnetic interference (EMI) of digital circuits, by intentionally sweeping the clock frequency. In this way, the energy of each clock harmonic is spread over a larger bandwidth, thereby reducing the peak of the interfering spectrum. This paper describes an highly flexible all-digital spread-spectrum clock generator (SSCG) realized with a standard-cells design flow. The developed circuit supports discontinuous frequency modulation profiles (with improved EMI reduction capability) and features reduced output jitter, due to delay interpolators and digital compensation of delay path asymmetries. The proposed SSCG is ideally suited for complex system on chips applications, having programmable spreading parameters, frequency synthesis capability and reduced recovery time to support local standby modes. The SSCG is implemented in bulk 28 nm CMOS technology, presents a maximum working frequency of 3.3 GHz and less than 3.2 ps rms output jitter. The measured peak level reduction of the clock power spectrum, at 1.0 GHz output frequency, is 27.0 dB with a 10% modulation depth. The power dissipation is 29.3 mW @ 3.3 GHz and the area occupation is 0.031 mm.
Spread spectrum clocking is an effective solution to reduce the electromagnetic interference produced by digital chips, using a clock signal with a frequency that is intentionally swept (frequency modulated) within a certain frequency range, with a predefined modulation profile. We present the implementation of an all-digital spread spectrum clock generator. The circuit is realized by using a design flow completely based on standard cells and is able to perform clock spreading with an arbitrary modulation profile and a modulation frequency up to 5 MHz. The circuit uses two digitally controlled delay lines driven by a digital modulator to synthesize the output waveform. A replica delay line is employed in a real-time measurement circuit to track process, voltage and temperature variations. A chip has been implemented in a 65 nm CMOS technology. The chip is able to generate signals up to 1.27 GHz. The measured peak level reduction of the clock spectrum, at 750 MHz output frequency, is 20.5 dB with a 6% modulation depth. The power dissipation is 44 mW @ 1.27 GHz.