In many FPGA-based systems, only sequential system control structures modeled by finite state machines are actually required. In order to deal with complexity, design time, and verification issues, which are weaknesses of traditional hardware description languages, it may be preferred to describe the control flow behaviorally in software. However, it is reported that high-level synthesis for FPGA often generates inferior results in terms of resources and performance when translating software-style control flow description to hardware. In this paper, the NanoSoftController is proposed as an open-source soft processor, which is optimized for minimal and efficient logic resource usage on FPGA platforms. It is targeted at processing sequential finite state machine functionality in software, featuring a compact ISA for control flow in embedded systems and a tiny accumulator-based data path. Furthermore, an efficient mapping of memory to small distributed LUT RAM instances enables its use as a system state machine controller in even very resource-constrained FPGA designs, requiring only 104 slice LUTs and 76 slice registers in total. However, despite all optimizations, in a case study with high-level synthesis results of three reference software-style control applications, i.e., electronic door lock, smart glucose sensor, and sequential sensor network node, a better resource efficiency could not be shown. We evaluate the negative results and provide lessons we learned from them.
The realization of autonomous, wearable, and implantable Systems-on-Chip (SoC) for health monitoring applications poses several challenges, such as achieving ultra-low size, cost, and power consumption, yet offering sufficient flexibility to re-program and adapt the autonomously operating SoC during the course of treatment. Commonly, programmability is not considered for ultra-low-power biomedical SoCs, since a dedicated finite state machine fixes the operation sequence. This work proposes a combination of a programmable NanoController with a SAR ADC, for potential use in ultra-low-power wearable or implantable biomedical SoCs, and integrates it as a prototype chip in a 22nm FDSOI-CMOS technology. The proposed NanoController achieves an extremely low power consumption of 505nW when programmed for 10-bit SAR ADC control and a sample rate of 100 kSamples/s. This result shows that energy efficiency and programmability during operation are not mutually exclusive for smart ultra-low-power biomedical platform chips. Based on the results of this prototype chip, a fully-integrated glucose sensor chip has been designed and submitted for fabrication, which will be evaluated in future work.
The efficiency of VLIW processors can be improved by reducing the energy consumption associated with accessing the register-file. This paper presents an energy-aware register allocation approach to reduce the dynamic energy consumption of a scheduled program by using an evolutionary algorithm approach and a register access transition energy model. Additionally, the influence of instruction scheduling on energy-aware register allocation is analyzed. By using the proposed energy-aware compiler backend on synthetic applications with random data in the register-file, the dynamic energy consumption of the total processor core for an exemplary VLIW processor can be reduced by up to 21% compared to a heuristic register allocation. In real-world applications from the domain of hearing aids, energy consumption can be reduced by up to 42.7% for single applications, and 17.6% average across a range of applications.
Precision livestock farming consists of technological tools and techniques to improve livestock management. Proper detection and classification of jaw movement (JM) events are indispensable for the estimation of dry matter intake, detection of health problems, and flag the onset of estrus, among other information. The analysis of acoustic signals is one of the most accepted ways to monitor the feeding behavior of free-grazing cattle. Different acoustic methods have been developed for recognizing JM-events in recent years. However, their operation is limited to off-line analysis on a personal computer. The lack of on-line acoustic monitoring systems is associated with the challenging operation requirements (low-power consumption, autonomy, portability, robustness and non-intrusive on the animal). In this paper, a fixed-point variant of the chew-bite energy-based algorithm is presented. This algorithm is implemented on a new low-power audio processor system for real-time recognition of JM-events. The system includes a Nanocontroller processor, which is always-on and detects JM-events; and a second transport-triggered architecture (TTA) based processor, which is mainly in power-down and classifies JM-events. The results demonstrate that the proposed fixed-point JM-events recognizer achieves a recognition rate of 91.4% and 90.2% in noiseless and noisy conditions, respectively. The recognition rate increases by 6.1% regarding a previous reference on-line system. Moreover, the proposed audio processor system chip consumes 4 μW on average, i.e., only 2.3% of the power of an always-on TTA-based processor system for the same audio sequence. An exemplary implementation of the proposed system in a 65 nm low-leakage CMOS technology is given.
In this paper, a heterogeneous controller system and its first-silicon ASIC implementation are presented, where the use of a programmable NanoController next to a general-purpose microcontroller enables more efficient and flexible power management strategies than typical timer-based, periodical power-up of a single microcontroller in state-of-the-art IoT devices. The NanoController features a compact, control-oriented 4-bit ISA, which is used to continuously pre-process data in order to decide when to power-up the microcontroller required for infrequent complex processing, e.g., encrypted wireless communication. Despite its programmability, the required silicon area and power consumption are very small and enable the use in the always-on domain of SoCs for energy harvesting platforms, instead of much simpler and constrained timer circuits. The first-silicon ASIC implementation of such a controller system using a 65nm UMC low-leakage process is presented and evaluated for a real home automation application intended to operate on harvested energy, i.e., electronic door lock, reducing the average power consumption of reference microcontrollers by up to 20x.
Distributed nodes in IoT and wireless sensor networks, which are powered by small batteries or energy harvesting, are constrained to very limited energy budgets. By intelligent power management and power gating strategies for the main microcontroller of the system, the energy efficiency can be significantly increased. However, timer-based, periodical power-up sequences are too inflexible to implement these strategies, and the use of a programmable power management controller demands minimum area and ultra-low power consumption from this system part itself. In this paper, the NanoController processor architecture is proposed, which is intended to be used as a flexible system state controller in the always-on domain of smart devices. The NanoController features a compact ISA, minimal silicon area and power consumption, and enables the implementation of efficient power management strategies in comparison to much simpler and constrained always-on timer circuits. For a power management control application of an electronic door lock, the NanoController is compared to small state-of-the-art controller architectures and has up to 86% smaller code size and up to 92% less silicon area and power consumption for 65 nm standard cell ASIC implementations.
An optimized instruction-set encoding can reduce the silicon area and power consumption of a processor architecture implementation. However, the design space of the input encoding problem is of factorial growth with the number of instruction patterns, so effective heuristics and an automated exploration tool are required to facilitate instruction-set encoding optimization in a processor design flow. This paper proposes a novel approach based on genetic algorithms to automatically optimize the instruction-set encoding of a specific processor architecture, reducing the silicon area and power consumption requirements for specific applications and hardware implementation technologies. Furthermore, an open-source tool, called VANAGA, is presented, which implements the proposed approach and allows flexible adaptation to custom instruction-set optimization scenarios. The tool flow is evaluated with an exemplary 65 nm standard cell ASIC implementation of a minimal controller architecture with 4-bit wide opcodes (NanoController). For different optimization scenarios, logic silicon area and total power consumption vary within a design space range of 6.3% and 0.46% for different instruction-set encodings, respectively.
In this paper, an operand masking approach is proposed to achieve lower energy consumption using approximate computing techniques in programmable high-performance processors, in this case horizontal and vertical SIMD vector processors for embedded computer vision applications. Contrary to state-of-the-art dedicated approximate arithmetic circuits, this mechanism enables programmable fine-grained accuracy control and switching energy reduction at runtime. An evaluation for a 45 nm ASIC technology shows a total effective energy reduction of up to 4.5% for a horizontal SIMD vector processor architecture executing approximate SIFT image feature extraction for an error-resilient egomotion estimation algorithm.
Microcontrollers to be used in harsh environmental conditions, e.g., at high temperatures or radiation exposition, need to be fabricated in robust technology nodes in order to operate reliably. However, these nodes are considerably larger than cutting-edge semiconductor technologies and provide less speed, drastically reducing system performance. In order to achieve low silicon area costs, low power consumption and reasonable performance, the processor architecture organization itself is a major influential design point. Parameters like data path width, instruction execution paradigm, code density, memory requirements, advanced control flow mechanisms etc., may have large effects on the design constraints. Application characteristics, like exploitable data parallelism and required arithmetic operations, have to be considered in order to use the implemented processor resources efficiently. In this paper, a design space exploration of five different architectures with MIPS- or ARM-compatible instruction set architectures, as well as transport-triggered instruction execution is presented. Using a 0.18 $$\upmu $$ μ m SOI CMOS technology for high temperature and an exemplary case study from the fields of communication, i.e., powerline communication encoder, the influence of architectural parameters on performance and hardware efficiency is compared. For this application, a transport-triggered architecture configuration has an 8.5 $$\times $$ × higher performance and 2.4 $$\times $$ × higher computational energy efficiency at a 1.6 $$\times $$ × larger total silicon area than an off-the-shelf ARM Cortex-M0 embedded processor, showing the considerable range of design trade-offs for different architectures.
Embedded automotive Computer Vision systems for real-time motion tracking and 3D scene reconstruction demand for high image feature extraction performance and have a heavily constrained energy budget unable to be met by general-purpose CPUs and GPUs. Due to the required programming flexibility for software updates and algorithmic extensions, the use of fully dedicated hardware accelerators is not advisable in most cases. In this paper, a vertical and a horizontal SIMD vector processor architecture are implemented and compared for accelerating the Scale-Invariant Feature Transform feature extraction algorithm, exploiting inherent data-level parallelism prevalent in this application and considering different programming code strategies for the different vectorization paradigms. An evaluation for a 45 nm ASIC technology shows an overall performance gain of up to 24.8x, and up to 151.3x higher total performance-area-energy efficiency compared to a reference scalar two-issue VLIW processor. Compared to other implementations on programmable ASIP and mobile GPU platforms, the proposed vertical SIMD vector processor achieves a performance gain of up to 5.1x and up to 31.3x higher performance-energy efficiency.
Numerous approximate adders have been proposed in the literature in response to the languishing benefits of technology scaling. However, they have been obtained with an ad-hoc and non-systematic methodology which does not fully exploit the design space possibilities. This paper provides a conceptual framework for the systematic design of approximate adders, including hybrid and non-equally segmented approaches as well as more robust error metrics. The framework discriminates the scenarios, where approximate processing does not provide significant benefits from those where it does; in this later case, it aids to obtain optimal configurations for the adders. Experimental results with a commercial technology assess the significant improvements of our systematic approach. Furthermore, a case study with a processor enhanced with an approximate accelerator highlights the usability of the methods.
The energy efficiency of application-specific processors for high-performance embedded Computer Vision systems can be increased by applying Approximate Computing mechanisms. Due to prevalent error resilience in typical feature extraction or image classification algorithms, the requirement of precise additions and multiplications can be relaxed to obtain more areaand energy-efficient ALU architectures for realtime operation on a limited energy budget. However, state-ofthe-art approximate adder and multiplier architectures do not consider the influence of pipelining in processor datapaths on the area-timing-energy (ATE) trade-offs. In this work, a pipeliningaware synthesis flow and ATE-accuracy trade-off exploration is presented, showing a reduction of up to 20% in silicon area and up to 11% in energy consumption when compared to unpipelined approximate units for the same target performance.
In this paper, the effects of application-specific instruction-set processor (ASIP) hardware optimizations on the performance of beamforming algorithms and on the hardware requirements (i.e., silicon area and power consumption) are studied. For that, the performance of three beamforming algorithms with different fixed-point implementations are compared using objective instrumental measures, i.e., PESQ, STOI, and iSNR. The proposed application-specific hardware optimizations are implemented in a VLIW-SIMD hearing aid processor, modifying the processor's datapath width, using a co-processor for the division operation and applying register file power optimizations. In total 24 different optimized processor configurations are studied. The result of this evaluation is that the same processor, running one of the beamformers, can be optimized, decreasing up to 2 times the silicon area requirements or up to 11 times the power consumption, thereby only slightly decreasing the overall algorithm performance (e.g., −2dB iSNR for a fixed beamformer).
ASICs for Stochastic Computing conditions are designed for higher energy-efficiency or performance by sacrificing computational accuracy due to intentional circuit timing violations. To optimize the stochastic gate-level circuit behavior of a specific design, iterative timing analysis campaigns have to be carried out for a variety of chip temperature- and supply voltage-dependent timing corner cases. However, the application of common event-driven logic simulators usually leads to excessive analysis runtimes, increasing design time for hardware developers. In this paper, a gate-level netlist-oriented FPGA-based timing analysis framework is proposed, offering a runtime-configuration mechanism for emulating different timing corner cases in hardware without requiring multiple FPGA bitstreams. For an exemplary timing analysis campaign of an existing chip design, speed-up factors of up to 267 are achieved while maintaining timing behavior deviations lower than 1.05% to timing simulations.
The demands of high-speed and power-efficient systems have resulted into the emergence of the approximate computing. Existing approximate circuits as well as stochastic techniques have shown promising advances in improving various figures of merit. However, a through fair comparison of arithmetic units still remains an issue which has not been studied. This paper reviews the prerequisites for a fair comparison of approximate arithmetic units. As one of the key components of arithmetic circuits, adders are the focus of this paper. For the first time in this paper, approximate and exact adders are studied together in the stochastic regime. Simulation results show that both the equal segmentation adder (ESA) and the error tolerant adder type II (ETAII) outperform exact adders working stochastically, if and only if the right configuration and sub-adder architectures are chosen. Otherwise, there is no reason to use the aforementioned architectures. In all, considering the cost-error trade-off, Lower-part OR adder (LOA) has the best behavior in the stochastic regime.