
A low voltage supply CMOS current conveyor circuit for digital input signals from 0.25 V up to 1.2 V is presented. The circuit is optimized and pre-layout simulated in a 65 nm CMOS process technology. At the target design voltage of 1.2 V, the current conveyor has a propagation delay of 2.86 ns, an energy consumption of only 80.9 pJ, and energy-delay product (EDP) of 231 pJns for resistive load of 10 kΩ. Superior performance of this work is demonstrated through comparison with other similar published work at a frequency of 5 MHz. It is shown that the proposed circuit is suitable for digital signaling. The developed CMOS circuit perfoms correctly until 50 MHz and its EDP is 31 pJns at 10 kΩ.
In the past few years, electrical energy harvesting from the ambient environment has attracted much of the research interests due to the increasing need for energy resources that suit the rapid development of power electronic devices. Recently, a novel mechanical energy harvesting and transduction technology has been introduced called the triboelectric nanogenerators (TENGs), with many advantages over the other harvesters. In this work, a new diagonal motion mode for TENGs is studied intensively in the attached electrode regime and simulated using COMSOL MULTIPHYSICS. The diagonal motion mode offers a new way of motion represented in the angle theta (theta)-between the left endpoint of the bottom dielectric and the left endpoint of the top dielectric-that ranges from 0 to 90 degrees, and allows for further energy optimization. A complete analytical model of the proposed TENG structure has been derived step by step, and its results were compared to COMSOL simulation results to validate its accuracy. The maximum error between the results is equal to 5.3%, 4.3% and 0.88% for the open circuit voltage (V-OC), short circuit charge (Q(SC)) and Capacitance (C) respectively. Based on the developed model, a performance optimization study has been performed on the dielectric material used in the device. The study revealed that, the tribo-pair of Nylon and Kapton corresponds to the best performance of the harvester amongst the other pairs used in the literature with a maximum energy of 0.40118 nJ. In addition to this, the accuracy of the analytical model was a motivation for constructing a Verilog-A model, by which the device was studied under different loading conditions such as simple resistive and capacitive loads. Finally, a MATLAB based CAD tool has been developed to allow rapid prototyping and testing of different device parameters on the device performance.
Exploiting the benefits of Virtualization in the world of Embedded technology has opened up new avenues for effective resource utilization, increased scalability, security and cost savings. With the above in perspective, the performance benchmarking of virtualized embedded systems is important. In this paper, we have assessed the performance of various types of virtualization techniques such as paravirtualization and hardware-assisted virtualization in a desktop environment. Microkernel based virtualization techniques are more suitable for embedded system environment, due to its low memory footprint and security advantages as only a small amount of trusted code is running at a high privileged level. We have used this implementation to analyze the performance of an OS on a microkernel based virtual environment and compared its performance with an OS in a nonvirtual environment on the same board. In addition to this, we have analyzed the performance of different types of virtualization techniques possible with a microkernel on a low power arm based embedded system with a benchmarking tool.
This paper presents a portable, low power electronic system for tactile sensory feedback for prosthetics. The system provides a prototype based on commercial off-the-shelf components having a wireless portable interface electronics and an array of 32 tactile sensors. The proposed prototype demonstrates to be able to transmit the touched signals with low power energy consumption achieving a battery lifetime of about 22 hours. Experimental results validate the correct functionality of the proposed prototype when different input stimuli have been tested. The proposed system overcomes similar state-of-the-art solutions having higher number of input channels featuring low power and real time operation.
We present a power-driven hierarchical framework for module/functional-unit selection, scheduling, and binding in high level synthesis. A significant aspect of algorithm design for large and complex problems is arriving at tradeoffs between quality of solution and timing complexity. Towards this end, we integrate an improved version of the very runtime-efficient list scheduling algorithm called modified list scheduling (MLS) with a power-driven simulated annealing (SA) algorithm for module selection. Our hierarchical framework efficiently explores the problem solution space by an extensive exploration of the power-driven module-selection solution space via SA, and for each module selection solution, uses MLS to obtain a scheduling and (integrated) binding (S&B) solution in which the binding is either a regular one (minimizing number of FUs and thus FU leakage power) or power-driven with mux/demux power considerations. This framework avoids the very runtime intensive exploration of both module selection and S&B within a conventional SA algorithm, but retains the basic prowess of SA by exploring only the important aspect of power-driven module-selection in a stochastic manner. The proposed hierarchical framework provides an average of 9.5% FU leakage power improvement over state of the art (approximate) algorithms that optimize only FU leakage power, and has a smaller runtime by factors of 2.5–3x. Further, compared to a sophisticated flat simulated annealing framework and an optimal 0/1-ILP formulation for total (dynamic and leakage) FU and architecture power optimization under latency constraints, PSA-MLS provides an improvement of 5.3–5.8% with a runtime advantage of 2x, and has an average optimality gap of only 4.7–4.8% with a significant runtime advantage of a factor of more than 1900, respectively.
In this paper, a new architecture of four-stage CMOS operational transconductance amplifier (OTA) based on an alternative differential AC boosting compensation called DACBC is proposed. The presented structure removes feedforward and boosts feedback paths of compensation network simultaneously. Moreover, the presented circuit uses a fairly small compensation capacitor in the order of 1 pF, which makes the circuit very compact regarding enhanced several small-signal and largesignal characteristics. The proposed circuit along with several state-of-the-art schemes from the literature have been extensively analysed and compared together. The simulation results show with the same capacitive load and power dissipation the unity-gain frequency (UGF) can be improved over 60 times than conventional nested Miller compensation. The results of the presented OTA with 15 pF capacitive load demonstrated 65° phase margin, 18.88 MHz as UGF and DC gain of 115 dB with power dissipation of 462 μW from 1.8 V.
This paper introduces a novel structure for the realization of a low voltage, low power current-mode analog to the digital converter (ADC) pipeline (12 bits). The proposed structure of the ADC is based on a novel design of a current comparator and Digital to Analog Converter (DAC) structure. This modification allows us to reach a higher speed, lower voltage, and lower power dissipation. ELDO simulators using 0.18 μm, CMOS and TSMC parameters are performed to confirm the workability of this architecture. The proposed ADC is powered with a 1 V supply voltage. It is characterized by wide conversion frequency (350 MHz) and low power consumption that is 2.76 mW.
In this paper the CMOS amplifier behaviour has been further investigated respect to the previous works in the literature. An exhaustive scenario for the EMI pollution has been considered: the injected interferences can indeed directly reach the amplifier pins or can be coupled from the PCB ground. This is a key point for evaluating also the susceptibility from the EMI coupled to the output pin, which is disclosed as a critical point. The investigated topologies are basically derived from the Miller and the Folded Cascode, which are well-known and widely used by the CMOS analog designers; all of them are re-designed in UMC 180 nm CMOS process in order to have a fair comparison.
In this paper, an extensive review of the available publications about comparing estimations versus measurements of power consumption in FPGA technology is carried out. This study reveals that the variety of experimental setups makes it difficult to elaborate solid studies departing from the results of different researchers using meta-analysis techniques. To mitigate this problem, we propose a procedure to standardize the setup of FPGA power estimation experiments. The goal is to make as close as possible power estimations and their corresponding actual on-chip measurements. The main idea is to use a fixed arrangement composed by a parameterized pattern generator block at the input, together with a set of interchangeable IP cores utilized as reference circuits. All the blocks are mapped together inside the FPGA sample, being the clock and reset lines the sole input signals. Thus, both power estimation and actual measurements are performed to the whole system in identical conditions. In order to illustrate the method, the paper includes some examples of the proposed methodology for different cores. A set of 25 circuits have been tested in two FPGA families, obtaining relative errors in power estimation between –61.5% and 9.2%.
Benthic microbial fuel cells (MFCs) are promising alternatives to conventional batteries for powering underwater low-power sensors. Regarding performances (10's μW at 100's mV for cm 2-scale electrodes), an electrical interface is required to maximize the harvested energy and boost the voltage. Because the MFCs electrical behavior fluctuates, it is common to refer to maximum power point tracking (MPPT). Using a sub-mW flyback converter, this paper compares the benefit of different MPPT strategies: either by maximizing the energy at the converter input or at the converter output, or by fixing the MFC operating point at its nominal maximum power point. A practical flyback has been validated and experimentally tested for these MPPT options showing a gain in efficiency in certain configurations. The results allow determining a power budget for MPPT controllers that should not exceed this gain. Eventually, considering typical MFC fluctuations, avoiding any MPPT controller by fixing the converter operating parameters may offer better performances for sub-mW harvesters.
In this work, we propose an approximate and energy-efficient CORDIC method, based on a trigonometric function spatial locality principle derived from benchmarks profiling. Successive sine/cosine computation requests cover more than 50% when the absolute phase difference is at most ten degrees. Consequently, this property suggests an optimized circuit implementation, both iterative or a succession of microrotation modules, where the last CORDIC requires fewer iterations, reducing the latency and the total energy budget at the same precision of two separate and independent instances. Thus, this simple design strategy allows significant area and energy dissipation in general-purpose VLSI architectures, but it introduces also dramatically optimizations in applicationspecific embedded systems used in the area of signal processing and radio frequency communication. In this contribution, we introduce a method, the hardware overhead and the energy budget per single cycle. Simulation results show the total energy saving in considered benchmarks is 40% in pipelined and iterative general purposes CORDIC. Furthermore, our application-specific systems (fast Fourier transform and digital oscillators for radiofrequency down conversions) show remarkable cycle savings when the successive sine/cosine computation requests are more than 70%. Finally, in this work, we extend the proposed approach to whichever phase difference less than 26.56° , as a variable for the second CORDIC number of angle rotations.
In a finite state machine (FSM), there is only one active state while the other states are in idle states simultaneously. Thus, only one state is required to power up, the other states can be switched off to save active power. Normally, a backward traversing algorithm is used to label the power gating cells and extract the enable signals for combinational logic gates in reducing the active power consumption. This conventional power gating technique uses the extracted enable signals to turn ON/OFF these inserted NMOS switches. Then, a power management unit is required to manage these enable signals and detect the idle periods. The proposed self-power saving technique uses internally generated enable signals from state transitions to control NMOS switches inserted under the ground rail of each state. All internal enable signals are created to activate/deactivate the machine states at the same time. Based on the next state of the FSM, a decoder creates the enable signals for each state to do power gating in an Automatic Teller Machine (ATM) application. The isolation cell is designed to isolate the current state and next state for retaining data. Simulation results show the power saving from 31.99% at a WAIT state to 82.87% at a LOCK state, depending on the current state of the finite state machine. On average, the power loss is saved up to 63.2% in the FSM. An overhead area is about 12% compared to the conventional technique while timing overhead is under 5%.
A high speed N × N bit multiplier architecture that supports signed and unsigned multiplication operations is proposed in this paper. This architecture incorporates the modified two's complement circuits and also N × N bit unsigned multiplier circuit. This unsigned multiplier circuit is based on decomposing the multiplier circuit into smaller-precision independent multipliers using Vedic Mathematics. These individual multipliers generate the partial products in parallel for high speed operation, which are combined by using high speed adders and parallel adder to generate the product output. The proposed architecture has regular-shape for the partial product tree that makes easy to implement. Finally, this multiplier architecture is implemented in UMC 65 nm technology for N = 8, 16 and 32 bits. The synthesis results shows that the proposed multiplier architecture improves in terms of speed and also reduces power-delay product (PDP), compared to the architectures in the literature.
In this paper, a low noise and low power analog front end for piezoelectric microphones used in hearing aid devices is presented. It consists of a Charge Amplifier, followed by a Variable Gain Amplifier and an Analog-to-Digital Converter. At the core of charge amplifier a two stage opamp with modified cascode current mirror is designed which achieves a gain of 93 dB and phase margin of 62°. Designed analog front end achieves an input referred noise of 0.12 μ Vrms and SNR of 74 dB. It consumes power of 430 μ W from 1.8 V supply and occupies an area of 1.2 mm × 0.22 mm. Proposed circuit is designed and fabricated in 0.18 μ m CMOS process. Designed circuit is interfaced with a sensor model of piezoelectric microphone, which mimics Ormia ochracea's auditory system, and its performance is successfully verified against simulation results.
IoT and autonomous systems are in charge of an increasing number of sensing, processing and communications tasks. These systems may be equipped with energy harvesting devices. Nevertheless, the energy harvested is uncertain and variable, which makes it difficult to manage the energy in these systems. Reinforcement learning algorithms can handle such uncertainties, however selecting the adapted algorithm is a difficult problem. Many algorithms are available and each has its own advantages and drawbacks. In this paper, we try to provide an overview of different approaches to help designer to determine the most appropriate algorithm according to its application and system. We focus on Q-learning, a popular reinforcement learning algorithm and several of these variants. The approach of Q-learning is based on the use of look up table, however some algorithms use a neural network approach. We compare different variants of Q-learning for the energy management of a sensor node. We show that depending on the desired performance and the constraints inherent in the application of the node, the appropriate approach changes.
CPU architecture has experienced great innovation in its architecture, from 8 bit to 64 bit, CISC to RISC, Single core to multi-core and single pipelined logic to deep multi-pipelined system. Today in an era of 64 bit architectures, 8 bits are still very relevant and has not lost its position and being used in many applications. Hence this research work deals with 8 bit CPU architecture and its features enhancement to make the 8 bit case very relevant in an era of 64 bit. The co-operative ALU, as name suggests, works in tandem with existing ALU and performs 16 bits operations. The specially designed instructions shares knowledge and efficiently handles existing ALU and Co-operative ALU to perform 8 bits and 16 bits operations. The Co-operative ALU is integrated with the 2 stage pipelined 8-bit RISC architecture ensuring that existing architecture is kept intact by way of applying new functionality in the form of an extension. The reconfigurable platform software tools are used for functionality verification and final deployment is done using reconfigurable platform hardware tools.
Intellectual Property threat is most challenging in the integrated chip manufacturing industry. Outsourcing the chip fabrication to untrusted foundries leads at the risk of intellectual property theft. Logical locking is the most popular countermeasure against the hardware intellectual property threats. In which the design is functionally locked with the added logic gates and thus the logically locked design is accessible only by applying the valid keys. This made the attacker to put all efforts to identify the valid key to unlock the integrated chip through various key-guessing attacks. Hence, an enhanced logical locking with locker-box is proposed to design a secured hardware. This locker-box is a logical non-linear component, which produces an ambiguous relationship between inputs and outputs. This non-linear mathematical function build a strong hardware against key-guessing attacks. Furthermore, the gates at which the logical locking is implanted plays a vital role as the design overheads like area, power and delay are concerned for any design. In this work, the gates with low observability values are chosen, such that an incorrect key will give an output corruption of 50% as a Hamming distance with minimal design overhead and implementation complexity. The experimental results are validated on ISCAS'85, ISCAS'89 and ITC'99 benchmark circuits.
The network-on-chip (NoC) paradigm has emerged as a high-performance communication architecture for integrating many IP cores on a single die. The benefit of performances arising out of adopting an NoC is often constrained by the interconnects that primarily consists of metallic communication channels. The reason is, as observed during the lifetime of the NoCs, the communication channels quickly become performance impediment while they experience a different kind of manufacturing faults. If it is attempted towards ensuring the desired reliability and yield of the NoCs, one must concern in tackling these faults on the interconnects of the NoCs. This paper presents a novel test-solution for detecting and locating different manufacturing faults and evaluates their severe impact on network performance. This paper for the first time presents a test strategy that can handle coexistent permanent faults occurring on both interswitch and local channels of a 2D mesh-based NoC system. Simulation results show the runtime evaluation of the proposed test strategy under realistic traffic scenarios in a 2D mesh-based network and provide deep insights into the underlying mechanism that affects multiple performance metrics in the presence of the basic manufacturing faults in the network. It is observed that the proposed scheme incurs low test cost in terms of low area overhead, low test time, low power consumption, low latency, etc. One advantage of the proposed strategy scales up seamlessly to large NoCs regardless of routing algorithms, topology, and channel width. Another advantage includes that the proposed test solution improves many basic quality parameters over the prior works, which are normally studied for system performance evaluation.
HAL is a multi-disciplinary open access archive for the deposit and dissemination of scientific research documents, whether they are published or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et à la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. An Optimum Inexact Design for an Energy Efficient Hearing Aid Sai Praveen Kadiyala, Aritra Sen, Shubham Mahajan, Quingyun Wang, Avinash Lingamaneni, James Sneed German, Hong Xu, Krishna Palem, Arindam Basu