
ABACUS parallel architecture was previously proposed as an alternate integer multiplication approach with column compression and parallel carry futures. This paper presents a VLSI implementation for ABACUS and benchmarks it against the conventional Wallace Tree Multiplier (WTM). Simulations are conducted with UMC180nm technology in Cadence environment. Although WTM implementation results in 26.6% fewer devices, ABACUS implementation has 8.6% less power dissipation with matched delay performance, due to 27.8% lower average activity.
Face detection is one of the most popular computer vision applications on mobile platforms. It is a compute-intensive task that consumes significant energy. In this paper, we present an energy efficient face detection implementation that offloads data- and compute-intensive portions of the application onto low-power mobile GPU to save overall power consumption without sacrificing performance. Our experiment on a state-of-the-art mobile processor demonstrates that our proposed approach saves power consumption up to 14.3% and improves performance by 87% over traditional CPU only execution.
Mobile data traffic has skyrocketed in recent years and will continue to grow and demand more capacity. Continued heterogeneous networks' evolution and the utilization of massive Femtocell deployments will facilitate the meeting of capacity demands as expected towards 5G and telecom 2020-vision. The CO2 footprint of power usage is becoming spotlighted and energy-aware solutions are required. This paper reviews recent Femtocell downlink power control frameworks and proposes a novel comprehensive one taking into consideration issues like intra-interference (Femto-Femto), dense-deployments and environmental impact. Simulation results show up to 68% user throughput improvement and up to 20.69kg/year CO2-emission reduction for one Femtocell.
Explosive growth of mobile devices usage and the quick increase of the mobile applications are facing many challenges in their resources as low computing power, battery life, limited bandwidth, and storage. Mobile Cloud Computing (MCC) has been introduced to be a potential technology for mobile services and to solve the mobile resources problem by moving the processing and the storage of data out from mobile devices to the cloud. The cloud enables the integration with additional development tool as graphical processing power (GPU) to increase the computational power. This paper presents a novel approach for real time face detection using GPU acceleration. The results of developed Applications demonstrate that the proposed Mobile GPU cloud computing increase both speed and accuracy of facial detection systems.
Typically, complex resource-interdependence and heterogeneous workload patterns can result in sub-optimal job allocation leading to performance loss or under-utilization of compute resources. A well behaved model can anticipate the demand patterns and proactively react to the dynamic stresses in a timely and well optimized manner. For a workload hosting environment, pool of available resources are optimally configured and utilized to sustain certain expectation of Quality-of-Service (QoS) in the presence of power, thermal and reliability constraints. The workload (or job) scheduling mechanism is expected to withstand dynamic variations in demand stresses while maximizing the resource utilization and minimizing the performance loss. Furthermore, workloads can be co-allocated to the clusters with least amount of resource contention. In this paper we introduce the methodology that facilitates the coordinated scheduling of the workloads to the systems with least contentious resources through phase-assisted dynamic characterization. We describe the method to perform optimal job scheduling by using phase model synthesized by learning and classifying the run-time behavior of workloads.
This paper presents a 24 Gbps SerDes transceiver circuit for on-chip high speed serial links for on-chip networks. The transceiver uses a proposed almost-differential self-timed 3-level signaling scheme, which works using a frequency of half the data rate for relaxing the design. Also, the third voltage level is created without the need for an external Vdd/2 supply source. Moreover, a 3-level inverter is proposed for the use in the front-end of both the TX and the RX. The transceiver is designed for a 5mm long lossy on-chip differential interconnect in GF 65nm CMOS technology. It serializes the parallel 3 Gbps 8-bit, and multiplexes them with the 12 GHz input clock. A simple RX extracts both the data and the clock from the same signals.
Energy consumption and energy modeling are important issues in designing and implementing of Wireless Sensor Networks (WSNs), which help the designers to optimize the energy consumption in WSN nodes. Good knowledge of the sources of energy consumption in WSNs is the first step to reduce energy consumption. Therefore, an accurate energy model is required for the evaluation of communication protocols. In this paper, we provide an energy model for WSNs considering the physical layer and MAC layer parameters by determining the energy consumed per payload bit transferred without error over AWGN channel. We show how the transmission power must be chosen in order to achieve energy-efficient communications over AWGN channel. We also find that, for each modulation scheme, there are optimal transmission power at which the energy consumption is minimized. Moreover, we investigated the energy saving gained from optimizing the constellation size.
All-Optical Clock and Data Recovery (OCDR) is an important function for future optical networks and optical signal processing. The OCDR realizes a long-distance optical data transmission system by restoring the incoming data and then retransmitting. The Self-Pulsating (SP) lasers are the promising technologies to enable fast and high-speed data recovery system in an optical domain. In this paper, we design and implement the OCDR based on two SP laser types, Amplified Feedback Laser (AFL) and Distributed Bragg Reflector Laser (DBRL). A comparative study and measurement of the network performance for the two types have been presented.
This paper presents characterization for coupling capacitance in through silicon Vias (TSV) arrays. Two scenarios are proposed to estimate the coupling capacitance between TSVs in TSVs array. First scenario is by using a closed form expression that accounts for the shielding effect resulted by TSVs. Second scenario is based on the existence of initial measured capacitance value at certain dimensions, thereafter the capacitance values can be obtained at other dimensions using scaling equations.
This paper presents a low input voltage and high step-up fully integrated DC-DC regulator in 0.18 μm standard CMOS technology for thermoelectric micro-power generation. The circuit avoids off-chip components, non-standard processes, and is thus suitable for ultra-low voltage low profile system-on-chip applications. The proposed system can deliver a regulated output voltage of 1.5 V at 31 μW output power with an input voltage as low as 0.2 V. The maximum simulated efficiency is 22% at the given step-up range. The design methodology of an integrated inductor layout and oscillator has been reported in detail for the standard process. At the ultra-low voltage range of interest, the regulator is estimated to have lower cost, higher integration, and improved efficiency compared to the alternatives reported in literature, including the 90 nm and 0.18 μm two-stage charge pump designs previously reported by our team.
This paper presents a fully self-powered interface circuit with a novel peak detector for piezoelectric energy harvesters (PEH). This circuit can be utilized to scavenge energy from low power environmental vibrations in 10s of μW range. Synchronous switching technique is used to extract maximum available power where switching instants are detected independently from excitation changes of the PEH. The proposed peak detector senses voltages higher than power supply for a wide frequency range of input vibration. The simulations with an output voltage range of 1 to 3.3 V show power conversion efficiency between 79% and 88% for an input power of 13.2 μW.
Recent focus in energy efficiency is motivated with diminishing conventional energy resources, and increasing demand in low power applications with shrinking platform sizes. In this work, various threshold logic technologies are compared with each other in terms of power-delay-product (PDP). Compound CMOS, complementary pass transistor, static NAND gate, full adder, capacitive and differential threshold logic technologies are compared within a developed comparison scenario. Results in UMC180nm technology indicate that complementary pass transistor based threshold logic proves at least 2.5% more efficient than the rest in terms of PDP, while NAND based implementation has 29.2% better in terms of delay performance.
This paper introduces a new algorithm and circuit design of Time-to-Digital Converter(TDC) with modified Successive Approximation Register(SAR) algorithm. This design enables continuous pulse disassemble. The input pulse is absolutely compared to pulses of widths proportional to Vfs/2, Vfs/4..Vfs/N, and each bit is evaluated independent of the previous bit result. Then bits correction is applied after the sample evaluation. A 4bit case study circuit is realized using TSMC CMOS 65nm design technology. The design demonstrated 3.67 Effective Number Of Bits (ENOB) for a sampling frequency of 666 MS\s.
This paper presents a performance enhancement feature for a novel power management circuit to generate 1.8 V from the low DC voltage rectified at the output of the vibration-based electromagnetic (EM) energy harvesters. The proposed 180 nm circuit utilizes a low voltage charge pump based boost converter with variable output-stages, and an autonomous regulator circuit with negative feedback topology. 2 and 3 stage charge pump options in the variable stage configuration has been validated to extend the supported input voltage range at the same load, or alternatively maintain higher efficiency operation at a higher load range. The simulation results showed that under no-load condition the output voltage reached to 1.8 V for input voltage of 0.65 V and 0.48 V with 2 and 3 stage outputs, respectively. The power conversion efficiency of the power management circuit can be kept stable around 55% by switching from 2 to 3 stages after 3.5 μA.
High cost of qualifying library standard cells on silicon wafer limits the number of test circuits on the test chip. This paper proposes a technique to share common load circuits among test circuits to reduce the silicon area. By enabling the load sharing, number of transistors for the common load can be reduced significantly. Results show up to 80% reduction in silicon area due to load area reduction.
Nowadays, FPGAs serve as Fields Programmable Systems on Chip (FPSoC) and are widely used to implement computationally intensive world applications. As the number of components in FPSoCs increases, the interconnect schemes based on Network on Chip (NoC) approach are increasingly used to overcome the problems of traditional bus based and point-to-point interconnect scheme. In this paper, we review several designs based on their contributions, architectures, implementations and future works. We also made our comparison between three of these routes to analyze the effect of varying the number of Virtual Channels (VCs), flit data width and buffer depth on the operating frequency, Logic Look-Up Tables (LUTs) and registers to help choosing the appropriate NoC based on system requirements.
In this paper, a 5-level buck converter is proposed. The circuit structure and the working principle are illustrated. The circuit is capable of providing five different voltage levels at the inductor input with the help of two flying capacitors. The 5-level buck converter can work at different operation regions covering wide range of output voltage values. By reducing the voltage difference at the inductor input, the 5-level buck converter can use smaller inductor compared to both 3-level and conventional buck converters which makes it more suitable for on-chip DC-DC conversion. A test circuit has been implemented in TSMC 65nm technology using 0.5nH on-chip spiral inductor and simulation results show better performance as compared to conventional and 3-level buck converters. For same switching frequency and inductor size, the 5-level buck converter achieves more than a 15% efficiency improvement over a 3-level buck converter at certain output voltage ranges.
With the advent of teraflop-scale computing on both a single coprocessor and many-core designs, there is tremendous need for techniques to fully utilize the compute power by keeping cores fed with data. Data prefetching has been used as a popular method to hide memory latencies by fetching data proactively before the processor needs the data. Fetching data ahead of time from the memory subsystem into faster caches reduces observable latencies or wait times on the processor end and this improves overall program execution times. We study two types of prefetching techniques that are available on a 61-core Intel Xeon Phi co-processor, namely software (compiler-guided) prefetching and hardware prefetching on a variety of workloads. Using machine learning techniques, we synthesize workload phases and the sequence of phase patterns using raw performance data from hardware counters such as memory bandwidth, miss ratios, prefetches issued, etc. Furthermore, we use performance data from workloads with different impacts and behaviors under various prefetcher settings. Our contribution can help in future prefetching design in the following ways: (1) to identify phases within workloads that have different characteristics and behaviors and help dynamically modify prefetch types and intensities to suit the phase; (2) to manage auto setting of prefetcher knobs without great effort from the user; (3) to influence software and hardware prefetching interaction designs in future processors; and (4) to use valuable insights and performance data in many areas such as power provisioning for the nodes in a large cluster to maximize both energy and performance efficiencies.
This paper presents the VHDL implementation of a novel power management algorithm for standalone PV-battery system. The algorithm performs two tasks, Maximum Power Point Tracking (MPPT) and dual load regulation. The MPPT is used to maximize the PV cells' output power, and is achieved by the “fractional open circuit voltage” method. The dual load regulation distributes the PV cells' output power among the loads, and delivers any surplus or deficit power to or from the battery. The proposed VHDL design has been synthesized on Xilinx using “5vlx50tff1136” as the target FPGA. The proposed design has utilized only 1% of the resources (slice registers and look-up-tables). This result has been found to be 23% lower than a previously implemented MPPT controller..
Low leakage power with maintained high throughput NoC is achieved. Traffic-based Virtual channel Activation (TVA) algorithm is presented to determine traffic load status at the NoC switch ports. Consequently adaptation signals are sent to activate or deactivate switch port VC groups. The algorithm is optimized to minimize power dissipation for a target throughput. TVA algorithm optimally utilizes VCs by deactivating idle VCs groups to guarantee high leakage power saving without affecting the NoC throughput. Network average leakage power has been reduced for different topologies (such as 2D-Mesh and 2D-Torus).