
There are increasing concerns on the security of IoT devices. Due to limited OS/firmware capability and insufficient external support in IoT devices, a new security architecture becomes imminent. “Root of Trust” and “cyber resilience” strengthened with hardware are two key security measures to make IoT devices less vulnerable to cyberattacks and have ability to recover by itself after attacks. Sec...
This paper presents a fast and efficient optimization engine with multi-directional, multi-objective algorithms based on a robust transistor sizing approach to improve digital circuit performance. However, such optimization processes are highly simulator-dependent and computationally expensive tasks. There-fore, we propose developing machine learning-based reliable models considering process and operating variations to speed up the optimization procedure by running them on developed Residual Neural Network (ResNN) models instead of running expensive circuit simulations. Results on 22nm Metal Gate High-K digital cells show a reduction in delay and leakage up to 36.7% and 18.8 %, respectively improving computational efficiency by several orders.
Semiconductor systems, whether cell-phones, laptops, servers or machine learning accelerators, require an increasing number of processing units, larger caches and faster interfaces while managing cost and power dissipation. As we continue to push the transistor, interconnect and memory scaling to smaller geometries, the industry is moving quickly towards 3D heterogeneous integration for higher den...
As we are moving into a new era with data-driven applications, not only datasets need to be analyzed in a smart way (with various kinds of neural models), but also new datasets have to be generated and observed from our living environment. As a result, life quality can be enhanced and improved with better usage of natural resources. In this paper, we will introduce new types of sensors which can b...
Heterogeneous 2.SD/3DIC integration is becoming more and more popular and important since device scaling and process shrink are becoming more and more difficult and expensive. This lecture will introduce status of advanced packaging, including FOWLP (Fan-out Wafer Level Packaging). Key technologies, applications, and achievements of modern EDA tools for 2.5D/3D IC design and analysis will also be presented
In this paper, we propose a very compact embedded CNN processor design that can easily fit into edge devices based on a modified logarithmic computing method using very low bit-width representation. For Yolov2, our processing circuit takes only 0.15 mm 2 using TSMC 40 nm cell library. The key idea is to apply low bit-width logarithmic representation and devise a unified, reusable CNN computing kernel that can significantly reduce computing resources. The proposed approach has been extensively evaluated on many popular image classification CNN models (AlexNet, VGG16, and ResNet-18/34) and object detection models (Yolov4). The hardware-implemented results show that our design achieves 20x performance improvement while consumes only minimal computing and storage resources, yet attains very high accuracy. The design is thoroughly verified on FPGAs, and the SoC integration is underway with promising results. With extremely efficient resource and energy usage, our design is excellent for edge computing purposes.
In this paper we propose a thermal quorum sensing (TQS) scheme that serves as a distributed thermal sensor network, based on whose feedback we can control the operation voltage and frequency of an integrated circuit (IC) within the safety margin, in order to extend the system reliability and lifetime. The TQS scheme and the existing dynamic voltage and frequency scaling (DVFS) modules can be integrated on a chip to serve as the secondary system (for online sensing and repair) in a symbiotic system. We also present a temperature simulation procedure for investigating the correlation between temperature change and sensor behavior of the proposed TQS scheme. Experimental result from a real IC design shows that the TQS-equipped symbiotic IC has a lifetime of 3.4x that of the original one without TQS, with an average clock frequency of 42 MHz as compared with the original 100 MHz clock. The hardware overhead for the extended lifetime is only about 0.37%.
This paper presents a model compression frame-work for both pruning and quantizing according to the channel distribution information. We apply the variational inference technique to train a Bayesian deep neural network, in which the parameters are modeled by probability distributions. According to the characteristic of the probability distribution, we can prune the redundant channels and determine the bit-width layer by layer. The experiments conducted on the CIFAR10 dataset with the VGG16 show that the number of parameters can be saved by 58.91x. The proposed compression approach can help implement hardware circuits for efficient edge and mobile computing.
Pattern-based weight pruning on CNNs has been proven an effective model reduction technique. In this paper, we first present how to select hardware-friendly pruning pattern sets that are universal to various models. We then propose a progressive pruning framework, which produces more globally optimized outcomes. Moreover, to the best of our knowledge, this is the first paper dealing with the pruning issue of the first and also the most sensitive layer of a CNN model through a two-staged pruning strategy. Experiment results show that the proposed framework achieves 2.25x/2x computation/model reduction while minimizing the accuracy loss.
Off-chip substrate routing for high-density packages is on the critical path for time to market. There are several substrate routing algorithms have been proposed in previously. Although routers can rapidly that produce routing results, these results might not be satisfied universally from expert's experiences. In other words, different routers tend to have strength and weakness from different SOC designs. In this paper, we propose a novel reroute framework to remedy the defect of substrate routers by using supervised machine learning. We build a classification model which extracts features from expert's experience. It will identify suboptimal routings that do not conform to manual routing style. Then, reroute these areas using different routers and produce diverse results, then feed to classification model until they are acceptable. Guided by the model, suboptimal routing areas are replaced by results that are closer to expert's manual routing. Experiments show that our rerouting framework achieves 36.5% improvement on the number of wire bends and 1.6% wirelength improvement, compared with initial results routed by recent related work.
High-frequency trading (HFT) systems require extremely low latency in response to market feeds to make profits. In order to reduce latency, the system is implemented with a 10 gigabit Ethernet physical transceiver with a low latency of 25 ns, custom network stack parsing and packaging, partial financial protocol decoding and encoding, order book handling and custom trading strategy. The hardware test and functional verification of the trading system use Taiwan futures trading environment provided by Yuanta Futures. The latency from the internal market packet analysis to the ordering packet triggered is approximately 433 ns.
The demand for low-power ADCs is increasing in battery-powered applications, and research has also been conducted on low-power continuous-time delta-sigma modulators (CTDSMs) in recent years. This talk explores various techniques to improve the energy efficiency of CTDSMs with some design examples.
This paper presents a structure pruning with design space exploration for DCNN accelerators. The design space exploration tool, called NNArch, generates optimized design scheduling for the row-stationary DCNN accelerators with fast and accurate analytical performance/energy models. Based on the NNArch, the proposed filter-segment pruning can efficiently compress the DCNNs with a simple filter index table, optimizing model accuracy, accelerator performance, and energy consumption. The experiment result shows that the proposed pruning scheme can achieve 2 times speedup and improve the energy-delay-product by 3.42 times on ResNet50 with the accuracy drop of 0.2%, on the accelerator of 168 PEs.
This article presents a true random number generator (TRNG) that achieves high entropy generation across wide voltage and temperature (VT) range (0.3–1.0 V, −40 °C to 110 °C) in a single latch-based entropy source (ES). In the ES, static inverter selection technique to minimize the mismatch between the paired inverters, and noise enhancement methods to increase the root mean square (rms) of noise voltage ( $\sigma _{n}$ ) are implemented for good randomness and robustness. In a 130-nm CMOS technology, the TRNG occupies 5343 $\mu \text{m} ^{\mathrm{ 2}}$ and consumes 0.116 pJ/bit at 0.3 V including an on-chip von Neumann post-processing circuit. The cryptographic quality of TRNG’s output is verified by National Institute of Standards and Technology (NIST) SP800-22 tests. Up to 325 mV $V$ pp noise injection attack tolerance is confirmed by power supply frequency injection attack. And an equivalent 20-year life at 0.3 V, 25 °C is verified by accelerated NBTI aging test.
The high power-efficiency requirement of the deep learning accelerator (DLA) draws attention to the in-memory computing technologies in the recent years. Several memory types such as MRAM, ReRAM, SRAM, PCM are potential candidates, and research groups make major progress in the TOPS/W index over the years. Different application scenarios of deep learning accelerators may impose constraints and lim...
A methodology for Artificial Intelligence (AI) edge Deep Convolutional Neural Network (DCNN) hardware design to increase computation parallelism and decrease latency is needed for a real time application. To increase the computation parallelism, a 1-bit by 1-bit high parallelism in-RRAM computing (IRC) macro is proposed. The goal of this testing macro is to test the characteristic of the RRAM and propose a co-training mechanism between DCNN algorithm and RRAM module to deal with the non-linearity issues of IRC.
Mixed-signal interfaces are the essential bridges between the physical world and the digital information processing backbone. In recent years, innovation in such interfaces has been increasingly fueled by application-level insight and the data-driven nature of modern systems. As a result, the traditional building block boundaries are blurring, and the extraction of information occurs through symbi...
Wide bandgap (WBG) power semiconductor devices, such as silicon carbide (SiC) and gallium nitride (GaN), offer the ability to reach higher voltage, higher frequency, and better temperature characteristics compared to silicon-based power devices. These potentials would revolutionize the way we deliver and manage power in the future and meet the demands for more efficient, smaller size and further cost reduction in the power electronics industry. In this paper, the opportunities, challenges, and potential transformative impacts on the power electronics industry would be discussed from both device and system perspectives.
Integrated circuits are known to be the key enabler of modern world of information and communication technologies (ICT). In the past several decades we have witnessed all major technology advancements which have enriched and enhanced everyone's everyday life. While human beings are enjoying technologies around them, the varieties and scales of information generated for technologies become increasi...
Ultra-low power attentive systems with always-on operation and signal monitoring with a disproportionately higher peak performance are now being in high demand, due to the convergence of AI and IoT. In this talk, circuits and architectures to enable exceptionally low power consumption in the common case while achieving high peak performance are discussed for next-generation intelligent systems. Se...