A 240-GHz wideband low-noise amplifier (LNA) incorporating high-speed customized transistors and dual-peak-G(max) cores is proposed in this work for 6G applications. The customized transistors, designed and modeled using an electromagnetic modeling approach, reduce the gate resistance and the drain-to-gate capacitance, enhancing f(max) from 288 to 394 GHz. The dual- peak-G(max) core utilizes a reciprocal embedding network consisting of two pre-embedding transmission lines and a DC-isolated Y-embedding transmission line to achieve maximum gain conditions at 221 and 261 GHz simultaneously, enabling the LNA to exhibit wideband characteristics efficiently. Implemented in a 40-nm digital CMOS technology, the proposed LNA shows a measured power gain of 16.2 dB at 220 GHz with a 3-dB bandwidth spanning from 208.6 to 223.6 GHz and a simulated noise figure of 11.5 dB while only consuming 34.7 mW from a 0.9-V supply. The measured output 1-dB compression point is -5.3 dBm at 220 GHz.
Computation in memory (CIM) overcomes the von Neumann bottleneck by minimizing the communication overhead between memory and process elements. However, using conventional CIM architectures to realize multiply-accumulate operations (MACs) with flexible input and weight bit precision is extremely challenging. This article presents a fully bit-flexible CIM design with a compact area and high energy efficiency. The proposed CIM macro employs a novel multi-functional computing bit cell design by integrating the MAC and the A/D conversion to maximize efficiency and flexibility. Moreover, an embedded input sparsity sensing and a self-adaptive dynamic range (DR) scaling scheme are proposed to minimize the energy-consuming A/D conversions in CIM. Finally, the proposed CIM macro implementation utilizes an interleaved placement structure to enhance the weight-updating bandwidth and the layout symmetry. The proposed CIM design fabricated in standard 28-nm CMOS technology achieves an area efficiency of 27.7 TOPS/mm2 and an energy efficiency of 291 TOPS/W, demonstrating a highly energy-area-efficient flexible CIM solution.
To enable energy-efficient computation for deep neural networks (DNNs) at edge, computing-in-memory (CIM) is proposed to reduce the energy costs during intense off-chip memory access. However, CIM is prone to multiply-accumulate (MAC) errors due to non-idealities of memory crossbars and peripheral circuits, which severely degrade the accuracy of DNNs. In this work, we propose a Data-Driven Non-ideality Aware Training (D-NAT) framework to compensate for the accuracy degradation. The proposed D-NAT framework has the following contributions: 1) We measured a fabricated SRAM-based CIM macro to obtain a data-driven MAC error model (D-MAC-EM). Based on the derived D-MAC-EM, we analyze the impact of the non-idealities on DNN’s accuracy. 2) To make DNNs robust to the non-idealities of CIM macros, we incorporate the measured D-MAC-EM into DNN’s training procedure. Moreover, we propose a statistical training mechanism to better estimate the gradients of the discrete D-MAC-EM. 3) We investigate trade-offs between quantization range and quantization errors. To mitigate the quantization errors in activations, we introduce extended PACT (E-PACT) that adaptively learns the upper and lower bounds of input activations for each layer. Simulation results show that our proposed D-NAT improves the accuracy of ResNet20, VGG8, ResNet34, and VGG16 by 78.98%, 71.8%, 72.04%, and 57.85%, respectively, which reaches the ideal upper bound of the quantized model. Lastly, the D-NAT framework is validated on an FPGA platform with the fabricated SRAM-based CIM macro chip. Based on the measurement results, D-NAT successfully recovers the accuracy under non-idealities of a real SRAM-based CIM macro.
The majority of digital sensors rely on von Neumann architecture microprocessors to process sampled data. When the sampled data require complex computation for 24×7, the processing element will a consume significant amount of energy and computation resources. Several new sensing algorithms use deep neural network algorithms and consume even more computation resources. High resource consumption prevents such systems for 24×7 deployment although they can deliver impressive results. This work adopts a Computing-In-Memory (CIM) device, which integrates a storage and analog processing unit to eliminate data movement, to process sampled data. This work designs and evaluates the CIM-based sensing framework for human pose recognition. The framework consists of uncertainty-aware training, activation function design, and CIM error model collection. The evaluation results show that the framework can improve the detection accuracy of three poses classification on CIM devices using binary weights from 33.3% to 91.5% while that on ideal CIM is 92.1%. Although on digital systems the accuracy is 98.7% with binary weight and 99.5% with floating weight, the energy consumption of executing 1 convolution layer on a CIM device is only 30,000 to 50,000 times less than the digital sensing system. Such a design can significantly reduce power consumption and enables battery-powered always-on sensors.
Mutli-camera vehicle tracking and re-identification (re-ID) have gradually gained attention due to their applications in the intelligent transportation system. However, these problems are fundamentally challenging. Specifically, for vehicle tracking, we observe that the results generated from single camera tracking algorithm usually recognize tracklets with same identity as different vehicles when the tracklets are occluded. Hence, we propose a Tracklet Reconnection technique to refine tracking results with pre-defined zone areas and GPS information. The proposed method can efficiently filter invalid tracklet pairs and reconnect the split tracklets into complete ones, which is important for the afterwards multi-target multi-camera tracking. As for re-ID, we also find that when a large-scale auxiliary dataset is used to assist the learning of main dataset for better model capability and generalization, there is a performance drop caused by data imbalance when the full auxiliary dataset is applied. To tackle this problem, we introduce Balanced Cross-Domain Learning to avoid the overemphasis on larger auxiliary dataset by a newly introduced training data sampler and loss function. The extensive experiments validate the empirical effectiveness of our proposed components.
Many of earlier attempts on mobile cloud integrations aim on increasing storage capacity of mobile devices, rather than its computation capacity. Traditional distributed computing model relies on static and reliable network connections to share workload among collaborative devices. The researches on pervasive and ubiquitous computing community enable collaborative computation to be conducted on connected computers. Similarly, it is limited to predefined computation services including predefined services and computation platforms. The computation resources in modern computation environment are heterogeneous and evolve over time. The aforementioned computation models do not make good use of such resources. We design and implement an extended OpenCL framework to federate the computation resources of mobile devices with cloud service so as to share its workload and shorten application response times. A virtual cloud core is attached to OpenCL context and can unify the computation between mobile devices and cloud services. The framework does not blindly off-load computation but take into account network capacity and load on connected servers so as to effectively share the load. It also allows a computation request to be conducted either on CPU on mobile device or GPU on connected servers. Our experiments show that the response time can be improved for up to 25 times with modern wireless network connection.
C. S. Shih合作论文数Department of Computer Science and Information Engineering, National Taiwan University3
Shao-Yi Chien (簡韶逸)合作论文数Department of Electrical Engineering, National Taiwan University1