Reinforcement learning (RL) algorithms can help solve continuous control tasks in robotics. Model-based RL in particular has shown promise for these applications. It can be orders of magnitude more sample-efficient than its model-free counter-part by utilizing Deep Neural Network (DNN) based dynamics models. Furthermore, model-based methods are more robust and more generalizable due to being reward-agnostic. However, the computational complexities involved in planning and control in model-based RL are much higher, causing challenges in real-time deployment on low-resource hardware at the edge. In our work, we focus on reducing the computational footprint of the dynamics models used in model-based RL. To make the algorithm more hardware-efficient, we introduce block floating point data types for DNNs with optimized block performance and different mixed-precision network configurations. The performance impact of these optimizations is assessed across benchmark continuous robotics control tasks. A total memory savings of ∼ 3.78× can be achieved over conventional FP32 networks while reaching comparable rewards in most scenarios.
A continuing rise in DNN usage in distributed and embedded use cases has demanded more efficient hardware execution in the field. Low-precision GeMMs with optimized data formats have played a key role in more memory and computationally-efficient networks. Recently trending formats are block-scaled representations stemming from tight HW-SW co-optimization, that compress network size by sharing exponents per data block. Prior work mostly focuses on deploying such block-scaled GeMM operations on domain-specific accelerators for optimum efficiency at the cost of flexibility and ease of deployment. In this work, we exploit and optimize the deployment of block-scaled GeMMs on fully-programmable in-order vector processors using ARM SVE. We define a systematic methodology for performing design space exploration to optimally match the workload specifications with processor vector-lengths, different microkernels, block sizes and shapes. We introduce efficient intrinsics-based microkernels with effective loop unrollings, and data-transfer efficient fused requantization strategies to maximize kernel performance, while also ensuring several deployment configurations. We enable generalized block-scaled kernel deployments through tunable block sizes and shapes, which helps in accommodating different accuracy-speed trade-off requirements. Utilizing 2D activation blocks instead of conventional 1D blocks, the static and dynamic BS-INT8 configurations yielded on average 3.8x and 2.9x faster speedups over FP32 models respectively, at no accuracy loss for CNN classification tasks on CIFAR10/100 datasets.
Reduced precision datatypes have become essential to the efficient training and deployment of Deep Neural Networks (DNNs). A recent development in the field has been the emergence of block-scaled datatypes: tensor representation formats derived from floating-point, that share a common exponent across multiple elements. While these formats are being broadly adopted and optimised for by DNN-specific inference accelerators, the potential benefits for training workloads on general-purpose (GP) vector processors has yet to be thoroughly explored. This work proposes a benchmarked implementation of block-scaled general matrix multiplications (GeMM) for DNN training at the edge using commercially available vector instruction sets (ARM SVE). Using this implementation, we highlight an accuracy-speed trade-off involving the shape of shared exponent blocks - vectors or squares. We exploit this result to optimize the training of fully connected networks by dynamically adapting the shared exponent block shapes during training. This strategy yields on average around 1.95x faster training with 2x lower memory footprint compared to standard IEEE 32-bit floating point (FP32), while achieving similar accuracy.
Recent years have seen a growing trend of deploying deep neural network-based applications on edge devices. Many of these applications, such as biometric identification, activity tracking, user preference learning, etc., require fine-tuning of the trained networks for user personalization. One way to prepare these models to handle new, unseen tasks, is to pretrain them on a distribution of known tasks. This observation has led to increasing research into meta-learning based few-shot learning techniques. However, basic meta-learning approaches do not account for the limited memory and computational resources during on-chip training. We propose a modified metalearning algorithm that enables quantized fine-tuning to optimally condition the models for on-chip few shot learning. The modification involves the inclusion of target hardware constraints upfront in the meta-learning process. Block floating point datatypes with low precision mantissa bits are utilized in the forward and backward passes, to allow hardware-friendly adaptation. Experiments show that our algorithm provides better initializations than conventional algorithms, more suitable for efficient quantized fine-tuning. This allows the few-shot learner to achieve better convergence, in terms of accuracy and speed. Extensive experiments are also performed to analyze the impact of initialization on quantized fine-tuning and further corroborate the benefits of our method.
We report a new optofluidic transmissive wavefront modulator optimized for gravity-neutral performance and miniaturized dimensions. The modulator is optimized for compensating gravity-induced aberrations by simulating its behavior with various design configurations. Miniaturization was achieved by a novel method for liquid filling and sealing of the modulator. Key elements of the new interface are intra-substrate channels fabricated by a selective laser-induced etching process. The modulator has a final thickness of 750 µm and can be operated in any orientation without significant loss of modulation quality.
An improved design of an electrothermally actuated two-terminal bistable microswitch is the focus of this paper. The proposed design has bimodal bistability which is obtained by using a pair of arches, a V-beam electrothermal actuator, and a novel initially retracting actuator. All these elements are monolithically integrated in a single planar releasable layer. The salient feature of the design is the usage of only a single pair of electrodes to switch between ON and OFF states, even though there are two actuators. In order to reduce the stress, the two actuators are mechanically decoupled but are electrically coupled to satisfy the two-terminal actuation. The switch design is experimentally verified by realizing on a silicon-on-insulator (SOI) wafer using a single-layer micro-fabrication technique. An actuation voltage of 11.8 V with 200-ms pulse-width, drives the switch from OFF to ON state and a 50-ms pulse of the same voltage across the same terminals, brings it back. [2018-0293]