The cellular neural network (CNN or CeNN) is known to be useful because of its suitability in real-time processing, parallel processing, robustness, flexibility, and energy efficiency. CeNNs have a large number of interconnected processing elements, which can be programmed to produce a wide range of patterns, including regular and irregular patterns, random patterns, and more. When implemented in memristive hardware, the pattern generator ability and inherent variability of memristive devices can be explored to create Physical Unclonable Functions (PUFs). This work reports a method of using memristive CeNNs to perform image processing tasks along with PUF image generation. The CeNN-PUF has dual mode capability combining data processing and encryption using PUF image watermarking. The proposed method provides unique device-specific image watermarks, following a two-stage process of (1) device-specific secret mask generation and (2) watermark embedding. The system is evaluated using multiple CeNN cloning templates and the robustness of the method is validated against ML attacks. A detailed analysis is presented to evaluate the uniqueness, randomness and reliability against different environmental changes. The experimental validation of the proposed model is done on FPGA Xilinx Zynq-7010 processor and benchmarked the system against quantization noise.
The errors in the memristive crossbar arrays due to device variations will impact the overall accuracy of neural networks or in-memory systems developed. For ensuring reliable use of memristive crossbar arrays, variability compensation techniques are essential to be part of the neural network design. In this paper, we present an input regulated variability compensation technique for memristive crossbar arrays. In the proposed method, the input image is split into non-overlapping blocks to be processed individually by small sized neural network blocks, which is referred to as imageSplit architecture. The memristive crossbar based Artificial Neural Network (ANN) blocks are used for building the proposed imageSplit. Circuit level analysis and integration is carried out to validate the proposed architecture. We test this approach on different datasets using various deep neural network architectures. The paper considers various device variations including $R_{OFF}/R_{ON}$ variations and aging using imageSplit. Along with hardware compensation techniques, algorithmic modifications like pruning and dropouts are also considered for analysis. The results show that splitting the input and independently training the smaller neural networks performs better in terms of output probabilistic values even with the presence of the significant amount of hardware variability.
Hyper Dimensional (HD) Computing when employed for machine learning tasks such as learning and classification involve the computation and comparison of large hypervectors within memory. The dimensionality of hypervectors in the order of thousands makes it difficult to implement HD computing in von Neumann systems. The in-memory computing capability of the memristor crossbar will speed up the vectormatrix multiplication to perform HD computing. The paper presents a memristive accelerator design for hyper-dimensional consumer text analytics. The hyper-dimensional computing offers simple encoding and data transformation techniques with higher accuracy in comparison with conventional techniques. The circuit implementation of the KNN-based HD classification using Ngrams is proposed in the paper. The effect of device variations on the performance of the proposed system is evaluated and the performance is compared with conventional classifiers. The trade-off between accuracy and dimensionality in HD computing is presented.
Counters are widely used for implementing counting, clocking, and computations. The digital counters are usually implemented with flip-flops using clocked mechanisms. Counters can consume a large area on chip, as the number of bits per count increases. Counting large numbers requires higher bits or use design methods making use of memories for scalability. Alternatively, in the paper, we present the use of memristor nodes for counting application. Each counter is implemented with a single memristor node and each count is encoded as a memristive conductive state. The conductance of the nodes in the crossbar is incremented or decremented to represent a count value. Also, the proposed method shows the possibility of implementing a larger number of counters in a single crossbar. The proposed counter design is evaluated using Voltage Threshold Adaptive Memristor (VTEAM) model. The on-chip area and power requirements of the proposed circuit is calculated and compared with counter array design using memristor crossbars.
Resistive random access memory is very well known for its potential application in in-memory and neural computing. However, they often have different types of device-to-device and cycle-to-cycle variability. This makes it harder to build highly accurate crossbar arrays. Traditional RRAM designs make use of various filament-based oxide materials for creating a channel that is sandwiched between two electrodes to form a two-terminal structure. They are often subjected to mechanical and electrical stress over repeated read-and-write cycles. The behavior of these devices often varies in practice across wafer arrays over these stresses when fabricated. The use of emerging 2D materials is explored to improve electrical endurance, long retention time, high switching speed, and fewer power losses. This study provides an in-depth exploration of neuro-memristive computing and its potential applications, focusing specifically on the utilization of graphene and 2D materials in RRAM for neural computing. The study presents a comprehensive analysis of the structural and design aspects of graphene-based RRAM, along with a thorough examination of commercially available RRAM models and their fabrication techniques. Furthermore, the study investigates the diverse range of applications that can benefit from graphene-based RRAM devices.
Intelligent sensor systems are essential for building modern Internet of Things applications. Embedding intelligence within or near sensors provides a strong case for analog neural computing. However, rapid prototyping of analog or mixed signal spiking neural computing is a non-trivial and time-consuming task. We introduce mixed-mode neural computing arrays for near-sensor-intelligent computing implemented with Field-Programmable Analog Arrays (FPAA) and Field-Programmable Gate Arrays (FPGA). The combinations of FPAA and FPGA pipelines ensure rapid prototyping and design optimization before finalizing the on-chip implementations. The proposed approach architecture ensures a scalable neural network testing framework along with sensor integration. The experimental set up of the proposed tactile sensing system in demonstrated. The initial simulations are carried out in SPICE, and the real-time implementation is validated on FPAA and FPGA hardware.
Synaptic stochasticity is an important feature of biological neural networks that is not widely explored in analog memristor networks. Synaptic Sampling Machine (SSM) is one of the recent models of the neural network that explores the importance of the synaptic stochasticity. In this paper, we present a memristive Echo State Network (ESN) with Extended-SSM (ESSM). The circuit-level design of the single synaptic sampling cell that can introduce stochasticity to the neural network is presented. The architecture of synaptic sampling cells is proposed that have the ability to adaptively reprogram the arrays and respond to stimuli of various strengths. The effect of stochasticity is achieved by randomly blocking the input with the probability that follows Bernoulli distribution, and can lead to the reduction of the memory capacity requirements. The blocking signals are randomly generated using Circular Shift Registers (CSRs). The network processing is handled in analog domain and the training is performed offline. The performance of the neural network is analyzed with a view to benchmark for hardware performance without compromising the system performance. The neural system was tested on ECG, MNIST, Fashion MNIST and CIFAR10 dataset for classification problem. The advantage of memristive CSR in comparison with conventional CMOS based CSR is presented. The ESSM-ESN performance is evaluated with the effect of device variations like resistance variations, noise and quantization. The advantage of ESSM-ESN is demonstrated in terms of performance and power requirements in comparison with other neural architectures.
Noisein test images can significantly reduce the inference accuracy of Convolution Neural Networks (CNNs). The learning alogrithm that optimise the CNN for inference often use training images that represent a limited number of variations of objects. Any variations from the training set makes it harder to recognise test images, thereby indicating lose of generalisation. We propose to use Gated Pixel Convolution Neural Network (PixelCNN) for generating training images to readjust the weights of pretrained CNN network. This way, the readjustment of weights using generated images, helps to improve the generalisation ability of the CNNs. Through this approach the CNN learning becomes continuous even with a limited training data and can be limited by the amount of generalisation and robustness to variability required for a given task. The proposed memristive loop generate architecture was designed using 22 nm CMOS technology and Knowm Meta-Stable Switch (MSS) memristor model ( $WO_{x}$ device parameters) with an on-chip area of $\text{1.54}\;\text{mm}^{2}$ and reduced power consumption.
The increasing demand for a high density of transistors in the chip with high yield often requires the manufacturing process to be mature. Defects are common when a new process technology is introduced. Manual testing for each wafer during a separate fabrication stage is a time-consuming process. To improve the production quality, the wafer map obtained after the electrical probe test should be categorized into corresponding defect classes. In this paper, we propose an automated wafer defect classification system using a Convolutional neural network (CNN)-memristor crossbar structure. The weights extracted from a pretrained neural network model are deployed into a crossbar structure. The output probability values from the softmax layer determine the class to which the wafer input belongs. The performance of CNN architecture is compared with other deep learning models. The area and power requirements of the proposed hardware architecture are also presented.
Miniaturization and energy efficiency are essential for building reliable edge AI computing devices using memristive crossbar accelerators. We propose that stochastic dropouts in inference stages are essential to improve the operational energy efficiency. Further, we show that the binarization of weights under stochastic processing improves the robustness in a crossbar based architecture. The inference dropouts can reduce the power consumption of memristive crossbars without compromising the performance accuracy for up to 10-15% of dropped neurons, saving significant energy in the crossbar. The architecture shows robustness to variability, hardware noise, aging, and drifting of memristor levels and memristor failures.
The first stage of tactile sensing is data acquisition using tactile sensors and the sensed data is transmitted to the central unit for neuromorphic computing. The memristive crossbars were proposed to use as synapses in neuromorphic computing but device intelligence at the sensor level are not investigated in literature. We propose the concept of Transistor Memristor Sensor (TMS)-crossbar by including sensor to memristor crossbar array configuration in the input layer of the neural network architecture. 2 possible cell configurations of TMS crossbar arrays: 1 Transistor 1 Memristor 1 Sensor (1T1M1S) and 2 Transistor 1 Memristor 1 Sensor (2T1M1S) are presented. We verified the proposed TMS-crossbar in the practical design of analog neural networks based Braille character recognition system. The proposed design is verified with SPICE simulations using circuit equivalent of FLX-A501 force sensor, TiO2 memristors and low power 22nm high-k CMOS transistors. The proposed analog neuromorphic computing system presents a scalable solution and is possible to encode 125 symbols with good accuracy in comparison with other Braille character recognition systems in the literature. The benefits of analog implementation of the TMS crossbar arrays is substantiated with results of accuracy, area and power requirements in comparison with the binary counterparts.
Pruning is a process of removing unwanted neurons from neural network computations. In Neural Networks, pruning creates sparse information processing, which can improve the overall generalisation and energy efficiency of the network. By excluding redundant weight values which are not contributing significantly to the system performance, hardware complexity can be reduced while maintaining the recognition accuracy. This paper evaluates the effectiveness of unstructured pruning on memristive crossbar based neural nodes in Artificial Neural Network (ANN) and Binary Weighted Neural Network (BWNN) architectures. The impact of pruning is analysed in terms of inference accuracy, energy consumption and area efficiency. The robustness of the pruned system is validated under the influence of conductance variability, bit errors, and multiply and accumulate errors.
In this paper, we present an input regulated modular neural network architecture for realising large neural networks. The proposed network become scalable in design for hardware by splitting an input image into non-overlapping blocks to be processed individually by small sized neural network blocks. Classification is done by fusing the class decisions detected from each individual block. We test this approach on three different datasets using various deep neural network architectures. The analog computation of the proposed splitting technique were evaluated using memristive-crossbar neural network architectures in SPICE tool. The results obtained show that splitting and processing an image in multiple small sized network gives higher accuracy as compared to processing the image as a whole in a larger single network. The area and power requirements of the neural network hardware architecture with the proposed splitting technique was computed and compared with the non-splitting case.
The network‐assisted device‐to‐device (D2D) communication as an underlay to cellular spectrum has attracted much attention for local area connectivity as a means to improve the cellular spectrum utilization. D2D communication as an underlay to cellular network open up various challenges including appropriate mode selection, spectral utilization, power control and efficient resource allocation. In this article, we study these issues to guarantee the quality‐of‐service requirements for the users, and a three‐step scheme is proposed. The proposed scheme first performs a mode selection procedure to choose the transmission mode of each user equipments (UEs). Then, a clustering scheme is developed to group the links that can share a common resource to improve the spectral efficiency. For the selection of suitable cellular UEs for each cluster whose resource can be shared, a cluster head selection algorithm is also developed. Finally, the expression for maximum number of links that the radio resource of shared UE can support is analytically derived. The performance of the proposed scheme is evaluated using a WINNER II A1 indoor office model. The results show that by proper management, D2D communication can effectively improve the total throughput than conventional cellular communication. From the results, it is clear that the proposed cluster head selection scheme is able to take advantage of both residual power requirement and the power requirement of UEs. Copyright © 2016 John Wiley & Sons, Ltd.
E-Learning enables the users to learn at anywhere at any time. In E-Learning systems, authenticating the E-Learning user has security issues. The usage of appropriate communication networks for providing the internet connectivity for E-learning is another challenge. WiMAX networks provide Broadband Wireless Access through the Multicast Broadcast Service so these networks can be most suitable for E-Learning applications. The authentication of E-Learning user is vulnerable to session hijacking problems. The repeated authentication of users can be done to overcome these issues. In this paper, session based Profile Caching Authentication is proposed. In this scheme, the credentials of E-Learning users can be cached at authentication server during the initial authentication through the appropriate subscriber station. The proposed cache based authentication scheme performs fast authentication by using cached user profile. Thus, the proposed authentication protocol reduces the delay in repeated authentication to enhance the security in ELearning. Keywords—Authentication, E-Learning, WiMAX, Security, Profile caching.
Device-to-device (D2D) communication has attracted much attention in the field of mobile networks for local area connectivity due to its spectral efficiency, high bit rate support and low power consumption. A group of D2D capable devices, called a cluster, can be connected through multiple links by sharing common resources. This may however result in co-channel interference between them. In this paper, we propose a novel orthogonal precoding vector selection method for reducing co-channel interference and thus maximizing the achievable data rate for each device in the cluster. The proposed method can be employed for uplink and downlink transmissions of both cellular and D2D communications. The analysis of the proposed method is carried out for the case where the cellular channel resource is being shared by single and multiple D2D links. Initially, the results via simulations are compared with the theoretical analysis and the performance of the proposed method is evaluated and compared for different resource sharing modes. The results show that our proposed method enhance the system throughput when compared with the conventional precoding vector allocation method. Finally, the paper illustrates that the introduction of cluster head in a cluster can save battery life of devices.
Direct communications between mobile terminals, also known as Device-to-Device (D2D) communications, is expected to be one of the key features supported by upcoming mobile networks. The D2D communication uses resources of the underlying mobile network which however results in mutual interference between the D2D and base station-to-terminal links. To reduce the mutual interference of these links, we propose a novel orthogonal Multiple Input Multiple Output (MIMO) precoding approach where the orthogonal precoding matrices are assigned to the D2D and base station-to-terminal communication links. The expression for outage probability of the proposed method is also presented in this paper. Numerical results show that our proposition achieves an enhancement of capacity gain in comparison with conventional precoding matrix allocation method and reduction in outage probability in comparison with conventional interference cancellation technique.
The concept of Device-to-Device (D2D) communication has been introduced for local area connectivity as a means to improve the cellular spectrum utilization and to reduce the energy consumption of user equipments. The performance of D2D communication is practically limited due to large distance and/or poor channel conditions between the D2D transmitter and receiver. To overcome these issues, a relay-assisted D2D communication has been introduced where a device relaying is an additional transmission mode along with the existing cellular and D2D transmission modes. In this paper, a transmission mode assignment algorithm based on the Hungarian algorithm is proposed to improve the overall system throughput. The proposed algorithm tries to solve two problems: a suitable transmission mode selection for each scheduled transmissions and a device selection for relaying communication between user equipments in the relay transmission mode. Simulation results show that our proposed algorithm could improve the system performance in terms of the overall system throughput and D2D data rate.
WiMAX networks are the most suitable for E-Learning through their Broadcast and Multicast Services at rural areas. Authentication of users is carried out by AAA server in WiMAX. In E-Learning systems the users must be forced to perform reauthentication to overcome the session hijacking problem. The reauthentication of users introduces frequent delay in the data access which is crucial in delaying sensitive applications such as E-Learning. In order to perform fast reauthentication caching mechanism known as Key Caching Based Authentication scheme is introduced in this paper. Even though the cache mechanism requires extra storage to keep the user credentials, this type of mechanism reduces the 50% of the delay occurring during reauthentication.