
The energy costs of Phase Change Memory (PCM) depends almost completely on the number of bits written per time unit. By using an encoding, we can reduce the number of bit flips when overwriting low-entropy data with low-entropy data. This is achieved by using a frequency table for bytes in classes of data to select the encoding. Using various corpora of mainly HTML files, we show that we can reduce the number of bit flips by about 0.5 bit flips per byte.
Any generic deep machine learning algorithm is essentially a function fitting exercise, where the network tunes its weights and parameters to learn discriminatory features by minimizing some cost function. Though the network tries to learn the optimal feature space, it seldom tries to learn an optimal distance metric in the cost function, and hence misses out on an additional layer of abstraction. We present a simple effective way of achieving this by learning a generic Mahalanabis distance in a collaborative loss function in an end-to-end fashion with any standard convolutional network as the feature learner. The proposed method DML-CRC gives state-of-the-art performance on benchmark fine-grained classification datasets CUB Birds, Oxford Flowers and Oxford-IIIT Pets using the VGG-19 deep network. The method is network agnostic and can be used for other similar classification tasks.
Complex numbers play a vital role in the implementation of a wide number of Digital Signal Processing (DSP) algorithms. Since they are represented as two separate components (real and imaginary), the algorithms increased the number of computations. In this letter, we present an improved approach that reduces the number of computations by integrating the computation of the radix-2 Fast Fourier Transform (FFT) with the Distributed Arithmetic (DA) and Complex Binary Number System (CBNS). Further, the proposed architecture replaces multipliers with the use of adders and shifters, which considerably decrease the design area. Although more adders are needed to implement the new DA-CBNS proposed architecture, when compared to the traditional radix-2 FFT and DA or CBNS -based FFT algorithms, results show that the proposed design yields an increase in operating frequency, and a reduction in size, memory usage, and power consumption.
Was it fair that Harry was hired but not Barry? Was it fair that Pam was fired instead of Sam? How can one ensure fairness when an intelligent algorithm takes these decisions instead of a human? How can one ensure that the decisions were taken based on merit and not on protected attributes like race or sex? These are the questions that must be answered now that many decisions in real life can be made through machine learning. However, research in fairness of algorithms has focused on the counterfactual questions "what if?" or "why?", whereas in real life most subjective questions of consequence are contrastive: "why this but not that?". We introduce concepts and mathematical tools using causal inference to address contrastive fairness in algorithmic decision-making with illustrative examples.
Recent integration of the cache-partition model has ushered in new paradigms of research in the predictability and schedulability analysis for real-time systems. However, simplicity in the analysis framework has prompted existing contributions to be biased towards a static cache-partitioning scheme. The dynamic scheme has largely been untackled despite its proficiency in schedulability, flexibility, and energy-efficiency. In this letter, we address the latter problem and make initial contributions to a dynamic cache-partition schedulability analysis for multicore partitioned scheduling of real-time fixed-priority sporadic tasks. We devise a sufficient schedulability test and then refine the upper-bound by proposing techniques to reduce the pessimism in the interference caused by cache-contention.
With more and more companies providing online services through the Internet, Web applications have been targeted by hackers. Although the existing signature-based Web Application Firewall (WAF) can well defend against known attack methods against Web applications, it is vulnerable to Web Advanced Persistent Threat (APT) using a large number of unknown Web attack methods to attack online services. In an effort to combat Web-based APT, we propose an unsupervised anomaly detection algorithm, Web-APT-Detect (WAD), which implements self-translation machine through an encoder-decoder using attention mechanism. Our attention mechanisms can improve the quality of self-translation machine used to detect malicious patterns in HTTP requests. Through experiments on the CSIC 2010 dataset, the F1-Score of our algorithm reaches 0.9844, which surpasses the known unsupervised algorithm and reaches the same level as the state-of-the-art supervised algorithm.
We propose a new multiplication technique for LWE based fully homomorphic encryption schemes. Unlike previous schemes, multiplication does not increase the size of the ciphertext. As a result, there is no need for relinearization. The noise associated with it increases only linearly.
Automatic analysis of phonocardiograms (PCGs) could be a useful tool assisting medical experts in diagnosing heart's functionality. This letter presents an algorithm processing PCGs for identifying existing abnormalities. It is based on a standardized feature set free of domain knowledge, the distribution of which is suitably approximated by a universal hidden Markov model capturing its temporal structure. At the same time, an integral part of the model is an adaptation module responsible for incorporating new data as soon as it is available without requiring complete model retraining. Extensive experiments following a standardized protocol show that the proposed algorithm reaches state of the art performance under noisy conditions in a subject/patient independent manner.
Reusable Intellectual property (IP) cores are increasingly being integrated in system-on-chips (SoCs) to reduce the SoC design complexity and satisfy the time to market constraint. However, globalization of design supply chain renders the IP cores such as digital signal processer (DSP) hardware accelerators vulnerable to piracy threat. Additionally, integrated circuits (ICs)/ IPs can be fraudulently claimed by a dishonest user. This letter presents a novel biometric fingerprint based hardware security approach using high level synthesis (HLS) framework to safeguard an IC/ IP against false ownership claim and piracy. The proposed approach embeds the IP vendor's biometric fingerprint into a hardware accelerator in the form of secret security constraints. Results show that the proposed approach outperforms a recent approach in terms of enhanced security.
Simulation based scientific applications generate increasingly large amounts of data on high-performance computing (HPC) systems. To allow data to be stored and analyzed efficiently, data compression is often utilized to reduce the volume and velocity of data. However, a question often raised by domain scientists is the level of compression that can be expected so that they can make more informed decisions, balancing between accuracy and performance. In this letter, we propose a deep neural network based approach for estimating the compressibility of scientific data. To train the neural network, we build both general features as well as compressor-specific features so that the characteristics of both data and lossy compressors are captured in training. Our approach is demonstrated to outperform a prior analytical model as well as a sampling based approach in the case of a biased estimation, i.e., for SZ. However, for the unbiased estimation (i.e., ZFP), the sampling based approach yields the best accuracy, despite the high overhead involved in sampling the target dataset.
An N-point Discrete Fourier Transform (DFT) has wide application such as speech signal amplitude/phase/frequency spectrum analysis and solving complex numerical problems etc. However a N-point DFT Application Specific Processor (ASP) can be prone to several hardware threats such as reverse engineering, counterfeiting, cloning and fraudulent ownership. This letter proposes a novel design flow of se...
A hardware accelerator is a pivotal component of a system-on-chip (SoC) employed in modern electronic systems. However, the design of hardware accelerators can be infected by inserting malicious logic (hardware Trojan) through reverse engineering (RE) by an adversary. Rising threats of RE and Trojan necessitates the security of hardware accelerator based SoCs. The proposed methodology secures the hardware accelerators using a novel structural obfuscation using key-driven transformation techniques such as key-based loop unrolling, key-based partitioning, key-based redundant operation elimination and key-based tree height transformation, which makes the design unobvious (non-interpretable) to an attacker. The results of the proposed approach on DSP hardware accelerators indicated 2.3x enhancement in strength of obfuscation (at gate level) compared to a recent approach (indicating enhanced security), at nominal design cost.
The performance of a digital signal processing (DSP) system is greatly affected by the performance of its multiplication operations. Simultaneous improvement in performance metrics such as delay, power, area, and energy efficiency is difficult to achieve and is a challenge to be addressed. To this end, an efficient carry save multiplier (CSM) that employs modified square root carry select adder (MSCA) for the vector-merging addition and improved full adder (IFA) in place of conventional full adder is proposed. Among $16\;\text{x}\;16$16x16 multipliers, the critical path delay (CPD), power, area, power delay product (PDP), and area delay product (ADP) of the proposed CSM are improved by 27.74, 19.4, 46.2, 41.4, and 60.87 percent respectively in comparison with improved booth multiplier and by 46.43, 31.46, 36.9, 63.05, and 65.96 percent respectively in comparison with low PDP booth multiplier. Cadence software with gpdk 45nm standard cell library is used for the design and implementation.
The increasing societal demand for data privacy has led researchers to develop methods to preserve privacy in data analysis. However, outlier analysis, a fundamental data analytics task with critical applications in medicine, finance, and national security, has only been analyzed for a few specialized cases of data privacy. This work is the first to provide a general framework for private outlier analysis, which is a two-step process. First, we show how to identify the relevant problem-specifications and then provide a practical solution that formally meets these specifications.
A novel light weight background subtraction algorithm proposed in this letter addresses the issue of executing real time motion detection in fog computing, which is a resource constrained environment. The proposed algorithm compresses the original high dimensional video frame into low dimensional space using random projection (RP) and the motion detection is identified using the structural similarity (SSIM) index. The objective of the proposed algorithm is to select the frames with motion and discard the remaining frame in order to reduce the storage and computational cost for the cloud users. The proposed algorithm is evaluated using CDNET 2014 video data and the results are verified in terms of selected frames per video(SFPV), percentage of reduction (POR) and running time. From the experiment, it is evident that the proposed algorithm reduces nearly 67 percent of storage cost with trifling computational expense.
Recent studies have shown that deep learning models are vulnerable to specifically crafted adversarial inputs that are quasi-imperceptible to humans. In this letter, we propose a novel method to detect adversarial inputs, by augmenting the main classification network with multiple binary detectors (observer networks) which take inputs from the hidden layers of the original network (convolutional kernel outputs) and classify the input as clean or adversarial. During inference, the detectors are treated as a part of an ensemble network and the input is deemed adversarial if at least half of the detectors classify it as so. The proposed method addresses the trade-off between accuracy of classification on clean and adversarial samples, as the original classification network is not modified during the detection process. The use of multiple observer networks makes attacking the detection mechanism non-trivial even when the attacker is aware of the victim classifier. We achieve a 99.5% detection accuracy on the MNIST dataset and 97.5% on the CIFAR-10 dataset using the Fast Gradient Sign Attack in a semi-white box setup. The number of false positive detections is a mere 0.12% in the worst case scenario.
The existing computational cognitive models in cyber security risk management have major limitations in the analysis of cyber security behaviors. To address this issue, we introduce Hilbert space and quantum cognition to cyber security risk management. We compare some key axioms and definitions of classical cognition and quantum cognition. We provide examples on how some unique principles of Hilbe...
Recently, the mobile segment observed the emergence of a new class of malware known as ransomware. In 2017, more than 468,830 unique mobile ransomware samples were discovered marking a 415 percent year-over-year increase in new ransomware. This trend presents a major concern for mobile users as they increasingly rely on their devices to safeguard sensitive information. Previous solutions have relied on high level bytecode and XML-based permission files to detect malicious applications. Unfortunately, attackers are resorting to obfuscation techniques that involve repackaging apps with malicious content directly in native machine code. As such, the aforementioned methods are insufficient for detecting modern mobile ransomware. To address these concerns, this work evaluates the effectiveness of using native instructions in detecting ransomware. We characterize different machine learning models and demonstrate that opcodes in native instructions can be used for detecting mobile ransomware with near ideal accuracy. In addition, we make the observation that the number of instruction opcodes that contribute to the detection of ransomware is significantly less than the full range of supported opcodes within a contemporary instruction set. Finally, we evaluate the robustness of our approach against six different ransomware families available in a state-of-the-art Android malware dataset.
ECDSA is a frequently used signature scheme that has attracted a great deal of software and hardware optimization efforts. In particular, the NIST P-256 curve is currently used for most of the TLS communication worldwide. This paper proposes some observations that lead to additional optimizations. The ECDSA verification includes two main bottlenecks: (a) modular inversion (modulo the group order);...
Energy storage systems (ESS) are effective solutions to reduce the cost of smart grid operations due to their ability to store and supply electricity on demand. Traditionally, ESS have been used to implement a plethora of cost reduction techniques for smart grid operations such as temporal supply-demand shifting, frequency regulation, voltage regulation etc. However, the stationary nature of ESS limits their flexibility, potentially leading to low utilization and risk being a stranded asset. In this work, we explore the use of Mobile ESS in implementing such cost reduction techniques in a grid consisting of multiple microgrids. We develop an algorithmic framework to assign multiple Mobile ESS to various microgrids of the smart grid across several days. Our framework attempts to maximize the cost reduction achieved due to Mobile ESS assignment minus the routing cost. We show that our algorithm is a polynomial time $1/e-$1/e-approximation for the NP-Hard problem of optimal assignment of MESS.