
Surveys show that more than 80% authentication systems are password based and these systems are increasingly under direct and indirect attacks. In an effort to protect the Positive Authentication System (PAS), the negative authentication concept was introduced [9]. Here, the representation space of password profile is called self-region; any element outside this self-region is defined as the non-self-region. Then anti-password detectors (clusters) are generated covering most of the non-self-region while leaving some space uncovered to reduce detector generation time and obfuscation. In this work, we investigate a Grid-based NAS approach, called G-NAS, where anti-password detectors are generated deterministically. This approach allows faster detector generation compared to previous NAS approaches. We reported some experimental results of G-NAS using different real-world password datasets. Results demonstrate the efficiency of the proposed approach and exhibited significant improvements compared to NAS approaches.. It appears to be more robust and scalable with respect to the size of password profiles and able to update of detector sets on-the-fly.
In this research, we explore the performances of two supervised learning techniques and two open-source network intrusion detection systems (NIDS) on backscatter darknet traffic. We employ Bro and Corsaro open-source systems as well as the CART Decision Tree and Naive Bayes machine learning classifiers. While designing our machine learning classifiers, we used different sizes of training/test sets and different feature sets to understand the importance of data pre-processing. Our results show that a machine learning base approach can achieve very high performance on such backscatter darknet traffic without using IP addresses and port numbers.
The Supervisory Control and Data Acquisition (SCADA) system discussed in this work manages a distributed control network for the Tunisian Electric & Gas Utility. The network is dispersed over a large geographic area that monitors and controls the flow of electricity/gas from both remote and centralized locations. The availability of the SCADA system in this context is critical to ensuring the uninterrupted delivery of energy, including safety, security, continuity of operations and revenue. Such SCADA systems are the backbone of national critical cyber-physical infrastructures. Herein, we propose adapting the Mean Failure Cost (MFC) metric for quantifying the cost of unavailability. This new metric combines the classic availability formulation with MFC. The resulting metric, so-called Econometric Availability (EA), offers a computational basis to evaluate a system in terms of the gain/loss ($/hour of operation) that affects each stakeholder due to unavailability.
Incremental Adaptive Corrective Learning is a method for testing ad-hoc wireless networks for vulnerabilities that adversaries can exploit. It is based on an evolutionary search for tests that define behaviors for adversary-controlled network nodes. The search incrementally increases the number of such nodes and first adapts each new node to the behaviors of the already existing attackers before improving the behavior of all attackers. Tests are evaluated in simulations and behaviors are corrected to fulfill all protocol induced obligations that are not explicitly targeted for an exploit. In this paper, we substantiate the claim that this is a general method by instantiating it for different vulnerability goals and by presenting an application for cooperative collision avoidance using VANETs. In all those instantiations, the method is able to produce concrete tests that demonstrate vulnerabilities.
This research summarizes the first attempt to incorporate Q-learning algorithm in software security. The Q-learning method is embedded as part of the software itself to provide a security mechanism that has ability to learn by itself to develop a temporary repair mechanism. The results of the experiment express that given the right parameters and the right setting the Q-learning approach rapidly learns to block all malicious actions. Data analysis on the Q-values produced by the software can provide security diagnostic as well. A larger scale experiment with extended parameter testing is expected to be seen in the future work.
The volume of online transactions has increased considerably in the recent years. Consequently, the number of fraud cases has also increased, causing billion dollar losses each year worldwide. Therefore, it is mandatory to employ mechanisms that are able to assist in fraud detection. In this work, it is proposed the use of Genetic Programming (GP) to identify frauds (charge back) in electronic transactions, more specifically in online credit card operations. A case study, using a real dataset from one of the largest Latin America electronic payment systems, has been conducted in order to evaluate the proposed algorithm. The presented algorithm achieves good performance in fraud detection, obtaining gains up to 17% with regard to the actual company baseline. Moreover, several classification problems, with considerably different datasets and domains, have been used to evaluate the performance of the algorithm. The effectiveness of the algorithm has been compared with other methods, widely employed for classification. The results show that the proposed algorithm achieved good classification effectiveness in all tested instances.
Industrial control systems are often large and complex distributed systems and therefore expose a large potential attack surface. Effectively minimizing this attack surface requires security experts and significant manpower during engineering and maintenance of the system. This task, which is already difficult for today's control systems, will become significantly more complex for tomorrow's systems, which can reconfigure themselves dynamically, e.g., if hardware failures occur.In this article, we present a dynamic security system which can automatically minimize the attack surface of a control system's communication network. This security system is specifically designed for next-generation industrial control systems, but can also be applied in current generation systems. The presented security system adapts the necessary parameters of network and security controls according to the underlying changes in the control system environment. This ensures a better cyber security resilience against system compromise and reduces the attack surface because security controls will only allow data transfer that is required by the control application. Our evaluations for a next generation industrial control system and a current generation substation automation system show that the attack surface can be reduced by up to 90%, depending on the size and actual configuration of the control system.
As any veteran of the editor wars can attest, Unix users can be fiercely and irrationally attached to the commands they use and the manner in which they use them.In this work, we investigate the problem of identifying users out of a large set of candidates (25-97) through their command-line histories. Using standard algorithms and feature sets inspired by natural language authorship attribution literature, we demonstrate conclusively that individual users can be identified with a high degree of accuracy through their command-line behavior. Further, we report on the best performing feature combinations, from the many thousands that are possible, both in terms of accuracy and generality.We validate our work by experimenting on three user corpora comprising data gathered over three decades at three distinct locations. These are the Greenberg user profile corpus (168 users), Schonlau masquerading corpus (50 users) and Cal Poly command history corpus (97 users). The first two are well known corpora published in 1991 and 2001 respectively. The last is developed by the authors in a year-long study in 2014 and represents the most recent corpus of its kind. For a 50 user configuration, we find feature sets that can successfully identify users with over 90% accuracy on the Cal Poly, Greenberg and one variant of the Schonlau corpus, and over 87% on the other Schonlau variant.
Recently, many Internet users, who seek anonymity, use Tor, which is one of the most popular anonymity software solutions. Tor provides this anonymity by hiding the identity of the user from the destination that the user aims to reach. It also hides the user activities into encrypted cells. In this work, we investigate up to what level we can define what the user in Tor is doing. To this end, we extended on the previous work to classify the user activities using information extracted from Tor circuits and cells. Moreover, we developed a classification system to identify user activities based on traffic flow features. Our results show that flow based classification can reach up to the accuracy of the cell level classification as well as being more flexible.
Anomaly detection refers to identifying the patterns in data that deviate from expected behavior. These non-conforming patterns are often termed as outliers, malwares, anomalies or exceptions in different application domains. This paper presents a novel, generic real-time distributed anomaly detection framework for multi-source stream data. As a case study, we have decided to detect anomaly for multi-source VMware-based cloud data center. The framework monitors VMware performance stream data (e.g., CPU load, memory usage, etc.) continuously. It collects these data simultaneously from all the VMwares connected to the network. It notifies the resource manager to reschedule its resources dynamically when it identifies any abnormal behavior of its collected data. We have used Apache Spark, a distributed framework for processing performance stream data and making prediction without any delay. Spark is chosen over a traditional distributed framework (e.g., Hadoop and MapReduce, Mahout, etc.) that is not ideal for stream data processing. We have implemented a flat incremental clustering algorithm to model the benign characteristics in our distributed Spark based framework. We have compared the average processing latency of a tuple during clustering and prediction in Spark with Storm, another distributed framework for stream data processing. We experimentally find that Spark processes a tuple much quicker than Storm on average.
Electronic mail has become the most popular, frequently-used and powerful medium for quicker personal and business communications. However, one of the common security issues and annoying problems faced by email users and organizations is receiving a large number of unsolicited email messages, known as spam emails, every day. A traditional countermeasure in most email systems nowadays is simple filtering mechanisms that can block or quarantine unwanted emails based on some keywords defined by the user. These filters require continual effort to keep them relevant and current with some extensions proposed to improve their performance. However, due to the gigantic volumes of received emails and the continual change in spamming techniques to bypass the implemented solutions, novel automated ideas and countermeasures need to be investigated. This paper explores a novel algorithm inspired by the immune system called dendritic cell algorithm (DCA). This algorithm is evaluated on a number of benchmark datasets to detect spam emails. The results demonstrate that this approach can be a promising solution for email classification and spam filtering.
In this work, a spread spectrum watermarking optimization algorithm is explored for digital color images using biobjective genetic algorithms and full-frame discrete-cosine transform. The aim of optimization is to generate the trade-off curve, a.k.a. optimal Pareto points, of watermark imperceptibility and robustness. The watermark imperceptibility is evaluated using the Structural SIMilarity (SSIM) index between the original image and the watermarked image whereas the watermark robustness is evaluated in terms of the Normalized Correlation Coefficient (NCC) between the original watermark and the recovered watermark. The watermarked image is susceptible to various types of attacks or processing distortions such as additive Gaussian noise, pepper-and-salt noise, JPEG compression, camera motion and median filtering. For the biobjective genetic algorithm, we used the fast elitist Non-dominated Sorting Genetic Algorithm (NSGA-II). We reviewed related work and investigated two color spaces (YCbCr and HSV) in addition to gray scale images where embedding is conducted in different frames and various distortions are applied before the extraction of the watermark. The results are compared for various cases under similar conditions.
Android mobile devices have reached a widespread use since the past decade, thus leading to an increase in the number and variety of applications on the market. However, from the perspective of information security, the user control of sensitive information has been shadowed by the fast development and rich variety of the applications. In the recent state of the art, users are subject to responding numerous requests for permission about using their private data to be able run an application. The awareness of the user about data protection and its relationship to permission requests is crucial for protecting the user against malicious software. Nevertheless, the slow adaptation of users to novel technologies suggests the need for developing automatic tools for detecting malicious software. In the present study, we analyze two major aspects of permission-based malware detection in Android applications: Feature selection methods and classification algorithms. Within the framework of the assumptions specified for the analysis and the data used for the analysis, our findings reveal a higher performance for the Random Forest and J48 decision tree classification algorithms for most of the selected feature selection methods.
In recent years, a large number of discrete chaotic cryptographic algorithms have been proposed. The chaotic based cryptograms are suitable for large-scale data encryption such as images, videos or audio data. This paper propose a novel higher dimensional chaotic system for audio encryption, in which variables are treated as encryption keys in order to achieve secure transmission of audio signals. Since the highly sensitive to the initial condition of a system and to the variation of a parameter, and chaotic trajectory is so unpredictable. As a result we obtain much higher security. The higher dimensional of the algorithm is used to enhance the key space and security of the algorithm. The security analysis of the algorithm is given. The experiments show that the algorithm has the characteristic of sensitive to initial condition, high key space; pixel distribution uniformity and the algorithm will not break in chosen/known-plaintext attacks.
In this paper, hybrid wireless sensor network model is envisaged over the power distribution grid for monitoring the health of the grid. The hybrid model is hierarchical. At the lower level, it uses a cluster topology at each tower to collect local information about the tower while at the higher level it uses linear chain topology to send the grid data to the base station (usually at the substation). Data is collected at each tower, aggregated over the linear chair network, and sent across to a base station for analysis. For analysis, a machine learning based model is employed. The model is designed to detect and classify anomalies in the sensory data and it ensures the security and stability of the smart grid. Initial topology model was investigated using a pilot simulation study followed by experimentation while the analysis is carried using the real time data collected using wireless sensor networks as an overlay network on the power distribution grid. Preliminary results show that detection mechanism is promising and is able to detect the occurrence of any anomalous event that may cause threat to the smart grid.
In this paper, we explore the effect of encircling behaviour on the topology of complex networks. We introduce the concept of topological encircling, which we define as an attacker making links to neighbours of a victim with the ultimate aim of undermining that victim. We introduce metrics to quantify topological encircling in complex networks, both at the network level and node pair (link) level. Using synthesized networks, we demonstrate that our measures are able to distinguish intentional topological encircling from preferential mixing. We discuss the potential utility of our measures and future research directions.
Having an idea of a user's location when he/she is using network services has been an area of interest ever since wireless networks became very popular. As the costs of wireless technologies decrease more and more, we observe the rise of an extremely diverse market of wireless capable devices. However, the field of indoor positioning is still wide open. In this field, most of the existing technologies are dependent on additional hardware and/or infrastructure, which increases the requirements for users. In this research, we investigate the ways of coupling indoor geo-fencing with access control including authentication and registration. To achieve this, we apply a classification based geo-fencing approach using received signal strength indicator. Consequently, we are mainly focusing on associating accurate geo-fencing with secure communication and computing. Experimental results show that we have achieved considerable positioning accuracy while providing a secure way of communication. Favouring diversity, our implementation does not mandate users to undergo any system software modification or adding new hardware components.
To detect and prevent network intrusions in Cloud computing environment, we propose a novel security framework hybrid-network intrusion detection system (H-NIDS). We use different classifiers (Bayesian, Associative and Decision tree) and Snort to implement this framework. This framework aims to detect network attacks in Cloud by monitoring network traffic, while ensuring performance and service quality. We evaluate the performance and detection efficiency of H-NIDS for ensuring its feasibility in Cloud. The results show that the proposed framework has higher detection rate and low false positives at an affordable computational cost.
Designing secure software systems is a non-trivial task as data on uncommon attacks is limited, costs are difficult to estimate, and technology and tools are continually changing. Consequently, a great deal of expertise is required to assess the security risks posed to a proposed system in its design stage. In this research we demonstrate how Evolutionary Algorithms (EAs) and Simulated Annealing (SA) can be used with Ordered Weighted Average (OWA) operators to provide a suitable aggregation tool for combining experts' opinions of individual components of an specific technical attack to produce an overall rating that can be used to rank attacks in order of salience. A set of thirty nine cyber security experts took part in an exercise in which they independently assessed a realistic system scenario. We show that using EAs and SA, OWA operators can be tuned to produce aggregations that are more stable when applied to a group of experts' ratings than those produced by the arithmetic mean, and that the difference between the solutions found by each of the algorithms is minimal. However, EAs do prove to be a quicker method of search when an equivalent number of evaluations is performed by each method.
Malware detection is one of the challenging tasks in Cyber security. The advent of code obfuscation, metamorphic malware, packers and zero day attacks has made malware detection a challenging task. In this paper we present a visualization based approach for malware detection. First the executable is converted to a gray-scale image called byteplot. Later we extract low level features like intensity based and texture based features. We apply computationally intelligent techniques for malware detection using these features. In this work we used Support Vector Machines (SVMs) and obtained an accuracy of 95% on a dataset containing 25000 malware and 12000 benign samples.