
Multilevel system modeling deals with the representation and implementation of relationships among types and among types and instances in software systems. It raises questions of interpretation and organization of types and instances in strict or interleaved layers. Reasoning over multilevel models requires a logic language that provides a uniform, extensible and flexible account for intra-type relationships, and type-instance relationships.FOML is a logic programming language that provides an expressive, executable formal basis for software models. It supports a wide variety of model-level activities, including reasoning about models, meta-modeling, and more. It is built as a semantic layer on top of PathLP, a compact logic programming language of guarded path expressions.In this paper we advocate the use of FOML as an underlying framework for the development and analysis of multilevel software modeling. We argue that FOML is suitable for multilevel modeling and for reasoning about such models. We show that FOML possesses major features needed for such tasks, including type-instance mixing in various organizational architectures, and demonstrate a sizable example using several approaches.
Honey pots are computer resources that are used to detect and deflect network attacks on a protected system. The data collected from honey pots can be utilized to better understand cyber-attacks and provide insights for improving security measures, such as intrusion detection systems. In recent years, attackers' sophistication has increased significantly, thus additional and more advanced analytical models are required. In this paper we suggest several unique methods for detecting attack propagation patterns using Markov Chains modeling and complex networks analysis. These methods can be applied on attack datasets collected from honey pots. The results of these models shed light on different attack profiles and interaction patterns between the deployed sensors in the honey pot system. We evaluate the suggested methods on a massive data set which includes over 167 million observed attacks on a globally distributed honey pot system. Analyzing the results reveals interesting patterns regarding attack correlations between the honey pots. We identify central honey pots which enable the propagation of attacks, and present how attack profiles may vary according to the attacking country. These patterns can be used to better understand existing or evolving attacks, and may aid security experts to better deploy honey pots in their system.
We propose a mathematical model for wireless industrial control systems and analyze the engineering challenge of co-scheduling networking and control in this context. We demonstrate the use of the model using a small case study and explain how our modelling approach allows for self-configuring and self-adapting systems. The main theoretical contribution of the paper is a mathematical formulation of the computational problem of synthesizing a co-design of scheduling and control for wireless industrial control systems and a mathematical proof that this problem is computationally hard. We also identify relevant special network topologies for which the problem can be efficiently solved.
Contests are one of the best ways to teach. It serves as a gamification of the learning process. In the cyber security field there are two additional unique obstacles: the first is that we don't want to teach criminal activities and the second is that we actually don't really know what the future cyber world will actually need. Both this problems are solved by asking to solve hard out-of-the-box computer programming tasks that are correlated to the current cyber security techniques.
We describe software-engineering lessons we learned by building, deploying, and operating a large-scale distributed wildlife tracking system. The design started four years ago, the system has been operational for the past two years, but kept evolving during this time. The paper describes the structure of the system and then a series of interesting and well-documented lessons we learned. Most of the lessons surprised us, in spite of some of us being fairly experienced, some are not so surprising, but we felt that they are interesting enough to document here. Some of the lessons are particularly interesting because they are specific to computer systems built by computer scientists for collecting or processing experimental science data. These issues mostly revolve around the difficulty of building and maintaining complex systems in small teams in which junior members often leave well before the project is over.
Object-Process Methodology (OPM) is a model-based systems engineering methodology which has been recognized as an ISO 19450: 2015 standard for automation systems and integration. It is domain-independent and intended for conceptual modeling of systems of various kinds. We collaborated remotely with a team that had implemented this methodology for modeling their proposed international standard, which concerned manufacturing systems lifecycle management. The goal of this collaboration was to achieve a satisfactory level of text-model coherence and model quality for the proposed standard that comply with ISO 19450. Achieving this goal proved to be an illstructured problem, without defined initial or final states and without a clear procedure for solving it. We compare three latter versions of the conceptual model using a quality assessment rubric that has been validated in previous studies. Utilizing OPM in this capacity enabled efficacious evaluation of model quality at every stage of authoring. The case described herein demonstrates how OPM can be utilized by interdisciplinary teams to tackle ill-structured problems more effectively.
Dynamic systems cause challenges for security because it is impossible to anticipate all the possible changes at design-time. Adaptive security is an applicable solution for this challenge. It is able to automatically select security mechanisms and their parameters at runtime in order to preserve the required security level in a changing environment. In this paper we describe an approach detailing key steps and activities to create an adaptive security solution. The architecture conforms to the MAPE-K reference model, which is a widely applied reference model in autonomic computing. Steps include: Establishing a control point to force control of the security decision process from the existing system. Identifying an environment parameter with which the security will trade off. Studying the interrelationship between security and the environment parameter in the existing system context by identifying measurable factors which influence their relationship. Formulation of a trade-off goal which is optimised based on multiple objectives, one of which is a security objective. Consequently, the proposed approach is capable of deploying mechanisms adaptive security to satisfy different security needs at different conditions.
Measuring and evaluating the Internet infrastructure is an important research area, with numerous proposals for sophisticated and novel ways to model traffic, infer networks' architectures and topology, or detect misconfigurations. An important requirement for such studies is the ability to reproduce the measurements in order to trace changes in the Internet over time, or to validate the reported results. Reproducibility is critical, for example, studying trends in configurations of services, adoption of patches and best practices, or fixes of misconfigurations. In this work we design automated tools for study of the architectures and configurations of the name servers in Domain Name System (DNS) and for evaluation of interoperability of unsigned domains with DNSSEC. Our measurement tools utilize wide-scale measurements of DNS servers in the forward and the reverse DNS trees, and conduct automated measurements, and provide reports with statistics. The tools are easy to use, with a convenient interface for clients and network operators. We provide access to the results collected with our tool, via a website.
Physical layer security can ensure secure communication over noisy channels in the presence of an eavesdropper with unlimited computational power. We adopt an information theoretic variant of semantic-security (SS) (a cryptographic gold standard), as our secrecy metric and study the open problem of the type II wiretap channel (WTC II) with a noisy main channel is, whose secrecy-capacity is unknown even under looser metrics than SS. Herein the secrecy-capacity is derived and shown to be equal to its SS capacity. In this setting, the legitimate users communicate via a discrete-memory less (DM) channel in the presence of an eavesdropper that has perfect access to a subset of its choosing of the transmitted symbols, constrained to a fixed fraction of the block length. The secrecy criterion is achieved simultaneously for all possible eavesdropper subset choices. On top of that, SS requires negligible mutual information between the message and the eavesdropper's observations even when maximized over all message distributions. A key tool for the achievability proof is a novel and stronger version of Wyner's soft covering lemma. Specifically, the lemma shows that a random codebook achieves the soft-covering phenomenon with high probability. The probability of failure is doubly-exponentially small in the block length. Since the combined number of messages and subsets grows only exponentially with the block length, SS for the WTC II is established by using the union bound and invoking the stronger soft-covering lemma. The direct proof shows that rates up to the weak-secrecy capacity of the classic WTC with a DM erasure channel (EC) to the eavesdropper are achievable. The converse follows by establishing the capacity of this DM wiretap EC as an upper bound for the WTC II. From a broader perspective, the stronger soft-covering lemma constitutes a tool for showing the existence of codebooks that satisfy exponentially many constraints, a beneficial ability for many other applications in information theoretic security.
The paper lists some of the quotations gathered during interviews and focus groups during a consulting engagement to help the client company improve its requirements engineering (RE) process. The paper describes in detail one of the phenomena observed, namely that of the problem of the lack of benefit of a document to its producer (PotLoBoaDtiP). It goes on to report other manifestations, including building and maintaining traces, of the PotLoBoaDtiP reported in the literature. It concludes by suggesting that the PotLoBoaDtiP cannot be solved with only technology, but that the motivational issue must be addressed.
This paper present a lightweight modeling technique that is suitable for attack description and reconstruction. It allows reconstruction of steps taken by the attacker during each stage using predefined attack ontology and traces left by the attacker. Simplicity and comprehensiveness of the proposed models makes them readable and appropriate for inclusion in incidence reports and investigation. At the same time given a predefined ontology the proposed modeling technique can be used to enhance reconstruction of attacks from forensic data.
In this paper, the author utilizes typical mistakes of third year undergraduate Computer Science and Software Engineering students in advanced software engineering courses to categorize and anticipates errors in engineering practice. In elementary courses, the students learn and apply techniques locally to relatively simple problems. In advanced courses the students are trained to select and integrate a number of techniques to solve several interdependent problems encountered in the development of complex systems. In this context, the work of Daniel Kahneman on System 1 (intuitive) and System 2 (rational) thinking is quite relevant to analysis of the patterns leading to cognitive errors. Implications for engineering practice are explored.
Mining closed frequent item sets is a key objective in the field of data mining due to its wide range of applications. Given a database of transactions, the task is to find closed subsets which appear frequently in different transactions. This subject has been studied thoroughly, and many efficient algorithms had been presented, however, most of them were designed for a non-distributed setting. The exponential growth of data in current times forces storing it in a distributed setting, meaning that most algorithms no longer apply. MapReduce is an acclaimed programming paradigm for processing large-scale, distributed data. In this paper we present an efficient algorithm for mining closed frequent item sets using the MapReduce paradigm. In addition to its novelty of running in a distributed setting, it also makes the duplication elimination step - a common step to all existing algorithms - redundant.
This paper focuses on a software framework to support face recognition, a specific area of image processing. For the processing approach, we use principal component analysis (PCA), a data dimensionality reduction approach. The goal of this study is to understand the entire face recognition process with PCA and to present a software framework supporting multiple variations, which can be used to help users create customized face recognition applications efficiently.
A highly available and robust control plane is a critical prerequisite for any Software-Defined Network (SDN) providing dependability guarantees. While there is a wide consensus that the logically centralized SDN controller should be physically distributed, today, we do not have a good understanding of how to design such a distributed and robust control plane. This is problematic, given the potentially large influence an SDN controller has on the network state compared to the distributed legacy protocols: the control plane can be an attractive target for a malicious attack. This paper initiates the study of distributed SDN control planes which are resilient to malicious controllers, for example controllers which have been compromised by a cyber attack. We introduce an adversarial control plane model and observe that approaches based on redundancy or threshold cryptography are insufficient, as incomplete or out-dated information about the network state introduces vulnerabilities. The approach presented in this paper is based on the insight that a control plane resilient to malicious behavior requires a basic notion of memory, and must be history-aware. In particular, we propose an in band approach, implemented on the SDN switch, to efficiently coordinate the different controller actions, and guarantee correct network updates even in the presence of malicious behavior. In our approach, the switch maintains a digest of the controller state and history, and only implements the update after verifying that a majority of controllers agree to the change. Our solution is not only robust but also, compared to existing consensus protocols such as Paxos, light-weight.
The Cyberspace presents enormous challenges to computer professionals. In addition to enhancing education in security, reliability, and privacy, it is important for students to have more understanding of ethics, law, policy, regulation, and responsible software development. This paper suggests the necessity of a course Law for Computer Professionals and a relevant curriculum that includes: legal aspects, ethical aspects, and professional responsibility. Student legal awareness of Privacy and Intellectual Property is very important because it affects every computer professional, and also brings one into contact with moral, philosophy and ethical issues. Curricular guidelines by the IEEE and ACM are addressed. Various study topics, for instance, on privacy and on patents, are demonstrated.
The proliferation of data often called Big Data has created problems with traditional approaches to data capture, storage, analysis and visualization, thus opening up new areas of research. Machine Learning algorithms are one area that has been used in Big Data for analysis. However, because of the challenges Big Data imposes, these algorithms need to be adapted and optimized to specific applications. One important decision made by software engineers is the choice of the language that is used in the implementation of these algorithms. This literature survey identifies and describes domain-specific languages and frameworks used for Machine Learning in Big Data with the intention of assisting software engineers in making more informed choices and providing beginners with an overview of the main languages used in this domain. This is the first survey that aims at better understanding how domain-specific languages for Machine Learning are used as a tool for research in Big Data.
Classifying human production of phonemes without additional encoding is accomplished at the level of about 77% using a version of reservoir computing. So far this has been accomplished with: (1) artificial data (2) artificial noise (designed to mimic natural noise) (3) natural human data with artificial noise (4) natural human data with its natural noise and variance albeit for certain phonemes. This mechanism, unlike most other methods is done without any encoding of the signal, and without changing time into space, but instead uses the Liquid State Machine paradigm which is an abstraction of natural cortical arrangements. The data is entered as an analogue signal without any modifications. This means that the methodology is close to "natural" biological mechanisms.
In this paper we explore how data mining can be applied to gastroenterology, and specifically to aid in the diagnosis of patients with high-risk lesions within Barrett's oesophagus (BE). BE is the only identifiable premalignant lesion for oesophageal adenocarcinoma (OA), a tumor whose incidence has been rising rapidly in the Western World.This paper makes two key contributions. First, as patient information is open to interpretation, we demonstrate that composite rules learned from multiple experts can be more accurate than that of one expert alone. Even expert doctors interpret endoscopy scans differently, potentially making it important to aggregate multiple opinions. Second, we demonstrate that decision trees can generate simple rules for dysplasia diagnosis. These rules can either be used to encapsulate the rules of the most accurate expert for training purposes or to help identify diagnostic errors.
This work deals with an Intelligent Tutoring System (ITS) for reading comprehension. Such a system could promote reading comprehension skills. An important step towards building a full ITS for reading comprehension is to build an automated ranking system that will assign a hardness level to questions used by the ITS. This is the main concern of this work. For this purpose we, first, had to define the set of criteria that determines the rate of difficulty of a question. Second, we prepared a bank of questions that were rated by a panel of experts using the set of criteria defined above. Third, we developed an automated rating software based on the criteria defined above. In particular, we considered and compared different machine learning techniques for the ranking system of the third part of the process: Artificial Neural Network (ANN), Support Vector Machine (SVM), decision tree and naïve Bayesian network. The definition of the criteria set for rating a question's difficulty, and the development of an automated software for rating a questions' difficulty, contribute to a tremendous advancement in the ITS domain for reading comprehension by providing a uniform, objective and automated system for determining a question's difficulty.