
This paper presents a testing approach for validating global adaptation in multi-embedded-agent systems. Those systems are gaining increasing attention due to their high adaptability and resilience. They differ from software multi-agent systems because embedded agents have additional constraints, like energy management that software agents don’t. Those constraints and other specificities, like the tight link with the physical environment, require the use of specific methods and tools for testing these systems. The proposed approach aims at validating at run-time the adaptation of those systems when the entities composing them, the agents, are able to change their global behaviors with self-organization processes. Self-organization processes are not specific to multi-agent systems but in their case, they allow agents to change their organization, i.e. their way of interacting, at runtime. The proposed approach and tool are designed to support lifelong monitoring of multi-embedded-agent systems. In such systems, agents have self-organization behaviors resulting in complex and ever adapting systems, which are challenging to test and monitor.
In recent times, the involvement of computer systems in our lives has been drastically increasing, as has the need of improving the resilience of these systems, e.g. so they can withstand errors and changes in their environment. Techniques such as testing and simulation are often used to ensure this, but in the case of complex, real-time systems, these techniques can only provide coverage for a limited set of possible system behaviours. Software model checking and stochastic verification are alternative techniques that formally and exhaustively verify whether software meets its functional requirements and establish the performance and dependability properties of software, respectively. The two formal techniques are often used in isolation, yet software must simultaneously ensure a combination of functional and non-functional requirements. The doctoral project described in this paper aims to bring these two areas of software verification together by enabling the joint analysis of functional and non-functional properties of software systems.
Within growing pervasive information systems, Systems of Systems (SoS) emerge as a new research frontier. A SoS is formed by a set of constituent systems that live on their own with well-established functionalities and requirements, and, in certain circumstances, they must collaborate to achieve a common mission. In this scenario, security is one crucial property that needs to be considered since the early stages of SoS lifecycle. Unfortunately, SoS security cannot be guaranteed by addressing the security of each constituent system separately. The aim of this paper is to discuss the challenges faced in addressing the security of SoS and to propose some research ideas centered around the notion of a mission to be carried out by the SoS.
A failure may occur at all architectural levels of the Internet of Things (IoT) applications: sensor and actuator nodes can be missed, network links can be down, and processing and storage components can fail to perform properly. That is the reason for which fault-tolerance (FT) has become a crucial concern for IoT systems. Our study aims at identifying and classifying the existing FT mechanisms that can tolerate the IoT systems failure. In line with a systematic mapping study selection procedure, we picked out 60 papers among over 2300 candidate studies. To this end, we applied a rigorous classification and extraction framework to select and analyze the most influential domain-related information. Our analysis revealed the following main findings: (i) whilst researchers tend to study fault-tolerant IoT (FT-IoT) in cloud level only, several studies extend the application to fog and edge computing; (ii) there is a growing scientific interest on using the microservices architecture to address FT in IoT systems; (iii) the IoT components distribution, collaboration and intelligent elements location impact the system resiliency. This study gives a foundation to classify the existing and future approaches for fault-tolerant IoT, by classifying a set of methods, techniques and architectures that are potentially capable to reduce IoT systems failure.
JARVIS is a Research & Development project, jointly developed by industrial SME partners and by the University of Florence, aimed at development of a hardware/software framework supporting integration among physical IoT devices, data analytic software agents, and human operators involved in operation and maintenance of resilient Industry 4.0 systems. At the heart of the JARVIS architecture, a suite of software digital twins deployed in a Java EE environment supports runtime monitoring and control of the hierarchy of hardware configuration items of the system, capturing their composition and representing their failure modes through a reflection architectural pattern enabling agile adaptation to the evolution of configurations. Besides, analytic modules can be deployed as micro-services leveraging both the knowledge base provided by digital twins and the data flowing from the ingestion layer. This enables agile development of advanced monitoring and control services supporting maintainability and resilience. We describe the JARVIS architecture, outlining responsibilities and collaborations among its modules, and we provide details on the structure of representation of digital twins, showing how this is exploited in a data analytic agent providing an executable representation of fault trees associated with failure modes of configuration items.
Resilience is an ability of the system to deliver its services in a dependable way despite the changes. In this paper, we propose a multi-agent based formal outlook on ensuring resilience of multi-robotic systems. We represent system functions as collaborative activities performed by the agents with different capabilities. Changes invoke either structural reconfigurations – forming different collaborations or compensative activities – introducing into the system agents with additional capabilities. We formalize the resilience mechanisms and demonstrate their use by a case study – a coordination of a swarm of drones.
Software systems are increasingly autonomous in making decisions on behalf of potential users. In these systems, the power of self goes beyond the ability of substituting human agents operating on software systems and exceeds the system boundaries invading the user prerogatives. Privacy and ethical issues are at the top of the research agenda in (big) data management and AI, that offer a wide range of techniques often used as key (black-box) components of autonomous systems. In this extended abstract, I discuss these issues from the software system developer perspective that uses such black-box components and outline a new approach based on a partially synthesized software exoskeleton that empowers the user by mediating her interactions in order to preserve her privacy and ethical preferences.
With the growing interest toward pervasive systems such as the Internet of Things or Cyber-Physical Systems, embedded multi-agent systems have been increasingly investigated. In these systems, agents cooperate to achieve their local goals and a global goal that would be impossible for an isolated agent to achieve. However, the dark side of this collaboration is that agents can easily be victim of malicious attacks coming from untrustworthy agents. Consequently, trust management systems are designed to help agents choosing trustworthy counterparts to cooperate based on available information. But gathering the necessary information may be too expensive in terms of energy for small embedded agents and not relevant in all contexts. We propose a solution that allows agents to manage the energy consumption associated with information gathering. Our solution uses a Multi-Armed Bandit algorithm, which is a reinforcement learning technique to allow the agents to adapt themselves and their energy consumption to the context.
Requirement change management (RCM) for global software development (GSD), facilitated by the cloud platform, faces communication, coordination and control issues especially when there is no effective information and knowledge-sharing mechanisms. This paper describes a reasonably effective requirement change management approach for cloud-based GSD. Objective: In this regard, we contribute a Reactive Middleware which facilitates a set of guidelines defined to manage change and traceability. Methods: This Reactive Middleware provides services for user management, requirement management, change management, and traceability of cloud-based GSD projects. We present (1) a process model for change management and traceability (CM-T) for cloud-based GSD, and then (2) detail our management approach for system engineering processes as part of the presented GSD guidelines. Results: To ensure that the defined CM-T process model complies with the CMMI Level 2 (Baseline) Capability, the process model is validated using an expert panel review process where a total average, 85.58% of the experts support the maturity of the process model. Also, we demonstrate the continual tight linkage of stakeholders’ requirements and system engineering processes towards change management and traceability, with an Airlock Control System case study.
Making informed choices when designing or contracting a system is yet a very challenging task. One of the biggest users’ concern is to select the most trustworthy solution. However, it is difficult to understand the trustworthiness of a system, because it encompasses a large diversity of properties such as security, privacy, performance, among others. Composing a measure that considers such a large number of properties, the relationship among them and their relevance in the composition requires a well defined model, such as a quality model. In this experience report, we study whether quality models can provide scores that are useful to characterize those properties, helping users to choose the most trustworthy of the available alternatives. Then, we have chosen a property that is on the top of users concerns: data privacy. Results showed a higher percentage of success of linkage attacks when the privacy score is lower, indicating the usefulness of quality models in measuring and improving data privacy and providing interesting insights to the users.
In recent years, we have observed the increasing interest in the system property resilience. We ascribe this increasing interest to the rapidly growing number of deployed, complex, socio-technical systems, which are facing uncertainty about changes they are expected to experience during their life-cycle and ways to deal with them. This paper contributes to current resilience research by focusing on the different definitions given for this system property, highlighting the risk that, using different terms in different communities, this contributes to create a “tower of Babel” problem, with the consequent difficulty in exchanging ideas and working together towards a common goal. We adopt an extended definition of dependability to define resilience. Based on that, we identify features of resilient systems, capture properties falling under the resilience umbrella, and define a conceptual framework for resilience characterization including how changes affect the system, strategies to design resilience, and discuss metrics for quantifying resilience at design and runtime.
Computer systems generate large amounts of event logs related to various operational aspects (positive and negative). Extracting from them useful information (e.g. targeted at dependability and resilience issues) is a challenging problem widely discussed in the literature and still needing deeper studies. We have developed a new holistic approach using enhanced event classification (based on original text mining algorithms) combined with multidimensional statistical analysis of various properties in vocabulary (words, phrases), time, spatial, local and global correlations. It has been incorporated in the developed tools and verified on event data sets collected from different computers.
The use of robots is gaining considerable traction in several domains, since they are capable of assisting and replacing humans for everyday tasks. To harvest the full potential of robots, it must be possible to define missions for robots that are domain-specific, resilient, and collaborative. Currently, robot vendors provide low-level APIs to program such missions, making mission definition a task-specific and error-prone activity. There is a need for quick definition of new missions, by users that lack programming expertise, such as farmers and emergency workers. In this paper, we extend the existing FLYAQ platform to support the high-level specification of adaptive and highly-resilient missions. We present an extensible specification language that allows users to declaratively specify domain-specific constraints as properties of missions, thus complementing the existing FLYAQ mission language. This permits to move at runtime, the actual generation of low-level operations to satisfy the declaratively specified mission. We show how this specification language can be automatically generated from a domain-specific FLYAQ mission language by using the generative ProMoBox approach. Next, we show how mission goals are achieved taking mission properties into account, and how missions may change due to unexpected circumstances.
Hierarchical State Machines (HSMs) are widely used in the design and implementation of spacecraft flight software. However, the traditional approach to using HSMs involves graphical languages (such as UML statecharts) from which implementation code is generated (e.g. in C or C++). This is driven by the fact that state transitions in an HSM can result in execution of action code, with associated side-effects, which is implemented by code in the target implementation language. Due to this indirection, early analysis of designs becomes difficult. We propose an internal Scala DSL for writing HSMs, which makes them short, readable and easy to work with during the design phase. Writing the HSM models in Scala also allows us to use an expressive monitoring framework (also in Scala) for checking temporal properties over the HSM behaviors. We show how our approach admits writing reactive monitors that send messages to the HSM when certain sequences of events have been observed, e.g., to inject faults under certain conditions, in order to check that the system continues to operate correctly. This work is part of a larger project exploring the use of a modern high-level programming language (Scala) for modeling and verification.
The paper lists three major issues: complexity, time and uncertainty, and identifies dependability as the permanent challenge. In order to enhance dependability, the paradigm shift is proposed where focus is on failure prediction and early malware detection. Failure prediction methodology, including modeling and failure mitigation, is presented and two case studies (failure prediction for computer servers and early malware detection) are described in detail. The proposed approach, using predictive analytics, may increase system availability by an order of magnitude or so.
Open Source software is increasingly used in a wide spectrum of applications. While the benefits of the open source components are unquestionable now, there is a great concern over security assurance provided by such components. Often open source software is a subject of frequent updates. The updates might introduce or remove a diverse range of features and hence violate security properties of the previous releases. Obviously, a manual inspection of security would be prohibitively slow and inefficient. Therefore, there is a great demand for the techniques that would allow the developers to automate the process of security assurance in the presence of frequent releases. The problem of security assurance is especially challenging because to ensure scalability, such main open source initiatives, as OpenStack adopt RESTful architecture. This requires new security assurance techniques to cater to stateless nature of the system. In this paper, we propose a model-driven framework that would allow the designers to model the security concerns and facilitate verification and validation of them in an automated manner. It enables a regular monitoring of the security features even in the presence of frequent updates. We exemplify our approach with the Keystone component of OpenStack.
Cyber-Physical Systems (CPS) are software and hardware systems that interact with the physical environment. Many CPSs have useful lifetimes measured in decades. This leads to unique concerns regarding security and longevity of software designed for CPSs which are exacerbated by the need for CPSs to adapt to ecosystem changes if they are to remain functional over extended periods. In particular, the software in long-lifetime CPSs must adapt to unanticipated trends in environmental conditions, aging effects on mechanical systems, and component upgrades and modifications. This paper presents the Toolkit for Evolving Ecosystem Envelopes (TEEE) system created to help address these challenges in CPSs. TEEE is able to detect environmental changes which have caused errors within the CPS without directly sensing the environmental change. TEEE uses dynamic profiling to detect the errors within the CPS, determine the root cause of the error, alert the user, and suggest a possible adaption.
An increasing openness and interconnectedness of safety-critical industrial control systems makes them vulnerable to security attacks. Hence, we should establish the integrated approaches enabling safety-security co-engineering. Such approaches should support an analysis of interdependencies between the mechanisms required for safety and security assurance. In this paper, we demonstrate how formal modelling can facilitate reasoning about the impact of certain security solutions on safety and vise versa. We rely on modelling and refinement in Event-B to systematically uncover mutual interdependencies and the constraints that should be imposed on the system to guarantee its safety even in the presence of security attacks. The approach is illustrated by a case study – a battery charging system of an electric car.
There has been much focus on applying graphical techniques to analyse various kinds of structural errors in knowledge bases as a method of verification and reliability estimation. The most commonly applied technique has been Petri nets, or variations thereof, in achieving this objective with much success. However, although this approach has been considerably useful for verifying rules in earlier generations of knowledge-based systems, it is unclear if this approach can continue to be as useful, or indeed accessible, for verifying current or later generations of KBS, which have significantly larger, more complex, probabilistic rule sets. It has recently been argued that stochastic Petri nets can be successfully applied to continue with knowledge base verification, although, this method has required extensive and complex modifications that has led into proposals for fuzzy Petri nets. It is the view of this paper that the stochastic activity network formalism can provide a potentially useful alternative for the verification of fuzzy rule sets and can be more efficient and effective than complex derivatives of Petri nets. We present a high-level discussion of how this approach could be applied and used to analyse knowledge bases in ensuring that there are free of structural errors.