Server replication is a common fault-tolerance strategy to improve transaction dependability for services in communications networks. In distributed architectures, fault-diagnosis and recovery are implemented via the interaction of the server replicas with the clients and other entities such as enhanced name servers. Such architectures provide an increased number of redundancy configuration choices. The influence of a (wide area) network connection can be quite significant and induce trade-offs between dependability and user-perceived performance. This paper develops a quantitative stochastic model using stochastic activity networks (SAN) for the evaluation of performance and dependability metrics of a generic transaction-based service implemented on a distributed replication architecture. The composite SAN model can be easily adapted to a wide range of client-server applications deployed in replicated server architectures. In order to obtain insight into the system behaviour, a set of relevant environment parameters and controllable fault-tolerance parameters are chosen and the dependability/performance trade-off is evaluated.
Multipoint-to-multipoint message broadcast is a demanding application scenario in ad-hoc networks. Adaptive management of wireless resources is necessary to support such applications in a safety critical context. In this work we study adaptation of transmission rate and power to varying densities of ad-hoc nodes. Our approach is to construct a cross-layer model building on existing models for physical and link layers. To enable optimization in relation to metrics of end-to-end delay and message reception probability a model of flooding broadcast is proposed as a part of the cross-layer model. In a simulation study we show that adaptation of transmission power and rate can be necessary to achieve delay requirements and maximize message reception probability. Compared to simulation our cross-layer model based optimization approach generates slightly more conservative parameter settings. It is further shown how correlated losses have a significant impact on the robustness of the broadcast technique.
Fault diagnosis is vital to initiate correct recovery actions in order to provide reliable end-to-end services in unreliable networking environments. In this paper we investigate end-node driven fault diagnosis assuming that no support functions in the network exist.Fault diagnosis in networking systems spanning wired and wireless links is complicated as faults are hidden in the network and observations are unreliable. To overcome this, we show how Bayesian Networks (BNs) can be applied for probabilistic fault diagnosis.We model a TCP end-to-end connection to estimate the state of the network and diagnose faults in the wireless and wired domain respectively. Estimation accuracy and diagnosis reactivity performance is evaluated from simulations and compared to a simple threshold based approach. We show how multiple unreliable cross-layer observations improve BN diagnosis performance and robustness to changes in the environment. Furthermore, we evaluate the BN approach and methods to extract features from network traffic to suggest improvements.
Organisation name of lead contractor for this deliverable FCUL Dissemination Level PU Public X PP Restricted to other programme participants (including the Commission Services) RE Restricted to a group specified by the consortium (including the Commission Services) CO Confidential, only for members of the consortium (including the Commission Services) Abstract: The main objectives of WP2 are to define a resilient architecture and to develop a range of middleware solutions (i.e. algorithms, protocols, services) to address resilience requirements in the design of highly available, reliable and trustworthy distributed solutions. This deliverable presents research results concerning the development of middleware services within the HIDENETS environment, the objective of which is to facilitate the construction of resilient, dependable car2car applications operating in ad-hoc environments in cooperation with infrastructure based services. The deliverable complements the work presented in deliverable D2.3 (Service level resilience solutions for the ad-hoc domain), focusing on ideas that typically reflect improvements of the services presented in D2.3. These improvements can be achieved due to the possibility of operating over infrastructure environments, in addition to the exclusive operation over ad-hoc environments. The deliverable provides three contributions. Firstly, it addresses the problem of fault detection from the perspective of end-to-end services, which is addressed under a probabilistic scope and which is suitable for IP based communication systems in general. Then, it provides a study on the dependability/performance trade-off for replicated servers in the infrastructure domain. Finally, an extension to the Intrusion Tolerant Agreement service is introduced, which exploits the assumed availability of a reliable server in the infrastructure domain to improve the performance of the basic (non-extended) service, as presented in D2.3. Added new input from BME. Made changes according to comments from the internal review team. Made changes according to comments from the external review team. 1 Executive Summary Objectives of the deliverable As stated in the HIDENETS Technical Annex, the main objectives of WP2 are " to define a resilient architecture and to develop a range of middleware solutions (i.e. algorithms, protocols, services) for resilience to be applied in the design of highly available, reliable, and trustworthy networking solutions ". The main objective of this deliverable is to complement the work presented in deliverable D2.3 (Service level resilience solutions for the ad hoc domain) [11], focusing on ideas that typically correspond to improvements, or would allow improving some of the services presented in D2.3. These ideas either exploit the possibility …
High-Availability as provided by fault-tolerance mechanisms comes at the price of increased overhead due to additional processing and communication, which may be a limiting factor to service performance as perceived by the clients. In order to quantify this impact and to understand the underlying mechanisms for performance degradation, this paper presents an approach for the analysis of client-centric performance metrics in cluster-based service deployment scenarios using High-Availability Middleware. The approach is based on a combination of measurement based empiric analysis under synthetically generated load patterns and simple queueing models, that allow for the extrapolation of empiric results and are used to gain insights into the underlying causes of the empiric performance behavior. The empiric and numerical results in the paper are based on an abstracted SIP-like call control service as deployed in future version of IP-based cellular networks, running on a two-node cluster system.
With the introduction of terminals that support multiple access technologies, macro-mobility and roaming are becoming even more important functionalities: it is expected that mobile terminals will handover between heterogeneous networks in order to extend coverage or to get connectivity to faster and/or cheaper access technologies. Access and session control must be guaranteed when roaming across different administrative domains. The 3GPP IP multimedia subsystem (IMS) is the first platform standardized towards network-independent access and session control. However, the current IMS specifications do not allow for any change of the user's IP address during a session, which prevents seamless mid-session macro-handover. In this paper, we present and compare a set of solutions, implemented respectively at the session and network layers, for shortening the handover delays and/or providing session continuity in case of macro-handover in IMS-based networks. Since one of the main functions of IMS is quality of service (QoS) negotiation and the control of legacy interfaces to the access networks for QoS provisioning, special attention is given to the discussions of the QoS re-negotiation procedures during such handover cases.
This document contains an update of the HIDENETS Reference Model, whose preliminary version was introduced in D1.1. The Reference Model contains the overall approach to development and assessment of end-to-end resilience solutions. As such, it presents a framework, which due to its abstraction level is not only restricted to the HIDENETS car-to-car and car-to-infrastructure applications and use-cases. Starting from a condensed summary of the used dependability terminology, the network architecture containing the ad hoc and infrastructure domain and the definition of the main networking elements together with the software architecture of the mobile nodes is presented. The concept of architectural hybridization and its inclusion in HIDENETS-like dependability solutions is described subsequently. A set of communication and middleware level services following the architecture hybridization concept and motivated by the dependability and resilience challenges raised by HIDENETS-like scenarios is then described. Besides architecture solutions, the reference model addresses the assessment of dependability solutions in HIDENETS-like scenarios using quantitative evaluations, realized by a combination of top-down and bottom-up modeling, as well as verification via test scenarios. In order to allow for fault prevention in the software development phase of HIDENETS-like applications, generic UML-based modeling approaches with focus on dependability related aspects are described. The HIDENETS reference model provides the framework in which the detailed solution in the HIDENETS project are being developed, while at the same time facilitating the same task for non-vehicular scenarios and applications
Reliable service provisioning in car-to-car networks is challenging because the environment is very dynamic and network topologies are changing rapidly, hence making communication unreliable. For service-level fault-tolerance, the service needs to be replicated onto several vehicles. For state-full services with dynamically changing state, a careful choice of the replica servers is necessary due to the dynamically changing properties of the communication paths between them. This paper proposes and analyzes a heuristic metric called geo-cost that aids the selection of replica candidates based on information about speed and direction of the cars. The analysis in simulation experiments shows that the proposed heuristic is performing equally well when compared to an existing approach based on a snapshot measurement of network delays. The difference between the existing metric and geo-cost is the feasibility of geo-cost compared to network delays that are often hard to determine.
The Session Initiation Protocol has been chosen for controlling multimedia sessions in the IMS part of UMTS infrastructures. In such networks, availability is crucial and the integration of SIP with a fault-tolerant solution, often based on a replication technique, has become necessary. Because the replicated stateful servers are deployed in distributed networks, state inconsistency may be introduced. Mechanisms have been proposed, which aim at keeping the inconsistency level below a certain threshold by introducing an adaptive delay before the states are committed. The effectiveness of those adaptive mechanisms depends on the accuracy of the inconsistency evaluation during the system operation. In this context, the careful definition of a practically measurable inconsistency metric is necessary in order to benefit from those mechanisms while minimizing their impacting on performance. This paper discusses the relevance of different inconsistency definitions and suggests a common model in which the inconsistency metrics are broken down into a set of measurable and/or analytically derivable contributing factors. We analyze the validity of this evaluation approach with results obtained in a prototype implementation of a 3GPP IMS call control system integrated in a distributed fault-tolerant architecture, so-called RSerPool, for the example of instant message sessions between users.
Third generation mobile networks are offering the user access to Internet services. The Session Initiation Protocol (SIP) is being deployed for establishing, modifying, and terminating those multimedia sessions. Mobile operators put high requirements on their infrastructure, in particular – and in focus of this paper – on availability and reliability of call control. One approach to minimize the impact of server failures is to implement redundant servers and to replicate the state between them in a timely manner. In this paper, we study two concepts for such a fault-tolerant architecture in their application to the highly relevant use-case of SIP call control. The first approach is implemented as a distributed set of servers gathered in a so-called pool, with fail-over functionality assigned to the pool access protocols. The second is a cluster-based solution that normally implies that the servers are confined in a local network, but on the other hand the latter solution is completely transparent to clients accessing the service deployed in the cluster. To evaluate these two approaches, both were implemented in an experimental testbed mimicking SIP call control scenarios in 3rd generation mobile networks. An approach for measurement of various dependability and performance parameters in this experimental setting is developed and concluded with a set of preliminary results.
Mohamed Kaâniche合作论文数Dependable Computing and Fault Tolerance research group;LAAS-CNRS2