This book provides an overview of the work of two successive ESPRIT Basic Research Projects on Predictably Dependable Computing Systems (PDCS), as well as their major achievements. The purpose of the projects has been "to contribute to making the process of designing and constructing dependable computing systems much more predictable and cost-effective". The book contains a carefully edited selection of papers on all four main topics in PDCS: fault prevention, fault tolerance, fault removal, and fault forecasting. Problems of real-time and distributed systems, system structuring, qualitative evaluation, and software dependability modelling are emphasized. The book reports on the latest research on PDCS from a team including many of Europe's leading researchers.
Our society is faced with an ever increasing dependence on computing systems, which lead to question ourselves about the limits of their dependability. In order to respond this question, a global conceptual and terminological framework is needed, which is first given. The analysis of the limits in dependability which is then conducted identifies design faults as the major limiting factor, a consequence of which is the concluding recommendation of applying a fault tolerance approach to the improvement of the production process.
We say that a computer is fault-tolerant if it fulfills its intended function despite the presence or the occurrence of faults. Fault-tolerance is achieved through the introduction and the management of redundancy. A fault-tolerant computer may contain several forms of redundancy, depending on the types of faults it is designed to tolerate. For example, structural redundancy can be used to provide continued system operation even if some components have failed; information redundancy in the form of error control codes can allow the detection or correction of data errors; timing redundancy can be used to tolerate transient faults, etc.
Definition of resilience Resilience (from the Latin etymology resilire, to rebound) is literally the act or action of springing back. As a property, two strands can historically be identified: a) in social psychology [Claudel 1936], where it is about elasticity, spirit, resource and good mood, and b) and in material science, where it is about robustness and elasticity. The notion of resilience has then been elaborated: • in child psychology and psychiatry [Engle et al. 1996], referring to living and developing successfully when facing adversity; • in ecology [Holling 1973], referring to moving from a stability domain to another one under the influence of disturbances; • in business [Hamel & Välikangas 2003], referring to the capacity to reinvent a business model before circumstances force to; • in industrial safety [Hollnagel et al. 2006], referring to anticipating risk changes before damage occurrence. A common point to the above senses of the notion of resilience is the ability to successfully accommodate unforeseen environmental perturbations or disturbances. A careful examination of [Holling 1973] leads to draw interesting parallels between ecological systems and computing systems, due to: a) the emphasis on the notion of persistence of a property: resilience is said to " determine the persistence of relationships within a system and is a measure of the ability of these systems to absorb changes of state variables, driving variables, and parameters, and still persist " ; b) the dissociation between resilience and stability: it is noted that " a system can be very resilient and still fluctuate greatly, i.e., have low stability " and that " low stability seems to introduce high resilience " ; c) the mention that diversity is of significant influence on both stability (decreasing it) and resilience (increasing it). The adjective resilient has been in use for decades in the field of dependable computing systems, e.g. [Alsberg & Day 1976], and is more and more in use, however essentially as a synonym of fault-tolerant, thus generally ignoring the unexpected aspect of the phenomena the systems may have to face. A noteworthy exception is the preface of [Anderson 1985], which says " The two key attributes here are dependability and robustness. […] A computing system can be said to be robust if it retains its ability to deliver service in conditions which are beyond its normal domain of operation ". Fault-tolerant computing systems are known for exhibiting some robustness with respect …
This deliverable deals with the modelling and analysis of interdependencies between critical infrastructures, focussing attention on two interdependent infrastructures studied in the context of CRUTIAL: the electric power infrastructure and the information infrastructures supporting management, control and maintenance functionality. The main objectives are: 1) investigate the main challenges to be addressed for the analysis and modelling of interdependencies, 2) review the modelling methodologies and tools that can be used to address these challenges and support the evaluation of the impact of interdependencies on the dependability and resilience of the service delivered to the users, and 3) present the preliminary directions investigated so far by the CRUTIAL consortium for describing and modelling interdependencies. Keyword list:. critical infrastructures, power systems, interdependencies modelling, dependability and security evaluation Methodologies synthesis Page ii
operationaland abstract-denotational-semantics definitions and as well as logics and models for temporal-logic model checking. Slides for SOFTWARE TESTING P. Ammann, J. Offutt: http://cs.gmu.edu/~offutt/softwaretest/powerpoint/ This course is about concepts and techniques for testing software and assuring its quality. Topics cover software testing at the unit, module, subsystem, and system levels, automatic and manual techniques for generating and validating test data, testing process, static vs. dynamic analysis, functional testing, inspections, and reliability assessment.
In this paper we present and analyse a new set of software failure data which shows the failure behaviour, over a period of four years, of a single-user work station which was installed at the City University in March 1985. The details recorded in this data collection exercise allow us to subdivide the data into various subsets of inter-failure times. A sub-collection of these are chosen for more detailed analysis. Experience of applying reliability models in the past has shown that the relative predictive performance of the models depends entirely on the context. It has been found that there is no one model that performs well over all data sets. It has also been found that for some data sets all models applied are in error. In such cases two techniques for improving predictive accuracy have been shown to be beneficial: i) recalibrating the raw model predictions and ii) using the results of trend tests to apply the models. These two techniques may be used separately or in combination. This paper is mainly devoted to the first technique but we will also show the benefit to be gained by the application of the second technique. We apply a number of reliability models to the failure data and the recalibration technique and assess the performance of the resulting prediction systems. Introduction...... . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 1 1 Raw Reliability Models..... . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 2 Analysis of Predictive Quality and Recalibration .... . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 2.1 The u-Plot.... . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 2.2 The y-Plot.... . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 2.3 The Prequential Likelihood Ratio........................................ 4 2.4 The Recalibration Technique ..... . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 3 Preliminary Data Processing ..... . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 4 Data Collection Activity and Trend Analysis .... . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 4.1 Data Presentation ..... . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 4.2 Trend Test Analysis .... . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 14 5 Analysis of Resulting Prediction Systems...................................... 14 5.1 Data Set USBAR. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 16 5.2 Data Set TSW. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 17 5.3 Data Set TUSAB. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 5.4 General Comments ..... . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 Conclusions and Future Work ..... . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 20 References................................................................................... 21 Acknowledgments.......................................................................... 23 Figures....................................................................................... 24
The paper reports about a study conducted for RATP, the utility organisation for public transportation in Paris and region. RATP has developed since the mid eighties a mathematically formal approach for the development of safety-critical software, based on the B method. The question raised, in the context of evolutions in software development, was: Is it possible to demonstrate the same level of safety without resorting to mathematically formal approaches? In order to respond this question, several steps were considered: 1) reminding the infeasibility of quantifying safety-critical software, and its consequences on the development process, and on the system vision, 2) situating the current RATP approach with respect to other safety-critical domains, 3) examining and comparing alternate approaches for developing safety-critical software, 4) coming back to the RATP approach, for examining underlying assumptions. The conclusion was the recommendation to pursue the mathematically formal development approach.
The current state-of-knowledge and state-of-the-art reasonably enable the construction and operation of critical systems, be they safety-critical or availability-critical. The situation drastically worsens when considering large, networked, evolving, systems either fixed or mobile, with demanding requirements driven by their domain of application. There is statistical evidence that these emerging systems suffer from a significant drop in dependability and security in comparison with the former systems. The cost of failures in service is growing rapidly, as a consequence of the degree of dependence placed on computing systems, up to several million euros per hour of downtime for some businesses
Two fault-tolerant software techniques are investigated: recovery block and N-version programming. For each, the stable reliability model is transformed into a model that considers reliability growth via the transformation approach based on the hyperexponential model. Analytic and numeric processing of the transformed models identify the influence of fault removal on the reliability of the fault-tolerant software approaches. The modeling approach is based on the transformation of a Markov chain of the fault-tolerant software system in stable reliability into another, modified Markov chain which enables reliability growth to be considered. This approach has allowed reliability growth relative to the classes of faults (independent, related) affecting fault-tolerant software to be identifed and evaluated. The evaluations apply to systems of short successive mission durations with respect to the system life-time. Using generalized stochastic Petri nets to model the fault-tolerant software systems allows for an automatic application of the transformation technique. Analytic expressions are derived only to analyze explicitly the impact of fault-removal of each class. In practice, reliability measures can be directly evaluated by available tools for numerical processing of the Markov chains. Even though this work is a f is t attempt, the results are important since they show the influence of reliability growth on the reliability of fault-tolerant software systems. These results: a) confirm, from the reliability growth perspective, the importance of the faults whose occurrence can lead to common-mode failures, eg, decider faults and re2ded faults and b) enable the impact of these faults to be quantified.
This paper gives the main definitions relating to dependability, a generic concept including as special case such attributes as reliability, availability, safety, confidentiality, integrity, maintainability, etc. Basic definitions are given first. They are then commented upon, and supplemented by additional definitions, which address the threats to dependability (faults, errors, failures), and the attributes of dependability. The discussion on the attributes encompasses the relationship of dependability with security, survivability and trustworthiness.
Our society is faced with an ever increasing dependence on computing systems, which lead to question ourselves about the limits of their dependability, and about the challenges raised by those limits. In order to respond these questions, a global conceptual and terminological framework is needed, which is first given. The limits and challenges in dependability are then addressed, from technical and financial viewpoints. The recognition that design faults are the major limiting factor leads to recommending the extension of fault tolerance from products to their production process.
This paper gives the main definitions relating to dependability, a generic concept including a special case of such attributes as reliability, availability, safety, integrity, maintainability, etc. Security brings in concerns for confidentiality, in addition to availability and integrity. Basic definitions are given first. They are then commented upon, and supplemented by additional definitions, which address the threats to dependability and security (faults, errors, failures), their attributes, and the means for their achievement (fault prevention, fault tolerance, fault removal, fault forecasting). The aim is to explicate a set of general concepts, of relevance across a wide range of situations and, therefore, helping communication and cooperation among a number of scientific and technical communities, including ones that are concentrating on particular types of system, of system failures, or of causes of system failures.
Mohamed Kaâniche合作论文数Dependable Computing and Fault Tolerance research group;LAAS-CNRS15
Ian S. Welch合作论文数Victoria University;School of Mathematics;Statistics and Computer Science 5
Johan Karlsson合作论文数Chalmers University of Technology4
Bev Littlewood合作论文数Centre for Software Reliability,City University London3
Cinzia Bernardeschi合作论文数Universita' di Pisa;Dipartimento di Ingegneria dell'Informazione3