
A key feature in fault injection (FI) based validation is identifying the relevant test cases to inject. This problem is exacerbated at the protocol level where the lack of detailed fault distributions limits the use of statistical approaches in deriving and estimating the number of test cases to inject. In this paper we develop and demonstrate the capabilities of a formal approach to protocol validation, where the deductive and computational analysis capabilities of formal methods are shown to be able to identify very specific test cases, and analytically identify equivalence classes of test cases.
Network media redundancy is a clean and effective way of achieving high levels of reliability against temporary medium faults and availability in the presence of permanent faults. This is specially true of critical control applications such as those supported by the Controller Area Network (CAN). In our endeavor to provide CAN with media redundancy we ended-up devising a scheme which is extraordinarily simpler than previous approaches known for CAN or other LANs and field-buses.
We present Galileo, a dynamic fault tree modeling and analysis tool that combines the innovative DIF-Tree analysis methodology with a rich user interface built using package-oriented programming. DIFTree integrates binary decision diagram and Markov methods under the common notation of dynamic fault trees, allowing the user to exploit the benefits of both techniques while avoiding the need to learn additional notations and methodologies. Package-oriented programming (POP) is a software architectural style in which large-scale software packages are used as components, exploiting their rich functionality and familiarity to users. Galileo can be obtained for free under license for evaluation, and can be downloaded from the World-Wide Web.
Database management systems (DBMS) achieve high availability and fault tolerance usually by replication. However fault tolerance does not come for free. Therefore, DBMSs serving critical applications with real time requirements must find a trade of between fault tolerance cost and performance. The purpose of this study is two-fold. It evaluates the effectiveness of DBMS fault tolerance in the presence of corruption in database buffer cache, which poses serious threat to the integrity requirement of the DBMSs. The first experiment of this study evaluates the effectiveness of fault tolerance, and the fault impact on database integrity, performance, and availability on a replicated DBMS, ClustRa, in the presence of software faults that corrupt the volatile data buffer cache. The second experiment identifies the weak data structure components in the data buffer cache that give fatal consequences when corrupted, and suggest the need for some forms of guarding them individually or collectively.
This paper analyzes the effect of dormant faults on the mean time to failure (MTTF) of highly reliable systems. The analysis is performed by means of Markov models that allow quantifying the effect of dormant faults and other vital reliability parameters. It turns out that the presence of dormant faults can drastically reduce the MTTF of a system, particularly if the operating system allows a sporadic ("event-driven") change from a regular mode of operation to another mode. Virtually every practical system involves such a change, at least in case of emergency. It is demonstrated that on-line built-in self-test (BIST) is an effective means to overcome the deteriorating effect of dormant faults and re-establish a high MTTF. A very moderate test period may already be sufficient. The analysis Is performed for the example of a fail-silent communication system for safety-critical real-time applications.
Several user-level checkpointing libraries that checkpoint Unix processes have been developed. However they do not support multithreaded programs. This paper describes a user-level checkpointing library to checkpoint multithreaded programs that use the POSIX threads library provided by Solaris 2. Experiments with programs from the SPLASH-2 benchmark suite showed a 3% to 10% increase in execution time with checkpointing enabled, plus an additional overhead for saving the program's state. The checkpointing library described here is available at http://www.dcs.uky.edu//sup /spl sim//chkpt/.
In this paper, we describe an experimental study of Internet topological stability and the origins of failure in Internet protocol backbones. The stability of end-to-end Internet paths is dependent both on the underlying telecommunication switching system, as well as the higher level software and hardware components specific to the Internet's packet-switched forwarding and routing architecture. Although a number of earlier studies have examined failures in the public telecommunication system, little attention has been given to the characterization of Internet stability. We provide analysis of the stability of major paths between Internet Service Providers based on the experimental instrumentation of key portions of the Internet infrastructure. We describe unexpectedly high levels of path fluctuation and an aggregate low mean time between failures for individual Internet paths. We also provide a case study of the network failures observed in a large regional Internet backbone. We characterize the type, origin, frequency and duration of these failures.
As Windows NT workstations become more entrenched in enterprise-critical and even mission-critical applications, the dependability of the Windows 32-bit (Win32) platform is becoming critical. To date, studies on the robustness of system software have focused on Unix-based systems. This paper describes an approach to assessing the robustness for Win32 software and providing robustness wrappers for third party commercial off-the-shelf (COTS) software. The robustness of Win32 applications to failing operating system (OS) functions is assessed by using fault injection techniques at the interface between the application and the operating system. Finally, software wrappers are developed to handle OS failures gracefully in order to mitigate catastrophic application failures.
The goal of Winckp is to transparently checkpoint and recover applications on Windows NT. The definition of transparency is no modifications to applications at all, period. There is no need to get source code, or object code. It does not involve compilation, linking or generation of a different executable. We employ window message logging/replaying to recreate states that are otherwise difficult to recover by checkpointing alone. In the paper, we describe the design and implementation of Winckp, and present the challenges and limitations. The software is available for download from http://www.bell-labs.com/projects/swift.
Process replication is provided as the central mechanism for application level software fault tolerance in SwiFT and DOORS. These technologies, implemented as reusable software modules, support cold and warm schemes of passive replication. The choice of a scheme for a particular application is based on its availability and performance requirements. In this paper we analyze the performability of a server software which may potentially use these technologies. We derive closed form formulae for availability throughput and probability of loss of a job. Six scenarios of loss are modeled and for each, these expressions are derived. The formulae can be used either of time or online to determine the optimal replication scheme.
This paper presents a method to efficiently estimate the network reliability, in support of the "real time" needs of a geographically diverse set of users. The network subsystem considered here is a ring-structured, dual-backbone network. The measurement is usage-dependent in terms of location and time. For users at different locations and/or different times, different requirements are applied to the network, thus different network reliabilities are expected. Our method applies the hierarchical decomposition principle and takes advantages of the graph similarities. The real-time estimation for all locations, at one time step, takes O(nmk) time, where n is the number of site nodes that need to be connected, m is the number of backbone nodes for one backbone network and k is the number of locations considered.
The design of dependable systems requires that the assessment of fault tolerance solutions be carried out as early as possible. This paper addresses this issue by studying a framework aimed at analyzing the behavior of these systems in the presence of faults. The proposed approach combines functional and behavioral modeling, fault injection and simulation. The paper reports on the investigation of such an approach on four practical architectures derived from real-world industrial systems. In the experiments we used two design tools, namely: Statemate(TM) and RDD-100(TM). Finally, main results and lessons learnt are discussed, and subsequent directions to further exploit these results are proposed.
The development of two programmable memory BIST architectures is first reported. A memory synthesis framework which can automatically generate, verify and insert programmable as well as non-programmable BIST units is developed as a vehicle to efficiently integrate BIST architectures in today's memory-intensive systems. Custom memory test algorithms could be loaded in the developed programmable BIST unit and therefore any type of memory test algorithm could be realized. The flexibility and efficiency of the framework are demonstrated by showing that these memory BIST units could be generated, functionally verified and inserted in a short time.
A complete source and channel coding system is protected from both channel errors and errors emanating from internal hardware failures by introducing redundancy in the source encoding and decoding procedures as well as in frequently inserting parity symbols generated by a burst-detecting convolutional code. The combined protected system can detect errors in any significant subsystems whether from transmission errors or hardware faults. The arithmetic source coding procedures are augmented with a few checking operations which detect failure effects at each iteration of the algorithm. The normal input symbol sequence has a parity symbol inserted sparsely; every n/sup th/ symbol is determined by a high-rate burst-detecting convolutional code. These parity values provide end to end error detection for channel and hardware-based errors. The favorable probability of detection performance is evaluated and the low overhead costs as measured by increased compression length of the modified source coding procedure are determined.
We present a synthesis for testability approach to obtain EXOR-based circuits with inherently small delay. The starting point of our approach is a functional specification given in form of a so-called Kronecker Functional Decision Diagram (KFDD). The KFDD is transformed into a circuit by using a composition method based on Boolean matrix multiplication. Efficient algorithms working on the KFDD are applied during synthesis to avoid the creation of constant lines. Thereby first stuck-at fault testability is guaranteed by construction. Moreover tests for all faults can be derived efficiently from the graph of the KFDD. Thus, it is not necessary to apply automatic test pattern generation (ATPG) to compute test sets for the synthesized circuits or to check for redundancies. Area and delay of the circuits can be further improved by merging of equivalent gates. Altogether our approach makes it possible to combine high speed with fill testability for circuits derived from KFDDs. Finally, the efficiency of the proposed methods is demonstrated by experiments.
This paper describes the latest version of the software package DSPNexpress, a tool for modeling with deterministic and stochastic Petri nets (DSPNs). Novel innovative features of DSPNexpress 2.000 constitute an efficient numerical method for transient analysis of DSPNs with and without concurrent deterministic transitions. In particular, DSPNexpress 2.000 can perform transient analysis of DSPNs without concurrent deterministic transitions in three orders of magnitude less computational effort than the previously known method. Furthermore, DSPNexpress 2.000 contains an effective numerical method for steady-state analysis of DSPNs with concurrent deterministic transitions.
Some of the issues this panel will address include (1) the effect of processor bugs on the design of highly dependable systems, (2) COTS versus radiation hard components, (3) use of COTS to build highly available enterprise computing systems, and (4) new research issues in designing COTS-based high availability network systems.
Critical system designers are turning to off-the-shelf operating system (OS) software to reduce costs and time-to-market. Unfortunately, general-purpose OSes do not always respond to exceptional conditions robustly, either accepting exceptional values without complaint, or suffering abnormal task termination. Even though direct measurement is impractical, this paper uses a multiversion comparison technique to reveal a 6% to 19% normalized rate at which exceptional parameter values cause no error report in commercial POSIX OS implementations. Additionally 168 functions across 13 OSes are compared to reveal common mode robustness failures. While the best single OS has a 12.6% robustness failure rate for system calls, 3.8% of failures are common across all 13 OSes examined. However, combining C library calls with system calls increases these rates to 29.5% for the best single OS and 17.0% for common mode failures. These results suggest that OS implementations are not completely diverse, and that C library functions are both less diverse and less robust than system calls.
Incremental messages have been designed to efficiently and flexibly manage replicated copies of critical data. We describe this new type of messages and show how applications benefit from an extended message passing interface which provides support for creating updating and recovering data copies, combining the advantages of both kernel and user-level approaches. The paper describes the interface offered to applications, its functionality and how it can be used to implement traditional fault-tolerance mechanisms such as passive replicas, checkpoints and recovery. A fault-tolerant transactional service that exploits the advantages of incremental messages is also presented. Furthermore, the paper presents details about the implementation of incremental messages in VSTa, a micro-kernel based operating system, and the first experimental results obtained.
In asynchronous systems, the sender encodes a data word with a code word from an unordered code and transmits the code word on the parallel bus lines. In this paper, a transmission time analysis for the above parallel asynchronous communication scheme is presented. It is proven that the average transmission time for a code word is a strictly increasing function of the weight of the code word and it approaches the worst transmission time possible when the weight goes to infinity. This implies that fast parallel asynchronous systems can be designed using low weight codes. This paper also analyzes the transmission time performances of the proximity detecting codes and gives some efficient low constant weight code designs.