
A key feature in fault injection (FI) based validation is identifying the relevant test cases to inject. This problem is exacerbated at the protocol level where the lack of detailed fault distributions limits the use of statistical approaches in deriving and estimating the number of test cases to inject. In this paper we develop and demonstrate the capabilities of a formal approach to protocol validation, where the deductive and computational analysis capabilities of formal methods are shown to be able to identify very specific test cases, and analytically identify equivalence classes of test cases.
Network media redundancy is a clean and effective way of achieving high levels of reliability against temporary medium faults and availability in the presence of permanent faults. This is specially true of critical control applications such as those supported by the Controller Area Network (CAN). In our endeavor to provide CAN with media redundancy we ended-up devising a scheme which is extraordinarily simpler than previous approaches known for CAN or other LANs and field-buses.
We present Galileo, a dynamic fault tree modeling and analysis tool that combines the innovative DIF-Tree analysis methodology with a rich user interface built using package-oriented programming. DIFTree integrates binary decision diagram and Markov methods under the common notation of dynamic fault trees, allowing the user to exploit the benefits of both techniques while avoiding the need to learn additional notations and methodologies. Package-oriented programming (POP) is a software architectural style in which large-scale software packages are used as components, exploiting their rich functionality and familiarity to users. Galileo can be obtained for free under license for evaluation, and can be downloaded from the World-Wide Web.
Database management systems (DBMS) achieve high availability and fault tolerance usually by replication. However fault tolerance does not come for free. Therefore, DBMSs serving critical applications with real time requirements must find a trade of between fault tolerance cost and performance. The purpose of this study is two-fold. It evaluates the effectiveness of DBMS fault tolerance in the presence of corruption in database buffer cache, which poses serious threat to the integrity requirement of the DBMSs. The first experiment of this study evaluates the effectiveness of fault tolerance, and the fault impact on database integrity, performance, and availability on a replicated DBMS, ClustRa, in the presence of software faults that corrupt the volatile data buffer cache. The second experiment identifies the weak data structure components in the data buffer cache that give fatal consequences when corrupted, and suggest the need for some forms of guarding them individually or collectively.
This paper analyzes the effect of dormant faults on the mean time to failure (MTTF) of highly reliable systems. The analysis is performed by means of Markov models that allow quantifying the effect of dormant faults and other vital reliability parameters. It turns out that the presence of dormant faults can drastically reduce the MTTF of a system, particularly if the operating system allows a sporadic ("event-driven") change from a regular mode of operation to another mode. Virtually every practical system involves such a change, at least in case of emergency. It is demonstrated that on-line built-in self-test (BIST) is an effective means to overcome the deteriorating effect of dormant faults and re-establish a high MTTF. A very moderate test period may already be sufficient. The analysis Is performed for the example of a fail-silent communication system for safety-critical real-time applications.