This report describes an experiment in the design of a general purpose fault tolerant system, FTM. The main objective of the FTM design was to implement a “low-cost” fault tolerant system that could be used on standard workstations. At the operating system level, our goal was to provide a methodology for the design of modular reliable operating systems, while offering fault tolerance transparency to user applications. In other words, porting an application to FTM had only to require compiling the source code without having to modify it. These objectives were achieved using the Mach micro-kernel and a modular set of reliable servers which implement application checkpoints and provide continuous system functions despite machine crashes. At the architectural level, our approach relies on a high performance stable storage implementation, called Stable Transactional Memory (STM), which can be implemented either by hardware or software. We first motivate our design choices, then we detail the FTM implementation at both architectural and operating system level. We comment on the reasons for the evolution of our stable memory technology from hardware to software. Finally, we present a performance evaluation of the FTM prototype. We conclude with lessons learned and give some assessments. Key-words: Fault Tolerance, Blocking Consistent Checkpointing, Stable Memory, Modular Operating System, Micro-kernel. Leçons du projet FTM : une expérimentation dans la conception d’un système tolérant les fautes de faible coût Résumé :Ce document présente une expérimentation dans la conception d’un système tolérant les fautes à vocation générale, le FTM. Notre motivation principale était la conception d’un système de faible coût pouvant être utilisé sur des stations de travail standard. En ce qui concerne le système d’exploitation, notre objectif était de développer une méthodologie de conception de systèmes d’exploitation fiables offrant la transparence de la tolérance aux fautes aux applications utilisateurs. Autrement dit, le portage d’une application sur FTM ne doit nécessiter que la compilation du logiciel source sans avoir à modifier ce dernier. Nos objectifs ont été atteints en utilisant le micro-noyau Mach et un ensemble modulaire de serveurs fiables qui implément les points de reprises des applications et offrent un service système continu, malgré la défaillance d’une machine. Au niveau de l’architecture, notre approche a reposé sur la conception d’une mémoire stable rapide pouvant être mise en œuvre soit par matériel, soit par logiciel. Nous décrivons tout d’abord nos choix de conception, puis nous présentons la mise en œuvre du FTM en ce qui concerne l’architecture et le système d’exploitation. En particulier, nous décrivons l’évolution de la technologie mémoire stable depuis sa mise en œuvre par matériel jusqu’à son implémentation par logiciel. Enfin, nous présentons une évaluation des performances du prototype qui a été réalisé au cours de cette étude. Nous concluons en tirant les leçons de ce projet. Mots-clé : tolérance aux fautes, points de reprise cohérents, mémoire stable, système d’exploitation modulaire, micro-noyau. Lessons from FTM: an Experiment in the Design and Implementation of a Low Cost Fault
This paper describes an experiment in the design of a general purpose fault tolerant system, FTM. The main objective of the FTM design was to implement a low-cost fault-tolerant system that could be used on standard workstations, At the operating system level, our goal was to offer fault-tolerance transparency to user applications, In other words, porting an application to FTM need only require compiling the source code without having to modify it, These objectives were achieved using the Mach micro-kernel and a modular set of reliable servers which implement application checkpoints and provide continuous system functions despite machine crashes. At the architectural level, our approach relies on a high-performance stable storage implementation, called Stable Transactional Memory (STM), which can be implemented either by hardware or software, We first motivate our design choices, then we detail the FTM implementation at both architectural and operating system level. We discuss the reasons for the evolution of our stable memory technology from hardware to software; We evaluate the performance of the FTM prototype, We conclude with lessons learned and give some assessments.
The purpose of this note is to propose a model for building fault tolerant systems. We present an approach based on the object paradigm. To ensure system consistency in the event of failure we provide two basic mechamisms, the persistent state of an object and the recovery unit. Methods are called within recovery units. To define global checkpoints, we build dynamic distributed atomic actions from dependencies between recovery units. One implementation of this model is then briefly presented.
The main aspects of the FTM (fault tolerant machine) architecture, which has been built by combining stable transactional memory boards with processors of a standard machine, are reviewed, and the design principles are presented. The FTM design is based on GOTHIC, a fault-tolerant distributed system that relies on stable storage technology. A fast stable transactional memory (STM) board, which offers built-in atomic operations on groups of small data structures with very good response time, has been integrated into a multiprocessor architecture, each processor possessing its own STM. The FTM hardware architecture has been built from standard open machine using dynamic redundancy in the building of the processing elements. The FTM prototype is presented, and the STM functions are described in detail.<>
The purpose of the Fault Tolerant Multiprocessor project (FTM) is to design a fault tolerant machine based on a Stable Transactional Memory (STM). The STM allows manipulation of stable objects within atomic transactions. The building of the operating system has led us to define a C++ interface to the STM which provides stable classes. Robust object oriented programs which resist processor failures can be written in C++ using stable objects and transactions.
article Free Access Share on Stable transactional memories and fault tolerant architectures Authors: M. Banâtre IRISA Campus de Beaulieu, 35042 Rennes cedex (France) IRISA Campus de Beaulieu, 35042 Rennes cedex (France)View Profile , Ph. Joubert IRISA Campus de Beaulieu, 35042 Rennes cedex (France) IRISA Campus de Beaulieu, 35042 Rennes cedex (France)View Profile , Ch. Morin IRISA Campus de Beaulieu, 35042 Rennes cedex (France) IRISA Campus de Beaulieu, 35042 Rennes cedex (France)View Profile , G. Muller IRISA Campus de Beaulieu, 35042 Rennes cedex (France) IRISA Campus de Beaulieu, 35042 Rennes cedex (France)View Profile , B. Rochat IRISA Campus de Beaulieu, 35042 Rennes cedex (France) IRISA Campus de Beaulieu, 35042 Rennes cedex (France)View Profile , P. Sanchez IRISA Campus de Beaulieu, 35042 Rennes cedex (France) IRISA Campus de Beaulieu, 35042 Rennes cedex (France)View Profile Authors Info & Claims ACM SIGOPS Operating Systems ReviewVolume 25Issue 1Jan. 1991pp 68–72https://doi.org/10.1145/122140.122148Published:02 January 1991Publication History 0citation164DownloadsMetricsTotal Citations0Total Downloads164Last 12 Months4Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF