The management of modern distributed systems is complicated by scale and dynamics. Scalable, decoupled communication establishes flexible, loosely coupled component relationships, and these relationships help meet the present demands on management. However, traditional decoupled addressing mechanisms tend to focus the addressing on only one of the parties involved in communication while, in general, a communication relationship involves a sender, communicated content, and receivers. The state of all three are simultaneously relevant to correctness of a management relationship and its communications. We introduce selective notification, a scalable, decoupled event dissemination architecture supporting simultaneous and combined addressing of senders, receivers, and events. We demonstrate its application to programming dynamic, scalable management relationships. We then discuss its implementation, and present measurements of its effective capabilities.
Dealing with damage that arises during operation of networked information systems is essential if such systems are to provide the dependability required by modem critical applications. Extensive damage can arise from environmental factors, malicious actions and so on, and in most cases it is impractical to mask the effects of such damage using typical redundancy techniques. Reconfiguration is required of both the application and the underlying computing and communications fabric. Such reconfiguration is difficult to achieve because it requires communication with a significant number of nodes both to determine the problem and to effect a repair In this demonstration we present an approach to the implementation of such reconfiguration. The approach to reactive control includes formal description of the error states, synthesis of the implementation, a novel new communications mechanism for communication between the error detection system and the application, and a system for coordinating the effects of independent actions.
Abstract : The services provided by critical infrastructure systems are essential to the operation of modem society. These systems include the financial payments system, transportation systems, military command and control systems, the electric power grid, and telecommunications systems including the Internet. Widespread failure of any of these system might result in severe financial loss or perhaps human injury. Critical infrastructure systems rely heavily on distributed information systems for operation. These information systems must therefore be dependable; that is, they must "deliver service that can justifiably be trusted." Traditional dependability alone does not provide a rich enough model to deal with the faults in large, critical distributed systems operating in hostile environments. These systems require not simply dependability but instead require survivability. Informally, survivability is when a system has "the ability to continue to provide service (possibly degraded or different) in a given environment when various events cause major damage to the system or its operating environment." One means of achieving survivability is non-local fault tolerance, where faults that affect significant portions of the network must be detected and handled in a coordinated fashion. Our approach to doing this is with a survivability control system. This control system takes network sensor events as input, uses these to detect faults, and responds with application reconfiguration. This thesis presents TEDL, the Time-based Event Detection Language, for formal specification of the reactive policy of this control system. A translator is used to synthesize an executable implementation from this specification. The results from using TEDL to describe and execute several attack and failure scenarios for a simplified financial payments system are presented.
IntroductionA significant impediment to the development of large-scale survivable systems is the inability to accuratelymonitor these systems in real-time. Traditional methods of monitoring rely on system loggingand textual messaging to relay information from the system to the administrators. While this works forsmall systems consisting of at most of a few dozen nodes, its inherently serial human interface doesnot and cannot scale to national or global proportions. To overcome this...
Critical systems are becoming increasingly larger, distributed, and important. These systems and “systems of systems” enable many essential services including banking, electrical power, military defense, and telecommunications. At these scales, management and security become very difficult problems. One approach to these aspects is the utilization of emergent properties. With the use of emergent algorithms and behavior, the system is enacted with a generic specification and then self-organizes based on local action by its nodes. Hence, less information is known about specific system configuration and state during operation.
not and cannot scale to national or global proportions. To overcome this impediment, we must create visualization and sonification techniques that will allow us to easily monitor our survivable systems in order to quickly detect and respond to malfunctions or attacks. Here we must remember that survivabil-ity is not synonymous with reliability or availability, and that any steps we can take to decrease system degradation are important [4]. The primary step towards adequate security, and therefore survivability, is prevention by properly configuring all devices and implementing a robust network security architecture. However, even the best laid plans can be foiled, whether by advanced or unseen techniques, a known buffer overflo w software fla w, or common script kiddie tools. In many systems, attackers are only noticed after they have already breached defenses and possibly caused significant damage. If a system is methodically and consistently monitored, an administrator can recognize attacks as they occur and take action to defend against them. Some of the problems with prevention outlined in Mukherjee et al. include the impossibility of building an absolutely secure system using current software development techniques and technologies, the impracticality of of discarding current open systems (such as the Internet) in favor of new secure systems, the hindrance of overly cautious prevention mechanisms to a user's productivity, the fallibility of cryptographic techniques and methods, and the possibility of abuse by legitimate users [8]. For these and other reasons, we must not only fortify our defenses to prevent security breeches, but we must