2002•OpenGrey (Institut de l'Information Scientifique et Technique)Open access

Common Mechanisms for Supporting Fault Tolerance in DSM and Message Passing Systems

Ramamurthy Badrinath, Christine Morin, Centre National de la Recherche Scientifique (CNRS), 35 - Rennes (France). Inst . de Recherche en Informatique et Systemes Aleatoires (IRISA), Rennes-1 Univ., 35 (France). Inst. de Recherche en Informatique et Systemes Aleatoires (IRISA), Institut National des Sciences Appliquees de Rennes (INSA), 35 (France). Inst . de Recherche en Informatique et Systemes Aleatoires (IRISA), Institut National de Recherche en Informatique et en Automatique (INRIA), 35 - Rennes (France). Inst. de Recherche en Informatique et Systemes Aleatoires (IRISA)

Open full text 7 citations

Abstract

Backward error recovery involving checkpointing and restart of tasks is an important component of any system providing fault tolerance to applicati- ons distributed over a network. A central problem to checkpointing and recovery is the ability to track dependencies and arrive at a consistent global checkpoint. Traditionally literature treats one of either distributed shared memory (DSM) or message passing as the interprocess communication mechanism when considering the issue of fault tolerance. This paper describes preliminary investigation into common mechanisms that can be implemented to support a wide variety of protocols in both shared memory and message passing systems. In effect it can be used in a system that combines both these IPC mechanisms.

About this research paper

What this paper is about

Backward error recovery involving checkpointing and restart of tasks is an important component of any system providing fault tolerance to applicati- ons distributed over a network. A central problem to checkpointing and recovery is the ability to track dependencies and arrive at a consistent global checkpoint. Traditionally literature treats one of either distributed shared memory (DSM) or message passing as the interprocess communication mechanism when considering the issue of fault tolerance. This paper describes preliminary investigation into common mechanisms that can be implemented to support a wide variety of protocols in both shared memory and message passing systems. In effect it can be used in a system that combines both these IPC mechanisms.

Why it matters

OpenAlex reports 7 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Backward error recovery involving checkpointing and restart of tasks is an important component of any system providing fault tolerance to applicati- ons distributed over a network. A central problem to checkpointing and recovery is the ability to track dependencies and arrive at a consistent global checkpoint. Traditionally literature treats one of either distributed shared memory (DSM) or message passing as the interprocess communication mechanism when considering the issue of fault tolerance. This paper describes preliminary investigation into common mechanisms that can be implemented to support a wide variety of protocols in both shared memory and message passing systems. In effect it can be used in a system that combines both these IPC mechanisms.

Key concepts: Computer science, Fault tolerance, Message passing, Inter-process communication, Distributed computing, Shared memory, Component (thermodynamics), Variety (cybernetics)

Related papers

Back to paper searchBrowse research topicsOriginal source
Common Mechanisms for Supporting Fault Tolerance in DSM and Message Passing Systems — Research Paper | ScholarLens