An enhanced model-based checkpointing protocol
Jiang Wu, Yi Luo, Dakshnamoorthy Manivannan
Abstract
Jiang Wu, Yi Luo, Dakshnamoorthy Manivannan
Abstract
Checkpointing and rollback recovery are widely used techniques to handle failures in distributed computing systems. Usually we avoid taking checkpoints that are useless during the recovery process. Communication-Induced checkpointing algorithms guarantee the usefulness of all the checkpoints and provide considerable autonomy with relatively low overhead. In this paper, we propose an enhanced Communication-Induced checkpointing algorithm. Our algorithm is likely to have less checkpointing overhead than an existing algorithm in the literature.
OpenAlex reports 3 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Checkpointing and rollback recovery are widely used techniques to handle failures in distributed computing systems. Usually we avoid taking checkpoints that are useless during the recovery process. Communication-Induced checkpointing algorithms guarantee the usefulness of all the checkpoints and provide considerable autonomy with relatively low overhead. In this paper, we propose an enhanced Communication-Induced checkpointing algorithm. Our algorithm is likely to have less checkpointing overhead than an existing algorithm in the literature.
Key concepts: Computer science, Overhead (engineering), Distributed computing, Process (computing), Protocol (science), Rollback, Fault tolerance, Parallel computing