1993•Unpublished venueRequires access

Architecture for fault-tolerant program execution on large-scale parallel hypercubes

Saniya Ben Hassen

Open publisher page 0 citations

Abstract

Compared to traditional uniprocessor architectures, parallel systems provide superior processing power. However, their reliability may be significantly decreased due to the large number of processing elements and communication links. As large-scale parallel computer systems become more available, and for highly reliable applications that need to be implemented efficiently, reliability considerations require new solutions. We are interested in the implementation of fault-tolerant, efficient applications for large-scale hypercube computers. Our goal is three-fold: to achieve a low overhead during failure-free operating circumstances, to obtain graceful performance degradation as failures occur in the system, and to provide reasonably fast recovery procedures so that continuous service can be ensured. We first choose a system model and a programming methodology to write parallel programs that we think are suitable for a wide variety of applications. On top of that methodology, we implement fault tolerant programs and supply a reliable storage and transparent recovery mechanisms. The components of our system operate on some critical shared data structures. The efficient implementation of the fault-tolerant programs is achieved by allowing parallel non-blocking accesses to the shared data structures and the integration of the access primitives, the reliable storage management, and the communication procedures. Furthermore, our system is divided into clusters to limit the overhead. The analytical approximation of the response time of the system we have derived suggests that this implementation meets the performance goals for coarse to medium grain processing.

About this research paper

What this paper is about

Compared to traditional uniprocessor architectures, parallel systems provide superior processing power. However, their reliability may be significantly decreased due to the large number of processing elements and communication links. As large-scale parallel computer systems become more available, and for highly reliable applications that need to be implemented efficiently, reliability considerations require new solutions. We are interested in the implementation of fault-tolerant, efficient applications for large-scale hypercube computers. Our goal is three-fold: to achieve a low overhead during failure-free operating circumstances, to obtain graceful performance degradation as failures occur in the system, and to provide reasonably fast recovery procedures so that continuous service can be ensured. We first choose a system model and a programming methodology to write parallel programs that we think are suitable for a wide variety of applications. On top of that methodology, we implement fault tolerant programs and supply a reliable storage and transparent recovery mechanisms. The components of our system operate on some critical shared data structures. The efficient implementation of the fault-tolerant programs is achieved by allowing parallel non-blocking accesses to the shared data structures and the integration of the access primitives, the reliable storage management, and the communication procedures. Furthermore, our system is divided into clusters to limit the overhead. The analytical approximation of the response time of the system we have derived suggests that this implementation meets the performance goals for coarse to medium grain processing.

Why it matters

A significance statement is not available in the OpenAlex record.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Compared to traditional uniprocessor architectures, parallel systems provide superior processing power. However, their reliability may be significantly decreased due to the large number of processing elements and communication links. As large-scale parallel computer systems become more available, and for highly reliable applications that need to be implemented efficiently, reliability considerations require new solutions. We are interested in the implementation of fault-tolerant, efficient applications for large-scale hypercube computers. Our goal is three-fold: to achieve a low overhead during failure-free operating circumstances, to obtain graceful performance degradation as failures occur in the system, and to provide reasonably fast recovery procedures so that continuous service can be ensured. We first choose a system model and a programming methodology to write parallel programs that we think are suitable for a wide variety of applications. On top of that methodology, we implement fault tolerant programs and supply a reliable storage and transparent recovery mechanisms. The components of our system operate on some critical shared data structures. The efficient implementation of the fault-tolerant programs is achieved by allowing parallel non-blocking accesses to the shared data structures and the integration of the access primitives, the reliable storage management, and the communication procedures. Furthermore, our system is divided into clusters to limit the overhead. The analytical approximation of the response time of the system we have derived suggests that this implementation meets the performance goals for coarse to medium grain processing.

Key concepts: Computer science, Uniprocessor system, Distributed computing, Fault tolerance, Overhead (engineering), Reliability (semiconductor), Parallel computing, Embedded system

Related papers

Back to paper searchBrowse research topicsOriginal source
Architecture for fault-tolerant program execution on large-scale parallel hypercubes — Research Paper | ScholarLens