Architecture for fault-tolerant program execution on large-scale parallel hypercubes
Saniya Ben Hassen
Abstract
Saniya Ben Hassen
Abstract
Compared to traditional uniprocessor architectures, parallel systems provide superior processing power. However, their reliability may be significantly decreased due to the large number of processing elements and communication links. As large-scale parallel computer systems become more available, and for highly reliable applications that need to be implemented efficiently, reliability considerations require new solutions. We are interested in the implementation of fault-tolerant, efficient applications for large-scale hypercube computers. Our goal is three-fold: to achieve a low overhead during failure-free operating circumstances, to obtain graceful performance degradation as failures occur in the system, and to provide reasonably fast recovery procedures so that continuous service can be ensured. We first choose a system model and a programming methodology to write parallel programs that we think are suitable for a wide variety of applications. On top of that methodology, we implement fault tolerant programs and supply a reliable storage and transparent recovery mechanisms. The components of our system operate on some critical shared data structures. The efficient implementation of the fault-tolerant programs is achieved by allowing parallel non-blocking accesses to the shared data structures and the integration of the access primitives, the reliable storage management, and the communication procedures. Furthermore, our system is divided into clusters to limit the overhead. The analytical approximation of the response time of the system we have derived suggests that this implementation meets the performance goals for coarse to medium grain processing.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Compared to traditional uniprocessor architectures, parallel systems provide superior processing power. However, their reliability may be significantly decreased due to the large number of processing elements and communication links. As large-scale parallel computer systems become more available, and for highly reliable applications that need to be implemented efficiently, reliability considerations require new solutions. We are interested in the implementation of fault-tolerant, efficient applications for large-scale hypercube computers. Our goal is three-fold: to achieve a low overhead during failure-free operating circumstances, to obtain graceful performance degradation as failures occur in the system, and to provide reasonably fast recovery procedures so that continuous service can be ensured. We first choose a system model and a programming methodology to write parallel programs that we think are suitable for a wide variety of applications. On top of that methodology, we implement fault tolerant programs and supply a reliable storage and transparent recovery mechanisms. The components of our system operate on some critical shared data structures. The efficient implementation of the fault-tolerant programs is achieved by allowing parallel non-blocking accesses to the shared data structures and the integration of the access primitives, the reliable storage management, and the communication procedures. Furthermore, our system is divided into clusters to limit the overhead. The analytical approximation of the response time of the system we have derived suggests that this implementation meets the performance goals for coarse to medium grain processing.
Key concepts: Computer science, Uniprocessor system, Distributed computing, Fault tolerance, Overhead (engineering), Reliability (semiconductor), Parallel computing, Embedded system