Adaptive redundancy for fault-tolerant real-time systems
Chia-Mei Chen, Satish K. Tripathi, Sheng-Tzong Cheng
Abstract
Chia-Mei Chen, Satish K. Tripathi, Sheng-Tzong Cheng
Abstract
Reliability is an important aspect of real-time systems because the result of a real-time application may be valid only if the application functions correctly and its timing constraints are satisfied. There are two kinds of faults: hardware and software faults. In this paper, we consider hardware transient faults. Full replication or full hardware redundancy can achieve a high degree of reliability; however, it may waste resources. We propose a fault-tolerance approach, a hybrid method of rollback and replication, for the real-time systems which require both system reliability and the guarantee of meeting deadlines. We define that a task is fault-tolerant if it can be recovered from a transient error either by rollback or duplication. Our approach attempts to make as many tasks fault-tolerant as possible.
OpenAlex reports 4 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Reliability is an important aspect of real-time systems because the result of a real-time application may be valid only if the application functions correctly and its timing constraints are satisfied. There are two kinds of faults: hardware and software faults. In this paper, we consider hardware transient faults. Full replication or full hardware redundancy can achieve a high degree of reliability; however, it may waste resources. We propose a fault-tolerance approach, a hybrid method of rollback and replication, for the real-time systems which require both system reliability and the guarantee of meeting deadlines. We define that a task is fault-tolerant if it can be recovered from a transient error either by rollback or duplication. Our approach attempts to make as many tasks fault-tolerant as possible.
Key concepts: Redundancy (engineering), Rollback, Fault tolerance, Computer science, Software fault tolerance, Replication (statistics), Reliability engineering, Reliability (semiconductor)