Robust parallel resource management in shared memory multiprocessor systems
I‐Ling Yen, Farokh Bastani
Abstract
I‐Ling Yen, Farokh Bastani
Abstract
Parallel machines are being increasingly used for applications that require both quick response time and high reliability. This poses a challenge in programming these systems since it must be ensured that there is sufficient redundancy to cope with failures and that, at the same time, redundant components are used effectively during failure free periods to enhance the performance. Among the issues, resource management in such systems is highly critical to the robustness and efficiency of the system. A good resource management algorithm should allow the system to continue its operation even in the presence of a significant number of processor failures. Also, the incorporation of fault tolerance should not incur too much overhead. In this paper, we develop two robust resource management algorithms which simultaneously achieve the twin objectives of low overhead and high reliability.>
OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Parallel machines are being increasingly used for applications that require both quick response time and high reliability. This poses a challenge in programming these systems since it must be ensured that there is sufficient redundancy to cope with failures and that, at the same time, redundant components are used effectively during failure free periods to enhance the performance. Among the issues, resource management in such systems is highly critical to the robustness and efficiency of the system. A good resource management algorithm should allow the system to continue its operation even in the presence of a significant number of processor failures. Also, the incorporation of fault tolerance should not incur too much overhead. In this paper, we develop two robust resource management algorithms which simultaneously achieve the twin objectives of low overhead and high reliability.>
Key concepts: Fault tolerance, Computer science, Redundancy (engineering), Multiprocessing, Robustness (evolution), Distributed computing, Overhead (engineering), Resource management (computing)