Empirical study of parallelism throttling schemes on a massively parallel system
Yong Meng Teo, Jerry C. Yan
Abstract
Yong Meng Teo, Jerry C. Yan
Abstract
Exploitation of parallelism in massively parallel systems is intuitively appealing and is a promising avenue for achieving teraflop performance. However, parallelism is not free; apart from overheads in communication and synchronisation, having too much program parallelism can also raise serious resource management problems during program execution. The problem of resource management is particularly complicated by the distributed nature of massively parallel systems. We address the issue of managing parallelism in a massively parallel system. This is an extension of our previous work on hardware throttle for parallel systems. We propose two parallelism throttle schemes, and a scheme for monitoring and measurement of system workload. Processes in the system are initiated when runtime loading levels in the system permit. The aim is to match dynamic program parallelism to static machine parallelism. An experimental study is conducted using a simulated multi-cluster dataflow machine as a testbed. Our study shows the extent that control over runaway program parallelism is necessary, and that it is possible to have good distributed control of resource use in a massively parallel system.>
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Exploitation of parallelism in massively parallel systems is intuitively appealing and is a promising avenue for achieving teraflop performance. However, parallelism is not free; apart from overheads in communication and synchronisation, having too much program parallelism can also raise serious resource management problems during program execution. The problem of resource management is particularly complicated by the distributed nature of massively parallel systems. We address the issue of managing parallelism in a massively parallel system. This is an extension of our previous work on hardware throttle for parallel systems. We propose two parallelism throttle schemes, and a scheme for monitoring and measurement of system workload. Processes in the system are initiated when runtime loading levels in the system permit. The aim is to match dynamic program parallelism to static machine parallelism. An experimental study is conducted using a simulated multi-cluster dataflow machine as a testbed. Our study shows the extent that control over runaway program parallelism is necessary, and that it is possible to have good distributed control of resource use in a massively parallel system.>
Key concepts: Massively parallel, Computer science, Task parallelism, Parallel computing, Data parallelism, Parallelism (grammar), Testbed, Dataflow