Comparing SFMD and SPMD Computation for On-Chip Multiprocessing of Intermediate Level Image Understanding Algorithms
Steve Rehfuss, Dan W. Hammerstrom
Abstract
Steve Rehfuss, Dan W. Hammerstrom
Abstract
The SFMD model of computation is intermediate between SIMD and SPMD processing. It is designed to extend SIMD processing to tasks requiring much data-dependent branching, such as model matching and various intermediate level vision algorithms, while retaining the simple semantics of SIMD. In this paper, within the realm of multiple processors per chip, we review the comparative performance of SIMD and SFMD processing, and then compare the performance of SFMD and SPMD on tasks requiring extensive inter-processor communication (IPC). A variety of approaches show SFMD giving a 1.5 - 4x speedup over SIMD on many tasks of interest, with little increase in cost. We derive lower bounds on the slowdown of SFMD versus SPMD showing SFMD is at least 1/2 the speed of SPMD for a wide regime. Review of recent VLSI projections show that pin limitations may force SPMD to have only 1/2 the number of processors per chip as SFMD, implying that there is a wide regime, involving substantial IPC, where SFMD is competitive or superior in performance to SPMD. As SFMD is substantially easier to program than SPMD, it appears to be a viable architecture for on-chip multiprocessing.
OpenAlex reports 2 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The SFMD model of computation is intermediate between SIMD and SPMD processing. It is designed to extend SIMD processing to tasks requiring much data-dependent branching, such as model matching and various intermediate level vision algorithms, while retaining the simple semantics of SIMD. In this paper, within the realm of multiple processors per chip, we review the comparative performance of SIMD and SFMD processing, and then compare the performance of SFMD and SPMD on tasks requiring extensive inter-processor communication (IPC). A variety of approaches show SFMD giving a 1.5 - 4x speedup over SIMD on many tasks of interest, with little increase in cost. We derive lower bounds on the slowdown of SFMD versus SPMD showing SFMD is at least 1/2 the speed of SPMD for a wide regime. Review of recent VLSI projections show that pin limitations may force SPMD to have only 1/2 the number of processors per chip as SFMD, implying that there is a wide regime, involving substantial IPC, where SFMD is competitive or superior in performance to SPMD. As SFMD is substantially easier to program than SPMD, it appears to be a viable architecture for on-chip multiprocessing.
Key concepts: SPMD, SIMD, Computer science, Parallel computing, Multiprocessing, Speedup, Computation, Algorithm