2023Research SquareOpen access

Fine Grain Algorithm Parallelization on a Hybrid Control-flow and Dataflow Processor

Nenad Korolija, Borko Furht, Veljko Milutinović

Open full text 1 citations

Abstract

Abstract The execution time of a high performance computing algorithm depends on multiple factors: the algorithm scalability, the chosen hardware, the communication speed between processing elements, etc. This work is based on a hybrid processor consisting of both the control-flow and dataflow hardware, where the control-flow hardware includes manycore architecture. The algorithm decomposition into portions suitable for different architectures is presented. The results are presented by decomposing the Lattice-Boltzmann method implemented for both control-flow and dataflow hardware. The total acceleration factor of the decomposed Lattice-Boltzmann method comparing to the execution time using the control-flow and the dataflow hardware is obtained by the analysis based on the speed of the communication between both hardware equal to the speed of shared cache memories. Results indicate the advantage of using the proposed hybrid architecture and the capability of accelerating considerably even the suitable for dataflow architectures. The analysis of the problem size needed to justify the utilization of control-flow hardware is also presented. The main benefit is in accelerating algorithms for which only some algorithm portions are suitable for the dataflow hardware.

Open-access reader

About this research paper

What this paper is about

Abstract The execution time of a high performance computing algorithm depends on multiple factors: the algorithm scalability, the chosen hardware, the communication speed between processing elements, etc. This work is based on a hybrid processor consisting of both the control-flow and dataflow hardware, where the control-flow hardware includes manycore architecture. The algorithm decomposition into portions suitable for different architectures is presented. The results are presented by decomposing the Lattice-Boltzmann method implemented for both control-flow and dataflow hardware. The total acceleration factor of the decomposed Lattice-Boltzmann method comparing to the execution time using the control-flow and the dataflow hardware is obtained by the analysis based on the speed of the communication between both hardware equal to the speed of shared cache memories. Results indicate the advantage of using the proposed hybrid architecture and the capability of accelerating considerably even the suitable for dataflow architectures. The analysis of the problem size needed to justify the utilization of control-flow hardware is also presented. The main benefit is in accelerating algorithms for which only some algorithm portions are suitable for the dataflow hardware.

Why it matters

OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

Abstract The execution time of a high performance computing algorithm depends on multiple factors: the algorithm scalability, the chosen hardware, the communication speed between processing elements, etc. This work is based on a hybrid processor consisting of both the control-flow and dataflow hardware, where the control-flow hardware includes manycore architecture. The algorithm decomposition into portions suitable for different architectures is presented. The results are presented by decomposing the Lattice-Boltzmann method implemented for both control-flow and dataflow hardware. The total acceleration factor of the decomposed Lattice-Boltzmann method comparing to the execution time using the control-flow and the dataflow hardware is obtained by the analysis based on the speed of the communication between both hardware equal to the speed of shared cache memories. Results indicate the advantage of using the proposed hybrid architecture and the capability of accelerating considerably even the suitable for dataflow architectures. The analysis of the problem size needed to justify the utilization of control-flow hardware is also presented. The main benefit is in accelerating algorithms for which only some algorithm portions are suitable for the dataflow hardware.

Key concepts: Dataflow, Dataflow architecture, Computer science, Parallel computing, Control flow, Scalability, Speedup, Data flow diagram

Related papers

Back to paper searchBrowse research topicsOriginal source
Fine Grain Algorithm Parallelization on a Hybrid Control-flow and Dataflow Processor — Research Paper | ScholarLens