Analytical modeling of matrix–vector multiplication on multicore processors
Roman A. Gareev, Elena N. Akimova
Abstract
Roman A. Gareev, Elena N. Akimova
Abstract
The efficiency of matrix–vector multiplication is of considerable importance. No current approaches can optimize this sufficiently well under severe time constraints. All major existing methods are based on either manual‐tuning or auto‐tuning and can therefore be time‐consuming. We introduce an alternative model‐driven approach, which is used to map the implementation of matrix–vector multiplication to a target architecture and analytically obtain its parameters. The approach yields the performance that is competitive with optimized Basic Linear Algebra Subprograms (BLAS)‐like dense linear algebra libraries without the need for manual‐tuning or auto‐tuning. Our method provides competitive performance across hardware architectures and can be utilized to obtain single‐threaded and multi‐threaded implementations on multicore processors. We expect that this approach allows the community to progress from valuable engineering solutions to techniques with a broader application.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The efficiency of matrix–vector multiplication is of considerable importance. No current approaches can optimize this sufficiently well under severe time constraints. All major existing methods are based on either manual‐tuning or auto‐tuning and can therefore be time‐consuming. We introduce an alternative model‐driven approach, which is used to map the implementation of matrix–vector multiplication to a target architecture and analytically obtain its parameters. The approach yields the performance that is competitive with optimized Basic Linear Algebra Subprograms (BLAS)‐like dense linear algebra libraries without the need for manual‐tuning or auto‐tuning. Our method provides competitive performance across hardware architectures and can be utilized to obtain single‐threaded and multi‐threaded implementations on multicore processors. We expect that this approach allows the community to progress from valuable engineering solutions to techniques with a broader application.
Key concepts: Matrix multiplication, Multiplication (music), Linear algebra, Computer science, Implementation, Multi-core processor, Parallel computing, Matrix (chemical analysis)