Extra High Speed Matrix Multiplication on the Cray-2
David A. Bailey
Abstract
David A. Bailey
Abstract
The Cray-2 is capable of performing matrix multiplication at very high rates. Using library routines provided by Cray Research, Inc., performance rates of 300 to 425 MFLOPS can be obtained on a single processor, depending on system load. Considerably higher rates can be achieved with all four processors running simultaneously. This article describes how matrix multiplication can be performed even faster, at up to twice the above-listed rates. This can be achieved by: (1) employing Strassen’s matrix multiplication algorithm to reduce the number of floating-point operations performed and (2) utilizing local memory on the Cray-2 to avoid performance losses due to memory bank contention. The numerical stability and potential for parallel application of this procedure are also discussed.
OpenAlex reports 73 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The Cray-2 is capable of performing matrix multiplication at very high rates. Using library routines provided by Cray Research, Inc., performance rates of 300 to 425 MFLOPS can be obtained on a single processor, depending on system load. Considerably higher rates can be achieved with all four processors running simultaneously. This article describes how matrix multiplication can be performed even faster, at up to twice the above-listed rates. This can be achieved by: (1) employing Strassen’s matrix multiplication algorithm to reduce the number of floating-point operations performed and (2) utilizing local memory on the Cray-2 to avoid performance losses due to memory bank contention. The numerical stability and potential for parallel application of this procedure are also discussed.
Key concepts: Strassen algorithm, Matrix multiplication, FLOPS, Parallel computing, Computer science, Multiplication (music), Matrix (chemical analysis), Supercomputer