Study on GPU-accelerated extraction of interconnects parasitic using CUDA and MPI
Xiaoyu Xu, Guoqiang Liu, Hui Yan Qu, Wei Min Xu, Yang Zhang
Abstract
Xiaoyu Xu, Guoqiang Liu, Hui Yan Qu, Wei Min Xu, Yang Zhang
Abstract
Parallel computation is application-oriented, particularly for the GPU (Graphics Processing Unit) with the inherent parallelism. This paper shows the architecture of a GPU cluster based on MPI (Message Passing Interface) and CUDA (Compute Unified Device Architecture). Results show that the acceleration ratio is obviously improved but the acceleration effect seems decelerated in large-scale GPU cluster. The parallel algorithm is mainly focused on task partitioning sparse matrix-vector multiplications (SpVM) in GPUs.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Parallel computation is application-oriented, particularly for the GPU (Graphics Processing Unit) with the inherent parallelism. This paper shows the architecture of a GPU cluster based on MPI (Message Passing Interface) and CUDA (Compute Unified Device Architecture). Results show that the acceleration ratio is obviously improved but the acceleration effect seems decelerated in large-scale GPU cluster. The parallel algorithm is mainly focused on task partitioning sparse matrix-vector multiplications (SpVM) in GPUs.
Key concepts: CUDA, Computer science, Parallel computing, Graphics processing unit, GPU cluster, Message Passing Interface, Acceleration, Computational science