Accelerating the Production of Synthetic Seismograms by a Multicore Processor Cluster with Multiple GPUs
Ferdinando Alessi, Annalisa Massini, Roberto Basili
Abstract
Ferdinando Alessi, Annalisa Massini, Roberto Basili
Abstract
In this work we propose two different parallel versions of the software package COMPSYN, devoted to the production of syntethic seismograms. The first version consists in the parallelization of the code to run on a cluster of multicore processors and is obtained by exploiting the MPI paradigm and OpenMP API to the end of maximizing the performance on multicore processors. The second version exploits the set of GPU associated to the multicore processor cluster and uses CUDA to take advantage of the GPU's computational power. We analyze the application performance of the two different implementations by using a real case study. In particular, we obtain for the GPU version a speedup of 10x over the parallelized version running on the cluster of multicore processors. Furthermore, we can estimate about at least 100x the speedup of the GPU version using a single node of the cluster with respect to the original sequential version.
OpenAlex reports 1 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
In this work we propose two different parallel versions of the software package COMPSYN, devoted to the production of syntethic seismograms. The first version consists in the parallelization of the code to run on a cluster of multicore processors and is obtained by exploiting the MPI paradigm and OpenMP API to the end of maximizing the performance on multicore processors. The second version exploits the set of GPU associated to the multicore processor cluster and uses CUDA to take advantage of the GPU's computational power. We analyze the application performance of the two different implementations by using a real case study. In particular, we obtain for the GPU version a speedup of 10x over the parallelized version running on the cluster of multicore processors. Furthermore, we can estimate about at least 100x the speedup of the GPU version using a single node of the cluster with respect to the original sequential version.
Key concepts: Speedup, Computer science, Parallel computing, Multi-core processor, CUDA, GPU cluster, Instruction set, General-purpose computing on graphics processing units