High-speed FDTD simulation algorithm for GPU with compute unified device architecture
Naoki Takada, Tomoyoshi Shimobaba, Nobuyuki Masuda, Tomoyoshi Ito
Abstract
Naoki Takada, Tomoyoshi Shimobaba, Nobuyuki Masuda, Tomoyoshi Ito
Abstract
In this paper, we propose a high-speed FDTD algorithm for GPU. Our algorithm includes two important techniques: coalesced global memory access on a GPU board and the improved cache block algorithm for GPU which resembles that for the central processing unit (CPU). Our algorithm achieved an approximately 20- fold improvement in computational speed compared with a conventional CPU at the maximum computational speed of the GPU.
OpenAlex reports 19 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
In this paper, we propose a high-speed FDTD algorithm for GPU. Our algorithm includes two important techniques: coalesced global memory access on a GPU board and the improved cache block algorithm for GPU which resembles that for the central processing unit (CPU). Our algorithm achieved an approximately 20- fold improvement in computational speed compared with a conventional CPU at the maximum computational speed of the GPU.
Key concepts: Computer science, Parallel computing, CUDA, Central processing unit, Speedup, Block (permutation group theory), Graphics processing unit, Computational science