2017Unpublished venueRequires access

Performance Optimization of OpenFOAM* on Clusters of Intel® Xeon Phi (TM) Processors

Ravi Ojha, Prasad Pawar, Sonia Gupta, Michael Klemm, Manoj Nambiar

Open publisher page 3 citations

Abstract

OpenFOAM* is a software package for solving partial differential equations and is very popular for computational fluid dynamics in the automotive segment. Intel has recently launched the Intel® Xeon Phi" processor which offers an opportunity of porting an application that is already running on Intel® Xeon® processors. The new Intel Xeon Phi processor offers 512 bit wide SIMD vector registers and high bandwidth on-package memory. This paper presents the work done to optimize the performance of OpenFOAM for both single node and clusters of Intel Xeon Phi processors. Feasibility and scope of techniques such as vectorization have been studied closely and implemented for top hotspots to gain improvements. For large scale runs, it becomes critical to deploy the problem on multiple nodes so that the solver scales and provides best possible performance. Thus, getting close to optimal domain decomposition becomes more important. We have studied the scalability of the resulting data distribution of such domain decomposition algorithms. To further improve the scalability we have designed our own algorithm (partition & reshuffle) which reshuffles an already decomposed mesh in order to minimize inter-node communication costs, i.e., network traffic. Our algorithm performs better or is at par when compared to existing decomposition algorithms for a motorbike benchmark with 20 million cells. A compelling study has also been done against Xeon processors to emphasize the speedup that could be achieved.

About this research paper

What this paper is about

OpenFOAM* is a software package for solving partial differential equations and is very popular for computational fluid dynamics in the automotive segment. Intel has recently launched the Intel® Xeon Phi" processor which offers an opportunity of porting an application that is already running on Intel® Xeon® processors. The new Intel Xeon Phi processor offers 512 bit wide SIMD vector registers and high bandwidth on-package memory. This paper presents the work done to optimize the performance of OpenFOAM for both single node and clusters of Intel Xeon Phi processors. Feasibility and scope of techniques such as vectorization have been studied closely and implemented for top hotspots to gain improvements. For large scale runs, it becomes critical to deploy the problem on multiple nodes so that the solver scales and provides best possible performance. Thus, getting close to optimal domain decomposition becomes more important. We have studied the scalability of the resulting data distribution of such domain decomposition algorithms. To further improve the scalability we have designed our own algorithm (partition & reshuffle) which reshuffles an already decomposed mesh in order to minimize inter-node communication costs, i.e., network traffic. Our algorithm performs better or is at par when compared to existing decomposition algorithms for a motorbike benchmark with 20 million cells. A compelling study has also been done against Xeon processors to emphasize the speedup that could be achieved.

Why it matters

OpenAlex reports 3 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

OpenFOAM* is a software package for solving partial differential equations and is very popular for computational fluid dynamics in the automotive segment. Intel has recently launched the Intel® Xeon Phi" processor which offers an opportunity of porting an application that is already running on Intel® Xeon® processors. The new Intel Xeon Phi processor offers 512 bit wide SIMD vector registers and high bandwidth on-package memory. This paper presents the work done to optimize the performance of OpenFOAM for both single node and clusters of Intel Xeon Phi processors. Feasibility and scope of techniques such as vectorization have been studied closely and implemented for top hotspots to gain improvements. For large scale runs, it becomes critical to deploy the problem on multiple nodes so that the solver scales and provides best possible performance. Thus, getting close to optimal domain decomposition becomes more important. We have studied the scalability of the resulting data distribution of such domain decomposition algorithms. To further improve the scalability we have designed our own algorithm (partition & reshuffle) which reshuffles an already decomposed mesh in order to minimize inter-node communication costs, i.e., network traffic. Our algorithm performs better or is at par when compared to existing decomposition algorithms for a motorbike benchmark with 20 million cells. A compelling study has also been done against Xeon processors to emphasize the speedup that could be achieved.

Key concepts: Xeon Phi, Parallel computing, Computer science, Xeon, Scalability, SIMD, Vectorization (mathematics), Porting

Related papers

Back to paper searchBrowse research topicsOriginal source
Performance Optimization of OpenFOAM* on Clusters of Intel® Xeon Phi (TM) Processors — Research Paper | ScholarLens