2011Unpublished venueRequires access

The Effect of Correlated Failure on the Reliability of HPC Systems

Thanadech Thanakornworakij, Raja Nassar, Chokchai Leangsuksun, Mihaela Păun

Open publisher page 7 citations

Abstract

High Performance Computing (HPC) system utilization can be maximized and sustained if one understands the failure behavior. In general, Time to Failure (TTF) of HPC systems has been long studied and showed that the Wei bull distribution gives the best fit. In addition, in many cases, TTF of such systems exhibit correlations. In our previous study, we developed a reliability model of an HPC system where failures among nodes are independent. However, some studies have clearly shown that in some cases nodes do not fail independently of one another. Therefore, it is of importance to develop a reliability model for an HPC system based on the occurrence of simultaneous failures. In this paper, we develop such a model and derive expressions for the probability density function of time to failure, system reliability, system failure rate, and mean time to failure (MTTF). Results show that if the failure of the components (nodes) in the system possesses a degree of dependency, the system reliability decreases.

About this research paper

What this paper is about

High Performance Computing (HPC) system utilization can be maximized and sustained if one understands the failure behavior. In general, Time to Failure (TTF) of HPC systems has been long studied and showed that the Wei bull distribution gives the best fit. In addition, in many cases, TTF of such systems exhibit correlations. In our previous study, we developed a reliability model of an HPC system where failures among nodes are independent. However, some studies have clearly shown that in some cases nodes do not fail independently of one another. Therefore, it is of importance to develop a reliability model for an HPC system based on the occurrence of simultaneous failures. In this paper, we develop such a model and derive expressions for the probability density function of time to failure, system reliability, system failure rate, and mean time to failure (MTTF). Results show that if the failure of the components (nodes) in the system possesses a degree of dependency, the system reliability decreases.

Why it matters

OpenAlex reports 7 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

High Performance Computing (HPC) system utilization can be maximized and sustained if one understands the failure behavior. In general, Time to Failure (TTF) of HPC systems has been long studied and showed that the Wei bull distribution gives the best fit. In addition, in many cases, TTF of such systems exhibit correlations. In our previous study, we developed a reliability model of an HPC system where failures among nodes are independent. However, some studies have clearly shown that in some cases nodes do not fail independently of one another. Therefore, it is of importance to develop a reliability model for an HPC system based on the occurrence of simultaneous failures. In this paper, we develop such a model and derive expressions for the probability density function of time to failure, system reliability, system failure rate, and mean time to failure (MTTF). Results show that if the failure of the components (nodes) in the system possesses a degree of dependency, the system reliability decreases.

Key concepts: Mean time between failures, Reliability (semiconductor), Dependency (UML), Reliability engineering, Failure rate, Computer science, Function (biology), Probability density function

Related papers

Back to paper searchBrowse research topicsOriginal source
The Effect of Correlated Failure on the Reliability of HPC Systems — Research Paper | ScholarLens