The Effect of Correlated Failure on the Reliability of HPC Systems
Thanadech Thanakornworakij, Raja Nassar, Chokchai Leangsuksun, Mihaela Păun
Abstract
Thanadech Thanakornworakij, Raja Nassar, Chokchai Leangsuksun, Mihaela Păun
Abstract
High Performance Computing (HPC) system utilization can be maximized and sustained if one understands the failure behavior. In general, Time to Failure (TTF) of HPC systems has been long studied and showed that the Wei bull distribution gives the best fit. In addition, in many cases, TTF of such systems exhibit correlations. In our previous study, we developed a reliability model of an HPC system where failures among nodes are independent. However, some studies have clearly shown that in some cases nodes do not fail independently of one another. Therefore, it is of importance to develop a reliability model for an HPC system based on the occurrence of simultaneous failures. In this paper, we develop such a model and derive expressions for the probability density function of time to failure, system reliability, system failure rate, and mean time to failure (MTTF). Results show that if the failure of the components (nodes) in the system possesses a degree of dependency, the system reliability decreases.
OpenAlex reports 7 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
High Performance Computing (HPC) system utilization can be maximized and sustained if one understands the failure behavior. In general, Time to Failure (TTF) of HPC systems has been long studied and showed that the Wei bull distribution gives the best fit. In addition, in many cases, TTF of such systems exhibit correlations. In our previous study, we developed a reliability model of an HPC system where failures among nodes are independent. However, some studies have clearly shown that in some cases nodes do not fail independently of one another. Therefore, it is of importance to develop a reliability model for an HPC system based on the occurrence of simultaneous failures. In this paper, we develop such a model and derive expressions for the probability density function of time to failure, system reliability, system failure rate, and mean time to failure (MTTF). Results show that if the failure of the components (nodes) in the system possesses a degree of dependency, the system reliability decreases.
Key concepts: Mean time between failures, Reliability (semiconductor), Dependency (UML), Reliability engineering, Failure rate, Computer science, Function (biology), Probability density function