Multiple Comparisons in Microarray Data Analysis
Donghui Zhang, Li Liu
Abstract
Donghui Zhang, Li Liu
Abstract
Multiplicity is a challenging statistical issue in drug discovery, and a particular example is microarray study. The traditional approach of controlling of the family-wise error rate (FWER) is conservative when the number of tests is large. A more appropriate approach is to control the false discovery rate (FDR). Since the development of the Benjamini and Hochberg (BH) FDR procedure in 1995, many modifications have been proposed aimed at relaxing the requirement for independent test statistics or improving the power of the BH FDR procedure. Comparisons of these procedures in the current literature are not comprehensive and the conclusions on performances are inconsistent. The objectives of this article are three-fold: (a) to perform a more comprehensive comparison of extant multiple testing procedures using two real microarray datasets and various simulated data sets; (b) to explore potential reasons for the inconsistencies in published simulation results; and (c) to identify suitable FDR procedures under different scenarios according to covariance structure, percent of true null hypotheses among multiple tests, and sample size.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Multiplicity is a challenging statistical issue in drug discovery, and a particular example is microarray study. The traditional approach of controlling of the family-wise error rate (FWER) is conservative when the number of tests is large. A more appropriate approach is to control the false discovery rate (FDR). Since the development of the Benjamini and Hochberg (BH) FDR procedure in 1995, many modifications have been proposed aimed at relaxing the requirement for independent test statistics or improving the power of the BH FDR procedure. Comparisons of these procedures in the current literature are not comprehensive and the conclusions on performances are inconsistent. The objectives of this article are three-fold: (a) to perform a more comprehensive comparison of extant multiple testing procedures using two real microarray datasets and various simulated data sets; (b) to explore potential reasons for the inconsistencies in published simulation results; and (c) to identify suitable FDR procedures under different scenarios according to covariance structure, percent of true null hypotheses among multiple tests, and sample size.
Key concepts: False discovery rate, Multiple comparisons problem, Statistical hypothesis testing, Sample size determination, Computer science, Type I and type II errors, Data mining, Extant taxon