False Discovery Rate and Asymptotics
Thorsten-Ingo M. Sc. Dickhaus
Abstract
Thorsten-Ingo M. Sc. Dickhaus
Abstract
The false discovery rate (FDR) is a rather young error control criterion in multiple testing problems. Initiated by the pioneering paper by Benjamini and Hochberg from 1995, it has become popular in the 1990ies as an alternative to the strong control of the family-wise error rate, especially if a large system of hypotheses is at hand and the analysis has mainly explorative character. Instead of controlling the probability of one or more false rejections, the FDR controls the expected proportion of falsely rejected hypotheses among all rejections. One typical application with strong impact on the development of the FDR is the first step (screening phase) of a microarray experiment where the experimenter aims at detecting a few candidate genes or SNPs potentially associated with a disease, which are than further analyzed using more stringent error handling methods. Especially due to such nowadays' applications with families of ten thousands or even some hundred thousands of hypotheses at hand, asymptotic considerations (with the number of hypotheses to be tested simultaneously tending to infinity) become more and more relevant. In this work, the behavior of the FDR is mainly studied from a theoretical point of view. After some fundamental issues as a preparation in Chapter 1, focus is laid in Chapter 2 on the asymptotic behaviour of the linear step-up procedure originally introduced by Benjamini and Hochberg. Since it is well known that this procedure strongly controls the FDR under positive dependency, we investigate the asymptotic conservativeness of this procedure under various distributional settings in depth. The results imply that, depending on the strength of positive dependence among the test statistics and the proportion of true nulls, the FDR can be close to the pre-specified error level or can be very small. Typically, the latter case leads to low power of the linear step-up procedure which raises the possibility for improvements of the algorithm. One improvement of Benjamini and Hochberg's procedure is presented and discussed in Chapter 3. Instead of using critical values increasing linearly (or, in other words, a linear rejection curve), we derive a non-linear and in some sense asymptotically optimal rejection curve leading to the full exhaustion of the FDR level under some extreme parameter configurations. This curve is then implemented into some stepwise multiple test procedures which control the FDR asymptotically or (with slight modifications) for a finite number of hypotheses. For the proof of FDR control for procedures employing non-linear critical values, some new methodology of proof is worked out. Chapter 4 then compares the newly derived methods with the original linear step-up procedure and other improved procedures with respect to multiple power. The results in this comparisons section are based on computer simulations. It turns out that certain procedures perform better in certain distributional setups or in other words that one can choose the appropriate FDR controlling algorithm to serve the purpose of detecting the most relevant alternatives most properly. Besides all these theoretical and methodological topics, we are also concerned with some practical aspects of FDR. We apply FDR controlling procedures to real life data and illustrate the functionality, assets and drawbacks of the different methods using these data sets.
OpenAlex reports 4 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The false discovery rate (FDR) is a rather young error control criterion in multiple testing problems. Initiated by the pioneering paper by Benjamini and Hochberg from 1995, it has become popular in the 1990ies as an alternative to the strong control of the family-wise error rate, especially if a large system of hypotheses is at hand and the analysis has mainly explorative character. Instead of controlling the probability of one or more false rejections, the FDR controls the expected proportion of falsely rejected hypotheses among all rejections. One typical application with strong impact on the development of the FDR is the first step (screening phase) of a microarray experiment where the experimenter aims at detecting a few candidate genes or SNPs potentially associated with a disease, which are than further analyzed using more stringent error handling methods. Especially due to such nowadays' applications with families of ten thousands or even some hundred thousands of hypotheses at hand, asymptotic considerations (with the number of hypotheses to be tested simultaneously tending to infinity) become more and more relevant. In this work, the behavior of the FDR is mainly studied from a theoretical point of view. After some fundamental issues as a preparation in Chapter 1, focus is laid in Chapter 2 on the asymptotic behaviour of the linear step-up procedure originally introduced by Benjamini and Hochberg. Since it is well known that this procedure strongly controls the FDR under positive dependency, we investigate the asymptotic conservativeness of this procedure under various distributional settings in depth. The results imply that, depending on the strength of positive dependence among the test statistics and the proportion of true nulls, the FDR can be close to the pre-specified error level or can be very small. Typically, the latter case leads to low power of the linear step-up procedure which raises the possibility for improvements of the algorithm. One improvement of Benjamini and Hochberg's procedure is presented and discussed in Chapter 3. Instead of using critical values increasing linearly (or, in other words, a linear rejection curve), we derive a non-linear and in some sense asymptotically optimal rejection curve leading to the full exhaustion of the FDR level under some extreme parameter configurations. This curve is then implemented into some stepwise multiple test procedures which control the FDR asymptotically or (with slight modifications) for a finite number of hypotheses. For the proof of FDR control for procedures employing non-linear critical values, some new methodology of proof is worked out. Chapter 4 then compares the newly derived methods with the original linear step-up procedure and other improved procedures with respect to multiple power. The results in this comparisons section are based on computer simulations. It turns out that certain procedures perform better in certain distributional setups or in other words that one can choose the appropriate FDR controlling algorithm to serve the purpose of detecting the most relevant alternatives most properly. Besides all these theoretical and methodological topics, we are also concerned with some practical aspects of FDR. We apply FDR controlling procedures to real life data and illustrate the functionality, assets and drawbacks of the different methods using these data sets.
Key concepts: False discovery rate, Multiple comparisons problem, Type I and type II errors, Computer science, Word error rate, Control (management), Mathematics, Statistics