A Proposed Solution to the Base Rate Problem in the Kappa Statistic
Edward L. Spitznagel
Abstract
Edward L. Spitznagel
Abstract
Because it corrects for chance agreement, kappa (kappa) is a useful statistic for calculating interrater concordance. However, kappa has been criticized because its computed value is a function not only of sensitivity and specificity, but also the prevalence, or base rate, of the illness of interest in the particular population under study. For example, it has been shown for a hypothetical case in which sensitivity and specificity remain constant at .95 each, that kappa falls from .81 to .14 when the prevalence drops from 50% to 1%. Thus, differing values of kappa may be entirely due to differences in prevalence. Calculation of agreement presents different problems depending on whether one is studying reliability or validity. We discuss quantification of agreement in the pure validity case, the pure reliability case, and those studies that fall somewhere between. As a way of minimizing the base rate problem, we propose a statistic for the quantification of agreement (the Y statistic), which can be related to kappa but which is completely independent of prevalence in the case of validity studies and relatively so in the case of reliability.
OpenAlex reports 491 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Because it corrects for chance agreement, kappa (kappa) is a useful statistic for calculating interrater concordance. However, kappa has been criticized because its computed value is a function not only of sensitivity and specificity, but also the prevalence, or base rate, of the illness of interest in the particular population under study. For example, it has been shown for a hypothetical case in which sensitivity and specificity remain constant at .95 each, that kappa falls from .81 to .14 when the prevalence drops from 50% to 1%. Thus, differing values of kappa may be entirely due to differences in prevalence. Calculation of agreement presents different problems depending on whether one is studying reliability or validity. We discuss quantification of agreement in the pure validity case, the pure reliability case, and those studies that fall somewhere between. As a way of minimizing the base rate problem, we propose a statistic for the quantification of agreement (the Y statistic), which can be related to kappa but which is completely independent of prevalence in the case of validity studies and relatively so in the case of reliability.
Key concepts: Kappa, Statistic, Cohen's kappa, Base (topology), Statistics, Econometrics, Mathematics, Psychology