More than Just the Kappa Coefficient: A Program to Fully Characterize Inter-Rater Reliability between Two Raters
Michael P. Cunningham
Abstract
Michael P. Cunningham
Abstract
The kappa coefficient is a widely used statistic for measuring the degree of reliability between raters. SAS ® procedures and macros exist for calculating kappa with two or more raters, but none address situations when the kappa coefficient alone does not sufficiently describe the level of reliability. When the prevalence of a rating in the population is very high or low, the value of kappa may indicate poor reliability even with a high observed proportion of agreement. Researchers have recommended reporting several other values in addition to the kappa to address this and another paradox of the kappa statistic. This program, developed in SAS ® 9.1, calculates kappa, but also outputs the observed and expected proportions of agreement, the prevalence and bias indices, and the prevalence adjusted bias adjusted kappa (PABAK) for two raters. Designed for input of the rater responses in the familiar 2x2 table format using the SAS %WINDOW statement, users with minimal SAS experience will be able to report these statistics to more fully characterize the extent of inter-observer agreement between two raters.
OpenAlex reports 62 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The kappa coefficient is a widely used statistic for measuring the degree of reliability between raters. SAS ® procedures and macros exist for calculating kappa with two or more raters, but none address situations when the kappa coefficient alone does not sufficiently describe the level of reliability. When the prevalence of a rating in the population is very high or low, the value of kappa may indicate poor reliability even with a high observed proportion of agreement. Researchers have recommended reporting several other values in addition to the kappa to address this and another paradox of the kappa statistic. This program, developed in SAS ® 9.1, calculates kappa, but also outputs the observed and expected proportions of agreement, the prevalence and bias indices, and the prevalence adjusted bias adjusted kappa (PABAK) for two raters. Designed for input of the rater responses in the familiar 2x2 table format using the SAS %WINDOW statement, users with minimal SAS experience will be able to report these statistics to more fully characterize the extent of inter-observer agreement between two raters.
Key concepts: Cohen's kappa, Kappa, Inter-rater reliability, Statistics, Reliability (semiconductor), Statistic, Population, Mathematics