Interrater reliability: the kappa statistic
Marry L. McHugh
Abstract
Open-access reader
Marry L. McHugh
Abstract
Open-access reader
Many situations in the healthcare industry rely on multiple people to collect research or clinical laboratory data .The question of consistency, or agreement among the individuals collecting data immediately arises due to the variability among human observers .Well-designed research studies must therefore include procedures that measure agreement among the various data collectors .Study designs typically involve training the data collectors, and measuring the extent to which they record the same scores for the same phenomena .Perfect agreement is seldom achieved, and confidence in study results is partly a function of the amount of disagreement, or error introduced into the study from inconsistency among the data collectors .The extent of agreement among data collectors is called, "interrater reliability" .Interrater reliability is a concern to one degree or another in most large studies due to the fact that multiple people collecting data may experience and interpret the phenomena of interest differently .
OpenAlex reports 19388 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Many situations in the healthcare industry rely on multiple people to collect research or clinical laboratory data .The question of consistency, or agreement among the individuals collecting data immediately arises due to the variability among human observers .Well-designed research studies must therefore include procedures that measure agreement among the various data collectors .Study designs typically involve training the data collectors, and measuring the extent to which they record the same scores for the same phenomena .Perfect agreement is seldom achieved, and confidence in study results is partly a function of the amount of disagreement, or error introduced into the study from inconsistency among the data collectors .The extent of agreement among data collectors is called, "interrater reliability" .Interrater reliability is a concern to one degree or another in most large studies due to the fact that multiple people collecting data may experience and interpret the phenomena of interest differently .
Key concepts: Inter-rater reliability, Kappa, Cohen's kappa, Statistics, Statistic, Reliability (semiconductor), Mathematics, Test (biology)