Published Studies of Interrater Reliability Often Overestimate Reliability: Computing the Correct Coefficient
Xitao Fan, Michael Chen
Abstract
Xitao Fan, Michael Chen
Abstract
It is erroneous to generalize the interrater reliability coefficient estimated from two or more raters rating only a (small) portion of the sample to the rest of the sample data for which only one rater is used for scoring, although such generalization is often made implicitly in practice. If the interrater reliability estimate from part of a sample is available, the score reliability for the rest of the sample data for which only one rater is used for scoring can be estimated both within the framework of classical reliability theory and that of generalizability theory. As intuitively expected, score reliability when only one rater is used for scoring is lower than the score reliability for which two raters are used. The authors provide a sample of published studies in different disciplines that inappropriately generalized reliability coefficients involving several raters to scores generated by a single rater.
OpenAlex reports 22 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
It is erroneous to generalize the interrater reliability coefficient estimated from two or more raters rating only a (small) portion of the sample to the rest of the sample data for which only one rater is used for scoring, although such generalization is often made implicitly in practice. If the interrater reliability estimate from part of a sample is available, the score reliability for the rest of the sample data for which only one rater is used for scoring can be estimated both within the framework of classical reliability theory and that of generalizability theory. As intuitively expected, score reliability when only one rater is used for scoring is lower than the score reliability for which two raters are used. The authors provide a sample of published studies in different disciplines that inappropriately generalized reliability coefficients involving several raters to scores generated by a single rater.
Key concepts: Generalizability theory, Inter-rater reliability, Reliability (semiconductor), Statistics, Sample (material), Generalization, Psychology, Sample size determination