2000Educational and Psychological MeasurementRequires access

Published Studies of Interrater Reliability Often Overestimate Reliability: Computing the Correct Coefficient

Xitao Fan, Michael Chen

Open publisher page 22 citations

Abstract

It is erroneous to generalize the interrater reliability coefficient estimated from two or more raters rating only a (small) portion of the sample to the rest of the sample data for which only one rater is used for scoring, although such generalization is often made implicitly in practice. If the interrater reliability estimate from part of a sample is available, the score reliability for the rest of the sample data for which only one rater is used for scoring can be estimated both within the framework of classical reliability theory and that of generalizability theory. As intuitively expected, score reliability when only one rater is used for scoring is lower than the score reliability for which two raters are used. The authors provide a sample of published studies in different disciplines that inappropriately generalized reliability coefficients involving several raters to scores generated by a single rater.

About this research paper

What this paper is about

It is erroneous to generalize the interrater reliability coefficient estimated from two or more raters rating only a (small) portion of the sample to the rest of the sample data for which only one rater is used for scoring, although such generalization is often made implicitly in practice. If the interrater reliability estimate from part of a sample is available, the score reliability for the rest of the sample data for which only one rater is used for scoring can be estimated both within the framework of classical reliability theory and that of generalizability theory. As intuitively expected, score reliability when only one rater is used for scoring is lower than the score reliability for which two raters are used. The authors provide a sample of published studies in different disciplines that inappropriately generalized reliability coefficients involving several raters to scores generated by a single rater.

Why it matters

OpenAlex reports 22 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

It is erroneous to generalize the interrater reliability coefficient estimated from two or more raters rating only a (small) portion of the sample to the rest of the sample data for which only one rater is used for scoring, although such generalization is often made implicitly in practice. If the interrater reliability estimate from part of a sample is available, the score reliability for the rest of the sample data for which only one rater is used for scoring can be estimated both within the framework of classical reliability theory and that of generalizability theory. As intuitively expected, score reliability when only one rater is used for scoring is lower than the score reliability for which two raters are used. The authors provide a sample of published studies in different disciplines that inappropriately generalized reliability coefficients involving several raters to scores generated by a single rater.

Key concepts: Generalizability theory, Inter-rater reliability, Reliability (semiconductor), Statistics, Sample (material), Generalization, Psychology, Sample size determination

Related papers

Back to paper searchBrowse research topicsOriginal source
Published Studies of Interrater Reliability Often Overestimate Reliability: Computing the Correct Coefficient — Research Paper | ScholarLens