Effects of Anchor Item Methods on Differential Item Functioning Detection with the Likelihood Ratio Test
Wen‐Chung Wang, Ya-Li Yeh
Abstract
Wen‐Chung Wang, Ya-Li Yeh
Abstract
Through simulations, this study investigates the effects of anchor item methods on Type I error and power of detecting differential item functioning (DIF) using the likelihood ratio test within the framework of item response theory. Four anchor item methods were compared: the all-other, 1-item, 4-item, and 10-item methods. The results showed that it is the average signed area between the reference and focal groups rather than the percentage of DIF items in a test that determines the Type I error of the all-other method. The all-other method yields good control over Type I error and reasonable power only when the average signed area approaches zero. The all-other method is not recommended for practical DIF analysis because it is only adequate under very stringent conditions. The other three methods perform appropriately under all the simulated conditions. The more anchor items are used, the higher the power of DIF detection.
OpenAlex reports 142 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Through simulations, this study investigates the effects of anchor item methods on Type I error and power of detecting differential item functioning (DIF) using the likelihood ratio test within the framework of item response theory. Four anchor item methods were compared: the all-other, 1-item, 4-item, and 10-item methods. The results showed that it is the average signed area between the reference and focal groups rather than the percentage of DIF items in a test that determines the Type I error of the all-other method. The all-other method yields good control over Type I error and reasonable power only when the average signed area approaches zero. The all-other method is not recommended for practical DIF analysis because it is only adequate under very stringent conditions. The other three methods perform appropriately under all the simulated conditions. The more anchor items are used, the higher the power of DIF detection.
Key concepts: Differential item functioning, Type I and type II errors, Item response theory, Statistics, Equating, Item analysis, Test (biology), Mathematics