A Note on the Type I Error Rate of the PARSCALE G 2 Statistic for Long Tests
Kyong Hee Chon, Sandip Sinharay
Abstract
Kyong Hee Chon, Sandip Sinharay
Abstract
The PARSCALE G 2 statistic is arguably the most popular item fit statistic in operational testing. For long tests, the Type I error rates of the statistic have often been found to be satisfactory. However, the Type I error rates of the statistic have only been studied for sample sizes of up to several thousands. The authors examined the Type I error rates of the PARSCALE G 2 statistic in a simulation study using sample sizes much larger than those considered in the literature. For any fixed test length, the Type I error rate of the PARSCALE G 2 statistic is found to increase to 1 as the sample size increases. The findings contradict the claim in the PARSCALE software manual that the PARSCALE G 2 statistic leads to a large-sample test and also contradict the common belief that the statistic has reasonable Type I error rates for long tests. Thus, this simulation study conveys the important practical message that the use of the PARSCALE G 2 statistic cannot always be recommended even for long tests. The Type I error rates of the item fit statistics of Orlando and Thissen were found to be close to the nominal level for all simulation conditions considered here.
A significance statement is not available in the OpenAlex record.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
The PARSCALE G 2 statistic is arguably the most popular item fit statistic in operational testing. For long tests, the Type I error rates of the statistic have often been found to be satisfactory. However, the Type I error rates of the statistic have only been studied for sample sizes of up to several thousands. The authors examined the Type I error rates of the PARSCALE G 2 statistic in a simulation study using sample sizes much larger than those considered in the literature. For any fixed test length, the Type I error rate of the PARSCALE G 2 statistic is found to increase to 1 as the sample size increases. The findings contradict the claim in the PARSCALE software manual that the PARSCALE G 2 statistic leads to a large-sample test and also contradict the common belief that the statistic has reasonable Type I error rates for long tests. Thus, this simulation study conveys the important practical message that the use of the PARSCALE G 2 statistic cannot always be recommended even for long tests. The Type I error rates of the item fit statistics of Orlando and Thissen were found to be close to the nominal level for all simulation conditions considered here.
Key concepts: Statistic, Type I and type II errors, Statistics, PRESS statistic, Ancillary statistic, F-test, Mathematics, Sample size determination