1993ETS Research Report SeriesRequires access

A SIMULATION STUDY OF METHODS FOR ASSESSING DIFFERENTIAL ITEM FUNCTIONING IN COMPUTER‐ADAPTIVE TESTS

Rebecca Zwick, Dorothy T. Thayer, Marilyn S. Wingersky

Open publisher page 13 citations

Abstract

ABSTRACT Simulated data were used to investigate the performance of modified versions of the Mantel‐Haenszel and standardization methods of differential item functioning (DIF) analysis in computer‐adaptive tests (CATs). Each “examinee” received 25 items out of a 75‐item pool. A three‐parameter logistic item response model was assumed, and examinees were matched on expected true scores based on their CAT responses and on estimated item parameters. Both DIF methods performed well. The CAT‐based DIF statistics were highly correlated with DIF statistics based on nonadaptive administration of all 75 pool items and with the true magnitudes of DIF in the simulation. DIF methods were also investigated for “pretest items,” for which item parameter estimates were assumed to be unavailable. The pretest DIF statistics were generally well‐behaved and also had high correlations with the true DIF. The pretest DIF measures, however, tended to be slightly smaller in magnitude than their CAT‐based counterparts. Also, in the case of the Mantel‐Haenszel approach, the pretest DIF statistics tended to have somewhat larger standard errors than the CAT DIF statistics.

About this research paper

What this paper is about

ABSTRACT Simulated data were used to investigate the performance of modified versions of the Mantel‐Haenszel and standardization methods of differential item functioning (DIF) analysis in computer‐adaptive tests (CATs). Each “examinee” received 25 items out of a 75‐item pool. A three‐parameter logistic item response model was assumed, and examinees were matched on expected true scores based on their CAT responses and on estimated item parameters. Both DIF methods performed well. The CAT‐based DIF statistics were highly correlated with DIF statistics based on nonadaptive administration of all 75 pool items and with the true magnitudes of DIF in the simulation. DIF methods were also investigated for “pretest items,” for which item parameter estimates were assumed to be unavailable. The pretest DIF statistics were generally well‐behaved and also had high correlations with the true DIF. The pretest DIF measures, however, tended to be slightly smaller in magnitude than their CAT‐based counterparts. Also, in the case of the Mantel‐Haenszel approach, the pretest DIF statistics tended to have somewhat larger standard errors than the CAT DIF statistics.

Why it matters

OpenAlex reports 13 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

ABSTRACT Simulated data were used to investigate the performance of modified versions of the Mantel‐Haenszel and standardization methods of differential item functioning (DIF) analysis in computer‐adaptive tests (CATs). Each “examinee” received 25 items out of a 75‐item pool. A three‐parameter logistic item response model was assumed, and examinees were matched on expected true scores based on their CAT responses and on estimated item parameters. Both DIF methods performed well. The CAT‐based DIF statistics were highly correlated with DIF statistics based on nonadaptive administration of all 75 pool items and with the true magnitudes of DIF in the simulation. DIF methods were also investigated for “pretest items,” for which item parameter estimates were assumed to be unavailable. The pretest DIF statistics were generally well‐behaved and also had high correlations with the true DIF. The pretest DIF measures, however, tended to be slightly smaller in magnitude than their CAT‐based counterparts. Also, in the case of the Mantel‐Haenszel approach, the pretest DIF statistics tended to have somewhat larger standard errors than the CAT DIF statistics.

Key concepts: Differential item functioning, Statistics, Computerized adaptive testing, Item response theory, Psychology, Mathematics, Psychometrics

Related papers

Back to paper searchBrowse research topicsOriginal source
A SIMULATION STUDY OF METHODS FOR ASSESSING DIFFERENTIAL ITEM FUNCTIONING IN COMPUTER‐ADAPTIVE TESTS — Research Paper | ScholarLens