1985•ETS Research Report SeriesRequires access

AN ASSESSMENT OF THE RELATIONSHIP BETWEEN THE ASSUMPTION OF UNIDIMENSIONALITY AND THE QUALITY OF IRT TRUE‐SCORE EQUATING1,2,3

Linda L. Cook, Neil J. Dorans, Daniel R. Eignor, Nancy S. Petersen

Open publisher page 20 citations

Abstract

ABSTRACT A strong assumption made by most commonly used item response theory (IRT) models is that the data are unidimensional, i.e., statistical dependence among item scores can be explained by a single ability dimension. One of the major practical applications of item response theory models to testing has been in the area of score equating. This research assesses the relationship between violations of the assumption of unidimensionality and the quality of IRT true‐score equating. First‐order and second‐order factor analyses were conducted on correlation matrices among item parcels. The item parcels were constructed to yield correlation matrices that were amenable to linear factor analyses. The first‐order analyses were employed to assess the effective dimensionality of the item parcel data. Second‐order analyses were employed to test meaningful hypotheses about the structure of the data, hypotheses that were suspected to be pertinent to the quality of equating results. Parcels were constructed for three SAT‐verbal forms and three forms of the Mathematics Level II Achievement test. The quality of IRT true‐score equating was assessed by score scale drift. Scale drift is said to have occurred if the results of equating test form D directly to test form A is not the same as that obtained by equating test form D to test form A through intervening forms B and C. Scale drift was less evident in the Mathematics Level II chain of equatings than it was in the SAT‐verbal chain. The factor analyses uncovered structural similarities and differences across test forms that were consistent with the scale drift results. The dimensionality analyses revealed that the Mathematics Level II item parcels were more nearly unidimensional than the SAT‐verbal item parcels. In addition, the dimensionality analyses revealed that one SAT‐verbal test form and one Mathematics Level II form were each less parallel to the other two forms in their respective equating chains than these other forms were to each other. Refinements in the dimensionality methodology and a more systematic dimensionality assessment are logical extensions of the present research.

About this research paper

What this paper is about

ABSTRACT A strong assumption made by most commonly used item response theory (IRT) models is that the data are unidimensional, i.e., statistical dependence among item scores can be explained by a single ability dimension. One of the major practical applications of item response theory models to testing has been in the area of score equating. This research assesses the relationship between violations of the assumption of unidimensionality and the quality of IRT true‐score equating. First‐order and second‐order factor analyses were conducted on correlation matrices among item parcels. The item parcels were constructed to yield correlation matrices that were amenable to linear factor analyses. The first‐order analyses were employed to assess the effective dimensionality of the item parcel data. Second‐order analyses were employed to test meaningful hypotheses about the structure of the data, hypotheses that were suspected to be pertinent to the quality of equating results. Parcels were constructed for three SAT‐verbal forms and three forms of the Mathematics Level II Achievement test. The quality of IRT true‐score equating was assessed by score scale drift. Scale drift is said to have occurred if the results of equating test form D directly to test form A is not the same as that obtained by equating test form D to test form A through intervening forms B and C. Scale drift was less evident in the Mathematics Level II chain of equatings than it was in the SAT‐verbal chain. The factor analyses uncovered structural similarities and differences across test forms that were consistent with the scale drift results. The dimensionality analyses revealed that the Mathematics Level II item parcels were more nearly unidimensional than the SAT‐verbal item parcels. In addition, the dimensionality analyses revealed that one SAT‐verbal test form and one Mathematics Level II form were each less parallel to the other two forms in their respective equating chains than these other forms were to each other. Refinements in the dimensionality methodology and a more systematic dimensionality assessment are logical extensions of the present research.

Why it matters

OpenAlex reports 20 citations for this work. Citation counts describe recorded attention and do not establish research quality.

Key contribution

A contribution statement is not available in the OpenAlex record.

Method / approach

Method details are not available in the OpenAlex metadata.

Main findings

Findings are not separately available in the OpenAlex metadata.

Limitations

Limitations are not available in the OpenAlex metadata.

Applications

Application details are not available in the OpenAlex metadata.

Available abstract

ABSTRACT A strong assumption made by most commonly used item response theory (IRT) models is that the data are unidimensional, i.e., statistical dependence among item scores can be explained by a single ability dimension. One of the major practical applications of item response theory models to testing has been in the area of score equating. This research assesses the relationship between violations of the assumption of unidimensionality and the quality of IRT true‐score equating. First‐order and second‐order factor analyses were conducted on correlation matrices among item parcels. The item parcels were constructed to yield correlation matrices that were amenable to linear factor analyses. The first‐order analyses were employed to assess the effective dimensionality of the item parcel data. Second‐order analyses were employed to test meaningful hypotheses about the structure of the data, hypotheses that were suspected to be pertinent to the quality of equating results. Parcels were constructed for three SAT‐verbal forms and three forms of the Mathematics Level II Achievement test. The quality of IRT true‐score equating was assessed by score scale drift. Scale drift is said to have occurred if the results of equating test form D directly to test form A is not the same as that obtained by equating test form D to test form A through intervening forms B and C. Scale drift was less evident in the Mathematics Level II chain of equatings than it was in the SAT‐verbal chain. The factor analyses uncovered structural similarities and differences across test forms that were consistent with the scale drift results. The dimensionality analyses revealed that the Mathematics Level II item parcels were more nearly unidimensional than the SAT‐verbal item parcels. In addition, the dimensionality analyses revealed that one SAT‐verbal test form and one Mathematics Level II form were each less parallel to the other two forms in their respective equating chains than these other forms were to each other. Refinements in the dimensionality methodology and a more systematic dimensionality assessment are logical extensions of the present research.

Key concepts: Equating, Item response theory, Statistics, Test score, Test (biology), Scale (ratio), Curse of dimensionality, Econometrics

Related papers

Back to paper searchBrowse research topicsOriginal source
AN ASSESSMENT OF THE RELATIONSHIP BETWEEN THE ASSUMPTION OF UNIDIMENSIONALITY AND THE QUALITY OF IRT TRUE‐SCORE EQUATING1,2,3 — Research Paper | ScholarLens