The Vodka is Potent, but the Meat is Rotten1: Evaluating Measurement Equivalence across Contexts
Zachary Elkins, John M. Sides
Abstract
Zachary Elkins, John M. Sides
Abstract
Valid measurement in comparative research depends on the equivalence of constructs and indicators across contexts, but thus far the pace of comparative research has outstripped attention to measurement equivalence. We describe different types and sources of equivalence, as well as methods of diagnosing non-equivalence. We emphasize the need to develop and test theories about threats to equivalence. We derive hypotheses about non-equivalence related to two concepts, democracy and development. Our empirical tests of items measuring these concepts demonstrate non-equivalence, though of differing magnitude and form. We conclude with general guidelines for the prevention, diagnosis, and treatment of non-equivalence. 1 Russian (mis)translation of “The spirit is strong but the flesh is weak, ” from a cross-national survey item (Smith 2003).Scholars of comparative politics continue to venture into more and more jurisdictions across longer stretches of time. Some cross-national datasets contain the universe of independent states since 1800, and cross-national survey projects now include much of the developed and developing worlds. The benefits of this expansion are clear: added cases can produce more variation in variables of interest, provide more powerful tests of extant theory, and illuminate empirical puzzles that lead to new theories. However, comparative inquiry grinds to halt if scholars cannot develop comparable concepts and measures across diverse national and historical contexts. This is true whether the purpose of comparative inquiry is descriptive or inferential. If an observed attribute is a product of not only the underlying construct but also a measurement irregularity particular to time or place, then inferences are compromised. Thus the importance of equivalent measurement, the challenge we address in this paper. 2 Concern about measurement non-equivalence is not new. Most of chapter 2 of The Civic Culture, for example, seeks to justify the equivalence of the authors ’ measures (Almond and Verba 1963; see also
OpenAlex reports 5 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
Valid measurement in comparative research depends on the equivalence of constructs and indicators across contexts, but thus far the pace of comparative research has outstripped attention to measurement equivalence. We describe different types and sources of equivalence, as well as methods of diagnosing non-equivalence. We emphasize the need to develop and test theories about threats to equivalence. We derive hypotheses about non-equivalence related to two concepts, democracy and development. Our empirical tests of items measuring these concepts demonstrate non-equivalence, though of differing magnitude and form. We conclude with general guidelines for the prevention, diagnosis, and treatment of non-equivalence. 1 Russian (mis)translation of “The spirit is strong but the flesh is weak, ” from a cross-national survey item (Smith 2003).Scholars of comparative politics continue to venture into more and more jurisdictions across longer stretches of time. Some cross-national datasets contain the universe of independent states since 1800, and cross-national survey projects now include much of the developed and developing worlds. The benefits of this expansion are clear: added cases can produce more variation in variables of interest, provide more powerful tests of extant theory, and illuminate empirical puzzles that lead to new theories. However, comparative inquiry grinds to halt if scholars cannot develop comparable concepts and measures across diverse national and historical contexts. This is true whether the purpose of comparative inquiry is descriptive or inferential. If an observed attribute is a product of not only the underlying construct but also a measurement irregularity particular to time or place, then inferences are compromised. Thus the importance of equivalent measurement, the challenge we address in this paper. 2 Concern about measurement non-equivalence is not new. Most of chapter 2 of The Civic Culture, for example, seeks to justify the equivalence of the authors ’ measures (Almond and Verba 1963; see also
Key concepts: Equivalence (formal languages), Pace, Mathematics, Logical equivalence, Dynamic and formal equivalence, Econometrics, Psychology, Mathematical economics