Linda Darling‐Hammond, Audrey Amrein‐Beardsley, Edward H. Haertel, Jesse Rothstein
Abstract
Practitioners, researchers, and policy makers agree that most current teacher evaluation systems do little to help teachers improve or to support personnel decision making. There's also a growing consensus that evidence of teacher contributions to student learning should be part of teacher evaluation systems, along with evidence about the quality of teacher practices. (VAMs), designed to evaluate student test score gains from one year to the next, are often promoted as tools to accomplish this goal. Value-added models enable researchers to use statistical methods to measure changes in student scores over time while considering student characteristics and other factors often found to influence achievement. In large-scale studies, these methods have proved valuable for looking at factors affecting achievement and measuring the effects of programs or interventions. Using VAMs for individual teacher evaluation is based on the belief that measured achievement gains for a specific teacher's students reflect that teacher's This attribution, however, assumes that student learning is measured well by a given test, is influenced by the teacher alone, and is independent from the growth of classmates and other aspects of the classroom context. None of these assumptions is well supported by current evidence. Most importantly, research reveals that gains in student achievement are influenced by much more than any individual teacher. Others factors include: * School factors such as class sizes, curriculum materials, instructional time, availability of specialists and tutors, and resources for learning (books, computers, science labs, and more); * Home and community supports or challenges; * Individual student needs and abilities, health, and attendance; * Peer culture and achievement; * Prior teachers and schooling, as well as other current teachers; * Differential summer learning loss, which especially affects low-income children; and * The specific tests used, which emphasize some kinds of learning and not others and which rarely measure achievement that is well above or below grade level. measure achievement that is well above or below grade level. However, value-added models don't actually measure most of these factors. VAMs rely on statistical controls for past achievement to parse out the small portion of student gains that is due to other factors, of which the teacher is only one. As a consequence, researchers have documented a number of problems with VAM models as accurate measures of teachers' effectiveness. 1. Value-added models of teacher effectiveness are inconsistent. Researchers have found that teacher effectiveness ratings differ substantially from class to class and from year to year, as well as from one statistical model to the next, as Table 1 shows. TABLE 1. Percent of teachers whose effectiveness rankings change BY 1 OR MORE BY 2 OR MORE BY 3 OR MORE DECILES DECILES DECILES Across 56-80% 12-33% 0-14% models (a) Across 85-100% 54-92% 39-54% courses (b) Across 74-93% 45-63% 19-41% years (b) Note: a Depending on pair of models compared. b Depending on the model used. Source: Newton, Darling-Hammond, Haertel, & Thomas (2010). A study examining data from five school districts found, for example, that of teachers who scored in the bottom 20% of rankings in one year, only 20% to 30% had similar ratings the next year, while 25% to 45% of these teachers moved to the top part of the distribution, scoring well above average. (See Figure 1.) The same was true for those who scored at the top of the distribution in one year: A small minority stayed in the same rating band the following year, while most scores moved to other parts of the distribution. …