Estimating Teacher Impacts on Student Achievement: An Experimental Evaluation
Kane TJ, Staiger DO · 2008
grade Brctdeveloper-involvedreplicatednumbers spot-checked
Sample
78 randomized pairs of classrooms (156 classrooms, 3,194 students, 140 unique teachers), grades 2-5, school years 2003-04 and 2004-05. Pre-experimental value-added was estimated on 1,950 teachers in the same schools using 1999-2000 to 2002-03 data.
Population
Los Angeles Unified elementary classrooms (grades 2-5), within-school teacher pairs. Teachers were National Board certification applicants plus same-grade, same-calendar-track comparison teachers with at least 3 years of experience — novices are excluded by design, so the measured spread of teacher quality is narrower than the district-wide spread.
Design
Teachers were randomly assigned to classrooms within schools, so non-experimental value-added estimated in prior years can be tested against experimental outcomes. This is the one genuinely randomized validation of VA. NBER working paper — the DOI recorded is the NBER series DOI, and NBER working papers are explicitly not peer-reviewed. Sample attrition to note: 151 pairs were randomized, but 42 were dropped for missing prior VA estimates, 12 for administrative reasons, and 19 because principals withdrew on the day of the roster switch, leaving 78 analysed pairs (52% of those randomized). Withdrawal was balanced on whether the roster had in fact been switched (10 switched vs 9 not). Analysis is intention-to-treat; about 15% of students had a different teacher by testing time, and roughly 10% of students were missing follow-up scores, but Table 5 shows neither attrition nor teacher-switching correlated with pre-experimental VA. Independence downgraded to developer-involved: Kane and Staiger are the leading methodologists and policy advocates of the value-added estimation being validated here (the paper cites their own Gordon, Kane & Staiger 2006 as the source of its approach), so this is a self-validation, not a third-party audit. Mitigating: an external advisory board of Hanushek, Goldhaber and Ballou advised on study design.
Key findings
The only experimental test of whether value-added measures what it claims. Randomly assigned classrooms performed about as prior value-added predicted, which is the strongest single reason to treat multi-year VA as a real measurement rather than a sorting artefact. Three things the full text adds that the record previously oversold or omitted. (1) The unbiasedness result is specification-dependent, not general: only specifications conditioning on prior test scores came out unbiased. Raw uncontrolled value-added overstates teacher differences by about half (coefficient ~0.4-0.5, significantly below 1) and student-fixed-effects value-added understates them by about half (~1.9-2.1, significantly above 1). "VA is unbiased" is true of the right VA only. (2) Accuracy is moderate, not high — the best model explains about 53% of true teacher-level variation. (3) The measured teacher effect fades by roughly 50% per year: at t+2 the math coefficients (0.03-0.07) are indistinguishable from zero. The authors flag this themselves as "raising important concerns about whether unbiased estimates of the short-term teacher impact are misleading in terms of the long-term impacts of a teacher." This paper validates VA as a measure of a teacher's effect on this year's test score, and simultaneously documents that most of that effect is gone within two years.
Genetic confound
Random assignment of teachers to classrooms; student composition cannot differ systematically between arms.
Replication notes
Agrees with the Chetty-Friedman-Rockoff switching validation; the two together are why VA survived the Rothstein critique. Two caveats on how independent that record is: CFR is quasi-experimental (teacher-switching), not a replication of this randomized design, and the subsequent random-assignment validation (the MET project) again has Kane and Staiger as authors. The paper's own conclusion asks for exactly this: "while our results need to be replicated elsewhere, these findings from Los Angeles schools suggest that recent concerns about bias in teacher value added estimates may be overstated in practice."
DOI / URL
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Non-experimental VA (levels, prior-score + peer controls) predicting experimental outcomes | forecast coefficient (1.0 = unbiased) | math 0.852 (SE 0.177) and 0.905 (0.180) with school FE; ELA 0.987 (0.277) and 1.089 (0.289) with school FE. All significantly different from zero and none statistically distinguishable from one — i.e. unbiased (Table 6). Gains specifications behave the same: math 0.794-0.865, ELA 0.765-0.886 | standardized | end of the experimental school year | active-alternative | end-of-treatment | domain-skill |
| Non-experimental VA with NO student/peer controls (raw mean scores) | forecast coefficient | math 0.511 (SE 0.108), ELA 0.418 (0.155) — significantly greater than zero but also significantly LESS than one, so uncontrolled value-added overstates true teacher differences by roughly half (Table 6). Consistent with tracking on unmeasured student characteristics | standardized | end of the experimental school year | active-alternative | end-of-treatment | domain-skill |
| Non-experimental VA estimated with student fixed effects | forecast coefficient | math 1.859 (SE 0.470), ELA 2.144 (0.635) — significantly greater than one, so the increasingly popular student-fixed-effects specification UNDERSTATES true teacher variation by about half (Table 6). The authors link this to Rothstein's (2008) point that student FE is biased under dynamic tracking | standardized | end of the experimental school year | active-alternative | end-of-treatment | domain-skill |
| How much of true teacher-effect variation VA actually captures | R2 / share of maximum attainable R2 | best specification R2 = 0.226 (math) and 0.169 (ELA); against a maximum attainable R2 of 0.425 given sampling noise, that is about 53% of teacher-level variation in each subject. VA is a real but roughly half-accurate measurement, not a precise one | standardized | end of the experimental school year | none | end-of-treatment | domain-skill |
| Magnitude of teacher effects (pre-experimental, preferred specification) | SD of teacher effects in student-level SD units | 0.219 in math and 0.175 in ELA with prior-score + peer controls and school FE; 0.448 / 0.453 with no controls (Table 3). Table 10 finds a similar 0.16-0.19 (math) and 0.13-0.16 (ELA) spread in New York and Boston, and the correlation between teacher effect and student baseline achievement is only 0.04-0.12 in all three cities | standardized | pre-experimental years 1999-2003 | none | not-applicable | domain-skill |
| Persistence of the randomly assigned teacher's effect, one year later | forecast coefficient at t+1 | math 0.359 and 0.390 (both p<0.10 only), ELA 0.477 and 0.569 (not significant) — roughly half the first-year coefficient (Table 6). The authors' headline fade-out estimate is about 50% per year | standardized | one year after the experimental year | active-alternative | 1-2yr | domain-skill |
| Persistence of the randomly assigned teacher's effect, two years later | forecast coefficient at t+2 | math 0.034 and 0.07 — statistically indistinguishable from zero; ELA 0.476 and 0.541 (the latter p<0.10). Only about 25% of the original effect remains, and in math nothing measurable does (Table 6). Compensatory teacher assignment was ruled out as the explanation (correlation between experimental-year and next-year teacher VA = -0.01) | standardized | two years after the experimental year | active-alternative | over-2yr | domain-skill |
Cited by
- Teacher quality — selection over credentials and workshopsstrong supportconf: highgc: low