The Evidence on Teaching

Getting Beneath the Veil of Effective Schools: Evidence from New York City

Dobbie, W., & Fryer, R. G. · 2013

grade Clongitudinalindependentunclearnumbers spot-checked
Sample
39 surveyed NYC charter schools (29 with usable admissions-lottery records; 16,179 lottery applicants); separate out-of-sample test on 59 charter schools with site-visit data
Population
NYC charter schools, grades 3-8, 2003-04 to 2010-11.
Design
Two-stage. Stage 1 estimates each school's effectiveness two ways: from admissions lotteries (available for only 29 of 39 schools) and from a matching-plus-regression observational model (all 39). Stage 2 regresses those school-level effectiveness estimates on survey-measured school practices — n=39 schools, cross-sectional, practice adoption NOT randomized. The headline numbers in Tables 5-9 all use the OBSERVATIONAL effectiveness estimates, not the lottery ones. The authors are explicit: "our estimates of the relationship between school inputs and school effectiveness are unlikely to be causal given the lack of experimental variation in school inputs. Unobserved factors, such as principal skill, student selection into lotteries, or the endogeneity of school inputs, could drive the correlations reported in the paper." The observational school-effectiveness estimates are themselves biased: regressing lottery on observational estimates gives a slope near 1 (0.946 math, 0.842 ELA), but the observational estimates are compressed (SD 0.099 vs 0.308 for lottery in math).
Key findings
The bridge from "charters work" to "here is what makes them work" — and the famous "five practices explain ~45% of effectiveness variance" number is the paper's weakest specification, not its strongest. That 45% comes from regressing observational (selection-prone) school-effectiveness estimates on practices in 39 schools. Swap in the lottery-identified effectiveness estimates and the same index explains 6.9% of math and 6.0% of ELA variance, with the ELA association no longer significant; the out-of-sample test on 59 schools explains under 10%. The genuinely robust half of the paper is the negative finding: traditional resource inputs — class size, per-pupil spending, teacher certification, advanced degrees — are not merely uncorrelated with effectiveness but significantly wrong-signed, explaining 14-23% of variance in the direction nobody wants.
Genetic confound
Mixed, and the record must not overstate it. The lottery-based school-effectiveness estimates are free of student selection, but they produce the 6-7% R2, not the 45% one. The 45% figure rests on OBSERVATIONAL effectiveness estimates that control for demographics and match cells only — student selection into charters, and into lotteries, remains. Separately, practice adoption is never randomized at either stage, so the practice-effectiveness link is correlational throughout: schools that run the bundle may differ in principal skill, staff quality and intake in unobserved ways.
Replication notes
The paper's own out-of-sample test (59 schools) is by the same authors in the same city and is not an independent replication; its R2 falls below 10%. Angrist, Pathak & Walters find that Massachusetts "No Excuses" charters outperform others, which is directionally convergent but a different design and a different question. Fryer (2014) transplanted the bundle experimentally into nine Houston public schools, which tests the practices rather than this correlation. We have not located an independent replication of the practice-effectiveness correlation itself, and cannot source the claim that none exists — hence unclear rather than unreplicated.
DOI / URL
10.1257/app.5.4.28

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Five-practice index (frequent teacher feedback, data-driven instruction, high-dosage tutoring, extended instructional time, high expectations) vs school effectiveness, OBSERVATIONAL effectiveness estimatesR2 and coefficient per 1 SD of indexR2 = 0.444 (math) and 0.475 (ELA), n=39 schools; 1 SD of index associated with +0.053 SD (SE 0.010) annual math gains and +0.039 SD (SE 0.008) annual ELA gains. This is the "explains ~45% of the variation" headline.standardizedannualnonenot-applicabledomain-skill
The SAME five-practice index against LOTTERY-identified school effectiveness (the only causally identified effectiveness estimates in the paper)R2 and coefficient per 1 SD of indexR2 = 0.069 (math) and 0.060 (ELA), n=29 schools; index coefficient +0.045 (SE 0.020, p<.05) math and +0.031 (SE 0.021, NOT significant) ELA. Footnote 6 states it plainly: "the measure only explains 6.9 percent of the variance in math effectiveness and 6.0 percent of the variation in ELA effectiveness in the lottery sample." Authors attribute the drop to imprecision in the lottery estimates (only 7 of 29 schools have significant lottery effects). The 45% headline is an artifact of using compressed, selection-prone observational effectiveness estimates.standardizedannualnonenot-applicabledomain-skill
Out-of-sample test of the practice index on 59 charter schools with site-visit rather than survey dataR2 and coefficient per 1 SD of indexIndex +0.027 SD (SE 0.009) math and +0.013 SD (SE 0.006) ELA; R2 = 0.102 (math) and 0.061 (ELA) — "the index explains less than 10 percent of the variation in math and ELA." Same sign, half the coefficient, a fifth of the variance. Same authors, same city, so not an independent replication.standardizedannualnonenot-applicabledomain-skill
Traditional resource inputs (class size, per-pupil expenditure, teacher certification, teachers with MA)R2 and coefficient per 1 SD of indexNot zero — significantly WRONG-SIGNED. Index coefficient -0.030 (SE 0.011, p<.01) math and -0.025 (SE 0.010, p<.05) ELA; R2 = 0.140 (math) and 0.228 (ELA). The paper: "An index of the four dichotomous measures explains 14.0 to 22.8 percent of the variance in charter school effectiveness but in the unexpected direction." Schools with >=89% certified teachers have annual math gains 0.041 SD (SE 0.023) lower. On lottery estimates the resource index is -0.035 (ns) math and -0.041 (p<.10) ELA.standardizedannualnonenot-applicabledomain-skill

Cited by