The Evidence on Teaching

Preschool at scale — Head Start, state pre-K, and what universal provision delivers

At scale, pre-K doesn't durably raise test scores — Tennessee went negative — yet Boston shows real attainment gains beside a test-score zero. It buys trajectory, not ability.

mixedconf: highgc: low

early-childhood · ages 45 · structure

Effect summary

Scaled pre-K does not durably raise achievement. The national Head Start RCT is null by grade 3; Tennessee's lottery-randomised state program went NEGATIVE by grade 6 (−0.13 to −0.18 SD, more special-ed, more discipline); Quebec's universal expansion produced persistent negative non-cognitive effects. Yet Boston's lottery shows a precise test-score ZERO alongside real attainment gains (+8.3pp on-time college). Head Start beats home care (~+0.37 SD) but is indistinguishable from the other preschools it displaces.

Practical takeaway

Do not expect a preschool programme to raise later achievement — the best randomised evidence says it won't, and can backfire. The defensible case rests on attainment/behaviour and on serving children whose alternative is genuinely home care. Check your actual counterfactual before spending: if families already have decent preschool, the measured return is near zero.

Who this applies to

Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.

Verdict

This is the evidence that should actually govern a preschool decision, and it is uncomfortable for every side of the debate.

Scaled pre-K does not durably raise achievement. Three of the four best designs say so, and one says worse than that:

  • Head Start (national RCT, n=4,667): small end-of-year gains in narrow language/literacy, gone by kindergarten for 4-year-olds, and by grade 3 essentially nothing survives multiplicity correction. Both arms remain ~0.5 SD below national norms.
  • Tennessee VPK (lottery RCT, n=2,990): positive at end of pre-K, gone by kindergarten, negative by grade 6 (−0.128 ELA, −0.178 math, −0.132 science), with more special-education placements (11.7% vs 8.4%) and more disciplinary infractions.
  • Quebec (universal expansion): persistent negative non-cognitive effects into school age, worse adult health and higher crime.
  • Boston (lottery): a precisely estimated zero on achievement across grades 3-10 — ruling out effects above ~0.12 SD.

And yet the attainment story is real. Boston, on the same precise zero, shows +8.3pp on-time college enrolment and +6.0pp high-school graduation. Three independent designs (Boston lottery, the Head Start rollout, the Ludwig-Miller RD) find attainment gains without test-score gains. This is the same decoupling seen in Perry — and it is now established in a modern, at-scale, lottery-based programme, which is what makes it credible.

So: mixed, at high confidence. Not "preschool doesn't work" — but "preschool does not do the thing it is usually sold as doing."

What the evidence shows

Source Design Grade Key finding
Durkin 2022 (TN-VPK) lottery RCT, n=2,990 A End of pre-K +0.24 → K +0.02 → grade 6 −0.13 to −0.18; IEPs 11.7% vs 8.4%
Gray-Lobe 2023 (Boston) lottery A MCAS grades 3-10 +0.005/+0.029 (precise zero); on-time college +8.3pp; HS grad +6.0pp
Puma 2012 (HSIS grade 3) national RCT A One favourable result (Reading +0.11) failing multiplicity; one unfavourable; both arms ~0.5 SD below norms
Kline & Walters 2016 HSIS decomposition B vs home care +0.37 SD; vs other centre ~ZERO
Pages 2020 replication C Deming's +0.23 → +0.17 in his cohorts, −0.15 in later ones, ~0 combined
Watts 2023 (NC Pre-K) county-dose TWFE, n=1.2M C The dissent: ~+0.20 SD growing through grade 5 — but the weakest design, internally contradictory

The counterfactual is the ballgame. Kline & Walters is the most clarifying paper here: Head Start beats staying home (+0.37 SD) but is statistically indistinguishable from the other preschools it displaces. As alternative childcare has become near-universal, the measurable return to adding a programme has shrunk toward zero. Much of what looks like "fadeout at scale" is really this.

But substitution cannot explain everything. Boston's compliers came ~62% from other preschool (33% Head Start, 29% private) — which should have made its effects smaller — yet Boston is where the attainment gains appear. Tennessee's controls were 63% home care — which should have made its effects larger — and Tennessee went negative. The counterfactual story runs the wrong way here, and the divergence between Boston and Tennessee remains genuinely unexplained. Structural quality doesn't explain it either: Tennessee met 9 of 10 NIEER quality benchmarks.

On the "no scaled programme shows persistent gains" claim — that is false as stated, and the adversarial review caught it. North Carolina reports gains that grow through grade 5. But it is a population-level, county-funding-dose design with an IV-rescaled effect size, its two constituent papers contradict each other on special-ed and retention, and it is outvoted by two cleaner randomised designs. Record it as the live dissent, not the answer.

Head Start's long-run attainment claim has failed replication. Deming's +0.23 SD becomes −0.15 in later cohorts and ~zero combined. Anyone citing Deming without Pages is citing half the evidence.

Hereditarian-lens assessment

Risk: low — the verdict rests on lotteries and RCTs. The lens contributes three things:

  • Head Start's effect is floor remediation, not ability gain: >1 SD at the bottom of the distribution, ~0.25 SD at the top. Schools remediate the same floor for everyone within a couple of years, which is why the mean effect vanishes.
  • The attainment/ability split is architectural. Shared environment for ability collapses to ~0.10 by adulthood; shared environment for attainment stays much larger. The observed pattern (test scores revert, graduation doesn't) is exactly what that predicts — see fadeout & persistence.
  • No IQ test was ever administered in HSIS. PPVT is the g-loaded proxy, and notably it showed the smallest effect of any measure (+0.09) while parent-reported outcomes showed the largest (+0.31). That ordering is diagnostic.

Practical guidance

  • Audit your counterfactual first. If the families you serve already have adequate preschool, expect ~zero measurable academic return. If their alternative is genuinely home care, expect ~+0.37 SD at end of treatment — and expect most of it to fade.
  • Do not promise achievement gains. The best randomised evidence does not support them, and Tennessee shows a scaled, structurally high-quality programme can end up behind.
  • If you run one, aim it at trajectory — school readiness that keeps children out of special-education placement and retention, and the behavioural/attainment channel that Boston moved.
  • Watch for harm at scale. Tennessee and Quebec are both large, well-identified, and negative. This is not a "worst case it does nothing" intervention.
  • The health channel is real and undersold — Head Start's most robust short-run effect was dental care (+15-16pp), and the Ludwig-Miller RD finds substantial child-mortality reductions.

Open questions

  • Why Boston and Tennessee diverge is the central unanswered question in early childhood. Counterfactual, structural quality, and era all fail to explain it.
  • Whether Boston's attainment gains replicate anywhere else with lottery-grade identification.
  • Whether North Carolina's growing gains survive a stronger design.
  • Boston's college graduation effect (+5.2pp) is not significant — the durable-attainment claim rests on enrolment, which is weaker than usually reported.

Evidence (17 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore