The Evidence on Teaching

Perry Preschool and Abecedarian — what the famous studies actually establish

Perry and Abecedarian: n=123 and n=111, IQ gains gone by adolescence, attainment effects real but multiplicity-fragile. Too thin to carry the policy built on them.

mixedconf: mediumgc: low

early-childhood · ages 45 · structure

Effect summary

The two studies carrying most of early-childhood policy rest on n=123 and n=111, single sites, 1960s-70s, with a no-preschool counterfactual that no longer exists. Large end-of-treatment IQ gains vanished (Perry: 95 vs 84 at program end → 80 vs 81 at age 14). Attainment and crime effects persisted to midlife but are multiplicity-fragile: FDR correction across the full Perry outcome family leaves exactly TWO surviving effects. The famous 7-13% returns are model extrapolations, are male-driven, and collapse with a realistic discount rate.

Practical takeaway

Do not use Perry or Abecedarian to justify a modern preschool budget. They tested near-total deprivation against intensive intervention in the 1960s-70s; the comparison no longer exists. Their real lesson is the outcome pattern — ability reverted, life trajectory didn't — not their headline effect sizes or ROI figures.

Who this applies to

Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.

Verdict

Perry Preschool and Abecedarian are the most cited studies in early-childhood policy and among the thinnest evidence in this database. Both are real randomised trials with genuine long-run follow-up — that much deserves respect, and the effects are not fabricated. But four facts should govern how much weight they carry:

  1. They are tiny and singular. Perry n=123, Abecedarian n=111, one site each, 100% African-American, low-SES, 1962-77. Abecedarian's headline "BA+ 23% vs 6%" is twelve people versus three.
  2. The cognitive gains did not last. Perry's IQ advantage went 95 vs 84 at program end → ns by age 8 → 80 vs 81 at 14. Abecedarian's ~+4.4 points at 21 is the only durable IQ effect from a preschool RCT — and it sits at the ceiling that whole-childhood adoption produces, not above it.
  3. Multiplicity is brutal. Applying false-discovery correction across the full Perry outcome family leaves exactly two effects at q=.05: early male IQ and female high-school graduation. Six of eight outcomes clearing naive p<.10 fail to replicate in the sibling trials.
  4. The counterfactual is extinct. Perry's controls got nothing. Today's controls attend other preschool. The contrast that generated these effects cannot be recreated.

Hence mixed: real effects, genuinely durable on some life outcomes, but far too fragile and too era-specific to carry the policy weight placed on them.

What the evidence shows

Source Design Grade Key finding
Schweinhart 2005 (Perry, age 40) RCT, n=123 C IQ +11 → 0; HS graduation 65% vs 45% (females 84 vs 31; males ns); 5+ arrests 36% vs 55%
Anderson 2008 multiplicity reanalysis B Only TWO Perry effects survive FDR at q=.05; pooled adult ES females .27 (.09), males −.05
Campbell 2012 (Abecedarian, age 30) RCT, n=111 B BA+ 23.1% vs 6.1% (12 vs 3 people); crime NULL; income-to-needs NULL
Heckman 2010 (ROI) economic model D 7-10% — itself a downward revision of prior 16-17% claims
Heckman & Karapakula 2019 age-55 follow-up B Effects survive worst-case randomization inference — but strongest for MALES

Perry's randomization was compromised in documented ways: siblings were force-assigned to the same arm, and the number of treatment-to-control transfers is reported inconsistently as 2, 3-6, or 5 across HighScope's own publications. Baseline imbalance was substantial (mother employed at entry: 9% treatment vs 31% control). This is why the topic grades Schweinhart C rather than B.

An unresolved conflict worth flagging: Anderson finds the durable effects are female-driven (adult pooled: females .27, males −.05). Heckman & Karapakula find them male-strongest at age 55. Both cannot be the headline, and the field has not reconciled them.

On the ROI figures. The "7-10% return" (Perry) and "13.7%" (Abecedarian/CARE) are model outputs, not measured effects — graded D on that basis. They require extrapolating earnings decades past the last observation. Three specifics a founder should know: Heckman's own paper revised earlier 16-17% claims downward; the benefit-cost ratio swings by 2.5× purely on how murder is priced, and there are four murders in the entire record; and the Abecedarian female return (10%, SE 8%) is not statistically distinguishable from zero. Benefit-cost also falls from 17.4 to 2.9 depending only on the discount rate chosen.

Hereditarian-lens assessment

Risk: low for the causal claims (these are RCTs). This is the domain where the lens earns its keep most clearly, and where it calibrates rather than debunks:

  • The pattern is exactly what heritability predicts. Measured ability reverted to a genetically-anchored trajectory; attainment and behaviour — which load far more on shared environment — persisted. That is not a puzzle; it is the architecture.
  • Abecedarian's durable IQ gain is at the environmental ceiling, not above it. Its ~+4.4 points at 21 is essentially identical to the +4.41 points that a whole childhood of adoption into a better home produces (Kendler 2015). Roughly 3-5 points looks like the general ceiling for environmental manipulation of ability.
  • Bloom-style "anyone to the top 2%" equalisation claims are refuted by the same data that support the real effects.
  • One caveat against over-reading: Spitz showed Abecedarian's treatment group already led on the Bayley at six months by roughly the eventual gap — a baseline-imbalance explanation for its "persistent" IQ effect that has never been cleanly resolved.

Practical guidance

  • Don't cite these studies to justify a modern preschool programme. The counterfactual is gone. Use the modern at-scale evidence instead (see preschool at scale).
  • Don't cite the ROI figures at all without stating that they are projections, male-driven, and discount-rate-dependent.
  • Do take the outcome-class lesson: if early intervention pays off, it pays off through life trajectory, not through lasting ability gains. Design and measure accordingly.
  • If you are serving genuinely deprived children, note that both programmes bundled health care and nutrition with education — Abecedarian included two daily meals — so they are not clean tests of "preschool."

Open questions

  • Female-driven (Anderson) versus male-driven (Heckman-Karapakula) — unreconciled.
  • Whether Abecedarian's IQ effect is treatment or six-month baseline imbalance.
  • Whether any of this generalises to a counterfactual of ordinary modern preschool. The direct test — Tennessee — went the other way.

Evidence (8 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore