The Evidence on Teaching

Class-size reduction

Smaller classes buy real K-1 gains that fade on tests yet persist in attainment — at roughly triple tutoring's cost, and diluted to nothing when scaled fast.

mixedconf: highgc: low

class-size · ages 418 · structure

Effect summary

Real but expensive and front-loaded: the STAR RCT shows ~0.2 SD in the first K-1 exposure year (2x for disadvantaged students), fading on tests but persisting in college attainment (+1.8 to +2.7pp). The founding international RD replicated to a precise ZERO two decades later; at scale (California) mass hiring diluted teacher quality and ate the gains. Cost per 0.1 SD is ~3x high-dosage tutoring's — but class size holds the database's only registry-proven long-run wage evidence, so the two ledgers disagree.

Practical takeaway

Default to ~22-25 students with excellent teachers and spend the marginal dollars on tutoring, coaching, and teacher selection. Consider ≤17 only in K-1 — and never by lowering the teacher hiring bar (the California lesson: class size and teacher quality compete for the same dollars and labor pool).

Who this applies to

Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.

Verdict

Class-size reduction is the most intuitive structural lever and among the most expensive per unit of learning. The grade-A core (STAR): assignment to a ~15 (vs ~22) class in K-3 buys ~0.2 SD in the first exposure year — roughly double for Black and free-lunch students — then little more; scores fade by grade 8, but college attainment durably rises (+1.8pp college attendance at 20; +2.7pp enrollment and +1.6pp degrees,

2x for Black students and +7.3pp in the poorest third of schools). The age-25-27 earnings estimate is underpowered, not null — its confidence interval contains exactly the gain the score effects predict.

Three complications keep this at mixed:

  • Context-dependence: the founding Israeli RD's ~0.25 SD became a precisely estimated zero in the same country two decades later — class-size effects are not a constant of nature.
  • The scale trap: California's statewide CSR forced mass hiring that diluted teacher quality most in poor schoolsthe equilibrium response ate the gains.
  • The two ledgers: on score-per-dollar, CSR (~$1-1.75k per 0.1 SD/yr) loses ~3x to high-dosage tutoring; but on long-run registry outcomes, class size has proof tutoring lacks — Sweden: +0.7% adult wages per pupil removed (IRR 18.6%), and the database's only credible case of an instructional lever durably moving IQ-type general-ability measures (small, at 18, for low-income sons). The direct wage effect ran 3-6x what score-imputation predicts, so per-score-dollar rankings may understate CSR.

Hereditarian-lens assessment

Risk: low (RCT + RD designs). The ATI ledger is genuinely two-sided and worth stating precisely: STAR's score gains favored disadvantaged students; Sweden's age-13 skill gains were income-neutral, its age-18 general-ability gains went to low-income sons, and its wage gains concentrated among the advantaged. No tidy equity story survives — which is itself the finding.

Practical guidance

  • Buy teacher quality first (a 1 SD better teacher ≈ the effect of a 10-student reduction at a fraction of the cost), then tutoring, then — if money remains — small K-1 classes.
  • If reducing, do it K-1 only, and only at a hiring-quality-neutral pace.
  • Never judge the lever on first-year test scores alone: this is the canonical fadeout-with-attainment-persistence case.

Evidence (7 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore