Class-size reduction
Smaller classes buy real K-1 gains that fade on tests yet persist in attainment — at roughly triple tutoring's cost, and diluted to nothing when scaled fast.
mixedconf: highgc: lowclass-size · ages 4–18 · structure
Real but expensive and front-loaded: the STAR RCT shows ~0.2 SD in the first K-1 exposure year (2x for disadvantaged students), fading on tests but persisting in college attainment (+1.8 to +2.7pp). The founding international RD replicated to a precise ZERO two decades later; at scale (California) mass hiring diluted teacher quality and ate the gains. Cost per 0.1 SD is ~3x high-dosage tutoring's — but class size holds the database's only registry-proven long-run wage evidence, so the two ledgers disagree.
Default to ~22-25 students with excellent teachers and spend the marginal dollars on tutoring, coaching, and teacher selection. Consider ≤17 only in K-1 — and never by lowering the teacher hiring bar (the California lesson: class size and teacher quality compete for the same dollars and labor pool).
Who this applies to
Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.
Verdict
Class-size reduction is the most intuitive structural lever and among the most expensive per unit of learning. The grade-A core (STAR): assignment to a ~15 (vs ~22) class in K-3 buys ~0.2 SD in the first exposure year — roughly double for Black and free-lunch students — then little more; scores fade by grade 8, but college attainment durably rises (+1.8pp college attendance at 20; +2.7pp enrollment and +1.6pp degrees,
2x for Black students and +7.3pp in the poorest third of schools). The age-25-27 earnings estimate is underpowered, not null — its confidence interval contains exactly the gain the score effects predict.
Three complications keep this at mixed:
- Context-dependence: the founding Israeli RD's ~0.25 SD became a precisely estimated zero in the same country two decades later — class-size effects are not a constant of nature.
- The scale trap: California's statewide CSR forced mass hiring that diluted teacher quality most in poor schools — the equilibrium response ate the gains.
- The two ledgers: on score-per-dollar, CSR (~$1-1.75k per 0.1 SD/yr) loses ~3x to high-dosage tutoring; but on long-run registry outcomes, class size has proof tutoring lacks — Sweden: +0.7% adult wages per pupil removed (IRR 18.6%), and the database's only credible case of an instructional lever durably moving IQ-type general-ability measures (small, at 18, for low-income sons). The direct wage effect ran 3-6x what score-imputation predicts, so per-score-dollar rankings may understate CSR.
Hereditarian-lens assessment
Risk: low (RCT + RD designs). The ATI ledger is genuinely two-sided and worth stating precisely: STAR's score gains favored disadvantaged students; Sweden's age-13 skill gains were income-neutral, its age-18 general-ability gains went to low-income sons, and its wage gains concentrated among the advantaged. No tidy equity story survives — which is itself the finding.
Practical guidance
- Buy teacher quality first (a 1 SD better teacher ≈ the effect of a 10-student reduction at a fraction of the cost), then tutoring, then — if money remains — small K-1 classes.
- If reducing, do it K-1 only, and only at a hiring-quality-neutral pace.
- Never judge the lever on first-year test scores alone: this is the canonical fadeout-with-attainment-persistence case.
- grade AProject STAR and Krueger's reanalysis (Experimental Estimates of Education Production Functions)Krueger, A. B. (Tennessee STAR RCT, 1985-89) · 1999 · rct
- grade AHow Does Your Kindergarten Classroom Affect Your Earnings? Evidence from Project StarChetty R, Friedman JN, Hilger N, Saez E, Schanzenbach DW, Yagan D · 2011 · rct
- grade AExperimental Evidence on the Effect of Childhood Investments on Postsecondary Attainment and Degree CompletionDynarski S, Hyman J, Schanzenbach DW · 2013 · rct
- grade BMaimonides' Rule and its Redux: the class-size RD replication failureAngrist, J., & Lavy, V. (1999); Angrist, J., Lavy, V., Leder-Luis, J., & Shany, A. (2019) · 2019 · replication
- grade BLong-Term Effects of Class Size (Swedish RD, grades 4-6)Fredriksson, P., Ockert, B., & Oosterbeek, H. · 2013 · natural-experiment
- grade BCalifornia's 1996 class-size reduction at scale (CSR Consortium; Jepsen & Rivkin)CSR Research Consortium (2002); Jepsen, C., & Rivkin, S. (2009) · 2009 · quasi-experiment
- grade BThe Promise of Tutoring for PreK-12 Learning: A Systematic Review and Meta-Analysis of the Experimental EvidenceNickow, A., Oreopoulos, P., & Quan, V. · 2024 · meta-analysis
Related decisions
- Does educational technology raise learning — CAI, adaptive software, devices, screens, and AI tutors?mixedconf: mediumgc: low
- Teacher quality — selection over credentials and workshopsstrong supportconf: highgc: low
- Tutoring — the honest effect, the Bloom 2-sigma myth, and what survives scalestrong supportconf: highgc: low