The Evidence on Teaching

Mastery learning (teach → test → reteach to criterion → advance)

Mastery learning moves tests of what it taught (~0.25) and barely moves independent measures (~0.05) — and time-to-mastery gaps widen, converting ability differences into time.

mixedconf: highgc: low

mastery · ages 618 · method

Effect summary

The cleanest measure-inflation case in education: ~+0.2-0.3 on tests of the objectives taught, ~+0.0-0.1 on independent standardized measures at equal time (Slavin's median +0.04; the pro-mastery meta's own within-study standardized effects 0.04-0.09; modern at-scale RCT +0.07). Gains concentrate in low-aptitude students; time-to-mastery inequality GROWS (2.5:1 → 4.2:1), so mastery converts ability differences into time inequality rather than reducing them.

Practical takeaway

Use the cheap ingredients of mastery — frequent checks, feedback, targeted reteaching of what a student actually missed — without the strict gate-everyone-to-criterion machinery, which taxes fast learners' time for near-zero measurable benefit on broad measures. If you do gate, gate the FOUNDATIONS (decoding, math facts) where coverage loss costs least.

Who this applies to

Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.

Verdict

Mastery learning — don't advance until the current unit is mastered, retest and recorrect until criterion — sounds like the obviously right architecture, and Bloom claimed ~1 sigma for it. The evidence says something much more specific: mastery learning reliably improves performance on tests of exactly the objectives it gated, and does approximately nothing on independent standardized measures at equal instructional time. This is the sharpest researcher-vs-standardized measure gap in the entire database — sharper than the generic 2× rule — and it was established by Slavin in 1987 and confirmed by the pro-mastery meta's own within-study data and by the modern at-scale RCTs. Both camps' numbers agree; they only dispute which measure is "fair."

The second finding matters as much for a school builder: mastery pacing does not shrink ability differences — it converts them into time inequality. Four years of time-to-mastery records show the slowest:fastest ratio growing from 2.5:1 to 4.2:1. Group-based mastery is a "Robin Hood" redistribution of teacher time toward the bottom of the class, with fast learners paying in waiting.

What the evidence shows

Source Design Grade Key effect
Slavin 1987 best-evidence synthesis B Standardized tests, equal time: median +0.04 (0/7 significant); experimenter-made: +0.24
Kulik 1990 108-study meta (pro-mastery) C Headline 0.52 — but own within-study standardized effects 0.04/0.09/0.07/−0.05; low-aptitude 0.61 vs high 0.40
Arlin 1983-84 experiment + 4-yr records C Retention per hour favored controls (−1.17); time-inequality grows 2.5:1 → 4.2:1
EEF Maths Mastery 2 cluster RCTs (10,114 pupils) B Pooled +0.073 on independent standardized tests
Cognitive Tutor Algebra 73-school RCT (adaptive mastery software) B Year 1 ~0; year 2 +0.21 (fragile); companion geometry RCT negative

The within-study comparisons are decisive. In every study that used both test types, the same program produced substantial effects on the experimenter's aligned test and trivial effects on the standardized one (the starkest: +0.64 vs +0.04 in the same study). The advocate meta-analyses (Guskey's ~0.78) got their numbers by counting formative-retake scores as outcomes (mastery students got retakes, controls didn't) and using class-mean SDs. Extra-corrective-time gains (+0.31) vanish within 4–12 weeks.

The time accounting kills the strong claim. Bloom-lab studies gave mastery arms 20–33% extra time; when Arlin held a fair clock, huge aligned-test effects became negative learning per hour. Individualized-pacing data show the slowest students need 200–600% more time — a cost that never converges, contra Bloom's prediction that mastering prerequisites shrinks differences.

Hereditarian-lens assessment

Risk: low for the verdict (experimental base). This topic is where the lens is most clarifying: Bloom's model implicitly claims environment (time + correction) can equalize learning, and the falling aptitude-achievement correlation in his lab (r=.60 → .25 under tutoring/mastery) was read as ability differences "washing out." The genetics review's verdict: that is a restricted-range/ ceiling artifact on aligned tests of 3-week units — on cumulative material over years, aptitude variance reasserts, heritability of achievement rises with age, and Arlin's growing time ratios are exactly what heritable differences in learning rate predict. Mastery learning is best understood as holding achievement constant and letting time absorb the ability distribution. That can be a legitimate choice for foundational skills — but it is a trade-off, not an equalization, and the fast-learner tax (gains concentrate at the bottom, near-zero at the top, with waiting imposed) is a real cost the advocacy literature omits.

What critics say / limits

  • Anderson & Burns argued Slavin's standardized-test criterion is itself a values choice ("mastery teaches what it teaches; broad tests measure what wasn't taught"). True — but for a school answerable to external outcomes, the broad measure IS the relevant one, and the coverage-vs-mastery dilemma (material not covered is material not learned) is precisely the point.
  • Modern caveats cut both ways: the EEF trials omitted strict gating and CTAI's gating was widely bypassed by teachers — so no modern at-scale RCT cleanly tests enforced mastery thresholds. Enforcement in real classrooms appears close to impossible, which is itself a finding.

Practical guidance

  • Keep the diagnostic core, drop the gate: frequent low-stakes checks (retrieval practice), immediate feedback, and targeted reteaching of the specific gaps — these carry mastery's real ingredients without the whole-class criterion machinery.
  • If you gate anything, gate foundations: decoding, number facts, prerequisite procedures — domains where a gap genuinely blocks later learning and where coverage loss is cheapest.
  • Don't hold fast learners hostage to group criterion cycles; give them forward paths (acceleration/enrichment) rather than waiting.
  • Budget honestly: mastery's costs are time and coverage; expect ~0 on broad measures unless the extra corrective time comes from somewhere real.

Open questions

  • Whether strictly enforced mastery gating (which no at-scale RCT has achieved) would do better — or simply reproduce Arlin's time-inequality at larger scale.
  • Whether modern adaptive platforms can deliver individualized mastery pacing without the group-time tax (current evidence: 0 to +0.2, fragile, one negative trial).

Evidence (6 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore