The Evidence on Teaching

Tutoring — the honest effect, the Bloom 2-sigma myth, and what survives scale

High-dosage tutoring is education's most reliable lever: ~0.29 SD in trials, ~0.2 well-scaled — not Bloom's 2σ. Groups of 3–4 work; 1:1 is unnecessary.

strong supportconf: highgc: low

tutoring · ages 518 · method

Effect summary

High-dosage tutoring is education's most reliable intervention — and its honest size is ~0.29 SD in efficacy RCTs, 0.16-0.21 at well-run scale, 0.03-0.09 at national scale, NOT Bloom's 2.0 (a compound design artifact). The recipe is known: paid consistent tutors, groups up to 1:3-4 (1:1 unnecessary), >=3 sessions/week DURING the school day, 30+ delivered hours. Cost $1,400-4,800/student/yr.

Practical takeaway

Buy the expensive boring version: small-group (2:1-4:1), paid trained tutors, embedded in the school day as a scheduled period, 3-5x/week, curriculum-aligned. Expect ~0.2 SD on real tests for the students who get it — a large effect by education standards. Never buy opt-in, after-school, once-weekly, or rotating-tutor models; they reliably fail.

Who this applies to

Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.

Verdict

Tutoring is the best-replicated positive intervention in this database — and the gap between its folklore and its evidence is the most instructive in education. Bloom's 2-sigma is a myth: a compound artifact of researcher-built tests on 3-week novel units, tutoring fused with retest-to-criterion mastery (the mastery arm alone was ~1.1σ), specially trained tutors, extra time, and a higher criterion in the tutored arm — never replicated in 160+ controlled studies since, and Bloom's own cited source put tutoring at d=0.40. The real numbers form a clean gradient:

  • Efficacy RCTs (small, well-run): pooled ≈0.29 (published Nickow meta; 0/96 studies reach 2.0)
  • Well-run at scale (US, standardized tests, 400+ students): 0.16–0.21
  • Best replicated program (Saga Chicago, daily in-school 2:1): 0.18–0.40, with rare measured persistence (+0.23 SD two years later)
  • Typical district/national scale (Nashville, ESSER districts, UK's £1B programme): 0.00–0.09

Even the honest number is large by education standards — one year of Saga roughly doubles a year of normal high-school math learning — which is why the verdict is strong-support. But the at-scale gradient is the school builder's central lesson: the model is not what fails at scale; implementation is. Programs that kept the bundle intact lost only ~18% of their effect at scale; programs in general lost ~42%; programs that went opt-in/after-school/low-dosage lost everything.

What the evidence shows

Source Design Grade Key effect
Nickow 2024 RCT-only meta (~90 trials) B Pooled 0.288; teacher 0.39 > para 0.30 > volunteer 0.17; ≥3x/week required
Guryan 2023 (Saga) 2 large RCTs + replication A TOT 0.18–0.40 math; persistence +0.23 SD at 11th grade; $3,200–4,800/student
Kraft, Schueler & Falken 2024 265-RCT scale meta B Full sample 0.42 → target-aligned at scale 0.16–0.21; bundle halves attenuation
UK NTP Years 2-3 national evaluation C 0.03–0.09 KS2 maths, ~0 GCSE, ~no 1-yr persistence; school-led beat vendors
Carbonari 2026 8 ESSER districts C Mostly null as implemented; only tiny high-dosage programs (+0.22–0.33) worked
VanLehn 2011 tutoring-systems review C Human tutoring 0.79, step-based software 0.76 — parity

Design facts that survive across all of it:

  • 1:1 is unnecessary. No advantage over small groups anywhere (Elbaum; Nickow; Saga's 2:1; Bhatt's 4:1 tech-rotation kept 0.23 at 30% lower cost). The binding ingredient is daily scheduled in-school dosage with an accountable adult, not the ratio.
  • Frequency beats total hours: ≥3 sessions/week; once-weekly programs repeatedly fail.
  • Paid beats volunteer (0.30–0.39 vs 0.17), though structured supervised volunteers manage ~0.30 on drilled subskills (Ritter).
  • In-school beats after-school (~2×); opt-in models fail on take-up — the students who need it don't show up.
  • Virtual is ~1/4–1/2 as effective at ~1/3 the cost (OnYourMark: 0.05–0.12 at $1,400) — a legitimate option where staffing fails, not a substitute.
  • Software parity: step-based tutoring systems match human tutors in efficacy settings — but collapse to ~0.13 on broad standardized measures, and adaptive systems can harm low achievers.

Durability is tutoring's soft spot: Saga's ~70% retention at 1–2 years is the exception; the UK national programme showed ~no persistence, and the flagship 1:1 program (Reading Recovery) shows contested evidence of long-term reversal. Assume decay and plan re-dosing.

Hereditarian-lens assessment

Risk: low (randomized designs throughout). Three lens findings worth a founder's attention:

  1. Bloom's equalization claim is refuted, not just his effect size. The falling aptitude-achievement correlation he celebrated is a ceiling artifact; heritability of achievement rises with age, and time-to-mastery ratios grow. Tutoring raises the mean for those tutored; it does not abolish ability differences. "Move anyone to the top 2%" was never real.
  2. Human-tutoring RCTs cannot tell you tutoring "closes gaps" — ~90% sample only low performers, so no high-ability arm exists. What's measurable: it reliably raises low performers' domain skills.
  3. Software tutoring can widen gaps (meta evidence of negative effects for low achievers, g=−0.18, while helping the general population) — a caution against assuming tech variants inherit human-tutoring's equity profile.
  4. Sample-SD units on low-performer samples overstate population-SD effects somewhat — one more reason the honest planning number is ~0.2, not 0.4.

What critics say / limits

  • Fryer's well-powered NYC high-dosage reading RCT was null at the mean — tutoring is more reliable in math than reading, and reading effects concentrate in early grades.
  • Peer-reviewed studies pool higher than grey literature (0.45 vs 0.22) despite clean funnel tests — selective reporting can't be ruled out.
  • Only 9 RCTs exist of programs serving ≥1,000 students; the at-scale bin is imprecise.
  • One district program produced significantly negative math effects (plausibly displacing core instruction) — tutoring time comes from somewhere.

Practical guidance

  • The recipe (each element evidence-moderated): groups of 2–4 · paid, trained, consistent tutors · scheduled period during the school day · 3–5 sessions/week · ≥30 actually-delivered hours · curriculum-aligned, data-informed sessions.
  • Budget $1,400–4,800/student/yr depending on labor model; tech-rotation 4:1 (~$2,500–3,000) is the best current cost-effect compromise; structured virtual ($1,400) where staffing fails.
  • Target it: biggest, most reliable effects for students behind grade level, math, and (for reading) early grades. It is the best-evidenced remediation tool in existence — not a whole-population strategy at 20× the cost of curriculum changes.
  • Measure delivered dosage, not enrolled students — implementation is where every failed scale-up died.
  • Assume fadeout without re-dosing; buy persistence by keeping students in the system across years (the Saga model) rather than one-shot pull-outs.

Open questions

  • Whether AI tutors change the cost curve: step-based systems already match human efficacy on aligned measures, but no published causal evidence yet shows chatbot tutoring moving standardized achievement (claims of "2-sigma AI tutoring" recapitulate Bloom's error).
  • Whether Saga-style persistence survives independent (non-developer) replication at scale.

Evidence (18 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore

Tutoring — the honest effect, the Bloom 2-sigma myth, and what survives scale · The Evidence on Teaching