Tutoring — the honest effect, the Bloom 2-sigma myth, and what survives scale
High-dosage tutoring is education's most reliable lever: ~0.29 SD in trials, ~0.2 well-scaled — not Bloom's 2σ. Groups of 3–4 work; 1:1 is unnecessary.
strong supportconf: highgc: lowtutoring · ages 5–18 · method
High-dosage tutoring is education's most reliable intervention — and its honest size is ~0.29 SD in efficacy RCTs, 0.16-0.21 at well-run scale, 0.03-0.09 at national scale, NOT Bloom's 2.0 (a compound design artifact). The recipe is known: paid consistent tutors, groups up to 1:3-4 (1:1 unnecessary), >=3 sessions/week DURING the school day, 30+ delivered hours. Cost $1,400-4,800/student/yr.
Buy the expensive boring version: small-group (2:1-4:1), paid trained tutors, embedded in the school day as a scheduled period, 3-5x/week, curriculum-aligned. Expect ~0.2 SD on real tests for the students who get it — a large effect by education standards. Never buy opt-in, after-school, once-weekly, or rotating-tutor models; they reliably fail.
Who this applies to
Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.
Verdict
Tutoring is the best-replicated positive intervention in this database — and the gap between its folklore and its evidence is the most instructive in education. Bloom's 2-sigma is a myth: a compound artifact of researcher-built tests on 3-week novel units, tutoring fused with retest-to-criterion mastery (the mastery arm alone was ~1.1σ), specially trained tutors, extra time, and a higher criterion in the tutored arm — never replicated in 160+ controlled studies since, and Bloom's own cited source put tutoring at d=0.40. The real numbers form a clean gradient:
- Efficacy RCTs (small, well-run): pooled ≈0.29 (published Nickow meta; 0/96 studies reach 2.0)
- Well-run at scale (US, standardized tests, 400+ students): 0.16–0.21
- Best replicated program (Saga Chicago, daily in-school 2:1): 0.18–0.40, with rare measured persistence (+0.23 SD two years later)
- Typical district/national scale (Nashville, ESSER districts, UK's £1B programme): 0.00–0.09
Even the honest number is large by education standards — one year of Saga roughly doubles a year of normal high-school math learning — which is why the verdict is strong-support. But the at-scale gradient is the school builder's central lesson: the model is not what fails at scale; implementation is. Programs that kept the bundle intact lost only ~18% of their effect at scale; programs in general lost ~42%; programs that went opt-in/after-school/low-dosage lost everything.
What the evidence shows
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Nickow 2024 | RCT-only meta (~90 trials) | B | Pooled 0.288; teacher 0.39 > para 0.30 > volunteer 0.17; ≥3x/week required |
| Guryan 2023 (Saga) | 2 large RCTs + replication | A | TOT 0.18–0.40 math; persistence +0.23 SD at 11th grade; $3,200–4,800/student |
| Kraft, Schueler & Falken 2024 | 265-RCT scale meta | B | Full sample 0.42 → target-aligned at scale 0.16–0.21; bundle halves attenuation |
| UK NTP Years 2-3 | national evaluation | C | 0.03–0.09 KS2 maths, ~0 GCSE, ~no 1-yr persistence; school-led beat vendors |
| Carbonari 2026 | 8 ESSER districts | C | Mostly null as implemented; only tiny high-dosage programs (+0.22–0.33) worked |
| VanLehn 2011 | tutoring-systems review | C | Human tutoring 0.79, step-based software 0.76 — parity |
Design facts that survive across all of it:
- 1:1 is unnecessary. No advantage over small groups anywhere (Elbaum; Nickow; Saga's 2:1; Bhatt's 4:1 tech-rotation kept 0.23 at 30% lower cost). The binding ingredient is daily scheduled in-school dosage with an accountable adult, not the ratio.
- Frequency beats total hours: ≥3 sessions/week; once-weekly programs repeatedly fail.
- Paid beats volunteer (0.30–0.39 vs 0.17), though structured supervised volunteers manage ~0.30 on drilled subskills (Ritter).
- In-school beats after-school (~2×); opt-in models fail on take-up — the students who need it don't show up.
- Virtual is ~1/4–1/2 as effective at ~1/3 the cost (OnYourMark: 0.05–0.12 at $1,400) — a legitimate option where staffing fails, not a substitute.
- Software parity: step-based tutoring systems match human tutors in efficacy settings — but collapse to ~0.13 on broad standardized measures, and adaptive systems can harm low achievers.
Durability is tutoring's soft spot: Saga's ~70% retention at 1–2 years is the exception; the UK national programme showed ~no persistence, and the flagship 1:1 program (Reading Recovery) shows contested evidence of long-term reversal. Assume decay and plan re-dosing.
Hereditarian-lens assessment
Risk: low (randomized designs throughout). Three lens findings worth a founder's attention:
- Bloom's equalization claim is refuted, not just his effect size. The falling aptitude-achievement correlation he celebrated is a ceiling artifact; heritability of achievement rises with age, and time-to-mastery ratios grow. Tutoring raises the mean for those tutored; it does not abolish ability differences. "Move anyone to the top 2%" was never real.
- Human-tutoring RCTs cannot tell you tutoring "closes gaps" — ~90% sample only low performers, so no high-ability arm exists. What's measurable: it reliably raises low performers' domain skills.
- Software tutoring can widen gaps (meta evidence of negative effects for low achievers, g=−0.18, while helping the general population) — a caution against assuming tech variants inherit human-tutoring's equity profile.
- Sample-SD units on low-performer samples overstate population-SD effects somewhat — one more reason the honest planning number is ~0.2, not 0.4.
What critics say / limits
- Fryer's well-powered NYC high-dosage reading RCT was null at the mean — tutoring is more reliable in math than reading, and reading effects concentrate in early grades.
- Peer-reviewed studies pool higher than grey literature (0.45 vs 0.22) despite clean funnel tests — selective reporting can't be ruled out.
- Only 9 RCTs exist of programs serving ≥1,000 students; the at-scale bin is imprecise.
- One district program produced significantly negative math effects (plausibly displacing core instruction) — tutoring time comes from somewhere.
Practical guidance
- The recipe (each element evidence-moderated): groups of 2–4 · paid, trained, consistent tutors · scheduled period during the school day · 3–5 sessions/week · ≥30 actually-delivered hours · curriculum-aligned, data-informed sessions.
- Budget $1,400–4,800/student/yr depending on labor model; tech-rotation 4:1 (~$2,500–3,000) is the best current cost-effect compromise; structured virtual ($1,400) where staffing fails.
- Target it: biggest, most reliable effects for students behind grade level, math, and (for reading) early grades. It is the best-evidenced remediation tool in existence — not a whole-population strategy at 20× the cost of curriculum changes.
- Measure delivered dosage, not enrolled students — implementation is where every failed scale-up died.
- Assume fadeout without re-dosing; buy persistence by keeping students in the system across years (the Saga model) rather than one-shot pull-outs.
Open questions
- Whether AI tutors change the cost curve: step-based systems already match human efficacy on aligned measures, but no published causal evidence yet shows chatbot tutoring moving standardized achievement (claims of "2-sigma AI tutoring" recapitulate Bloom's error).
- Whether Saga-style persistence survives independent (non-developer) replication at scale.
- grade BThe Promise of Tutoring for PreK-12 Learning: A Systematic Review and Meta-Analysis of the Experimental EvidenceNickow, A., Oreopoulos, P., & Quan, V. · 2024 · meta-analysis
- grade ANot Too Late: Improving Academic Outcomes Among Adolescents (Saga/Match high-dosage tutoring, two Chicago RCTs)Guryan, J., Ludwig, J., Bhatt, M. P., Cook, P. J., et al. · 2023 · rct
- grade BWhat Impacts Should We Expect from Tutoring at Scale? Exploring Meta-Analytic GeneralizabilityKraft, M. A., Schueler, B. E., & Falken, G. T. · 2024 · meta-analysis
- grade DThe 2 Sigma Problem: The Search for Methods of Group Instruction as Effective as One-to-One TutoringBloom BS · 1984 · review
- grade DTwo-Sigma Tutoring: Separating Science Fiction from Science Factvon Hippel, P. T. · 2024 · critique
- grade CEducational Outcomes of Tutoring: A Meta-Analysis of FindingsCohen, P. A., Kulik, J. A., & Kulik, C.-L. C. · 1982 · meta-analysis
- grade CThe Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring SystemsVanLehn, K. · 2011 · review
- grade CAcademic Interventions for Elementary and Middle School Students With Low Socioeconomic Status: A Meta-AnalysisDietrichson, J., Bog, M., Filges, T., & Jorgensen, A.-M. K. · 2017 · meta-analysis
- grade CThe Effectiveness of Volunteer Tutoring Programs: A Meta-Analysis (RCT-only)Ritter, G. W., Barnett, J. H., Denny, G. S., & Albin, G. R. · 2009 · meta-analysis
- grade CHow Effective Are One-to-One Tutoring Programs in Reading for At-Risk Students? A Meta-AnalysisElbaum, B., Vaughn, S., Hughes, M. T., & Moody, S. W. · 2000 · meta-analysis
- grade BCan Technology Facilitate Scale? Evidence from a Randomized Evaluation of High-Dosage Tutoring (Saga 4:1 variant)Bhatt, M. P., Guryan, J., Khan, S. A., LaForest-Tucker, M., & Mishra, B. · 2024 · rct
- grade BThe Scaling Dynamics and Causal Effects of a District-Operated Tutoring Program (Metro Nashville)Kraft, M. A., Edwards, D. S., & Cannata, M. · 2024 · quasi-experiment
- grade CUK National Tutoring Programme: Year 2 and Year 3 Impact Evaluations (NFER/DfE)National Foundation for Educational Research, for the Department for Education · 2024 · quasi-experiment
- grade CImpacts of Academic Recovery Interventions on Student Achievement, 2022-23 (Road to Recovery, 8 districts)Carbonari, M. V., DeArmond, M., Dewey, D., Goldhaber, D., Kane, T. J., Staiger, D. O., et al. · 2026 · quasi-experiment
- grade DA Blueprint for Scaling Tutoring and Mentoring Across Public Schools (cost model)Kraft, M. A., & Falken, G. T. · 2021 · review
- grade CHigh-Impact Tutoring: State of the Research and Priorities (design principles)Robinson, C. D., & Loeb, S. (with Kraft & Schueler companion brief) · 2021 · review
- grade BThe Effects of Virtual Tutoring on Young Readers: Results From a Randomized Controlled TrialRobinson, C. D., Pollard, C., Novicoff, S., White, S., & Loeb, S. · 2024 · rct
- grade CLong-Term Impacts of Reading Recovery through 3rd and 4th Grade: A Regression Discontinuity StudyMay, H., Blakeney, A., Shrestha, P., Mazal, M., & Kennedy, N. · 2023 · quasi-experiment
Related decisions
- Class-size reductionmixedconf: highgc: low
- Does educational technology raise learning — CAI, adaptive software, devices, screens, and AI tutors?mixedconf: mediumgc: low
- School spending, charters, and what school-level choices actually move outcomesmixedconf: highgc: low
- The parent as instructor — what survives when the person teaching is the person who raised the childmixedconf: mediumgc: medium