The Evidence on Teaching

The Promise of Tutoring for PreK-12 Learning: A Systematic Review and Meta-Analysis of the Experimental Evidence

Nickow, A., Oreopoulos, P., & Quan, V. · 2024

grade Bmeta-analysisindependentreplicatednumbers spot-checked
Sample
89-96 RCTs (WP 2020: 96; published AERJ 2024: 89)
Population
PreK-12, mostly low-performing US samples; ~80% literacy.
Design
RCT-only meta of supplemental human tutoring. Outcomes <=3 months post-treatment (fadeout unmeasurable by design). Published version revised the pooled estimate DOWN from the working paper.
Key findings
The central modern tutoring benchmark: ~0.29 SD pooled across ~90 RCTs — education's most reliable intervention, and 5-7x below Bloom. The moderator structure is the recipe: paid teacher/paraprofessional tutors, >=3 sessions/week, during the school day, small groups fine (1:1 unnecessary). Once-weekly and volunteer programs repeatedly fail.
Genetic confound
~90% low-performer samples: restricted-range SDs inflate d vs population units (Fitzgerald & Tipton).
Replication notes
Numbers verified incl. the published downward revision. Converges with Dietrichson (0.36 standardized-only) and Ritter (0.30).
DOI / URL
10.3102/00028312231208687

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Pooled tutoring effectg0.37 (2020 WP) -> 0.288 (SE .029, published AERJ 2024)mixed<=3 months postunclearunder-1yrdomain-skill
By tutor type (published)gteacher 0.39 > paraprofessional 0.30 > nonprofessional/volunteer 0.17mixed<=3 monthsunclearunder-1yrdomain-skill
Frequency / settingg>=3 sessions/wk needed; during-school ~2x after-school; 1:1 no better than 1:2-3mixed<=3 monthsunclearunder-1yrdomain-skill
By sample size (efficacy-scale gradient)gN<50: 0.45 -> N>400: 0.25mixed<=3 monthsunclearunder-1yrdomain-skill

Cited by