The Evidence on Teaching

What Impacts Should We Expect from Tutoring at Scale? Exploring Meta-Analytic Generalizability

Kraft, M. A., Schueler, B. E., & Falken, G. T. · 2024

grade Bmeta-analysisindependentreplicatednumbers spot-checked
Sample
265 RCTs / 340 interventions
Population
PreK-12, mostly US; 89% of studies target low performers; 59% evaluate programs serving <100 students.
Design
The scale-generalizability meta: bins programs by number of students served and restricts to the policy-relevant target (US, standardized tests, at scale). Only 9 RCTs of programs serving >=1,000 students.
Key findings
The honest at-scale expectation for tutoring: d=0.16-0.21 on standardized tests — one-third to one-half the efficacy-meta numbers. The attenuation is real (not publication bias) and partly explained by design drift at scale (bigger ratios, less delivered dosage). The high-impact bundle roughly HALVES the attenuation.
Genetic confound
Targeted low-performer samples shrink SD denominators.
Replication notes
Numbers verified. Corroborated by independent at-scale quasi-experiments in the US, UK, and Australia.
DOI / URL
10.26300/zygj-m525

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Full sample pooledd0.42 (prediction interval -0.31 to 1.16)mixedend of treatmentbusiness-as-usualend-of-treatmentdomain-skill
By program scaled<100 students 0.55 -> 1000+ students 0.14 (near-linear decline)mixedend of treatmentbusiness-as-usualend-of-treatmentdomain-skill
Target-aligned (US + standardized + at scale)d0.21 (400-999 students) to 0.16 (1000+)standardizedend of treatmentbusiness-as-usualend-of-treatmentdomain-skill
Researcher-generated testscoefficient+0.22 SD vs third-party standardizedmixedend of treatmentbusiness-as-usualend-of-treatmentdomain-skill
Bundled design features vs scale attenuation% decline-42% (all programs) vs -18% (full bundle: in-person, in-school-day, <=3:1, >=3x/wk, >=15 hrs, provided curriculum)mixedend of treatmentbusiness-as-usualend-of-treatmentdomain-skill

Cited by