Two-Sigma Tutoring: Separating Science Fiction from Science Fact
von Hippel, P. T. · 2024
grade Dcritiqueindependentnot-applicable
Sample
Forensic reading of the Anania/Burke dissertations + modern RCT record
Population
N/A.
Design
Education Next 24(2). Every load-bearing fact sourced from the primary dissertations and Walberg (1984).
Key findings
The best single documented takedown of 2-sigma: broken provenance, bundled treatment, aligned 3-week tests, higher mastery criterion in the tutored arm, extra time, replacement design. Modern replicated expectation: ~0.3 on broad tests for intense programs, 0-0.1 for typical at-scale vendor tutoring.
Genetic confound
Notes 2 sigma would exceed ~5 years of secondary-school learning — implausible on its face given stable individual differences.
Replication notes
No serious published rebuttal defends 2.0 as a tutoring effect. (Note: von Hippel quotes Cohen 1982's tutoring average as 0.33; the paper's tutee ES is 0.40 — the 0.33 is the tutor-as-learner effect.)
DOI / URL
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Honest ceiling for intense well-designed tutoring on broad tests | SD | ~0.33 (Saga 0.18-0.40; NCLB vendors 0.06) | standardized | end of treatment | business-as-usual | end-of-treatment | domain-skill |
Cited by
- Tutoring — the honest effect, the Bloom 2-sigma myth, and what survives scalestrong supportconf: highgc: low