The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems
VanLehn, K. · 2011
grade Creviewindependentreplicatednumbers spot-checked
Sample
10 human-tutoring comparisons; ~28 evaluations
Population
STEM-heavy; substantially college samples; experimenter tests.
Design
Tests the interaction-granularity hypothesis (finer interaction -> better): d=0.3 CAI / 1.0 ITS / 2.0 human predicted.
Key findings
The deflation from inside the tutoring community: human tutoring is d≈0.79, not 2.0 — and step-based tutoring SOFTWARE matches it (0.76). Effectiveness plateaus at step-level feedback; the human adds little beyond it. On Bloom: 'the d=2.0 effect seems due mostly to holding the tutees to a higher standard of mastery.'
Genetic confound
N/A (comparative review).
Replication notes
Kulik & Fletcher 2016 (50 ITS evaluations, median 0.66, collapsing to ~0.13 on broad standardized tests) independently confirms both the parity finding and the alignment inflation.
DOI / URL
10.1080/00461520.2011.611369
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Human tutoring vs no tutoring | d | 0.79 (10 comparisons; ~0.68 excluding Anania) | mixed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| Step-based intelligent tutoring systems | d | 0.76 — statistical parity with humans | mixed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| Answer-based CAI | d | 0.31 | mixed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
Cited by
- Does educational technology raise learning — CAI, adaptive software, devices, screens, and AI tutors?mixedconf: mediumgc: low
- Tutoring — the honest effect, the Bloom 2-sigma myth, and what survives scalestrong supportconf: highgc: low