The Evidence on Teaching

The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems

VanLehn, K. · 2011

grade Creviewindependentreplicatednumbers spot-checked
Sample
10 human-tutoring comparisons; ~28 evaluations
Population
STEM-heavy; substantially college samples; experimenter tests.
Design
Tests the interaction-granularity hypothesis (finer interaction -> better): d=0.3 CAI / 1.0 ITS / 2.0 human predicted.
Key findings
The deflation from inside the tutoring community: human tutoring is d≈0.79, not 2.0 — and step-based tutoring SOFTWARE matches it (0.76). Effectiveness plateaus at step-level feedback; the human adds little beyond it. On Bloom: 'the d=2.0 effect seems due mostly to holding the tutees to a higher standard of mastery.'
Genetic confound
N/A (comparative review).
Replication notes
Kulik & Fletcher 2016 (50 ITS evaluations, median 0.66, collapsing to ~0.13 on broad standardized tests) independently confirms both the parity finding and the alignment inflation.
DOI / URL
10.1080/00461520.2011.611369

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Human tutoring vs no tutoringd0.79 (10 comparisons; ~0.68 excluding Anania)mixedend of treatmentbusiness-as-usualend-of-treatmentdomain-skill
Step-based intelligent tutoring systemsd0.76 — statistical parity with humansmixedend of treatmentbusiness-as-usualend-of-treatmentdomain-skill
Answer-based CAId0.31mixedend of treatmentbusiness-as-usualend-of-treatmentdomain-skill

Cited by