A meta-analysis of the effectiveness of intelligent tutoring systems on K–12 students’ mathematical learning
Steenbergen-Hu, S., & Cooper, H. · 2013
grade Cmeta-analysisindependentreplicated
Sample
26 reports containing 34 independent samples, published 1997-2010
Population
K-12 mathematics in school settings, predominantly US; ITS defined broadly as "self-paced, learner-led, highly adaptive, and interactive learning environments operated through computers", which pulled in answer-based CAI systems (iLearnMath, Larson Pre-Algebra/Algebra, PLATO Algebra, PLATO Achieve Now) alongside true step-based tutors such as Cognitive Tutor.
Design
The K-12 half of the Steenbergen-Hu & Cooper pair, and the direct counterweight to Kulik & Fletcher (2016). Comparison conditions were mostly regular classroom instruction but the pool also admits one-to-one tutoring, homework and no-instruction controls, and Kulik & Fletcher document that several included studies had pretest gaps of 0.76-1.09 SD or gave no evidence of baseline equivalence - a real quality criticism of this meta that cuts AGAINST its own null (weak designs usually inflate, so the null survives the criticism awkwardly well). The definitional dispute is the substance of the disagreement with Kulik & Fletcher: Steenbergen-Hu & Cooper include answer-based CAI that ITS specialists would exclude; Kulik & Fletcher exclude any study with a non-conventional control or a failed implementation, which removes null results. NOT READ IN FULL - APA paywall, no OA copy located (unpaywall/OpenAlex/CORE/scholar.archive.org all negative). Numbers below come from the paper's own ERIC abstract plus two independent secondary sources that read the tables (Kulik & Fletcher 2016 pp. 45, 69; Leite et al. 2025 preprint), which agree with each other.
Key findings
Intelligent tutoring systems do essentially nothing for K-12 mathematics: average effect sizes range from g = 0.01 to g = 0.09 across model specifications - "no negative and perhaps a small positive effect" in the authors' own words. Two moderators carry the interest. First, effects were LARGER in interventions shorter than one school year than in longer ones, which is the fadeout / hothouse signature rather than a dosage argument for the technology. Second, and squarely against the equity case for adaptive software, ITS worked better for general-population students than for low achievers (g = -0.18 for low-achieving samples vs g = 0.04 for general samples), raising the possibility that computerized instruction widens rather than narrows gaps. The contrast with the same authors' college meta-analysis one year later (g = 0.32 to 0.37) is the striking result across the pair: the same technology, the same reviewers, the same methods, four times the effect on self-selected college students.
Genetic confound
Medium - a mix of randomized and quasi-experimental school studies; several included studies had large baseline pretest differences (0.76-1.09 SD) documented by Kulik & Fletcher, so selection into conditions is not fully ruled out in the pool.
Replication notes
The near-zero K-12 finding is corroborated by the two largest independent bodies of evidence on the same question: Cheung & Slavin (2013), whose large randomized subset gives +0.06 (95% CI 0.00 to 0.13) on K-12 mathematics technology, and the federal Dynarski/Campuzano RCTs (+0.03). Kulik & Fletcher (2016) dispute it and report 0.40 for K-12 mathematics ITS - but their own standardized-measure-only K-12 mathematics subgroup is 0.10, which lands back on Steenbergen-Hu & Cooper. Leite et al. (2025, preprint) estimate g = 0.271 for US K-12 ITS with standardized measures at 0.200 and end-of-school-year measurement at 0.155, i.e. between the two camps and closer to this one once measure type and timing are held fixed.
DOI / URL
10.1037/a0032447
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| K-12 mathematics achievement, ITS vs control | g | 0.01 to 0.09 depending on model (the authors report a range, not a single point estimate); Kulik & Fletcher summarise it as "around 0.05 standard deviations, a trivial amount". The frequently quoted single figure is g = 0.09, p > .05. | mixed | end of intervention | business-as-usual | end-of-treatment | domain-skill |
| Low-achieving students vs general-population students | g | low-achieving K-12 samples g = -0.18 vs general samples g = 0.04. The authors flag the achievement-gap-widening implication explicitly. (Values as reported by Leite et al. 2025 from the paper's tables; the abstract states the direction.) | mixed | end of intervention | business-as-usual | end-of-treatment | domain-skill |
| Duration moderator | g | interventions shorter than one school year produced GREATER effects than longer ones - the opposite of a dose-response relationship, and the standard signature of short hothouse studies | mixed | end of intervention | business-as-usual | end-of-treatment | domain-skill |
| Grade-level breakdown | g | elementary 0.41, middle school 0.09, high school -0.09 (fixed-effects estimates, as reported by Leite et al. 2025 from this paper's tables) - a spread wide enough that the pooled near-zero is an average over sign changes, not a homogeneous null | mixed | end of intervention | business-as-usual | end-of-treatment | domain-skill |