Intelligent tutoring systems and learning outcomes: A meta-analysis
Ma, W., Adesope, O. O., Nesbit, J. C., & Liu, Q. · 2014
grade Cmeta-analysisindependentmixed
Sample
107 effect sizes from 73 reports, 14,321 participants; distribution of the 107 effect sizes M = 0.43, SD = 0.40
Population
All education levels from elementary school to postsecondary, mixed subjects (mathematics and language/literacy both g = 0.35); ITS used as principal instruction, as a supplement, as an integral component of a course, or as a homework aid.
Design
The meta-analysis organised around COMPARISON CONDITION rather than around the technology, which is what makes it useful here: it prices the same software against five different things a student could otherwise have been doing, and the answer changes sign. Pool is not restricted to high-quality experimental designs - it includes experimental, quasi-experimental and pre-experimental studies, which is the standing criticism of it (Leite et al. 2025) and the reason the grade stays C. Attrition was reported in only 21.5% of studies (33.6% explicitly reported none, 44.9% said nothing). Publication bias was assessed by fail-safe N only: the number of null studies needed to overturn the pooled effect exceeds Rosenthal's 5k+10 criterion, so the authors conclude bias is not a significant threat - no funnel plot, trim-and-fill or PET-PEESE. NOT READ IN FULL - APA paywall, no OA copy located (the author's SFU copy, ResearchGate, Academia.edu and Ovid all refused). Numbers below come from the paper's own ERIC/PsycNET abstract plus Kulik & Fletcher (2016, p. 45) and Leite et al. (2025), two independent readers of its tables. THE 95% CONFIDENCE INTERVALS FOR THE COMPARISON- CONDITION CONTRASTS COULD NOT BE RETRIEVED; the significance statements below are the paper's own.
Key findings
Across 107 comparisons the average ITS effect is g = 0.43, but the number decomposes entirely into what the control group was doing. ITS beat teacher-led LARGE-GROUP instruction by g = 0.42, beat non-ITS computer-based instruction by g = 0.57, and beat textbooks and workbooks by g = 0.35. Against the two conditions that give a student individual or near-individual human attention, the advantage disappears: g = 0.05 vs small-group instruction and g = -0.11 vs individualized human tutoring, neither significant. This is one of the two sources (with VanLehn 2011) behind the claim that intelligent tutoring software is "no worse than a human tutor" - and read precisely, the claim is that a non-significant point estimate favours the human by a tenth of a standard deviation, not that the software matches a tutor. Unusually for this literature, the measure-type moderator was essentially flat here (standardized g = 0.41 vs researcher-developed g = 0.42), which is the sharpest disagreement with Kulik & Fletcher (2016), whose same-moderator gap was 0.13 vs 0.73.
Genetic confound
Medium - the pool mixes randomized, quasi-experimental and pre-experimental designs without restricting on design quality, so selection into ITS conditions is not excluded; the human-tutoring contrasts in particular come mostly from small within-course randomizations where the confound risk is low but the samples are self-selected.
Replication notes
The "ITS is statistically indistinguishable from a human tutor" claim is replicated by VanLehn (2011), who found step-based ITS d = 0.76 against human tutoring d = 0.79 - the two sources the claim rests on. It is contradicted in direction by Steenbergen-Hu & Cooper (2014), who found ITS 0.25 SD BELOW human tutoring in college settings, though that gap is also not the two sigma of folklore. The ITS-vs-large-group contrast (g = 0.42) is corroborated by Kulik & Fletcher (2016, weighted g = 0.50 against conventional instruction) but not by the two large-scale K-12 literatures (Steenbergen-Hu & Cooper 2013 ~0.05; Cheung & Slavin 2013 large randomized +0.06). The measure-type null reported here (standardized 0.41 vs researcher-made 0.42) is itself unreplicated and is the outlier finding in this literature.
DOI / URL
10.1037/a0037123
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| ITS vs all comparison conditions, pooled | g | mean of the 107 effect sizes = 0.43 (SD 0.40); significant positive mean effects were found at every education level, in almost every subject domain, and whether the ITS was the principal instruction, a supplement, an integral course component or a homework aid | mixed | end of treatment | unclear | end-of-treatment | domain-skill |
| ITS vs teacher-led LARGE-GROUP instruction | g | +0.42, significant (reported as 0.44 by Leite et al. 2025 reading the paper's moderator table; CI not retrievable). The contrast that makes ITS look like a strong intervention. | mixed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| ITS vs SMALL-GROUP human instruction | g | +0.05, NOT significant - the advantage is already gone once the comparison group gets small-group teaching | mixed | end of treatment | active-alternative | end-of-treatment | domain-skill |
| ITS vs INDIVIDUALIZED HUMAN TUTORING | g | -0.11, NOT significant - the point estimate favours the human tutor. This, with VanLehn (2011), is the entire evidential basis of the "ITS is as good as a human tutor" claim; the paper's 95% CI could not be retrieved (paywalled), so the width of the equivalence claim is unverified here. | mixed | end of treatment | active-alternative | end-of-treatment | domain-skill |
| ITS vs textbooks or workbooks | g | +0.35, significant (0.36 in Leite et al.'s reading of the table) | mixed | end of treatment | active-alternative | end-of-treatment | domain-skill |
| ITS vs non-ITS computer-based instruction | g | +0.57, significant (0.577) - the largest contrast in the review, i.e. ITS mostly beats ordinary CAI rather than beating teaching | mixed | end of treatment | active-alternative | end-of-treatment | domain-skill |
| Outcome-measure type moderator | g | standardized measures 0.41 (p < .05) vs researcher-developed measures 0.42 (p < .05) - essentially NO alignment gap, in flat contradiction of Kulik & Fletcher (2016) (0.13 vs 0.73) and of Xu et al. (2019) on reading-comprehension ITS (researcher-developed measures 0.73 SD larger). Values via Leite et al. (2025) reading this paper's tables; the archive treats this as the least replicated finding in the file. | mixed | end of treatment | unclear | end-of-treatment | domain-skill |
| Grade level, prior knowledge and subject moderators | g | elementary 0.31 vs middle school 0.41 vs high school 0.41; low prior domain knowledge 0.38, medium 0.28, varying 0.48 (all positive); mathematics 0.35 and language/literacy 0.35. Publication bias assessed by fail-safe N only, which exceeded Rosenthal's 5k+10 criterion. | mixed | end of treatment | unclear | end-of-treatment | domain-skill |