The Evidence on Teaching

Intelligent tutoring systems and learning outcomes: A meta-analysis

Ma, W., Adesope, O. O., Nesbit, J. C., & Liu, Q. · 2014

grade Cmeta-analysisindependentmixed
Sample
107 effect sizes from 73 reports, 14,321 participants; distribution of the 107 effect sizes M = 0.43, SD = 0.40
Population
All education levels from elementary school to postsecondary, mixed subjects (mathematics and language/literacy both g = 0.35); ITS used as principal instruction, as a supplement, as an integral component of a course, or as a homework aid.
Design
The meta-analysis organised around COMPARISON CONDITION rather than around the technology, which is what makes it useful here: it prices the same software against five different things a student could otherwise have been doing, and the answer changes sign. Pool is not restricted to high-quality experimental designs - it includes experimental, quasi-experimental and pre-experimental studies, which is the standing criticism of it (Leite et al. 2025) and the reason the grade stays C. Attrition was reported in only 21.5% of studies (33.6% explicitly reported none, 44.9% said nothing). Publication bias was assessed by fail-safe N only: the number of null studies needed to overturn the pooled effect exceeds Rosenthal's 5k+10 criterion, so the authors conclude bias is not a significant threat - no funnel plot, trim-and-fill or PET-PEESE. NOT READ IN FULL - APA paywall, no OA copy located (the author's SFU copy, ResearchGate, Academia.edu and Ovid all refused). Numbers below come from the paper's own ERIC/PsycNET abstract plus Kulik & Fletcher (2016, p. 45) and Leite et al. (2025), two independent readers of its tables. THE 95% CONFIDENCE INTERVALS FOR THE COMPARISON- CONDITION CONTRASTS COULD NOT BE RETRIEVED; the significance statements below are the paper's own.
Key findings
Across 107 comparisons the average ITS effect is g = 0.43, but the number decomposes entirely into what the control group was doing. ITS beat teacher-led LARGE-GROUP instruction by g = 0.42, beat non-ITS computer-based instruction by g = 0.57, and beat textbooks and workbooks by g = 0.35. Against the two conditions that give a student individual or near-individual human attention, the advantage disappears: g = 0.05 vs small-group instruction and g = -0.11 vs individualized human tutoring, neither significant. This is one of the two sources (with VanLehn 2011) behind the claim that intelligent tutoring software is "no worse than a human tutor" - and read precisely, the claim is that a non-significant point estimate favours the human by a tenth of a standard deviation, not that the software matches a tutor. Unusually for this literature, the measure-type moderator was essentially flat here (standardized g = 0.41 vs researcher-developed g = 0.42), which is the sharpest disagreement with Kulik & Fletcher (2016), whose same-moderator gap was 0.13 vs 0.73.
Genetic confound
Medium - the pool mixes randomized, quasi-experimental and pre-experimental designs without restricting on design quality, so selection into ITS conditions is not excluded; the human-tutoring contrasts in particular come mostly from small within-course randomizations where the confound risk is low but the samples are self-selected.
Replication notes
The "ITS is statistically indistinguishable from a human tutor" claim is replicated by VanLehn (2011), who found step-based ITS d = 0.76 against human tutoring d = 0.79 - the two sources the claim rests on. It is contradicted in direction by Steenbergen-Hu & Cooper (2014), who found ITS 0.25 SD BELOW human tutoring in college settings, though that gap is also not the two sigma of folklore. The ITS-vs-large-group contrast (g = 0.42) is corroborated by Kulik & Fletcher (2016, weighted g = 0.50 against conventional instruction) but not by the two large-scale K-12 literatures (Steenbergen-Hu & Cooper 2013 ~0.05; Cheung & Slavin 2013 large randomized +0.06). The measure-type null reported here (standardized 0.41 vs researcher-made 0.42) is itself unreplicated and is the outlier finding in this literature.
DOI / URL
10.1037/a0037123

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
ITS vs all comparison conditions, pooledgmean of the 107 effect sizes = 0.43 (SD 0.40); significant positive mean effects were found at every education level, in almost every subject domain, and whether the ITS was the principal instruction, a supplement, an integral course component or a homework aidmixedend of treatmentunclearend-of-treatmentdomain-skill
ITS vs teacher-led LARGE-GROUP instructiong+0.42, significant (reported as 0.44 by Leite et al. 2025 reading the paper's moderator table; CI not retrievable). The contrast that makes ITS look like a strong intervention.mixedend of treatmentbusiness-as-usualend-of-treatmentdomain-skill
ITS vs SMALL-GROUP human instructiong+0.05, NOT significant - the advantage is already gone once the comparison group gets small-group teachingmixedend of treatmentactive-alternativeend-of-treatmentdomain-skill
ITS vs INDIVIDUALIZED HUMAN TUTORINGg-0.11, NOT significant - the point estimate favours the human tutor. This, with VanLehn (2011), is the entire evidential basis of the "ITS is as good as a human tutor" claim; the paper's 95% CI could not be retrieved (paywalled), so the width of the equivalence claim is unverified here.mixedend of treatmentactive-alternativeend-of-treatmentdomain-skill
ITS vs textbooks or workbooksg+0.35, significant (0.36 in Leite et al.'s reading of the table)mixedend of treatmentactive-alternativeend-of-treatmentdomain-skill
ITS vs non-ITS computer-based instructiong+0.57, significant (0.577) - the largest contrast in the review, i.e. ITS mostly beats ordinary CAI rather than beating teachingmixedend of treatmentactive-alternativeend-of-treatmentdomain-skill
Outcome-measure type moderatorgstandardized measures 0.41 (p < .05) vs researcher-developed measures 0.42 (p < .05) - essentially NO alignment gap, in flat contradiction of Kulik & Fletcher (2016) (0.13 vs 0.73) and of Xu et al. (2019) on reading-comprehension ITS (researcher-developed measures 0.73 SD larger). Values via Leite et al. (2025) reading this paper's tables; the archive treats this as the least replicated finding in the file.mixedend of treatmentunclearend-of-treatmentdomain-skill
Grade level, prior knowledge and subject moderatorsgelementary 0.31 vs middle school 0.41 vs high school 0.41; low prior domain knowledge 0.38, medium 0.28, varying 0.48 (all positive); mathematics 0.35 and language/literacy 0.35. Publication bias assessed by fail-safe N only, which exceeded Rosenthal's 5k+10 criterion.mixedend of treatmentunclearend-of-treatmentdomain-skill

Cited by