Effectiveness of Intelligent Tutoring Systems: A Meta-Analytic Review
Kulik, J. A., & Fletcher, J. D. · 2016
grade Cmeta-analysisindependentmixednumbers spot-checked
Sample
50 controlled evaluations (63 independent comparisons) drawn from 550 candidate reports; 35 published, 15 dissertations/technical reports
Population
Mixed K-12 and postsecondary (23 elementary/high-school evaluations, 27 postsecondary); mathematics 18, other subjects 32; 39 US, 11 other. Systems include Cognitive Tutor (15 evaluations), AutoTutor, ALEKS, Wayang Outpost and others.
Design
THE archive's canonical measure-inflation case for ed-tech, because the same 50 studies produce a famous number and a boring one depending only on which test was used. Primary estimator is Glass's ES, unweighted, with evaluation study as the unit and 90% Winsorization of outliers; Hedges's g and weighted random-effects estimates are reported as supplements and are systematically lower (weighted g = 0.50, 95% CI 0.40-0.59). Inclusion rules are strict in ways that matter and are themselves selective: control groups had to receive CONVENTIONAL instruction (studies with human-tutor, CAI-tutor or no-instruction controls were excluded); pretest gaps >= 0.5 SD excluded; overaligned outcome measures excluded; and - critically - "the treatment had to be implemented without major failures", which removes exactly the implementation failures that characterise ordinary school deployment. Only 13 of 50 evaluations randomized; 32 used intact groups. Publication bias was handled only as a published/unpublished moderator (0.67 vs 0.46); no funnel plot, trim-and-fill or PET-PEESE was performed. Fletcher is at the Institute for Defense Analyses and Kulik at Michigan; neither is a vendor, but the paper is written from inside the ITS research community and its discussion argues that studies finding nulls "were weak in design or execution".
Key findings
The median ITS effect across 50 evaluations is 0.66 SD - the number everyone quotes - and it is almost entirely an artefact of outcome-measure alignment. Test type is the strongest moderator in the paper (r = -.63): local, researcher-developed tests give 0.73 (k = 38), a mix gives 0.45 (k = 3), and independent standardized tests give 0.13 (k = 9; 0.09 under Hedges g homogeneity analysis). A 5-6x gap. Within Cognitive Tutor alone the median is 0.86 on local tests and 0.16 on standardized. The same deflation shows up on every other quality axis: studies with <= 80 participants 0.78 vs > 250 participants 0.30; conventional controls 0.60 vs controls given ITS-derived materials 0.18; strong implementations 0.44 vs weak implementations -0.01. The paper's own K-12 mathematics subgroup on standardized measures only is 0.10 - the same null the large RCT literature reports.
Genetic confound
Medium - only 13 of 50 evaluations randomized (mean 0.55) and 32 used intact groups (mean 0.67); pretest gaps up to 0.5 SD were tolerated. Not the main threat here; measure alignment is.
Replication notes
The headline 0.66 is contradicted by Steenbergen-Hu & Cooper (2013), who found ~0.05 on K-12 mathematics; Kulik & Fletcher devote a discussion section to the disagreement and attribute it to definition and control-group standards. It is ALSO contradicted from inside this paper: the same 50 studies give 0.13 on standardized tests, and the paper's own K-12 mathematics subgroup gives 0.10 on standardized-only measures - a number identical to the large randomized literature (Cheung & Slavin 2013 large randomized +0.06; Pane 2014 ~0 to +0.21). The measure-type gap itself is heavily replicated (Cohen/Kulik/Kulik 1982: 0.84 vs 0.27; Cheung & Slavin 2016; Xu 2019 reading-comprehension ITS: researcher-developed measures 0.73 SD larger).
DOI / URL
10.3102/0034654315581420
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| ITS vs conventional instruction, overall (the headline) | ES (Glass) | median 0.66, mean 0.65 (SD 0.56); 90% Winsorized mean 0.61; 58 of 63 comparisons favoured ITS. Weighted random-effects Hedges g = 0.50 (95% CI 0.40 to 0.59, p < .001) by study, g = 0.49 (95% CI 0.40 to 0.58) by comparison - the weighted estimate is a third lower than the quoted median. | mixed | end of treatment (30 minutes to a full school year) | business-as-usual | end-of-treatment | domain-skill |
| ITS effect measured on INDEPENDENT STANDARDIZED tests | ES (Glass) | 0.13 (SD 0.17, k = 9 evaluations); 0.09 under the Hedges g homogeneity analysis. This is the number the archive treats as the headline for this meta. | standardized | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| ITS effect measured on LOCAL researcher-developed tests | ES (Glass) | 0.73 (SD 0.32, k = 38); 0.62 under Hedges g homogeneity analysis. Local and standardized together 0.45 (k = 3). Test type is the strongest moderator in the review (r = -.63, p < .001). | researcher-designed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| Cognitive Tutor evaluations by test type (15 studies) | ES (Glass) | median 0.86 on local tests vs 0.16 on standardized (means 0.76 vs 0.12); overall median 0.35. Same product, same review, 5x apart depending on the test. | mixed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| K-12 mathematics subgroup, standardized measures only | ES (Glass) | 0.10 (8 studies); vs 0.72 in 7 studies using local tests and 0.45 in 3 using both. Overall K-12 mathematics 0.40 (18 studies). Reported in the discussion when reconciling with Steenbergen-Hu & Cooper (2013). | standardized | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| Sample-size gradient | ES (Glass) | <= 80 participants 0.78 (k = 26); 81-250 0.53 (k = 10); > 250 0.30 (k = 13); r = -.55, p < .001. Confounded with test type (r = .60) - the big studies are the ones using standardized tests. | mixed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| Control-condition quality (11 experiments in 6 excluded reports) | ES (Glass) | median 0.24, mean 0.18 when the control group read ITS-derived "canned text" or watched recorded tutoring sessions, vs mean 0.60 with conventional classroom controls. Giving controls the same CONTENT without the software erases two-thirds of the effect. | researcher-designed | end of treatment | active-alternative | end-of-treatment | domain-skill |
| Implementation adequacy (4 studies reporting both strong and weak implementations) | ES (Glass) | median 0.44 for stronger implementations vs -0.01 for weaker ones. In Koedinger & Anderson (1993) the experienced teacher got 0.96 and teachers new to the system got -0.23; the inexperienced teachers treated the ITS as a replacement for themselves and graded papers while it ran. | researcher-designed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| Design and publication moderators | ES (Glass) | random assignment 0.55 (k = 13) vs intact groups 0.67 (k = 32); published 0.67 (k = 35) vs unpublished 0.46 (k = 15); postsecondary 0.75 (k = 27) vs elementary/high school 0.44 (k = 23); mathematics 0.40 (k = 18) vs other subjects 0.72 (k = 32); constructed-response tests 0.84 vs objective tests 0.53. Assignment and publication were non-significant in the ANOVA but significant in the Hedges homogeneity analysis. NO funnel plot, trim-and-fill or PET-PEESE was run. | mixed | end of treatment | business-as-usual | end-of-treatment | domain-skill |