A meta-analysis of the effectiveness of intelligent tutoring systems on college students’ academic learning
Steenbergen-Hu, S., & Cooper, H. · 2014
grade Cmeta-analysisindependentmixed
Sample
35 reports containing 39 studies, covering 22 different ITS
Population
College and other higher-education students; most-studied systems AutoTutor, ALEKS (Assessment and Learning in Knowledge Spaces), xTES (eXtended Tutor-Expert System) and Web Interface for Statistics Education. Subject domains mixed (physics, computer literacy, statistics, mathematics).
Design
Companion to the same authors' 2013 K-12 mathematics meta-analysis, and the pair is the archive's cleanest natural experiment on POPULATION rather than technology: identical reviewers, identical methods, four times the effect once the students are self-selected adults in voluntary courses. The control-group moderator is the load-bearing result and it is exactly the "what did the control group do with that hour" question - ITS beat a no-instruction control by 0.86, beat a conventional class by 0.37, and LOST to human tutoring by 0.25. Effects were also significantly larger in earlier studies than in more recent ones, the standard decay signature. Typical college ITS evaluations are short (hours to weeks) and use instructor- or researcher-built posttests aligned to the tutored material, which is the main reason the college numbers should not be read as comparable to the K-12 standardized-test numbers. NOT READ IN FULL - APA paywall; unpaywall lists a Duke institutional-repository copy (hdl.handle.net/10161/9536) that is behind bot protection and could not be retrieved. Numbers below come from the paper's ERIC abstract plus Kulik & Fletcher (2016, p. 45), which reports the control-condition breakdown from the tables.
Key findings
Intelligent tutoring systems raise college students' academic learning by g = 0.32 to 0.37 overall - moderate, and roughly four times what the same reviewers found for K-12 mathematics. The number that matters is the control-condition breakdown, because it decomposes the headline into three different quantities: +0.86 against a control that received no instruction at all, +0.37 against conventional classroom instruction, and -0.25 against human tutoring. ITS outperformed every non-human comparison (classroom teaching, printed or computerized reading material, ordinary CAI, laboratory and homework assignments, no treatment) and underperformed one-to-one humans. Effects in earlier studies were significantly larger than in more recent studies. Differences across ITS types, subject domains and modes of implementation were not significant.
Genetic confound
Medium - college ITS studies are typically within-course randomizations of volunteers, so between-arm genetic differences are unlikely, but selection INTO the courses and the samples (self-selected postsecondary students) is what makes these estimates non-transportable to K-12.
Replication notes
The postsecondary-beats-K-12 pattern replicates in Kulik & Fletcher (2016), whose postsecondary subgroup averages 0.75 vs 0.44 for elementary/high school, and in Ma et al. (2014). The control-condition ordering (no instruction > conventional class > human tutoring) replicates in Ma et al. (2014), which likewise finds ITS beating large-group instruction (g = 0.42) and losing slightly to individual human tutoring (g = -0.11). What does NOT replicate across the pair is the magnitude: the same reviewers found g = 0.01-0.09 on K-12 mathematics one year earlier. Flagged mixed because the college estimate is not corroborated on any standardized third-party measure - this literature is dominated by course-aligned instructor-built posttests.
DOI / URL
10.1037/a0034752
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| College academic learning, ITS vs all comparison conditions pooled | g | 0.32 to 0.37 depending on model - the authors report a range rather than a single point estimate; Kulik & Fletcher summarise it as "approximately 0.35 standard deviations" | mixed | end of course or module | unclear | end-of-treatment | domain-skill |
| ITS vs a control receiving NO instruction | g | +0.86 - the largest contrast in the review, and the one with the least practical meaning | mixed | end of course or module | none | end-of-treatment | domain-skill |
| ITS vs conventional classroom instruction | g | +0.37 - the contrast that answers the actual substitution question for a college course | mixed | end of course or module | business-as-usual | end-of-treatment | domain-skill |
| ITS vs human tutoring | g | -0.25 - ITS scored a quarter of a standard deviation BELOW human-tutored controls. The "software matches a human tutor" claim is not supported by this meta-analysis, though the gap is far smaller than Bloom's two sigma. | mixed | end of course or module | active-alternative | end-of-treatment | domain-skill |
| Study-recency moderator | g | effectiveness in earlier studies was significantly greater than in more recent studies - the same decay-over-time pattern the archive sees across ed-tech. No significant differences by ITS type, subject domain, or how the ITS was integrated into instruction. | mixed | end of course or module | unclear | end-of-treatment | domain-skill |