The Evidence on Teaching

The effectiveness of educational technology applications for enhancing mathematics achievement in K-12 classrooms: A meta-analysis

Cheung, A. C. K., & Slavin, R. E. · 2013

grade Cmeta-analysisindependentreplicatednumbers spot-checked
Sample
74 studies, 56,886 K-12 students (45 elementary studies, N = 31,555; 29 secondary studies, N = 25,331); over 700 candidate reports screened, 1980-2011
Population
K-12 mathematics, any country if reported in English but overwhelmingly US; programmes grouped as supplemental CAI (Jostens, PLATO, Larson Pre-Algebra, SRA Drill), computer-managed learning (Accelerated Math only) and comprehensive models (Cognitive Tutor, I Can Learn).
Design
Best Evidence Encyclopedia review (Johns Hopkins CRRE, IES-funded), notable for inclusion standards that pre-empt most of the inflation this literature is famous for. Criterion 7 excluded measures "inherent to the program"; only standardized/state/district tests and comprehensive experimenter-made measures fair to controls were admitted, and the authors describe the pool as "primarily standardized tests". Criterion 5 required randomization or matching with pretest adjustment; pretest gaps > 0.5 SD were excluded; minimum duration 12 weeks; at least two teachers per arm; programmes had to be replicable in ordinary schools. Random-effects model (heterogeneity Q = 345.80, df = 73, p < .001); fixed-effects average was only +0.10. Abstract and discussion quote +0.15 while Table 2 gives the random-effects estimate as +0.16 - a small internal inconsistency in the paper, not a transcription error here. Independence: neither author develops mathematics technology; the review is unusual in that its conclusions cut against the sector it reviews.
Key findings
Educational technology in K-12 mathematics moves standardized achievement +0.16 SD (random effects; +0.10 fixed) - and the entire effect is carried by weak designs. Restricting to randomized experiments halves it to +0.08; restricting to LARGE randomized experiments takes it to +0.06 (95% CI 0.00 to 0.13, p = .07), i.e. statistically indistinguishable from the +0.03 of the federal Dynarski/Campuzano RCTs. Small studies returned twice the effect of large ones (+0.26 vs +0.12) inside every design category. By programme type the ordering is the reverse of the marketing: cheap supplemental drill CAI +0.19, computer-managed learning +0.09, and the flagship comprehensive adaptive models (Cognitive Tutor, I Can Learn) +0.06 and not significant. Publication bias is NOT the explanation - published and unpublished reports both average +0.15.
Genetic confound
Low for the randomized subset (26 studies); medium overall, since 48 of 74 studies are matched or post-hoc-matched quasi-experiments where selection into technology-using classrooms is plausible. The design moderator is itself the estimate of that contamination.
Replication notes
The central pattern - effects shrinking toward zero as design quality and sample size rise - replicates in the authors' own companion reading meta-analysis (Cheung & Slavin 2012, large randomized +0.07) and independently in the two federal Dynarski/Campuzano RCTs (+0.03 in math), which are themselves inside this pool. Cheung & Slavin (2016, 'How methodological features affect effect sizes in education') generalises the same design-quality gradient across subjects. No independent re-analysis of this specific study pool is known to us.
DOI / URL
10.1016/j.edurev.2013.01.001

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Mathematics achievement, all 74 studiesES+0.16 (random effects, 95% CI 0.11 to 0.20, Z = 7.14, p < .001); +0.10 fixed effects (95% CI 0.09 to 0.12); the abstract and discussion quote +0.15standardizedend of programme (minimum 12 weeks; most one school year)business-as-usualend-of-treatmentdomain-skill
Supplemental CAI programmes (Jostens, PLATO, Larson, SRA Drill)ES+0.19 (95% CI 0.14 to 0.24, p < .001); 55 effect sizes from 37 studies; the largest of the three programme categoriesstandardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Computer-managed learning (Accelerated Math)ES+0.09 (95% CI 0.00 to 0.18, p = .05); 10 effect sizes from 7 studiesstandardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Comprehensive adaptive models (Cognitive Tutor, I Can Learn)ES+0.06 (95% CI -0.04 to 0.15, p = .26) - NOT significant; 9 effect sizes from 8 studies. The most expensive category is the weakest.standardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Randomized experiments only vs quasi-experimentsESrandomized (k = 26) +0.08 (95% CI 0.04 to 0.16) vs quasi-experimental (k = 48) +0.20 (95% CI 0.14 to 0.25); QB = 7.20, p = .01. Quasi-experiments return 2.5x the randomized estimate.standardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Large randomized studies only (N > 250 and random assignment)ES+0.06 (95% CI 0.00 to 0.13, Z = 1.81, p = .07), k = 15 - the honest headline. Small randomized +0.17 (k = 11); large matched +0.16 (k = 29); small matched +0.31 (k = 19).standardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Sample size gradientESlarge studies (N > 250, k = 44) +0.12 (95% CI 0.08 to 0.17) vs small studies (N < 250, k = 30) +0.26 (95% CI 0.16 to 0.36); QB = 6.13, p = .01. The doubling holds within every design category.standardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Publication bias checksESpublished (k = 18) +0.15 vs unpublished (k = 56) +0.15, QB = 0.01, p = .94 - no publication bias detectable. Classic fail-safe N = 3,506-3,629 null studies; Orwin fail-safe N = 701-702. No funnel/trim-and-fill or PET-PEESE was run.standardizednot applicablenonenot-applicabledomain-skill
Implementation quality moderatorEShigh implementation +0.26 vs low/medium +0.12 (significant) - but 41% of studies gave no implementation information and the authors rated those themselves, and they warn the rating is contaminated (null results invite 'poor implementation' as the explanation).standardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Dose (intended minutes per week) and grade levelESintensity: <30 min/wk +0.06, 30-75 min/wk +0.20, >75 min/wk +0.14 (QB = 5.85, p = .05). Grade: elementary +0.17 vs secondary +0.14, not significant (p = .51). SES: low +0.12, high +0.25, not significant.standardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill

Cited by