The Evidence on Teaching

How features of educational technology applications affect student reading outcomes: A meta-analysis

Cheung, A. C. K., & Slavin, R. E. · 2012

grade Cmeta-analysisdeveloper-involvedreplicatednumbers spot-checked
Sample
84 studies, 60,553 K-12 participants (8 kindergarten studies, N = 2,068; 59 elementary, N = 34,200; 18 secondary, N = 24,285)
Population
K-12 reading, any country if reported in English, overwhelmingly US. Four programme categories - supplemental CAI (Destination Reading, PLATO Focus, Waterford, WICAT), computer-managed learning (Accelerated Reader only), innovative technology applications (Fast ForWord, Reading Reels, Lightspan) and comprehensive models (READ 180, Writing to Read, Voyager Passport).
Design
Companion to the 2013 mathematics review; same Best Evidence Encyclopedia inclusion standards. Criterion 7 excluded measures of objectives inherent to the programme and also excluded phonemic awareness, oral vocabulary and writing measures, so the pool is dominated by standardized reading tests; comprehensive experimenter-made reading measures were admitted only where fair to controls. Random assignment or pretest-adjusted matching required; pretest gaps > 0.5 SD excluded; 12-week minimum; two teachers per arm. Random effects (Q = 362.52, df = 83, p < .001). INDEPENDENCE CAVEAT, and this is why the flag is developer-involved rather than independent: the "innovative technology applications" category the review calls promising contains Reading Reels, the embedded-multimedia component of Slavin's own Success for All model, and the two qualifying Reading Reels trials (Chambers et al. 2006, 2008; ES +0.17 and +0.27) are co-authored by both reviewers. The conflict touches 2 of 84 studies and one of four programme categories; it does not touch the deflationary headline, which runs against the reviewers' interest. Version read: the April 2012 Johns Hopkins CRRE/BEE technical report ("The Effectiveness of Educational Technology Applications for Enhancing Reading Achievement in K-12 Classrooms: A Meta-Analysis"), whose abstract, N, and all reported effect sizes match the published Educational Research Review 7(3) 198-215 article; the published version's title is the one recorded above.
Key findings
Educational technology moves K-12 reading +0.16 SD overall, and again the effect is a function of study quality rather than of technology: randomized experiments give +0.08, large randomized experiments +0.07, and small studies twice what large ones give (+0.25 vs +0.13). The supplementary CAI that dominates actual classroom use - the category the federal Dynarski/Campuzano RCTs evaluated - returns +0.11, which the authors call not educationally meaningful. Unlike the mathematics companion, this review DOES detect publication bias: published articles +0.25 vs technical reports and dissertations +0.14 (p < .04). The categories that look better (comprehensive integrated models such as READ 180, +0.28; innovative applications, +0.18) contain no randomized studies at all for the comprehensive models, and the authors caution that non-randomized studies overstate effects.
Genetic confound
Medium - 59 of 84 studies are matched or post-hoc-matched quasi-experiments; only 25 are randomized. Within-study subgroup results by ability, gender and race are fixed-effects post-hoc comparisons on small k and should not be treated as causal.
Replication notes
The design-quality gradient (randomized +0.08, large randomized +0.07, small studies twice large studies) replicates exactly in the authors' companion mathematics meta-analysis (Cheung & Slavin 2013) and matches the federal Dynarski/Campuzano reading RCTs (+0.04 grade 1, +0.02 grade 4) that sit inside this pool. Cheung & Slavin (2016) generalises the gradient. The favourable categories (comprehensive models +0.28, innovative applications +0.18) rest on small, non-randomized study counts and are NOT replicated - the authors say so themselves.
DOI / URL
10.1016/j.edurev.2012.05.002

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Reading achievement, all 84 studiesES+0.16 (random effects; one-study-removal sensitivity range 0.12 to 0.21; Q = 362.52, df = 83, p < .001)standardizedend of programme (minimum 12 weeks; most one school year)business-as-usualend-of-treatmentdomain-skill
Supplemental CAI (the dominant classroom use)ES+0.11 across 56 studies - the authors' verdict is that this category does not produce "educationally meaningful effects"; 19 of the 57 supplemental studies were randomizedstandardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Comprehensive integrated models (READ 180, Writing to Read, Voyager Passport)ES+0.28 across 18 studies - the largest category effect, but NONE of the READ 180 or Voyager Passport studies was randomized, and these programmes bundle 90 min/day of small-group teacher instruction with the software, so they do not isolate technologystandardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Innovative technology applications (Fast ForWord, Reading Reels, Lightspan)ES+0.18 across 6 studies; includes the reviewers' own Reading Reels trials (+0.17, +0.27) - see the independence caveat in design_notesstandardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Computer-managed learning (Accelerated Reader)ES+0.19 across 4 studies; QB across the four programme categories = 7.15, df = 3, p < .07 (marginal)standardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Randomized experiments vs quasi-experimentsESrandomized (k = 25) +0.08 vs quasi-experimental (k = 59) +0.19 - quasi-experiments return 2.4x. By design and size - large randomized +0.07, small randomized +0.21, large matched +0.16, small matched +0.24 (QB = 12.37, p < .001).standardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Sample size gradientESlarge studies (N > 250, k = 49) +0.13 vs small (N < 250, k = 35) +0.25; QB = 4.66, p < .03. The report counts the small studies as 35 in one sentence and 40 in the next - an internal inconsistency in the source, not a transcription error here.standardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Publication biasESpublished articles +0.25 vs unpublished technical reports and dissertations +0.14; QB = 4.44, df = 1, p < .04 - publication bias IS present here, unlike in the companion mathematics review. Classic fail-safe N = 4,198; Orwin fail-safe N = 880. No funnel/trim-and-fill or PET-PEESE.standardizednot applicablenonenot-applicabledomain-skill
Implementation quality moderatorESlow implementation +0.01, medium +0.18, high +0.22 - no effect at all where implementation was rated low; 53% of studies gave insufficient implementation information, and the authors warn the rating is contaminated by post-hoc explanation of nullsstandardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Grade level and within-study ability subgroupsESkindergarten +0.15, elementary +0.10, secondary +0.31 (QB = 9.52, p < .01) - but only 2 of 18 secondary studies were randomized and the category is dominated by 8 READ 180 and 3 Accelerated Reader studies. Within-study subgroups (fixed effects, small k): low ability +0.37, middle +0.27, high +0.08; male +0.28 vs female +0.12; English learners +0.29 (3 studies).standardizedend of programmebusiness-as-usualend-of-treatmentdomain-skill

Cited by