Don't throw away your printed books: A meta-analysis on the effects of reading media on reading comprehension
Delgado, P., Vargas, C., Ackerman, R., & Salmerón, L. · 2018
grade Cmeta-analysisindependentreplicatednumbers spot-checked
Sample
54 studies (38 between-participant, 16 within-participant) yielding 56 + 18 effect sizes; 171,055 participants in total. Studies published 2000-2017.
Population
Readers compared on comparable texts across media. Educational level was a coded moderator — Grades 1-6 (k = 8), Grades 7-12 (k = 8), undergraduates (k = 36), graduates/professionals (k = 3, dropped from moderator analysis). Undergraduates therefore dominate the corpus; the school-age evidence is a minority of it.
Design
Meta-analysis of experimental within- and between-participant comparisons of reading the SAME text on paper versus on a screen, with reading comprehension as the outcome. Negative Hedges' g favours paper. This is the "screen as delivery medium" question and must not be confused with the "screen as competing attraction" question (phones, multitasking): here there is no distraction, no notification and no alternative activity — the only thing that changes is the surface the text is printed on. Random-effects models; moderator analysis by one-way ANOVA (QB) for categorical moderators and meta-regression for continuous ones. Independent academic work (Universities of Valencia and Technion), no device or publisher funding declared. WEAKNESSES: (1) heterogeneity is large and mostly unexplained — QW is significant in every moderator cell, and the moderators together explain only a fraction of it; (2) outcome measures are overwhelmingly researcher-designed comprehension tests administered immediately after reading, so nothing here speaks to durable learning; (3) the corpus is undergraduate-heavy, and the two school-age cells are k = 8 each; (4) the paper-advantage estimate is an average over wildly different digital conditions (42 computer effect sizes, 14 hand-held), and the hand-held subgroup is NOT significant (g = -.12, p = .11) — a fact usually dropped when the headline -.21 is cited; (5) the method of allocating participants to media conditions was tested as a moderator and was not significant, which is reassuring but does not make the underlying studies randomized trials. Displacement was NOT measured: no study here recorded what screen reading displaced (time, attention, re-reading), only how the comprehension score differed.
Key findings
Reading the same text on a screen costs about a fifth of a standard deviation of comprehension relative to paper (g = -.21, 95% CI [-.28, -.14]), and the same figure appears independently in within-participant designs (dc = -.21). The finding that matters most for policy is the time trend: a meta-regression on publication year is significant and NEGATIVE (beta = -.01 per year, QR = 4.95, p = .03, R2 = .64) — the paper advantage has GROWN by about .01 per year since 2000 rather than shrinking as digital-native cohorts arrived. That directly falsifies the standard "the kids will adapt" defence, though it rests on one meta-regression and should be treated as the paper's most interesting rather than its most secure result. The cost is concentrated where school actually happens: under time pressure (g = -.26 vs -.09 self-paced) and on informational text (g = -.27; narrative text shows no media effect at all, g = +.01). The authors benchmark -.21 against a year's growth in elementary reading comprehension (about .32) and against the mean reading intervention (about .45), i.e. roughly two-thirds of a school year of comprehension growth.
Genetic confound
Low as a threat to the medium contrast — most included comparisons randomly allocate the same readers or matched readers across media, so genes cannot differ between arms. The moderator analyses (education level, genre) are between-study comparisons and carry the usual confounding.
Replication notes
The core screen-inferiority effect is one of the better-replicated findings in this cluster. Delgado et al. themselves report it twice on non-overlapping evidence — between-participant designs (g = -.21, k = 56) and within-participant designs (dc = -.21, k = 18) — and it survives five different estimators (g between -.22 and -.20; dc between -.18 and -.24), removal of grey literature (g = -.19), and three publication-bias tests (Rosenthal fail-safe N above criterion; Egger p = .39 and p = .20). It agrees with the earlier meta-analyses it supersedes (Wang et al. 2007; Kong et al. 2018; Singer & Alexander 2017) and with Clinton (2019, Journal of Research in Reading, doi 10.1111/1467-9817.12269), which on an independently assembled corpus of 33 studies and 2,799 participants found g = -0.25 overall and reproduced the genre moderator almost exactly (expository -0.32, narrative -0.04). The MODERATORS are much less replicated than the main effect: the publication-year trend rests on a single meta-regression of 56 effect sizes (QR = 4.95, p = .03) and the narrative-text null on only k = 7, so those two should be read as the weakest, not the strongest, parts of the paper.
DOI / URL
10.1016/j.edurev.2018.09.003
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Reading comprehension, screen vs paper (between-participant designs) | Hedges g | -0.21, 95% CI [-0.28, -0.14], k = 56; robust to estimator choice (-0.22 to -0.20) and to dropping grey literature (-0.19, k = 38) | mixed | immediately after reading the text | active-alternative | end-of-treatment | domain-skill |
| Reading comprehension, screen vs paper (within-participant designs) | dc (corrected d for repeated measures) | -0.21, 95% CI [-0.37, -0.06], k = 18; -0.18 to -0.24 across estimators | mixed | immediately after reading the text | active-alternative | end-of-treatment | domain-skill |
| Change in the paper advantage over calendar time (publication-year meta-regression) | beta per year on Hedges g | -0.01 per year (QR = 4.95, p = .03, R2 = .64) — the paper advantage GREW from 2000 to 2017; single meta-regression, so the least robust of the three significant moderators | mixed | study publication year, 2000-2017 | active-alternative | not-applicable | domain-skill |
| Reading comprehension by time frame (moderator) | Hedges g | time-limited reading -0.26 [-0.35, -0.16] (k = 27) vs self-paced -0.09 [-0.22, 0.05] (k = 20); QB = 4.12, p = .04, R2 = .05 | mixed | immediately after reading the text | active-alternative | end-of-treatment | domain-skill |
| Reading comprehension by text genre (moderator) | Hedges g | informational -0.27 [-0.36, -0.18] (k = 34), mixed -0.30 [-0.40, -0.21] (k = 10), narrative +0.01 [-0.20, 0.20] (k = 7); QB = 7.00, p < .05, R2 = .31 | mixed | immediately after reading the text | active-alternative | end-of-treatment | domain-skill |
| Reading comprehension by device type (moderator, NOT significant) | Hedges g | computer -0.23 [-0.31, -0.15] (k = 42) vs hand-held -0.12 [-0.27, 0.03], p = .11 (k = 14); QB = 1.55, ns — the hand-held cell does not reach significance and this is routinely dropped when the headline -.21 is quoted | mixed | immediately after reading the text | active-alternative | end-of-treatment | domain-skill |
| Reading comprehension by school stage (moderator, NOT significant) | Hedges g | Grades 1-6 -0.19 [-0.35, -0.03] (k = 8); Grades 7-12 -0.15 [-0.29, -0.02] (k = 8); undergraduates -0.28 [-0.38, -0.18] (k = 36); QB = 2.33, ns, R2 = .00 | mixed | immediately after reading the text | active-alternative | end-of-treatment | domain-skill |