Distributed Practice in Verbal Recall Tasks: A Review and Quantitative Synthesis
Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. · 2006
grade Cmeta-analysisindependentreplicatednumbers spot-checked
Sample
317 experiments / 184 articles (out of 427 reviewed); 958 accuracy values, 839 assessments of distributed practice, 169 effect sizes. The headline massed-vs-spaced analysis rests on 271 comparisons from 254 studies, 14,811 participants.
Population
Overwhelmingly adult lab subjects (85% young adults); artificial verbal materials (word lists, paired associates, trivia). Recall tests only.
Design
Anchor spacing meta. Mostly within-subject spacing manipulations, study time held equal across massed and spaced arms. Analysis is largely by raw accuracy difference rather than d, because most primary studies did not report the variances needed for effect sizes. 80% of comparisons used a retention interval under 1 day and only ~4% used one over 1 month.
Key findings
Definitive demonstration that distributed beats massed for verbal retention: 36.7% vs 47.3% correct pooled (+10.6 percentage points, p < .001), with only 12 of 271 comparisons null or negative. Three qualifications the record previously lacked, all of which cut against reading this as strong support for classroom spacing. (1) The corpus is a SHORT-interval literature — 229 of the 271 comparisons test recall within 10 minutes of study, and in three of the four bins with retention intervals from 10 minutes to 7 days the spacing benefit is directionally positive but NOT statistically significant. (2) The retention-interval moderation is real and directional, not just "space more": optimal ISI grows with retention interval, and over-spacing HURTS (29-84 day ISI worse than 2-28 day ISI at long delays, p < .01). (3) The authors explicitly decline to generalise to children at educationally relevant delays — no usable data exist. All outcomes are recall of the trained verbal material: domain-skill, not transfer.
Genetic confound
Low: within-subject manipulations hold ability constant; a within-learner scheduling effect, not ability selection.
Replication notes
Effect >100 years old and very robust in direction; long (>1 month) intervals underrepresented (~4% of data). The authors note there is no way to count unpublished null findings, but argue the file-drawer problem is a non-issue given the effect's size.
DOI / URL
10.1037/0033-2909.132.3.354
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Spaced vs massed, pooled across ALL retention intervals (Table 1) | % correct | 36.7% massed vs 47.3% spaced = +10.6 percentage points, t(540) = 6.6, p < .001. NOTE the record previously reported "+15%" as the overall benefit; the paper's 15% figure is NOT the pooled effect, it is the average benefit in the four studies with a retention interval over one month (Discussion, "Educational Implications") | researcher-designed | pooled, retention intervals from 1 s to ~8 years, 80% under 1 day | active-alternative | unclear | domain-skill |
| Spaced vs massed at SHORT retention intervals (under 10 min) | % correct | 1-59 s bin 41.2 vs 50.1 (+8.9 pts, p < .001, 105 comparisons); 1-10 min bin 33.8 vs 44.8 (+11.0 pts, p < .001, 124 comparisons). Together these two bins are 229 of the 271 comparisons — the meta is overwhelmingly a short-retention-interval literature | researcher-designed | final test within 10 minutes of study | active-alternative | end-of-treatment | domain-skill |
| Spaced vs massed at EDUCATIONALLY RELEVANT retention intervals (Table 1) — cuts against the strength of the verdict | % correct | directionally positive in every bin but mostly NOT statistically significant: 10 min-1 day 40.6 vs 47.9, t(20) = 0.6, p = .535 (11 comparisons); 1 day 32.9 vs 43.0, t(28) = 1.2, p = .249 (15); 2-7 days 31.1 vs 45.4, t(16) = 1.4, p = .190 (9); 8-30 days 32.8 vs 62.2, t(10) = 2.3, p < .05 (6); 31+ days 17 vs 39 from a SINGLE comparison in a single study, no test reported | researcher-designed | final test 10 min to years after study | active-alternative | under-1yr | domain-skill |
| Retention-interval moderation of the LAG effect: optimal inter-study interval grows with retention interval (Table 3) | % correct at shorter vs longer ISI | no ISI effect when RI < 1 day (1-10 s vs 11-29 s ISI: 1.6 vs 3.9, p = .077; 30-59 s vs 1 day ISI: 3.4 vs 1.0, p = .397). At RI = 1 day, a 1-day ISI beat a 1-15 min ISI, 17.5 vs 6.4, p < .05. At RI = 2-28 days, 10.3 vs 1.5, p < .05. This is the paper's central and best-supported claim | researcher-designed | final test 1 day to 28 days after study | active-alternative | under-1yr | domain-skill |
| Over-spacing hurts: too-long ISIs at long retention intervals | % correct at shorter vs longer ISI | at RI 30-2900 days, a 29-84 day ISI performed WORSE than a 2-28 day ISI, -0.6 vs 9.0, t(17) = 3.0, p < .01; at RI 2-28 days a 2-28 day ISI was marginally worse than a 1-day ISI (3.5 vs 10.3, p = .091). "Space it as far apart as possible" is not what this meta supports | researcher-designed | final test 1 month to 8 years after study | active-alternative | over-2yr | domain-skill |
| Robustness | null/negative rate | only 12 of 271 massed-vs-spaced comparisons null or negative; most of the 12 were paired-associate tasks, the same task type as studies that did show the benefit | researcher-designed | lab | active-alternative | unclear | domain-skill |
| Evidence in CHILDREN at educationally relevant delays | coverage | 85% of the data come from young adults; the authors state that for retention intervals of one day or longer "no usable data" on children exist and "we cannot say for certain that children's long-term memory will benefit from distributed practice" | researcher-designed | not applicable | none | not-applicable | domain-skill |
Cited by
- Spaced (distributed) practicestrong supportconf: highgc: low