Retrieval practice (the testing effect)
Testing yourself beats rereading — robust in real classrooms; honest durable size ~0.1–0.3 SD, biggest after delay, thinnest on transfer.
strong supportconf: highgc: lowpractice · ages 8–18 · method
Retrieving information from memory (low-stakes quizzing, self-testing) produces more durable learning than rereading. Robust in real classrooms (222-study meta g≈0.50) but the honest standardized/durable magnitude is ~0.1-0.3 SD: the effect is largest on delayed tests, concentrated on material that was actually tested, thins on far transfer, and can REVERSE on complex procedural content.
Replace some rereading/review with low-stakes retrieval (quizzes, flashcards, brain-dumps, past-paper practice) ACROSS every subject — always with feedback. Space the retrieval out and expect the payoff on delayed tests, not the same day.
Who this applies to
Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.
Verdict
Retrieving information from memory — being quizzed, self-testing, doing practice questions — makes learning stick better than restudying the same material. This is one of the two best-supported cross-cutting methods in all of education (Dunlosky's team rates it, with spacing, as the only "high utility" study techniques), it replicates in real classrooms across subjects and ages, and it is genetic-confound-clean (the core studies randomize within students). It is the closest thing the learning sciences have to phonics: a robust core finding that must nonetheless be scoped carefully to avoid overclaiming.
The scope: (1) the benefit shows up on delayed tests, not immediately — at very short delays, restudy can win; (2) it is largest on the material actually tested and thins on far transfer; (3) headline effect sizes (g≈0.5) come from researcher-aligned tests, so the honest standardized/durable magnitude is roughly 0.1-0.3 SD; and (4) on complex procedural content it can even reverse (restudy wins). The practical verdict is still strongly positive — but "quiz more" is a reliable ~0.1-0.3 SD lever, not a magic bullet.
What the evidence shows
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Roediger & Karpicke 2006 | lab RCT (the crossover) | B | 5 min: restudy wins (~81% vs 75%); 1 week: testing wins (~61% vs 40%) |
| Yang 2021 | 222-study classroom meta (48k students) | B | Overall g=0.50; transfer to untested material 0.32; feedback 0.54 vs 0.37 |
| Adesope 2017 | meta (~217 studies) | C | vs restudy 0.51; vs no-activity 0.93 (mostly lab) |
| Rowland 2014 | testing-vs-restudy meta | C | g=0.50; 81% favor testing; effortful recall > recognition |
| Pan & Rickard 2018 | transfer meta (N=10,382) | C | Transfer 0.40 overall, but untested-but-studied material weakest (~0.15-0.20), near-zero without feedback |
| CVS classroom study 2024 | grade 5-6 classroom | C | Restudy BEAT retrieval on complex procedural content, immediate AND delayed |
| Bjork, Dunlosky & Kornell 2013 | narrative review (metacognition) | D | Only 18% self-test to learn; ~70% self-test only to check what they know |
The crossover is the load-bearing fact: retrieval is a durability method. It costs you a little now (retrieving is harder than rereading) and pays you back later. That is why students dislike it — rereading feels more effective in the moment (the same metacognitive illusion behind massing).
Added 2026-07-30 — the citation for a claim this file was already making. The sentence above ("that is why students dislike it") was asserted without a source until a citation-graph audit surfaced Bjork, Dunlosky & Kornell 2013. It does not move the verdict or the confidence — it is a narrative review (grade D) with no new effect size. What it supplies is the size of the belief problem: asked why they self-test, only 18% of students say "because I learn more that way," while about 70% say they do it to gauge what they know. Retrieval is thus widely used as an assessment device and rarely as a learning device, and most students believe rereading is the better strategy. In open-ended surveys only 11% name retrieval as a study strategy at all (rising to ~42% when offered as a forced choice), while 76% report rereading. The operational consequence is the one already in Practical guidance and now has evidence behind it: retrieval has to be built into the curriculum, because students will not choose it as a learning tool. Caveat worth stating plainly: this survey evidence is college students, not the 8–18 band this topic covers, so treat it as a plausible mechanism rather than a measured fact about children.
Two honest deflations from the adversarial review: the big numbers are partly measure alignment (Yang: aligned quiz/test formats 0.53 vs mismatched 0.40), and the edge over active elaborative controls (concept mapping, good note-taking) is only minimal — the large effects are mostly against rereading or doing nothing. And on complex, high-element-interactivity procedural content the effect can vanish or reverse (the 2024 CVS classroom study; the van Gog & Sweller boundary).
Hereditarian-lens assessment
Risk: low. The effect is on domain retention, never on g, and the core designs randomize within learners. One correction the genetics review forced, worth stating because it is widely repeated wrongly: retrieval practice does not show "expertise reversal." (That is a worked-examples phenomenon.) Retrieval's learner interaction is a working-memory-capacity one, and it is ability-robust to gap-narrowing: it works for lower-ability students when feedback is present (Agarwal et al. 2017 found the benefit slightly larger for lower-working-memory students; Moreira et al. 2019 found no moderation by IQ, working memory, decoding, or vocabulary). Without feedback it can be costly for low-working-memory learners (Zheng et al. 2023). So feedback is the equity lever — retrieval with feedback tends to help everyone and does not widen ability gaps.
Lab-to-classroom & durability
Unusually for a "desirable difficulty," this one does survive translation: Yang's 222-study classroom meta and Agarwal's 50-classroom review confirm real benefits on authentic course content. But those outcomes are researcher-aligned course tests, so apply the ~2× discount for a standardized equivalent. The benefit is durable by construction (it shows up at delay) — the open question is whether it transfers beyond the tested material, where the evidence (Pan & Rickard) says: weakly.
Boundaries & what critics say
- Complex procedural content: retrieval can lose to restudy (2024 CVS classroom reversal) — the element-interactivity boundary is now empirically supported, not just theoretical.
- Far transfer is weak: retrieval mostly re-strengthens what it directly tested.
- Active controls: it barely beats concept mapping / good note-taking; the big wins are vs rereading.
Practical guidance
- Build low-stakes, frequent retrieval into every subject: entry/exit quizzes, flashcards, cumulative "brain-dumps," past-paper practice. Keep stakes low so it's practice, not judgment.
- Always give feedback — it's what makes the effect robust and equitable.
- Space the retrieval (see spaced practice) and use effortful formats (short-answer/recall over multiple-choice) where feasible.
- Expect the payoff on delayed assessments; don't judge it by same-day performance.
- For brand-new, complex procedural skills, teach and let them restudy first — introduce retrieval once the basics are in place.
Open questions
- The true standardized effect size is not directly estimated (nearly all classroom evidence uses aligned course tests); ~0.1-0.3 SD is an inference.
- How far the complex-material reversal generalizes beyond procedural science skills is unresolved.
- grade BTest-Enhanced Learning: Taking Memory Tests Improves Long-Term RetentionRoediger, H. L., & Karpicke, J. D. · 2006 · rct
- grade BTesting (Quizzing) Boosts Classroom Learning: A Systematic and Meta-Analytic ReviewYang, C., Luo, L., Vadillo, M. A., Yu, R., & Shanks, D. R. · 2021 · meta-analysis
- grade CRethinking the Use of Tests: A Meta-Analysis of Practice TestingAdesope, O. O., Trevisan, D. A., & Sundararajan, N. · 2017 · meta-analysis
- grade CThe Effect of Testing Versus Restudy on Retention: A Meta-Analytic Review of the Testing EffectRowland, C. A. · 2014 · meta-analysis
- grade CRetrieval Practice Consistently Benefits Student Learning: A Systematic Review of Applied Research in Schools and ClassroomsAgarwal, P. K., Nunes, L. D., & Blunt, J. R. · 2021 · review
- grade DImproving Students' Learning With Effective Learning TechniquesDunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. · 2013 · review
- grade CTransfer of Test-Enhanced Learning: Meta-Analytic Review and SynthesisPan, S. C., & Rickard, T. C. · 2018 · meta-analysis
- grade DNot New, but Nearly Forgotten: the Testing Effect Decreases or even Disappears as the Complexity of Learning Materials Increasesvan Gog T, Sweller J · 2015 · critique
- grade CRetrieval Practice as a Tool for Teaching the Control-of-Variables Strategy: A Classroom ReversalKranz J, Möller J, Kaufmann E, Tempel T · 2024 · rct
- grade DSelf-Regulated Learning: Beliefs, Techniques, and IllusionsBjork RA, Dunlosky J, Kornell N · 2013 · review
Related decisions
- Feedback and formative assessmentmixedconf: highgc: low
- Mastery learning (teach → test → reteach to criterion → advance)mixedconf: highgc: low
- Spaced (distributed) practicestrong supportconf: highgc: low