Refutation Text Facilitates Learning: a Meta-Analysis of Between-Subjects Experiments
Schroeder NL, Kucera AC · 2022
grade Cmeta-analysisindependentreplicated
Sample
44 independent comparisons, 3,869 participants
Population
Readers of refutation vs non-refutation text on topics where a misconception is documented. Overwhelmingly post-secondary (k = 30 of 44); only k = 10 comparisons come from K-12 (primary k = 2, middle k = 4, secondary k = 4). Europe k = 21, North America k = 18.
Design
The best-reported estimate in this literature and the one to quote. Restricted to BETWEEN-SUBJECTS experiments, which removes the within-subject designs that inflate the older corpus; 39 of 44 comparisons randomly assigned participants. Two structural limits matter more than any moderator. FIRST, the control condition is almost always another text (expository/scientific text k = 30, unspecified text k = 10), so this is a TEXT-STRUCTURE contrast, not a test of instruction against business-as-usual teaching. SECOND, the outcome is in every case a bespoke comprehension or misconception measure built by the study authors around the very misconception the text refutes; the measure-type moderator the authors report is ITEM FORMAT (multiple choice, open-ended, true/false, Likert), not measure provenance. There is no standardized-instrument arm anywhere in this meta-analysis, so the researcher-designed-vs-independent inflation factor cannot be computed from it - which is itself the finding. Dose is a single reading session of a few hundred words.
Key findings
Refutation text beats non-refutation text at g = 0.41, 95% CI [0.30, 0.51], p < .001, k = 44, n = 3,869; heterogeneity Q(43) = 109.59, p < .001, I-squared = 60.76. Publication bias: Egger t(42) = 0.55, one-tailed p = 0.29 (no asymmetry), fail-safe N = 1,625, but trim-and-fill imputed 10 studies and cut the estimate to g = 0.28 [0.16, 0.39] - so the defensible range is roughly 0.28 to 0.41. THE DURABILITY RESULT IS THE HEADLINE FOR THIS TOPIC AND IT IS FLAT: same day g = 0.39 (k = 17), 2 days to 1 week g = 0.38 (k = 11), 8 days to 1 month g = 0.43 (k = 13), more than 1 month g = 0.56 (k = 2); Q-between = 0.48, p = 0.98. The refutation advantage does not decay over the horizons tested - but the longest cell holds two comparisons, so "more than a month" is not evidenced, it is merely not contradicted. Domain does NOT moderate (science g = 0.45 k = 33, mathematics 0.32 k = 5, social science 0.31 k = 6; Q-between = 1.09, p = 0.58): this is a general text-comprehension finding wearing science clothes. Age does not significantly moderate either (Q-between = 6.60, p = 0.25) although the point estimates run the wrong way for the adult-heavy corpus: primary 0.57 (k = 2), middle 0.71 (k = 4), secondary 0.49 (k = 4), post-secondary 0.33 (k = 30). Item format does not moderate (Q-between = 3.76, p = 0.58). The one significant moderator is PUBLICATION TYPE, Q = 9.01, p = 0.01: conference proceedings g = 0.73 (k = 2), journal articles 0.45 (k = 36), dissertations 0.11, NOT SIGNIFICANT (k = 6). The unpublished-thesis subset finds essentially nothing.
Genetic confound
Low. Randomised between-subjects contrasts on a knowledge outcome; genes cannot differ between arms.
Replication notes
Reproduces the direction of Guzzetti et al. 1993 and Tippett 2010 on a modern, between-subjects-only corpus, at roughly a third of a standard deviation rather than a large effect.
DOI / URL
10.1007/s10648-021-09656-z
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Refutation text vs non-refutation text, overall | Hedges g | 0.41 (95% CI 0.30-0.51), k = 44, n = 3,869 | researcher-designed | pooled across immediate and delayed post-tests | active-alternative | unclear | domain-skill |
| Overall effect after trim-and-fill correction for small-study bias | Hedges g | 0.28 (95% CI 0.16-0.39), 10 studies imputed | researcher-designed | pooled | active-alternative | unclear | domain-skill |
| Effect at same-day post-test | Hedges g | 0.39 (k = 17) | researcher-designed | same day | active-alternative | end-of-treatment | domain-skill |
| Effect at 2 days to 1 week | Hedges g | 0.38 (k = 11) | researcher-designed | 2 days to 1 week | active-alternative | under-1yr | domain-skill |
| Effect at 8 days to 1 month | Hedges g | 0.43 (k = 13) | researcher-designed | 8 days to 1 month | active-alternative | under-1yr | domain-skill |
| Effect at more than 1 month | Hedges g | 0.56 (k = 2) - two comparisons only; Q-between across all timing bands 0.48, p = 0.98 | researcher-designed | more than 1 month | active-alternative | under-1yr | domain-skill |
| Domain moderator - is this a science finding or a text finding? | Hedges g by subgroup | science 0.45 (k = 33), mathematics 0.32 (k = 5), social science 0.31 (k = 6); Q-between = 1.09, p = 0.58 (not significant) | researcher-designed | pooled | active-alternative | unclear | domain-skill |
| Age/education moderator | Hedges g by subgroup | primary K-5 0.57 (k = 2), middle 6-8 0.71 (k = 4), secondary 9-12 0.49 (k = 4), post-secondary 0.33 (k = 30); Q-between = 6.60, p = 0.25 (not significant) | researcher-designed | pooled | active-alternative | unclear | domain-skill |
| Publication-type moderator (the only significant one) | Hedges g by subgroup | conference proceedings 0.73 (k = 2), journal articles 0.45 (k = 36), dissertations 0.11 NOT SIGNIFICANT (k = 6); Q = 9.01, p = 0.01 | researcher-designed | pooled | active-alternative | unclear | domain-skill |
Cited by
- Science misconceptions and conceptual change — can naive intuitions be taught away?mixedconf: mediumgc: low