Quality of Research Design Moderates Effects of Grade Retention on Achievement: A Meta-Analytic, Multilevel Analysis
Allen CS, Chen Q, Willson VL, Hughes JN · 2009
grade Bmeta-analysisindependentreplicated
Sample
22 peer-reviewed North American studies, 207 effect sizes, screened from 199 candidates published 1990 to June 2007
Population
North American school children, kindergarten through secondary
Design
Educational Evaluation and Policy Analysis 31(4), 480-499. Two-level meta-analytic regression in Mplus. The contribution is the coding scheme: comparison-group quality (1 = all promoted students or matching on non-academic variables such as SES; 2 = low-achieving promoted; 3 = no significant pre-retention difference in academic ability; kappa = .93) crossed with statistical control (1 = none; 2 = distal covariate; 3 = proximal achievement covariate; 4 = pre-retention measure of the outcome itself; kappa = .91). Studies came out low quality in 4 cases (18%), medium in 12 (55%), high in 6 (27%). Critically, the authors note that NO study using random assignment exists in this literature. Heterogeneity Q = 874.52, p < .001; funnel plot showed no publication bias. An overall pooled mean across all 22 studies is never printed in the paper.
Key findings
The headline is a moderator, not a mean: 'the effect size for studies with the lowest design quality is -0.30, whereas the effect size for studies with medium and high design quality is 0.04' - and the 0.04 'is not practically or statistically significantly different from 0'. Design quality is a significant predictor, B = 0.341 (SE 0.05, p = .034), explaining R-squared = .395 of between-study variance; the medium-versus-high contrast is not significant (B = -0.218, p = .452), so the break is between studies that matched on prior achievement and studies that did not. The comparison-group problem is quantified separately: 15 of 22 studies (68%) used same-grade comparison and only 5 (23%) same-age, and the fade rate is about three times steeper under same-grade (-0.111 per year, p = .004) than same-age (-0.040 per year, p = .017) - the signature of an advantage that consists of being older and having covered the material twice. The authors state both halves of the honest conclusion: 'Results challenge the widely held view that retention has a negative impact on achievement' and 'the null hypothesis of no effect of retention on achievement cannot be rejected', but also that the results 'provide little support for proponents of grade retention. Given the expense... a finding of no significant difference calls into question the educational benefits of grade retention policies.' One further pattern worth recording: the four Chicago studies split by policy era, with small positive effects in the two conducted under the test-based promotion policy and moderate negative effects in the two earlier ones.
Genetic confound
This paper IS the genetic-confound argument, made empirically and without using the word. Retention is assigned on low achievement; achievement is substantially heritable; so a study that does not match on prior achievement is comparing children who differ in the trait the study is about. Move from studies that fail that test to studies that pass it and the measured harm of retention disappears. That is exactly what the archive's premise predicts a confounded literature will do under better designs.
Replication notes
Reproduced structurally by Goos, Pipa and Peixoto in 2021 on a larger, later and design-gated corpus (84 studies, pooled g = -0.04). It also recovers a result already present in the 1989 Holmes meta-analysis, whose own best-matched strata - five studies matched on both achievement and SES - gave -0.01, and +0.12 with one questionable study removed. The finding has therefore been independently arrived at three times across thirty years, and quoted almost never.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Achievement effect of retention in the lowest-design-quality studies | effect size | -0.30 | standardized | pooled across follow-ups | business-as-usual | unclear | domain-skill |
| Achievement effect of retention in medium- and high-design-quality studies | effect size | +0.04, described by the authors as neither practically nor statistically different from zero | standardized | pooled across follow-ups | business-as-usual | unclear | domain-skill |
| Design quality as a moderator | regression coefficient | B = 0.341 (SE 0.05), p = .034, R-squared = .395 for the lowest-versus-higher contrast; the medium-versus-high contrast is null (B = -0.218, p = .452) | standardized | study level | none | not-applicable | domain-skill |
| Fade of the retention advantage, same-grade comparisons | change per year post-retention | -0.111 (SE 0.04), p = .004, from 18 studies and 147 effect sizes | standardized | years after retention | business-as-usual | over-2yr | domain-skill |
| Fade of the retention advantage, same-age comparisons | change per year post-retention | -0.040 (SE 0.02), p = .017, from 7 studies and 60 effect sizes - about a third the same-grade rate | standardized | years after retention | business-as-usual | over-2yr | domain-skill |
| How often the flattering comparison is used | study count | 15 of 22 studies (68%) used same-grade comparison, 5 (23%) same-age, 2 (9%) both | standardized | study level | none | not-applicable | domain-skill |
Cited by
- Grade retention — holding a child back versus promoting, and test-based promotion policiesmixedconf: mediumgc: low