The Evidence on Teaching

Effectiveness of grade retention: A systematic review and meta-analysis

Goos M, Pipa J, Peixoto F · 2021

grade Bmeta-analysisindependentreplicated
Sample
84 studies yielding 1,328 effect sizes, published 2000-2019
Population
Kindergarten through grade 12, multiple countries
Design
Educational Research Review 34, 100401; read in the KU Leuven repository accepted manuscript. This is the design-gated meta-analysis the topic turns on. Design quality is an INCLUSION RULE rather than a moderator: a study had to use an experimental design or regression discontinuity, propensity scores, instrumental variables, difference-in-differences or factor-analytic methods, and studies 'without a credible control group were excluded'. Three-level meta-regression on 1,328 effect sizes nested in 84 studies. Because the gate removes the entire pre-2000 correlational literature, the pooled estimate here is not comparable with Holmes or Jimerson - it is what is left after the confounded studies are taken out.
Key findings
The pooled effect of grade retention is g = -0.04 (SE 0.04, 95% CI -0.11 to 0.03, t(1327) = -1.11, p = 0.266) - indistinguishable from zero. The vote count is a picture of a genuinely contested literature rather than a consensus: 35% of effects significantly negative, 41% non-significant, 24% significantly positive. Three moderators carry almost all the disagreement. (1) COMPARISON GROUP: same-grade +0.05 (95% CI -0.03 to 0.12) versus same-age -0.11 (-0.18 to -0.04), F(1,1326) = 23.60, p < .001, surviving joint estimation. (2) IDENTIFICATION STRATEGY: the regression-discontinuity subset gives +0.17 (0.03 to 0.31) while the instrumental-variables subset gives -0.15 (-0.31 to 0.00) - two designs the archive grades identically, pointing in opposite directions. (3) WHAT IS ATTACHED TO THE RETENTION: 'mere rehearsal' - repeating the year with no added support - gives g = -0.10 (-0.17 to -0.02), while 'retention plus' - repetition bundled with remediation - gives g = +0.20 (0.04 to 0.35), F = 11.57, p < .001. By outcome class the answers differ and must not be merged: achievement -0.07 (-0.14 to 0.01), psychosocial +0.08 (-0.01 to 0.17), school career (attainment and dropout) -0.10 (-0.19 to -0.01). Effects decay: g = 0.01 in the retention year, falling by about 0.01 per year after. Grade at retention (kindergarten -0.13, primary -0.02, secondary -0.04) is NOT a significant moderator here, F = 1.46, p = .232 - which sits awkwardly beside the individual RD studies, where the early/late contrast is the strongest single pattern. Peer-reviewed and grey studies do not differ.
Genetic confound
Low by construction of the inclusion rule, and that is the point of the paper. Retention is assigned on low measured achievement, which is substantially heritable, so the pre-2000 literature compared retained children with children who were better students to begin with. Gating on designs that break that link is what moves the pooled estimate from strongly negative to zero. Residual risk lives in the propensity-score subset (g = -0.09), where matching can only balance measured variables.
Replication notes
Independently reproduces Allen and colleagues' 2009 finding by a different route - Allen moderated ON design quality and found -0.30 for the weakest designs against +0.04 for the rest; Goos gated on design quality and found -0.04 overall. The authors say so themselves: 'Allen et al. (2009) obtained a similar zero average effect size when restricting the study sample to high-quality studies.' Two syntheses, different samples, different methods, same structural conclusion.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Overall effect of grade retention across design-gated studiesHedges g-0.04 (SE 0.04, 95% CI -0.11 to 0.03), t(1327) = -1.11, p = 0.266; 35% of effects significantly negative, 41% null, 24% significantly positivestandardizedpooled across follow-up horizonsbusiness-as-usualuncleardomain-skill
Same-grade versus same-age comparisonHedges g by comparison groupsame-grade +0.05 (95% CI -0.03 to 0.12, 41 studies); same-age -0.11 (-0.18 to -0.04, 52 studies); F(1,1326) = 23.60, p < .001standardizedpooledbusiness-as-usualuncleardomain-skill
Identification strategy as a moderatorHedges g by designregression discontinuity +0.17 (0.03 to 0.31, 17 studies); instrumental variables -0.15 (-0.31 to 0.00, 15 studies); propensity score -0.09 (-0.18 to 0.00); difference-in-differences -0.05 (-0.20 to 0.10)standardizedpooledbusiness-as-usualuncleardomain-skill
Retention alone versus retention bundled with supportHedges g'mere rehearsal' -0.10 (-0.17 to -0.02, 69 studies); 'retention plus' added support +0.20 (0.04 to 0.35, 15 studies); F = 11.57, p < .001standardizedpooledbusiness-as-usualuncleardomain-skill
Achievement outcomes onlyHedges g-0.07 (95% CI -0.14 to 0.01), 58 studies, 524 effect sizesstandardizedpooledbusiness-as-usualuncleardomain-skill
School-career outcomes (attainment and dropout)Hedges g-0.10 (95% CI -0.19 to -0.01), 29 studies, 446 effect sizes - the only outcome class whose confidence interval excludes zero, and it is negativestandardizedpooledbusiness-as-usualover-2yrattainment
Psychosocial outcomesHedges g+0.08 (95% CI -0.01 to 0.17), 22 studies, 326 effect sizesstandardizedpooledbusiness-as-usualunclearnon-cognitive
Decay of the effect over timechange in g per yearg = 0.01 in the retention year, diminishing by about 0.01 each year afterwardsstandardizedyear-by-yearbusiness-as-usualover-2yrdomain-skill

Cited by

Effectiveness of grade retention: A systematic review and meta-analysis · The Evidence on Teaching