The Power of Feedback Revisited: A Meta-Analysis of Educational Feedback Research
Wisniewski, B., Zierer, K., & Hattie, J. · 2020
grade Cmeta-analysisdeveloper-involvedfailednumbers spot-checked
Sample
435 studies / 994 effect sizes / >61,000 participants
Population
Students in educational contexts (school and post-secondary).
Design
Random-effects; education-only. NOT a naive meta-meta: the unit of analysis is the PRIMARY study. The authors searched the 32 meta-analyses used in the Visible Learning feedback synthesis, pulled the primary studies out of them, dropped duplicates and non-educational contexts, and re-meta-analysed at the study level with precision weighting — which is why it is graded as a meta-analysis of mixed-quality primary designs rather than rejected under the archive's no-meta-meta rule. Corpus is old (median publication year 1985; only 15% of effects from the preceding 15 years). Funnel asymmetry present for journal articles (Egger z = 9.75, p < 0.0001) but absent for dissertations (z = 1.03, p = 0.30) — i.e. the asymmetry is publication bias, not heterogeneity. Six moderators coded; "way of measuring the outcome" was one of two features the authors had to DROP for insufficient data, so measure alignment is uncoded.
Key findings
Hattie-coauthored revision that cuts the famous d = 0.79 to 0.48 (0.42 restricted to controlled designs). Information content dominates: high-information feedback 0.99, corrective 0.46, bare reinforcement/punishment 0.24. Note that 0.24 is positive with a CI excluding zero — this meta does NOT show praise or person-level feedback to be null or harmful on achievement, and has no self/person-level category at all; the negative signal appears only on motivational outcomes, where 21% of effects were negative and 86% of those came from uninformative feedback. Corpus is old (median 1985), publication-biased in journals but not dissertations, and never codes measure alignment because the authors had to drop that variable — so 0.48 is an UPPER bound; at-scale RCTs on independent exams run 0.0-0.10. The authors' own conclusion: feedback is not one treatment; the heterogeneity is the finding.
Genetic confound
Aggregates experimental comparisons; student-trait moderation uncoded.
Replication notes
Numbers verified against the full text. Hattie & Timperley's (2007) meta-synthesis figure of d = 0.79 — an unweighted fixed-effect sum of meta-analytic averages, with duplicated studies — is cut to 0.48 (~60% of it) by Hattie's own coauthored reanalysis of the underlying primary studies. The authors attribute the gap to de-duplication, precision weighting, and exclusion of studies that failed the inclusion criteria on inspection. Of 24 meta-analysis subsets recomputed, 21 confidence intervals overlapped the synthesised value and 3 were overestimates (Rummel & Feinberg 1988; Standley 1996; Miller 2003) — Standley's music-as-reinforcement effects of 1.74 to 11.98 could not be reconstructed at all and the author did not respond when contacted.
DOI / URL
10.3389/fpsyg.2019.03087
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Overall achievement/learning, all feedback types pooled | d | 0.48 [0.44, 0.51] after excluding 35 extreme values (3.5% of effects); 0.55 [0.48, 0.62] before outlier exclusion, of which 17% were negative. Fixed-effect equivalent would be 0.41 [0.40, 0.43]. I-squared = 83.4% — the heterogeneity, not the mean, is the finding. | mixed | end-of-treatment | active-alternative | end-of-treatment | domain-skill |
| Controlled (experimental) studies only | d | 0.42 [0.37, 0.46], k = 713 — against 0.63 [0.56, 0.69] for pre-post designs (k = 244). The design-restricted estimate is the defensible one. | mixed | end-of-treatment | business-as-usual | end-of-treatment | domain-skill |
| Published journal articles vs dissertations | d | 0.49 [0.45, 0.53] (k = 843) vs 0.36 [0.25, 0.46] (k = 116) — low and negative feedback effects are less likely to be published | mixed | end-of-treatment | unclear | end-of-treatment | domain-skill |
| High-information feedback (task + process + sometimes self-regulation level) | d | 0.99 [0.82, 1.15], k = 42 | mixed | end-of-treatment | active-alternative | end-of-treatment | domain-skill |
| Corrective feedback | d | 0.46 [0.39, 0.55], k = 238 | mixed | end-of-treatment | active-alternative | end-of-treatment | domain-skill |
| Reinforcement or punishment (uninformative feedback — praise, rewards, sanctions) | d | 0.24 [0.06, 0.43], k = 39 — small but POSITIVE and the CI excludes zero. This meta does not find praise/person-level feedback to be null or harmful on achievement; it finds it four times weaker than high-information feedback. There is no separate "self/person level" category in the coding scheme. | mixed | end-of-treatment | active-alternative | end-of-treatment | domain-skill |
| Cognitive outcomes (achievement, retention, test performance) | d | 0.51 [0.46, 0.55], k = 597 | mixed | end-of-treatment | active-alternative | end-of-treatment | domain-skill |
| Motivational outcomes (intrinsic motivation, locus of control, self-efficacy, persistence) | d | 0.33 [0.23, 0.42], k = 109. This is where the harm signal actually lives — 21% of motivational effect sizes were negative, and 86% of the interventions producing them were uninformative feedback (rewards or punishments). The authors state the results indicate not that feedback effects on motivation are low per se, but that effects of UNINFORMATIVE feedback on motivation are "low or even negative". That is a descriptive breakdown, not a computed subgroup estimate. | mixed | end-of-treatment | active-alternative | end-of-treatment | non-cognitive |
| Behavioural outcomes (classroom behaviour, discipline) | d | 0.48 [-0.09, 1.06], k = 30 — CI includes zero | mixed | end-of-treatment | active-alternative | end-of-treatment | behaviour |
| Direction of feedback | d | teacher-to-student 0.47 [0.43, 0.51] (k = 812); student-to-teacher 0.35 [0.13, 0.56] (k = 27, almost all higher education); student-to-student 0.85 [0.59, 1.11] (k = 16, only 8 studies) | mixed | end-of-treatment | active-alternative | end-of-treatment | domain-skill |
Cited by
- Feedback and formative assessmentmixedconf: highgc: low