Feedback and formative assessment
Feedback on the task helps modestly; feedback on the person backfires — a stable third of studies reverse. The famous 0.4–0.7 number has no computed source.
mixedconf: highgc: lowfeedback · ages 5–18 · method
Feedback helps on average but is the only major method with a stable ~1/3 BACKFIRE rate. Task-information feedback works (small: rigorous school-age estimate g≈0.07-0.17, near-zero for secondary students); person-directed feedback (praise, ego threat) nulls or reverses it; grades cancel comments. The famous formative-assessment claim (0.4-0.7) has no computed source — the developer's own at-scale RCT found +0.09 on national exams.
Give specific task-focused information ('here is what's wrong and what to do next'), never person-focused evaluation. On formative work, give comments WITHOUT grades. Cut marking volume drastically — written marking has no causal evidence. Treat 'formative assessment' programs as a cheap small bet (~0.1), not a transformation.
Who this applies to
Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.
Verdict
Feedback is the most double-edged tool in teaching. The famous numbers — Hattie's d=0.79, Black & Wiliam's 0.4–0.7 for formative assessment — are both dead: the first was cut to 0.48 by Hattie's own coauthored reanalysis and falls to g≈0.07–0.17 in the only risk-of-bias-graded school-age synthesis; the second was never computed from anything and deflates to +0.09 when the developer's own program was tested at scale on national exams. Meanwhile more than a third of measured feedback effects are negative — the only major method with a stable backfire rate.
What survives is a precise and useful moderator structure, stable across 30 years: feedback that carries task information helps; feedback that directs attention to the self hurts. That structure — not any headline effect size — is the actionable finding. The verdict is mixed because the method genuinely runs both directions; confidence is high because multiple grade-A/B at-scale RCTs and the rigor-graded synthesis all converge on the same picture.
What the evidence shows
Feedback per se:
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Kluger & DeNisi 1996 | 607-effect meta | C | d=0.41 overall, >38% of effects negative; praise 0.09, discouragement −0.14, ego threat 0.08 |
| Newman 2021 (EEF) | risk-of-bias-graded, K-12 only | B | Low/moderate-bias g=0.17; lowest-bias studies 0.07; secondary students ~0 |
| Wisniewski 2020 | 994-effect education meta | C | 0.48 (cuts Hattie's 0.79); high-information 0.99 vs bare reinforcement 0.24 |
| Butler 1988 / Koenka 2021 | RCT + 4 metas | C | Comments > grades on achievement AND motivation; grades+comments = grades (the grade cancels the comment) |
The rigor gradient is the story: serious-bias studies g=0.62 → moderate 0.20 → lowest-bias 0.07. Feedback in schools is real but small, positive mainly in early primary and math, and ~zero for secondary students. The backfire conditions are consistent: praise and person-comments, discouragement, ego-involving normative grades, and feedback on very complex tasks.
Formative assessment as a program:
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Black & Wiliam 1998 | narrative review | D | The 0.4-0.7 claim — never computed from anything (Bennett 2011) |
| Kingston & Nash 2011 | the only computed meta | C | 0.20 (13 usable studies of 300+; contested in both directions) |
| Speckesser 2018 (EEF EFA) | 140-school RCT, 25,393 pupils | A | +0.09 on GCSE (p=.09) — Wiliam's own program, 4-7x below the claim |
| Anders 2022 | peer-reviewed report of the same trial | A | 0.09 pre-registered; 0.11 on sensitivity/CACE; narrows the prior-attainment gap, not the FSM gap |
| Randel 2011 (CASL) | 67 schools / 9,596 pupils | B | Null on state math; teacher knowledge rose, teacher practice did not |
| Cordray 2012 (NWEA MAP) | 32 schools / 3,720 pupils | B | Null on the state test and on MAP's own composite; no change in differentiation |
| Konstantopoulos 2016 (Indiana) | 57 schools / ~25,000 pupils | B | ITT ns overall; scattered grade-specific positives (reading gr 3-4, math gr 5-6) |
| Konstantopoulos 2015 (same trial, quantiles) | quantile re-analysis | B | Math TOT 0.16-0.22, flat across ability — no compensatory gradient |
| Andersson & Palm 2017 | ~22 Year-4 teachers, one municipality | C | d≈0.30 — on a municipality-designed test aligned to the program |
| Boström & Palm 2023 | randomized replication, same developer | B | d=0.00 (p=.85) on an independent measure; null in every prior-attainment tercile |
The trajectory tracks rigor perfectly: 0.4–0.7 (asserted) → 0.20 (weak meta) → 0.09 (at-scale RCT) → ~0 (US commercial variants). The one mildly working variant (EEF's) cost £1.20/pupil/yr — so even 0.09 would be excellent value if real — and its distinguishing ingredient was monthly peer accountability, not the assessment techniques themselves.
Hereditarian-lens assessment
Risk: low — the verdict rests on randomized designs with national-exam outcomes. Two lens-relevant notes: (1) Butler's exception — high achievers anticipating more grades kept performing — plus the ego-threat moderator suggest grading regimes interact with stable (partly heritable) traits, plausibly harming fragile/low-ability students most; (2) the praise literature's perverse targeting loop (adults give the most inflated praise to exactly the low-self-esteem children it harms) is direction-plausible but replication-fragile (Mueller & Dweck vs Li & Bates) — treat person-praise avoidance as a free "do not," not a quantified lever.
What critics say / limits
- Added 2026-07-30, with a verification failure recorded honestly. A citation-graph audit flagged
Mandouit & Hattie 2023, "Revisiting 'The Power of Feedback' from the perspective of the learner" — Hattie
re-examining his own 2007 feed-up/feed-back/feed-forward model from the receiving end. We could not
read it. It is paywalled in Learning and Instruction with no open-access or repository copy and no
abstract deposited with Crossref or Semantic Scholar; the reported response frequencies ("how to improve"
51%, error flagging 50%, "done well" 41%, feed-forward 35%) come from secondary reporting and are not
verified against the paper, and we could not establish its sample size or design. It is recorded in the
archive as
disposition: unresolved— a known unread lead, deliberately not cited in the evidence table above, because a paper we could not read cannot be evidence. It is kept rather than dropped so the gap stays visible. It changes nothing here, and would not even if verified: it measures what students say they want, which is a preference, not an effect. Its only interest is directional corroboration from a different evidence type — students ask for exactly the task-and-improvement information that the Kluger & DeNisi moderator structure says works, and not for the person-evaluation that backfires. Note the irony worth keeping in view: this is the same author whose d = 0.79 headline was cut to 0.48 by his own coauthored reanalysis, so a Hattie paper agreeing with us is not independent support. - The at-scale nulls (CASL, MAP) were powered for ~0.25, so effects of ~0.1 can't be excluded — "null at scale" and "+0.09 at scale" are statistically compatible.
- McMillan et al. argue Kingston & Nash's 0.20 is not credible as an underestimate — but name no alternative number; uncertainty runs both ways within ~0.1–0.3.
- Durability is essentially unmeasured: virtually no feedback study has a ≥1-year follow-up.
Practical guidance
- Give task information: what was wrong, why, and what to do next. The three reliable ingredients: correct-solution information, progress-relative-to-prior-work, and high information content.
- Never person-evaluate: no intelligence praise, no ego-involving comparisons, no discouragement. Praise the work and the process sparingly if at all.
- Comments without grades on formative work. Adding the grade neutralizes the comment. Save grades for genuinely summative moments.
- Slash marking volume. Written marking consumes hours weekly with no causal evidence; mark less, mark better, and give class time to respond to it.
- Buy the cheap version of formative assessment (exit tickets, whiteboards, quizzing — which is really retrieval practice) and expect ~0.1, not 0.5. Skip commercial interim-testing data platforms; they change nothing.
Open questions
- Whether the EFA +0.09 is a real effect (p=.09, unreplicated) or noise.
- Whether comment-only marking's benefit survives at scale (the biggest marking RCT lost its outcome arm to COVID).
- grade CThe Effects of Feedback Interventions on Performance: A Historical Review, a Meta-Analysis, and a Preliminary Feedback Intervention TheoryKluger, A. N., & DeNisi, A. · 1996 · meta-analysis
- grade BThe Impact of Feedback on Student Attainment: A Systematic Review (EEF/EPPI-Centre)Newman, M., Bird, K., Kwan, I., Shemilt, I., Richardson, M., & Hoo, H. · 2021 · meta-analysis
- grade CThe Power of Feedback Revisited: A Meta-Analysis of Educational Feedback ResearchWisniewski, B., Zierer, K., & Hattie, J. · 2020 · meta-analysis
- grade CEnhancing and Undermining Intrinsic Motivation: Task-Involving and Ego-Involving EvaluationButler, R. · 1988 · rct
- grade CA Meta-Analysis on the Impact of Grades and Comments on Academic Motivation and AchievementKoenka, A. C., Linnenbrink-Garcia, L., Moshontz, H., et al. · 2021 · meta-analysis
- grade DA Marked Improvement? A Review of the Evidence on Written Marking (EEF/Oxford)Elliott, V., Baird, J., Hopfenbeck, T., et al. · 2016 · review
- grade CPraise for Intelligence Can Undermine Motivation and Performance (with the Li & Bates 2019 replication attempt)Mueller, C. M., & Dweck, C. S. (1998); Li, Y., & Bates, T. C. (2019) · 1998 · replication
- grade DAssessment and Classroom Learning / Inside the Black BoxBlack, P., & Wiliam, D. · 1998 · review
- grade CFormative Assessment: A Meta-Analysis and a Call for Research (with the Briggs et al. 2012 / McMillan et al. 2013 critiques)Kingston, N., & Nash, B. · 2011 · meta-analysis
- grade AEmbedding Formative Assessment: Evaluation report and executive summarySpeckesser S, Runge J, Foliano F, Bursnall M, Hudson-Sharp N, Rolfe H, Anders J · 2018 · rct
- grade AThe Effect of Embedding Formative Assessment on Pupil AttainmentAnders J, Foliano F, Bursnall M, Dorsett R, Hudson N, Runge J, Speckesser S · 2022 · rct
- grade BClassroom Assessment for Student Learning: Impact on Elementary School Mathematics in the Central RegionRandel B, Beesley AD, Apthorp H, Clark TF, Wang X, Cicchinelli LF, Williams JM · 2011 · rct
- grade BThe Impact of the Measures of Academic Progress (MAP) Program on Student Reading AchievementCordray D, Pion G, Brandt C, Molefe A, Toby M · 2012 · rct
- grade BEffects of Interim Assessments on Student Achievement: Evidence From a Large-Scale ExperimentKonstantopoulos S, Miller SR, van der Ploeg A, Li W · 2016 · rct
- grade BEffects of Interim Assessments Across the Achievement DistributionKonstantopoulos S, Li W, Miller SR, van der Ploeg A · 2015 · rct
- grade BThe effect of a formative assessment practice on student achievement in mathematicsBoström E, Palm T · 2023 · replication
Related decisions
- Does teaching music produce musical skill?moderate supportconf: mediumgc: medium
- Placement and mastery diagnosis — deciding what to teach next from evidence of current skillmoderate supportconf: mediumgc: medium
- Retrieval practice (the testing effect)strong supportconf: highgc: low