The Evidence on Teaching

Feedback and formative assessment

Feedback on the task helps modestly; feedback on the person backfires — a stable third of studies reverse. The famous 0.4–0.7 number has no computed source.

mixedconf: highgc: low

feedback · ages 518 · method

Effect summary

Feedback helps on average but is the only major method with a stable ~1/3 BACKFIRE rate. Task-information feedback works (small: rigorous school-age estimate g≈0.07-0.17, near-zero for secondary students); person-directed feedback (praise, ego threat) nulls or reverses it; grades cancel comments. The famous formative-assessment claim (0.4-0.7) has no computed source — the developer's own at-scale RCT found +0.09 on national exams.

Practical takeaway

Give specific task-focused information ('here is what's wrong and what to do next'), never person-focused evaluation. On formative work, give comments WITHOUT grades. Cut marking volume drastically — written marking has no causal evidence. Treat 'formative assessment' programs as a cheap small bet (~0.1), not a transformation.

Who this applies to

Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.

Verdict

Feedback is the most double-edged tool in teaching. The famous numbers — Hattie's d=0.79, Black & Wiliam's 0.4–0.7 for formative assessment — are both dead: the first was cut to 0.48 by Hattie's own coauthored reanalysis and falls to g≈0.07–0.17 in the only risk-of-bias-graded school-age synthesis; the second was never computed from anything and deflates to +0.09 when the developer's own program was tested at scale on national exams. Meanwhile more than a third of measured feedback effects are negative — the only major method with a stable backfire rate.

What survives is a precise and useful moderator structure, stable across 30 years: feedback that carries task information helps; feedback that directs attention to the self hurts. That structure — not any headline effect size — is the actionable finding. The verdict is mixed because the method genuinely runs both directions; confidence is high because multiple grade-A/B at-scale RCTs and the rigor-graded synthesis all converge on the same picture.

What the evidence shows

Feedback per se:

Source Design Grade Key effect
Kluger & DeNisi 1996 607-effect meta C d=0.41 overall, >38% of effects negative; praise 0.09, discouragement −0.14, ego threat 0.08
Newman 2021 (EEF) risk-of-bias-graded, K-12 only B Low/moderate-bias g=0.17; lowest-bias studies 0.07; secondary students ~0
Wisniewski 2020 994-effect education meta C 0.48 (cuts Hattie's 0.79); high-information 0.99 vs bare reinforcement 0.24
Butler 1988 / Koenka 2021 RCT + 4 metas C Comments > grades on achievement AND motivation; grades+comments = grades (the grade cancels the comment)

The rigor gradient is the story: serious-bias studies g=0.62 → moderate 0.20 → lowest-bias 0.07. Feedback in schools is real but small, positive mainly in early primary and math, and ~zero for secondary students. The backfire conditions are consistent: praise and person-comments, discouragement, ego-involving normative grades, and feedback on very complex tasks.

Formative assessment as a program:

Source Design Grade Key effect
Black & Wiliam 1998 narrative review D The 0.4-0.7 claim — never computed from anything (Bennett 2011)
Kingston & Nash 2011 the only computed meta C 0.20 (13 usable studies of 300+; contested in both directions)
Speckesser 2018 (EEF EFA) 140-school RCT, 25,393 pupils A +0.09 on GCSE (p=.09) — Wiliam's own program, 4-7x below the claim
Anders 2022 peer-reviewed report of the same trial A 0.09 pre-registered; 0.11 on sensitivity/CACE; narrows the prior-attainment gap, not the FSM gap
Randel 2011 (CASL) 67 schools / 9,596 pupils B Null on state math; teacher knowledge rose, teacher practice did not
Cordray 2012 (NWEA MAP) 32 schools / 3,720 pupils B Null on the state test and on MAP's own composite; no change in differentiation
Konstantopoulos 2016 (Indiana) 57 schools / ~25,000 pupils B ITT ns overall; scattered grade-specific positives (reading gr 3-4, math gr 5-6)
Konstantopoulos 2015 (same trial, quantiles) quantile re-analysis B Math TOT 0.16-0.22, flat across ability — no compensatory gradient
Andersson & Palm 2017 ~22 Year-4 teachers, one municipality C d≈0.30 — on a municipality-designed test aligned to the program
Boström & Palm 2023 randomized replication, same developer B d=0.00 (p=.85) on an independent measure; null in every prior-attainment tercile

The trajectory tracks rigor perfectly: 0.4–0.7 (asserted) → 0.20 (weak meta) → 0.09 (at-scale RCT) → ~0 (US commercial variants). The one mildly working variant (EEF's) cost £1.20/pupil/yr — so even 0.09 would be excellent value if real — and its distinguishing ingredient was monthly peer accountability, not the assessment techniques themselves.

Hereditarian-lens assessment

Risk: low — the verdict rests on randomized designs with national-exam outcomes. Two lens-relevant notes: (1) Butler's exception — high achievers anticipating more grades kept performing — plus the ego-threat moderator suggest grading regimes interact with stable (partly heritable) traits, plausibly harming fragile/low-ability students most; (2) the praise literature's perverse targeting loop (adults give the most inflated praise to exactly the low-self-esteem children it harms) is direction-plausible but replication-fragile (Mueller & Dweck vs Li & Bates) — treat person-praise avoidance as a free "do not," not a quantified lever.

What critics say / limits

  • Added 2026-07-30, with a verification failure recorded honestly. A citation-graph audit flagged Mandouit & Hattie 2023, "Revisiting 'The Power of Feedback' from the perspective of the learner" — Hattie re-examining his own 2007 feed-up/feed-back/feed-forward model from the receiving end. We could not read it. It is paywalled in Learning and Instruction with no open-access or repository copy and no abstract deposited with Crossref or Semantic Scholar; the reported response frequencies ("how to improve" 51%, error flagging 50%, "done well" 41%, feed-forward 35%) come from secondary reporting and are not verified against the paper, and we could not establish its sample size or design. It is recorded in the archive as disposition: unresolved — a known unread lead, deliberately not cited in the evidence table above, because a paper we could not read cannot be evidence. It is kept rather than dropped so the gap stays visible. It changes nothing here, and would not even if verified: it measures what students say they want, which is a preference, not an effect. Its only interest is directional corroboration from a different evidence type — students ask for exactly the task-and-improvement information that the Kluger & DeNisi moderator structure says works, and not for the person-evaluation that backfires. Note the irony worth keeping in view: this is the same author whose d = 0.79 headline was cut to 0.48 by his own coauthored reanalysis, so a Hattie paper agreeing with us is not independent support.
  • The at-scale nulls (CASL, MAP) were powered for ~0.25, so effects of ~0.1 can't be excluded — "null at scale" and "+0.09 at scale" are statistically compatible.
  • McMillan et al. argue Kingston & Nash's 0.20 is not credible as an underestimate — but name no alternative number; uncertainty runs both ways within ~0.1–0.3.
  • Durability is essentially unmeasured: virtually no feedback study has a ≥1-year follow-up.

Practical guidance

  • Give task information: what was wrong, why, and what to do next. The three reliable ingredients: correct-solution information, progress-relative-to-prior-work, and high information content.
  • Never person-evaluate: no intelligence praise, no ego-involving comparisons, no discouragement. Praise the work and the process sparingly if at all.
  • Comments without grades on formative work. Adding the grade neutralizes the comment. Save grades for genuinely summative moments.
  • Slash marking volume. Written marking consumes hours weekly with no causal evidence; mark less, mark better, and give class time to respond to it.
  • Buy the cheap version of formative assessment (exit tickets, whiteboards, quizzing — which is really retrieval practice) and expect ~0.1, not 0.5. Skip commercial interim-testing data platforms; they change nothing.

Open questions

  • Whether the EFA +0.09 is a real effect (p=.09, unreplicated) or noise.
  • Whether comment-only marking's benefit survives at scale (the biggest marking RCT lost its outcome arm to COVID).

Evidence (18 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore