The Evidence on Teaching

Does teaching a growth mindset raise achievement?

Growth-mindset interventions change beliefs almost everywhere and change achievement almost nowhere: the two independent national-scale trials measured standardized tests and found exactly zero.

mixedconf: highgc: low

character · ages 418

Effect summary

The effect is real, tiny, and confined to teacher-assigned grades. The best trial ever run — 65 US high schools, 12,490 students, individually randomised against an active control — moved ninth-grade GPA by d = 0.11 for lower-achieving students and by d = 0.01 (CI −0.03 to 0.06) for everyone else. It has no independent standardized test outcome. Every trial that DOES use one reports zero: England's 101-school EEF effectiveness trial found −0.01, −0.00 and −0.00 on KS2 maths, reading and grammar including for disadvantaged pupils, and Argentina's 202-school trial found +0.015 maths and −0.008 reading, ruling out anything above 0.07 SD. Quality-graded meta-analysis walks the pooled effect down from d = 0.08 (Sisk) to 0.05 (all studies) to 0.04 (manipulation check succeeded) to 0.02 (highest quality), and to non-significance after trim-and-fill. Authors with a financial stake in a mindset product were 2.5× more likely to report a positive result, and their published effects (d = 0.18) were nine times their own unpublished ones (d = 0.02).

Practical takeaway

Do not buy a growth-mindset programme, and do not spend curriculum time on one. It costs almost nothing, which is the only honest argument for it, and it delivers almost nothing measurable on any test a school does not write itself. Do keep the free half: praise process and effort rather than the person, tell students that difficulty is normal and improvement is possible, and never say 'you're just not a maths person.' That costs zero minutes. What you must not do is treat mindset work as an achievement lever, a substitute for teaching, or an explanation for why a child is behind.

Who this applies to

Group size
independentsoftwarewhole-class
Delivered by
softwareteacherself
Ages studied
1018(narrower than the 418 this topic is filed under — outside it is extrapolation)
Dose
Two 25-minute online sessions about 20 days apart produced the largest well-identified effect (National Study of Learning Mindsets). Dose is NOT the binding constraint: England's 8-week teacher-delivered programme at up to 2.5 hours a week produced exactly zero on national tests.
Cost
low
Moves
non-cognitive
Needs first
For the one credible positive, the student must already be lower-achieving, in a medium- or low-achieving school, with peer norms that make challenge-seeking socially safe — and the outcome must be a course grade.
Not for
Raising scores on an independent standardized test. Two national-scale independent trials measured exactly that and found zero, including for disadvantaged pupils. Also not for higher-achieving students: the best trial's estimate for them is d = 0.01 (95% CI −0.03 to 0.06).

Verdict

mixed, and the word is doing precise work. Growth mindset is not a fraud and it is not brain training. There is a real effect, it has been measured under an excellent design, and the archive's rules require reporting it: in the National Study of Learning Mindsets, two 25-minute online sessions raised ninth-grade core GPA by 0.10 grade points (d = 0.11) among lower-achieving students.

Three facts hold that effect in place, and each of them is what turns moderate-support into mixed:

  1. It appears on teacher-assigned grades and nowhere else. Every positive result in the literature — Blackwell 2007, Yeager 2019, Yeager 2022, Porter 2022 — has grades as its outcome. Every study whose primary outcome is an externally set and externally marked test reports zero.
  2. It is zero for the majority of students. d = 0.01, CI −0.03 to 0.06, for higher achievers in the trial that established the effect.
  3. It does not survive transport. England's 101-school effectiveness trial and Argentina's 202-school trial both tested it independently, on national tests, in exactly the populations the theory nominates, and both returned zero.

Under this archive's replication rule, a finding with failed independent replications cannot exceed mixed regardless of the original effect size, and mindset is one of the two or three canonical cases the rule exists for.

What the evidence shows

Source Design Grade Key effect
Foliano 2019 (EEF) Independent 101-school cluster RCT, 5,018 pupils, KS2 national tests A Maths −0.01 (−0.04–0.01), reading −0.00, GPS −0.00. Zero months' progress, and zero for FSM pupils
Ganimian 2020 Independent 202-school RCT, Argentina, government national assessment A Maths +0.015, reading −0.008; rules out effects > 0.07 SD; self-efficacy fell for girls and low-income students
Yeager 2019 (NSLM) 65-school, 12,490-student RCT, active control, preregistered A Lower achievers +0.10 GPA (d = 0.11); higher achievers d = 0.01 (ns); D/F rate −5.3pp; advanced maths +3pp. No standardized test
Macnamara & Burgoyne 2023 Preregistered meta, quality ladder, 63 studies / 97,672 B 0.05 → 0.04 → 0.02; non-significant after trim-and-fill; COI authors 2.5× more likely positive
Sisk 2018 Two metas, k = 273 / k = 43 B r = .09; intervention d = 0.08; effect present only where the manipulation check failed or was absent
Burnette 2023 Meta, 53 samples — the steelman C d = 0.14 conditional on at-risk subsample AND high fidelity; prediction interval −0.08 to 0.35
Gazmuri 2025 Structured review, 24 RCTs, quality-weighted B Strongest studies: d = −0.01 to +0.065; of 4 trustworthy conflict-free trials, 2 null and 2 very small
Bahník & Vranka 2017 N = 5,653 applicants, admissions aptitude test C Mindset ↔ test score r = −.03; no relation to persistence; does not predict score change
Rienzo 2015 (EEF pilot) 6- and 30-school pilots B +2 months (ns) and −2 months English (ns). The "promise" that vanished at scale
Blackwell 2007 n = 91 RCT + n = 373 longitudinal C b = 0.53 on maths grades; targeting interaction only marginal (p < .10)
Porter 2022 1,996 students, teacher-randomised, Mindset Works product C Struggling students β = 0.27 on report-card grades; preregistration badge withdrawn; vendor employees on the author list
Yeager 2022 NSLM moderator analysis, 9,167 students B Works only where teachers are growth-minded (0.11 vs −0.02)
Rege 2021 NSLM reanalysis + Norwegian replication, N = 6,541 B Challenge-seeking and advanced-maths enrolment replicate. No achievement claim made
WWC 2022 Federal review, 6 qualifying studies (postsecondary) B GPA "potentially positive" (+13); enrolment and progression: no discernible effects

The measurement fault line is almost perfectly confounded with the result, and that is the finding. Sort the literature by outcome type rather than by author and it stops being controversial. Grades → small positive. Externally marked test → zero. This is the archive's standing effect-size rule (measure type matters more than the number) producing an unusually clean partition.

The quality ladder does the rest. Macnamara and Burgoyne pre-registered a best-practice screen and the estimate falls monotonically as quality rises: 0.05 across all 63 studies, 0.04 where the manipulation check actually succeeded, 0.02 in the six highest-quality studies, and nothing at all after trim-and-fill imputes the ten missing small or negative studies. Sisk had already found the same pattern in miniature — the achievement effect was significant only in studies where the manipulation check was absent or had failed, which is the signature of an artefact rather than a mechanism.

The conflict-of-interest analysis is the most quotable result in this tranche. Thirty percent of mindset documents had at least one author with a financial incentive. Those articles were 2.5 times as likely to report a significant positive effect (56% vs 21%, p = .008). And among financially incentivised authors, published effects averaged d = 0.18 against their own unpublished effects of d = 0.02 — a gap absent among authors without a stake. Porter 2022 is that finding in a single paper: two co-authors employed by Mindset Works, evaluating Mindset Works' product, with the journal's preregistration badge later withdrawn.

And the pilot-to-scale record is the archive's cleanest example of effect deflation. England's 2015 EEF pilot found +2 months on 286 pupils and was written up as "evidence of promise." Scaled seventeen-fold and measured on national tests, it produced −0.01, −0.00 and −0.00.

Hereditarian-lens assessment

Risk: low — the verdict rests on randomised trials, and genes cannot differ between arms.

But the premise is doing quiet work in the background, and it explains why the correlational leg was always weak. Mindset theory proposes that beliefs about the malleability of ability drive achievement. Bahník and Vranka put that to the strongest available test — 5,653 applicants, a high-stakes admissions aptitude test rather than a teacher's mark — and found r = −.03, no relation to a behavioural persistence measure, and no prediction of score change. Sisk's pooled correlation across 365,915 people is r = .09, about one percent of variance. That is what a belief measure looks like when the outcome is substantially heritable and the belief is not the bottleneck.

The archive's premise (3) is worth restating here, because it cuts against over-dismissal too: heritability of individual differences does not mean instruction is useless. The reason to reject mindset is not that traits are heritable — it is that two well-powered independent trials measured the thing and found nothing.

One further note. Yeager 2019 found significant cross-school variation in the achievement effect but no significant variation in the belief-change effect. The intervention changed minds everywhere and changed grades only somewhere. Whatever produced the GPA movement, it was not simply the belief.

Boundaries & what critics say

  • The steelman is Burnette 2023, published in the same issue as the deflationary meta, reporting d = 0.14. Read what it actually is: not an average treatment effect but an estimate conditional on both a targeted subsample and high implementation fidelity, with no bias correction and a 95% prediction interval of −0.08 to 0.35. The authors' own interval says a new trial of this kind is not predicted to do anything reliably.
  • Tipton et al. 2023 is the developer team's methodological reply, and its case — that a single average effect is the wrong estimand for a heterogeneous literature — is a serious statistical point, not a dodge. But the reply comes from the NSLM's own authors and statisticians, and the critics' response charges that the reanalysis changed effect sizes, moderator coding and study inclusion. The disagreement is analytic; both sides hold the same studies.
  • The behavioural and course-taking results really do replicate. Rege 2021 reproduced challenge-seeking on a behavioural task and advanced-maths enrolment in Norway. Note the framing that concession implies — the paper explicitly positions itself as going beyond grades or performance. Getting more students to opt into harder courses is a genuine outcome and should not be denied; it is simply a different claim from raising achievement.
  • The two teacher-mindset moderation results contradict each other outright. Yeager 2022 finds the effect only where teachers are growth-minded; Porter 2022 finds it largest where teachers are fixed-minded. Both are developer-side papers from the same year. When the moderators disagree this badly, the moderator literature is not yet evidence.
  • The praise finding is separate and survives better than the intervention finding. Person praise ("you're so clever") versus process praise is a zero-cost distinction that the archive already records as plausible-but-fragile — the famous demonstration (Mueller & Dweck 1998) failed its only large direct replication and the dispute is unresolved. Treat it as a free "do not," not as an effect size.
  • Cost is the honest pro-mindset argument. England's programme cost about £4 per pupil and Argentina's about $2.82 per student. At that price a d = 0.02 is not obviously a bad trade. What is a bad trade is the eight weeks of curriculum time England's version consumed to produce zero.

Practical guidance

  • Do not purchase a growth-mindset programme or curriculum. The one trial of a commercial product in this file was run by the vendor's own employees and lost its preregistration badge.
  • Do not give up teaching time for it. The cheap online version is the one with evidence; the eight-week classroom version is the one with a national-scale zero.
  • Keep the free behaviours. Praise the work, not the person. Say that struggle is expected and that the subject is learnable. Never tell a child they are "not a maths person." These cost nothing and nothing in this file argues against them.
  • If you are going to try it, target it and be honest about the outcome. The only credible positive is for lower-achieving adolescents in medium-achieving schools, on course grades, at about d = 0.11 — and it did not reproduce in England or Argentina.
  • Refuse mindset as an explanation for a struggling child. A child who is behind in reading needs reading instruction. The archive's fadeout and persistence and phonics findings are what move that child; a belief module is not.
  • Interrogate every positive mindset claim on two questions: was the outcome a test somebody else wrote and marked, and does any author have a financial stake? Those two questions sort this literature almost perfectly.

Open questions

  • Whether the NSLM's GPA effect is a real learning effect or a grading effect is untested. Nobody has run the NSLM design with an independent standardized test as the primary outcome, which is remarkable given how much rests on it.
  • The challenge-seeking and course-enrolment effects are the most robust things in the literature and the least studied downstream. Does taking advanced maths because of a mindset module produce anything at the end of it?
  • Why belief change is homogeneous across schools while achievement change is heterogeneous has no accepted explanation, and it is the strongest internal evidence that the proposed mechanism is not the operative one.
  • Nothing here is genetically informative about differential responsiveness to belief interventions — the same untested trainability assumption the archive flags in talent and trainability.

Evidence (18 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore