The Evidence on Teaching

Subject

Teaching character — executive function, self-control, mindset and social-emotional learning

Bottom line. Do not buy character as a curriculum. This is the literature with modern psychology's worst replication record and the largest recent school spending, and the pattern is almost perfectly consistent: interventions move the ratings and beliefs they are evaluated on, and do not move the attainment they are sold on. Growth mindset is exactly zero on national tests in two independent trials. Executive-function training decays to 0.008 at follow-up. Grit is conscientiousness renamed, adding 0.4% of attainment variance. SEL's famous +0.27 is now 0.10 and has never been reproduced independently on an externally marked test. The traits themselves are 37–100% heritable with shared-environment variance near zero, which is why the correlational case for character education is worthless and why the randomised case is so thin. What a school should actually run is a well-managed room — small, replicated, cheap — and teach the subjects.

The program (what a school should actually do)

Teach the subjects; run the room; do not buy the trait.

Character is where this archive's premise, its replication rule, and its measurement rules all point the same way at once, and where the money has gone anyway. Four separate literatures were built on the same idea — that a school can install a durable disposition which then pays out in achievement — and all four have deflated in the same direction, on the same fault line, over roughly the same fifteen years.

Do not buy a growth-mindset programme.

  • The best-designed trial is real and small: 12,490 students, active control, preregistered, +0.11 SD on ninth-grade GPA for lower-achieving students and d = 0.01 (CI −0.03 to 0.06) for everyone else. It has no independent standardized test outcome anywhere in it.
  • Every trial that does use one reports nothing. England: 101 schools, 5,018 pupils, −0.01, −0.00, −0.00 on KS2 maths, reading and grammar, and zero for disadvantaged pupils. Argentina: 202 schools, +0.015 maths, effects above 0.07 SD ruled out.
  • Quality-graded meta-analysis walks it from 0.05 to 0.02 and to non-significance after bias correction, and finds authors with a financial stake 2.5× more likely to publish a positive result — with their published effects (0.18) nine times their own unpublished ones (0.02). (evidence: growth mindset)

Do not buy an executive-function curriculum, a working-memory trainer, or school mindfulness.

  • Training improves the trained task (g = 0.38) and nothing survives to follow-up: the pooled follow-up estimate across all EF interventions is 0.008 after bias correction; EF-specific curricula are 0.12 at post-test and −0.03 at follow-up.
  • Tools of the Mind — the field's flagship — is null in a 60-classroom independent US trial where every significant effect favoured the control group through first grade, null on every credible interval in a registered Canadian trial with an active control, and rated "no discernible effects" by the federal clearinghouse, which also found the famous 2007 Science paper did not meet evidence standards.
  • School mindfulness on EF is 0.25 against passive controls and 0.04 against active ones; the 85-school MYRIAD trial returned 0.005, 0.02 and 0.02 with five secondary outcomes favouring the control arm.
  • The latent executive-function factor is 99–100% heritable with shared environment at zero in two independent twin cohorts. (evidence: executive function training)

Stop measuring grit; stop citing the marshmallow test.

  • Grit correlates with conscientiousness at ρ = .84, shares 95% of its variance with it, and has a genetic correlation of .86 with it. Its incremental validity over conscientiousness is .000–.004 in R², and 0.5 percentage points of GCSE variance.
  • The marshmallow effect halves in a proper sample and dies under controls (β 0.236 → 0.050, ns), and the original cohort's own preregistered mid-life follow-up finds preschool delay predicts none of eleven capital-formation outcomes (average r = 0.02).
  • Self-control is different and genuinely predicts — 13% vs 43% criminal conviction across the Dunedin extremes, surviving a discordant-twin design — but it is 60% heritable with shared environment ≈ 0, and forty-one randomised training trials have moved rating scales and never once measured an objective long-run outcome. (evidence: self-control and grit)

Do not buy an SEL curriculum, least of all on its attainment claim.

  • The founding +0.27 rested on 35 of 213 interventions. The field's own current estimates are 0.109 and 0.101, with significant publication bias and 53% of outcome data child self-report.
  • England measured PATHS on externally marked national tests across 45 schools: +0.026, −0.029, −0.025, and −0.106 on Year 6 English, which the evaluator judged attributable to the programme.
  • The US federal government randomised 84 schools across seven flagship programmes with an independent analyst and found 2 of 60 impacts significant against 3 expected by chance, 0 of 18 in growth-curve analysis, and two significant negative domain effects. (evidence: SEL programmes)

Run the room instead. Ordinary classroom management returns g = 0.22–0.23, replicated across two meta-analyses nine years apart, and classroom order is causally worth a great deal — one disruptive peer in a class of 25 costs classmates 3% of their earnings at ages 24–28. That is the part of this domain with real evidence, and it lives in behaviour management.

What to refuse to spend money or time on

  • Growth-mindset programmes and Brainology — zero on national tests in two independent national-scale trials, and the only trial of the commercial product was run by the vendor's employees and lost its preregistration badge.
  • Executive-function curricula, including Tools of the Mind — null under independent evaluation, with every significant effect favouring the control group in the largest such trial, and fidelity unrelated to outcomes.
  • Working-memory and brain-training software — see brain training; far transfer is 0.001 with zero true heterogeneity.
  • Universal school mindfulness for attention, behaviour or mental health — the definitive 8,376-student trial is null on all three co-primary outcomes and five secondaries went the wrong way.
  • Grit measurement of any kind — it adds four-thousandths of R² over ordinary personality measurement.
  • Branded SEL curricula bought on an attainment business case — the number in the business case is a 2011 estimate that the field has since more than halved and that no independent large trial has reproduced on an externally marked test.
  • Any character claim evidenced only by a rating scale the vendor wrote. Berry's UK PATHS trial is the demonstration: at 12 months the developer's own instrument was significant on 6 of 11 subscales and the standard instrument on 0 of 7, in the same children on the same day. By 24 months both were null.

The measurement rule that sorts this entire domain

If you read only one thing from this subject, read this. Sort every study by who wrote and scored the outcome, and the controversy disappears.

Outcome type What the character literature reports
Vendor's own rating scale Large effects (SEL skills 0.57; PATHS developer scale 6 of 11 subscales)
Teacher rating / teacher-assigned grade Small effects (mindset +0.11 GPA; SEL rated competence ≈ 0.22)
Externally set and marked test Zero (KS2 −0.01/−0.00/−0.00; PATHS −0.106 to +0.026; GBG 0.03; Argentina +0.015)
Administrative record (arrests, earnings, suspensions) Effects exist, but for structural interventions, not character ones

Two corollaries a founder should internalise. First, an active control halves or erases most of what is left — mindfulness on EF falls from 0.35 to 0.04 the moment the comparison group also does something. Second, follow-up erases most of the rest — EF training goes to 0.008, PATHS's developer-scale gains vanish by 24 months, Second Step's 42% aggression reduction is gone by year 3.

The hereditarian bottom line for a founder

The traits this subject proposes to teach are among the most heritable things measured in children, and — the part that matters more — their shared-environment variance is repeatedly indistinguishable from zero:

Trait Heritability Shared environment
Common executive function (latent) 99–100% ≈ 0
Self-control (meta of 31 twin studies) 60% ≈ 0
Grit, perseverance of effort 37% not significant

Three consequences.

  1. The observational case for character education is worthless. Rimfeld's twin analysis finds 88% of the grit–GCSE correlation is genetically mediated; the marshmallow association dies on controls for early ability and the home; Jacob and Parkinson find 1 of 13 EF–achievement associations survives control for background and IQ. Every gradient in this domain is doing the job the archive predicts it will do.
  2. "Shared environment ≈ 0" is a statement about schools. As Rimfeld puts it, current differences between families and schools explain little variance in the development of these traits. A programme aimed at a trait with near-zero between-school variance is aiming at a small target, which is what the effect sizes look like.
  3. But heritability is not a reason to dismiss a randomised trial, and this subject deliberately does not do that. Alan, Boneva and Ertac randomised 52 Istanbul schools, taught that effort is productive and failure informative for twelve weeks, and raised standardized mathematics by 0.31 SD immediately and 0.23 SD at 2.5 years — while moving teacher grades not at all. It is developer-led and unreplicated, and it is the single most interesting result in this tranche. (background: talent and trainability)

What survives replication, and what does not

Survives:

  • Behavioural challenge-seeking and advanced-course enrolment from mindset interventions (replicated US → Norway).
  • Rated social-emotional competence from SEL programmes (g ≈ 0.22, and larger under independent evaluation — the one place this archive's developer-bias prior was disconfirmed).
  • Near transfer from EF training, in proportion to task similarity.
  • Self-control as a predictor, including under a discordant-twin design.
  • Classroom management as a practice: 0.22 in 2016, 0.23 in 2025.

Does not survive:

  • Growth mindset on standardized achievement — zero in England and Argentina.
  • The praise-backfire effect — failed its only large direct replication.
  • Tools of the Mind — failed independently twice, with the Science paper rejected by WWC.
  • The marshmallow test's predictive claim — halved, then dead under controls, then nothing at mid-life.
  • Grit as a construct distinct from conscientiousness — refuted by three methods.
  • Durlak's +0.27 on achievement — 0.109 and 0.101 on re-estimation, never independently reproduced.
  • School mindfulness for mental health or executive control — nulls at 8,376 and 460 students.
  • Second Step's aggression reduction — gone by year 3 of the same trial.
  • The Good Behavior Game's Baltimore results, when transported and independently evaluated.

Open questions a founder should watch

  • Alan et al. 2019 is the exception that needs testing. An effort-beliefs curriculum that raises a standardized maths score at 2.5 years without touching teacher grades is not what any other result in this subject looks like. It is developer-led and has never been independently replicated. If it replicates, this subject's bottom line changes.
  • Why independent SEL evaluations produce larger rated-competence effects than developer-led ones (0.69 vs 0.21) is unexplained and cuts against the archive's own prior.
  • Whether the mindset GPA effect is learning or grading has never been tested — nobody has run the National Study of Learning Mindsets design with an independent standardized test.
  • Nothing in this subject is genetically informative about differential responsiveness to any of these interventions, which is the same untested trainability assumption the archive flags throughout.

Evidence topics

← Home