The Evidence on Teaching

Are self-control and grit teachable levers on achievement?

Grit is conscientiousness renamed and adds 0.4% to grade prediction; self-control genuinely predicts life outcomes but is 60% heritable with zero shared environment, and training moves ratings, not lives.

mixedconf: highgc: medium

character · ages 418

Effect summary

Separate the trait from the outcome and the picture resolves. GRIT is not a distinct construct: it correlates with conscientiousness at ρ = .84 meta-analytically, shares 95% of its variance with conscientiousness in factor analysis, and has a genetic correlation of .86 with it in twins. Its incremental validity over conscientiousness is .000–.004 in R², and 0.5 percentage points of GCSE variance. Grit↔achievement is r = .18–.19, driven entirely by perseverance (.21–.26) not consistency (.08–.10), and the association is WEAKER against standardized than teacher-assigned outcomes. SELF-CONTROL predicts more robustly — the Dunedin gradient gives 13% vs 43% criminal conviction across the extremes, and the effect survives a discordant-twin design (B = −0.13 on school performance, halving to −0.07 under sibling IQ control). But it is 60% heritable with shared environment ≈ 0 across 31 twin studies, and the marshmallow test itself does not survive: β = 0.236 → 0.050 (ns) under full controls, and the original cohort's mid-life follow-up finds preschool delay predicts NONE of 11 capital-formation outcomes (average r = 0.02). TRAINING moves rater-report scales (d = 0.27–0.32) and, after bias correction, adult self-control training gives 0.13–0.24 with an author-allegiance effect. Exactly one RCT — developer-led, Istanbul, 52 schools — moved objective standardized maths (+0.31 SD, +0.23 SD at 2.5 years) while moving teacher grades not at all.

Practical takeaway

Stop measuring grit — it is conscientiousness with a new name and adds four-thousandths of R² over ordinary personality measurement, and no school decision should turn on it. Stop using the marshmallow test as a rationale for anything: the effect halves in a proper sample, vanishes under controls, and the original cohort's own mid-life follow-up finds the preschool measure predicts none of eleven adult outcomes. Self-control is different — it genuinely predicts, and it survives a discordant-twin test — but it is 60% heritable with zero detectable shared-environment variance, which means differences between families and schools explain almost none of it. If you want to try anything, the only trial that raised objective achievement taught children that effort is productive and that failure is informative, in twelve weeks, for pennies, and it raised standardized maths by 0.19–0.31 SD without touching teacher grades. That is worth replicating independently. It is not worth buying a character curriculum for.

Who this applies to

Group size
whole-class
Delivered by
teacher
Ages studied
418
Dose
The one trial with objective achievement gains ran a 12-week teacher-delivered curriculum of animated videos, goal-setting and effort-productivity beliefs, in Grade 4, with follow-up at 1.5 and 2.5 years. The rater-report training literature is typically 8–20 sessions in children under 10.
Cost
low
Moves
non-cognitive
Needs first
An objective outcome measure specified in advance. Every effect in this literature that rests on a rating scale is scored by someone who knows the child was in the programme.
Not for
Predicting or selecting individual children. Grit adds 0.4 percentage points of attainment variance beyond ordinary personality measurement, and preschool delay of gratification predicts nothing at all at mid-life. It is also not for explaining why a child is behind: the observational gradients are substantially genetic, and treating them as diagnoses of deficient character is exactly the error this archive exists to prevent.

Verdict

mixed, because two constructs with very different evidence get sold as one thing.

Grit does not survive as a construct at all. It correlates with conscientiousness at ρ = .84 meta-analytically, shares 95% of its variance with conscientiousness in confirmatory factor analysis, and has a genetic correlation of .86 with it in twins. Three methods, three research groups, one answer: grit is conscientiousness renamed. Its incremental contribution over conscientiousness to predicting academic performance is .000 to .004 in R², and 0.5 percentage points of GCSE variance in the twin sample. This is the archive's cleanest example of the jangle fallacy — a new name for an old measure, with a decade of policy attention behind it.

Self-control is a real and robust predictor. The Dunedin gradient is not a small effect: 13% versus 43% criminal conviction across the extremes of childhood self-control, with 96% cohort retention and court records as the outcome. And crucially it survives the design this archive demands — comparing dizygotic co-twins raised in the same home, the less self-controlled twin does worse at school and is more antisocial at 12.

But self-control is not obviously a lever. It is 60% heritable with shared environment indistinguishable from zero across 31 twin studies and 100,000 people. The marshmallow test — the single most famous demonstration in this literature — does not replicate at full strength and does not survive to mid-life. And the training literature moves rater-report scales without ever having demonstrated a change in an objective outcome.

One trial breaks the pattern, and the archive's rules require reporting it prominently: Alan, Boneva and Ertac randomised 52 Istanbul primary schools, raised standardized mathematics by 0.31 SD short-run and 0.23 SD at 2.5 years, and moved a behavioural effort task. It is developer-led, and what it taught was arguably effort-productivity beliefs rather than a trait.

What the evidence shows

Grit is conscientiousness

Source Design Grade Key effect
Credé 2017 Meta, 584 ES / 88 samples / 66,807 people B Grit↔conscientiousness ρ = .84; grit↔achievement ρ = .18; incremental R² = .000 / .002 / .004
Rimfeld 2016 Twin study, 4,642 UK 16-year-olds, GCSE records B h² = 37% (perseverance), C ≈ 0; grit adds 0.5pp of GCSE variance beyond Big Five; 88% of the grit–GCSE link is genetically mediated
Lam & Zhou 2022 Meta, 137 studies / N = 285,331 B r = .19; perseverance .21, consistency .08; association weaker on standardized than nonstandardized outcomes
Schmidt 2018 Hierarchical CFA, two German samples C Perseverance shares 95% of its variance with conscientiousness; grit "can be fully integrated into" it
Duckworth 2007 6 correlational studies, researcher-built scale D Grit explains ~4% of outcome variance; grit↔conscientiousness r = .77; grit↔ability r = .02
Credé 2018 Critique D No basis for combining perseverance and passion; no evidence grit responds to intervention

Note what Duckworth's own founding paper reports: r = .77 with conscientiousness, in Study 2, in 2007. The deflation did not require new data so much as somebody asking the question.

Note also Lam and Zhou's one surviving moderator, because it is the same measurement fault line that runs through growth mindset: the perseverance–achievement association is stronger for nonstandardized than for standardized achievement. The correlation weakens when the outcome stops being a teacher's judgement of the same child whose diligence is being rated.

Self-control predicts — and is heritable

Source Design Grade Key effect
Moffitt 2011 Dunedin birth cohort (n = 1,037, 96% retention) + E-Risk discordant-twin FE B Conviction 13% vs 43% across extremes; within-family school performance B = −0.13 → −0.07 under sibling IQ control; antisocial and smoking hold
Willems 2019 Meta of 31 twin studies, >100,000 people B h² = 60% (MZ .58 / DZ .28); C ≈ 0; h² = 75% parent-report, 53% self-report, 41% observation
Watts 2018 Conceptual replication, n = 918 (552 primary), SECCYD C Age-15 achievement β 0.236 → 0.081 → 0.050 (p = .140); behaviour outcomes null even unadjusted
Benjamin 2020 Preregistered ~45-year follow-up of the original Bing cohort (n = 113) B Preschool delay predicts none of 11 mid-life outcomes, average r = 0.02; removing it does not change the composite index
Shoda 1990 The original longitudinal claim D SAT-V r = .42, SAT-Q r = .57in one cell of n = 35; other cells −.12 to −.40; 46% response rate, responders delayed 62s longer
Falk 2020 Preregistered reanalysis + Monte Carlo — the steelman B 55% censoring attenuates a true .42 to .349; with predetermined controls only, β 0.236 → 0.137, still p < .001
Doebel 2020 Commentary D Argues the controls are constituents of delay ability, not confounds
Watts & Duncan 2020 Reply D Concedes partial replication of achievement correlations; maintains behavioural correlations are null

Willems's informant moderator deserves to be read twice. Heritability of self-control is 75% when a parent rates both twins, 53% by self-report, and 41% by observation. The genetic estimate is largest exactly where a single rater supplies both scores — which is a warning about the measurement, not a refutation of the finding, since even the observational estimate is 41% with shared environment near zero.

And Benjamin et al. is the most under-cited paper in this file. Mischel, Shoda and Peake are among its authors. They followed their own cohort to age 45–50 with a preregistered plan, and the preschool delay measure predicted none of eleven capital-formation outcomes, average r = 0.02. Removing it from their composite index did not reduce that index's predictive power at all.

Training: ratings move, lives are unmeasured

Source Design Grade Key effect
Alan 2019 School-randomised, AEA-registered, 52 schools, 3,200+ children, 2.5-yr follow-up A Standardized maths +0.31 SD, +0.19 at 1.5yr, +0.23 at 2.5yr; behavioural effort task moved; teacher grades unchanged
Piquero 2016 Meta of 41 RCTs, children under 10 C Self-control 0.32, problem behaviour 0.27 — all rater-report at post-test; no objective long-run outcome anywhere
Friese 2017 Meta, 33 studies / 158 ES, bias-corrected B g = 0.30 → 0.13–0.24 corrected; larger against inactive controls; larger when strength-model proponents were authors
Duckworth 2016 (stitch in time) Two small field RCTs C Outcome is self-reported goal attainment over one week. No achievement measure
Duckworth 2016 (situational strategies) Theoretical review D No trial, no effect size — what the leading advocates offer in place of one

Alan et al. is the exception and it should be treated seriously. It is registered, school-level randomised, internally replicated across two independent samples, and its achievement gains persist to 2.5 years on a researcher-administered standardized test. One detail makes it more credible rather than less: teacher-assigned grades did not move. In a literature where every other positive result lives on teacher judgement, this one lives on the test and not on the grade — the reverse of the growth-mindset pattern, and evidence against a grading artefact.

Three things bound it. It is developer-led. What it taught — that effort is productive and that failure is informative — is closer to an effort-beliefs curriculum than to grit as measured. And the authors themselves note that dropping their covariate set costs the sample-2 long-run effect its significance.

Hereditarian-lens assessment

Risk: medium, and this topic is where the archive's premise earns its keep most directly.

The observational literature relating self-control and grit to outcomes is confounded exactly as the premise predicts, and here we can put numbers on it rather than assert it:

  • 88% of the grit–GCSE correlation is genetically mediated (Rimfeld). The phenotypic association is r = 0.17; almost all of it is genes shared between the trait and the attainment.
  • Shared environment is zero. Not small — zero, and not significant, for grit and for every Big Five trait in a UK-representative twin sample, and approximately zero across 31 twin studies of self-control. Rimfeld states the implication plainly: current differences between families and schools explain little variance in the development of grit.
  • The marshmallow association dies on exactly the controls the premise nominates. Watts et al. find β = 0.236 falling to 0.081 with family background and home environment, and to 0.050 (p = .140) once concurrent 54-month cognitive and behavioural measures are added. Those controls are early ability and the home — heritable traits supplied by the same parents.

But the premise must not be allowed to do more work than it can. Two findings resist it and are recorded as such.

First, Moffitt's discordant-twin analysis. Comparing dizygotic co-twins in the same household, the less self-controlled twin at 5 does worse at school and is more antisocial at 12. That is a within-family design; shared genes and shared home are differenced out. The school-performance coefficient halves under sibling IQ control (−0.13 → −0.07) but survives, and the behavioural coefficients hold. Something about self-control is doing causal work.

Second, randomised programme evaluations are graded on their design, not on the premise. Alan et al. is an A-grade school-randomised trial with a 2.5-year objective follow-up, and heritability of the trait is not an argument against it. The correct reading of the twin evidence is that it demolishes the observational case for character education, not that it forecloses instruction. That distinction is the archive's premise (3), and this topic is where ignoring it would be most tempting.

The honest synthesis: the trait is substantially heritable and the correlational literature about it is nearly worthless; whether teaching effort beliefs raises achievement is a separate, open, and partly encouraging question.

Boundaries & what critics say

  • Falk, Kosse & Pinger 2020 is the strongest defence of the marshmallow test and it is a good one. Watts et al. censored the delay task at 7 minutes, truncating 55% of children against 15% in the original, and Monte Carlo simulation shows this alone attenuates a true .42 to .349 — statistically indistinguishable from the observed .30. Under predetermined controls only, the age-15 association falls by a third rather than two thirds and stays significant at p < .001.
  • The disagreement is about what counts as a confound, and both sides are coherent. Doebel et al. argue that early executive function and verbal ability are constituents of delay ability, so controlling them removes the effect of interest. This archive controls them, because the question it asks is whether there is a teachable lever left once heritable early ability is accounted for. Falk et al.'s own closing line concedes the point that matters: settling this requires targeted intervention studies. Those studies are the training literature, and they measure ratings.
  • Watts & Duncan concede partial replication. The achievement correlations are smaller but real. What did not replicate at all — even unadjusted — is the association with adolescent behaviour, which was the more sociologically consequential half of the original claim.
  • Moffitt's within-family result is the best case for self-control as a lever, and this topic does not dismiss it. Its limits are that the outcome horizon is age 12, the achievement coefficient halves under IQ control, and the predictor is a rater composite.
  • Piquero's meta-analyses look supportive and are not. Forty-one randomised trials with d = 0.32 sounds decisive until you check the instruments: the Kansas Reflection-Impulsivity Scale, the Kendall & Wilcox Self-Control Rating Scale, the Child Behavior Checklist — all rater report, at post-test, scored by adults who know which children were in the programme. No objective long-run outcome appears anywhere in either review.
  • Friese's author-allegiance moderator applies to this whole literature. Self-control training effects are larger when proponents of the strength model wrote the paper. Combined with growth mindset's financial-incentive finding, allegiance is now a documented moderator in two separate character literatures.
  • Nothing here argues against expecting effort. Attendance, homework completion, and persistence through difficulty are worth requiring on their own terms. The claim under test is that they are a trainable trait that transfers.

Practical guidance

  • Do not administer a grit scale. It measures conscientiousness, and it adds four-thousandths of R² over measuring conscientiousness. Nothing a school does should turn on the result.
  • Retire the marshmallow test as a rationale. The original is a 35-child cell with a 46% response rate; the replication halves it and controls kill it; and the original authors' own mid-life follow-up finds it predicts nothing.
  • Never use a self-control gradient as a diagnosis. It is 60% heritable with zero shared-environment variance. Reading a child's impulsivity as a character failing the family produced is precisely the inference the twin data forbid.
  • If you spend anything here, spend it on effort beliefs, not on trait training — and only as a cheap experiment. The Istanbul trial is twelve weeks, teacher-delivered, and moved a standardized maths test at 2.5 years. It is one developer-led trial and needs independent replication before it is a recommendation.
  • Demand an objective outcome from any character programme. If the evidence is a rating scale completed by the teacher who delivered the programme, there is no evidence.
  • Separate conduct from attainment when you set the goal. Wanting an orderly, diligent classroom is legitimate and is addressed by behaviour management, which has better evidence than anything in this file.

Open questions

  • Alan et al. is the single most important unreplicated result in this tranche. An independent replication with a standardized outcome would change this topic's verdict in either direction.
  • No study has established what the Moffitt within-family effect is. It survives sibling comparison, so it is not simply shared genes and shared home — but the mechanism is unidentified and the horizon is age 12.
  • Nobody has run a self-control training trial with an objective long-run outcome. Forty-one RCTs exist and not one measures achievement, earnings or convictions.
  • Whether the perseverance facet is separable from conscientiousness in any way that matters practically remains contested; the twin genetic correlation of .86 suggests not.

Evidence (20 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore