The Evidence on Teaching

Deliberate practice and the 10,000-hour rule

The 10,000-hour rule is dead: practice explains ~14% of performance variance, is itself heritable, and identical twins 20,000 hours apart didn't differ. Necessary, insufficient.

mixedconf: highgc: medium

deliberate-practice · ages 618 · method

Effect summary

The 10,000-hour rule is dead as stated. Accumulated deliberate practice explains ~14% of performance variance overall (games 24%, music 23%, sports 20%, education 5%, professions <1%) — and that figure is an UPPER BOUND, not a causal estimate, because practice is itself 40-70% heritable and MZ co-twins differing by up to 20,000 hours did not differ in ability. Practice is necessary and insufficient.

Practical takeaway

Practice is necessary — nobody reaches expertise without it — but it is not sufficient and there is no hour threshold. Reject both 'anyone can be great with 10,000 hours' and 'talent is everything.' Design for accumulated quality practice while expecting large, irreducible individual differences in rate and ceiling.

Who this applies to

Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.

Verdict

Ericsson's framework — that expertise is "largely" explained by accumulated structured, effortful, feedback-rich practice — was popularised by Gladwell into a threshold rule that the original paper never claimed (10,000 hours was the best violinists' mean by age 20, not a cutoff). The honest number is far smaller than either version implies, and the deflation runs deeper than the headline meta-analysis because of a genetic confound the meta cannot address.

What the evidence shows

Source Design Grade Key effect
Macnamara 2014 meta, 5 domains C 12% (corrigendum 14%) of variance; games 24%, music 23%, sports 20%, education 5%, professions <1% (n.s.)
Macnamara 2016 sports meta C 18–20%; 19% sub-elite → 1% (n.s.) among elite
Hambrick 2014 chess/music cases C Masters ranged 832–24,284 h; 4 players exceeded the master-group mean (10,530 h) and stayed intermediate
Macnamara & Maitra 2019 preregistered replication B Best violinists logged fewer hours (8,224) than the merely good (9,844)
Mosing 2014 twin (N=10,539) B Practice is 40–70% heritable; within MZ pairs differing up to 20,228 h, the twin who practised more was not better

Three findings kill the threshold reading. Chess masters' practice hours span a nearly 30-fold range; a third of masters had less practice than the mean of the tier below; and in a preregistered, double-blind replication of Ericsson's own 1993 study the best violinists had fewer accumulated hours than the merely good ones, with both groups passing 10,000 hours by age 20.

The variance-explained figure tracks artifacts as much as reality. It rises with domain predictability (24% high / 12% moderate / 6% low) and falls sharply with measurement quality — interview 20% > questionnaire 12% > objective logs 5%. The headline numbers rest substantially on retrospective recall contaminated by current skill.

Ericsson's rebuttal is testable and fails. Every definitional restriction he demanded still leaves the sports estimate at 17–22%; his own best-case reanalysis retains 14 of 88 effects and assumes reliabilities to reach 61% — and even there, teacher-designed practice (r=.56) is statistically indistinguishable from self-directed practice (r=.51).

Hereditarian-lens assessment

Risk: medium, and this is the most important thing on the page. The R² figures are correlations between practice and attainment, and the adversarial pass established that they cannot be read as practice's causal contribution:

  • Practice is itself heritable (h² ≈ .69 men / .41 women), so people who practise more are genetically non-random.
  • The practice–ability covariance is predominantly genetic (rA .33–.57) with a nonshared- environment correlation of ~0.
  • In the decisive test, MZ co-twins discordant by up to 20,228 practice hours did not differ in ability (all p > .2).

So 14% is an upper bound containing gene-environment correlation; the true causal environmental effect is smaller. Note the confound runs in the deflationary direction — it makes the verdict more secure, not less, which is why confidence stays high at medium risk.

Two honest counterweights. The elite r ≈ .11 is computed after selection on both practice and performance, so range restriction inflates the collapse — the artifact licenses neither "practice stops mattering at the top" nor a hereditarian reading. And Mosing's outcome is auditory discrimination (a capacity), not instrument performance, which the authors concede likely does improve with practice.

Boundaries & what critics say

  • This is not "talent is everything." Practice is necessary, non-trivial, and the largest single identified predictor in several domains.
  • Domain matters enormously: 24% in games vs <1% in professions. Highly predictable, well-structured activities are where practice buys most.
  • Retrospective measurement inflates the estimates; objective logs give 5%.

Practical guidance

  • Provide structured, feedback-rich practice with a coach — it is necessary, and the alternative is worse.
  • Do not promise thresholds. No hour count guarantees expertise; the range is enormous at every level.
  • Expect and plan for large differences in rate, rather than treating slow progress as insufficient effort — the strongest causal test finds no within-genotype effect of extra hours.
  • Prefer domains and framings where practice actually pays (structured, predictable skills).

Open questions

  • What the causal environmental effect of practice actually is, once gene-environment correlation is removed, is unestimated — only bounded above.
  • Mosing's design needs replication on performance outcomes (instrument playing, sport skill) rather than discrimination capacity.

Evidence (8 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore