Deliberate Practice and Proposed Limits on the Effects of Practice on the Acquisition of Expert Performance: Why the Original Definition Matters and Recommendations for Future Research
Ericsson, K. A., & Harwell, K. W. · 2019
grade Dcritiquedeveloper-ledunclearnumbers spot-checked
Sample
14 independent effect sizes across 12 studies, retained from the 191 non-aggregated effect sizes (88 studies) in Macnamara et al. (2014)
Population
Post-hoc subset of the Macnamara meta-analytic database, predominantly music, chess/games and sports; zero effect sizes from professions, one from education
Design
No new empirical data. Re-analysis of Macnamara et al.'s (2014) own dataset after three post-hoc exclusionary criteria (performance measure diagnostic of domain skill and capturing reproducibly superior performance; practice designed to improve the targeted performance; separate estimate of solitary practice, teacher-guided or self-directed). Figure 3's own flow diagram: 191 non-aggregated effect sizes evaluated -> 188 non-duplicates -> 80 pass Criterion 1 -> 43 pass Criterion 2 -> 14 pass Criterion 3, minus 2 plus 2 = 14 independent effect sizes across 12 studies. That is a ~93% discard of the effect sizes, applied after the results were known and never prespecified. The retained 14 are then disattenuated using ASSUMED reliabilities (.60 for practice estimates, .80 for performance) that the authors argue for from unrelated literatures rather than measure in the included studies.
Key findings
The most favourable defensible estimate for deliberate practice — built on a post-hoc 14-of-191 effect-size subset plus disattenuation with assumed rather than measured reliabilities. Take it at face value and practice still leaves ~39% unexplained on the most generous correction and ~71% on the raw correlation. Ericsson's own numbers show teacher-designed practice (r=.56) is no better than self-directed practice (r=.51), Q(1)=0.22, p=.64 — the same null Macnamara & Maitra (2019) found experimentally. For an education archive the decisive number is the domain breakdown: after the filtering, ZERO effect sizes from professions and exactly ONE from education survive (r=.31, Spelling Bee). The paper concedes hours are a crude proxy and that future work should measure practice QUALITY.
Genetic confound
Two-sided, and the second side must be recorded. The paper does recommend combining detailed training histories WITH genome-wide analyses to assess relative contributions and interactions of hereditary and environmental factors — conceding that correlational DP data cannot separate them. But its main argument runs against the hereditarian premise: the abstract asserts that "genetic effects have so far accounted for remarkably small amounts of variance – with exception of genetic influences of height and body size," and the paper argues that heritability estimates from general-population twin samples should not be extrapolated to expert performers ("what is" vs "what could be"), that Hambrick & Tucker-Drob's music heritability goes non-significant under Ericsson's reanalysis, and that Mosing et al. (2014) never published the professional-musician ACE analysis despite repeated requests. These are arguments, not data.
Replication notes
What exists is methodological rebuttal, not a replication attempt. Hambrick & Macnamara (2020, Front. Psychol. 11:1134, "Is the Deliberate Practice View Defensible?") criticise the post-hoc subsetting and the speculative disattenuation; Macnamara & Hambrick (2020, Psychological Research) make the related charge on Moxley, Ericsson & Tuffiash (2017), and Ericsson (2020, Psychological Research) replied. Nobody has attempted to reproduce the 14-effect-size reanalysis independently, so we record unclear rather than failed — the criteria were applied after results were known and never prespecified, which is a design objection, not a failed replication.
DOI / URL
10.3389/fpsyg.2019.02396
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Performance, narrowly-defined DP subset (uncorrected) | r; % variance | r=.54, 95% CI [.44, .63], p<.001, k=14; ~29% of variance (vs Macnamara's 14% overall). Still leaves ~71% unexplained. | mixed | attained expertise; cross-sectional; field | none | not-applicable | domain-skill |
| Same subset after assumed-reliability disattenuation | % variance | ~61% of variance — the strongest number the DP camp has produced, contingent on assumed reliabilities of .60/.80 and on keeping only 14 of 191 effect sizes | mixed | attained expertise; cross-sectional; field | none | not-applicable | domain-skill |
| Deliberate (teacher-guided) vs purposeful (self-directed) practice | r | DP r=.56 vs purposeful practice r=.51, Q(1)=0.22, p=.64 — statistically indistinguishable, undercutting the claim that teacher-designed practice is the active ingredient | mixed | attained expertise; field | none | not-applicable | domain-skill |
| Domain breakdown of the surviving 14 effect sizes — education and professions all but vanish | r; k | Games r=.50 (k=5), music r=.71 (k=3), sports r=.58 (k=5), education r=.31 (k=1, the Duckworth Spelling Bee study), professions r=n/a (k=0 — no effect size from the professions met the criteria). On Ericsson's own most-favourable filtering, the deliberate-practice case rests on games, music and sport, and there is essentially no education evidence left. | mixed | attained expertise; cross-sectional; field | none | not-applicable | domain-skill |
| Claim (argued, not estimated): genetic contributions to expert performance are small and general-population heritability does not transfer to experts | assertion, no new data | The abstract states "genetic effects have so far accounted for remarkably small amounts of variance – with exception of genetic influences of height and body size." Supported by argument and selective reanalysis (Hambrick & Tucker-Drob music heritability reported as non-significant under the authors' re-specification; GWAS hit-rates; the unpublished Mosing professional-musician ACE analysis), not by any genetically informative study the authors ran. Recorded here so the record is not silent on the paper's central claim against this archive's premise. | mixed | not-applicable | none | not-applicable | domain-skill |
Cited by
- Deliberate practice and the 10,000-hour rulemixedconf: highgc: medium