The Evidence on Teaching

Talent and trainability — what is heritable, and what that does not license

Physical capacity is about as heritable as cognitive ability (~60%, no shared environment). Whether trainability is a stable trait remains unproven — the skeptics are winning.

strong supportconf: highgc: low

talent · ages 418 · input

Effect summary

Physical capacity is heritable at roughly the level of cognitive ability: measured VO2max h2 ≈ 60% (twin meta 59-72%), youth fitness and motor tests 52-79%, athlete status ~66% — with NO detectable shared-environment component anywhere. Whether TRAINABILITY (heritable differences in response to training) is a real trait is contested and the sceptical side has strengthened: HERITAGE's 47% is an uncontrolled family-design upper bound, control-arm meta-analyses find most response variance is measurement error, 'non-response' falls 69%→0% as training DOSE rises, and the flagship molecular result (a 21-SNP panel 'explaining 49%' of trainability) is an unreplicated in-sample stepwise fit.

Practical takeaway

Expect large, irreducible differences in physical capacity between children, and expect good coaching to widen rather than narrow them. Do not infer prediction from heritability: the molecular record is null and expert consensus — co-signed by HERITAGE's own principal investigator — is that no child should be genetically screened for athletic potential. Before labelling any child a 'non-responder', raise the training dose: in the one dose-varying study, non-response disappeared entirely at 4+ sessions a week.

Who this applies to

Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.

Verdict

This topic is the physical-domain counterpart to the project's founding premise, and the adversarial pass forced a sharper statement of it than the one this project started with. There are two claims here, and they have very different evidential standing.

Capacity is heritable — strongly supported. Measured aerobic capacity, youth fitness, motor coordination, and athlete status all show substantial heritability in twin designs, and — strikingly — no detectable shared-environment component in any of them. Household does not explain aerobic capacity.

Trainability is contested — and this project previously overstated it. The widely-cited HERITAGE figure of 47% heritability of the training response is an explicit "maximal heritability": an upper bound from a two-generation family design (not twins) that cannot separate genes from shared household, with no non-exercise control arm. Control-arm meta-analyses since find that most variance in change scores is measurement error. "Response to training is heritable" is a hypothesis, not an established fact.

What the evidence shows

Source Design Grade Key effect
Schutte 2016 twin + meta (n≈1,088) B Measured VO₂max h²=60% (47–69); meta 59% absolute / 72% per-kg; no detectable C
Silventoinen 2024 child twin study C Youth fitness/motor tests 52–79% heritable
De Moor 2007 twin + linkage C Athlete status h²=66%; no genome-wide significant locus
HERITAGE (Bouchard 1999) family training study C Response ranged ~0 → >1,000 ml/min; 2.5× more variance between than within families; "maximal" h²=47% (upper bound, no control arm)
Renwick 2024 meta, control-arm only B No strong evidence true VO₂max trainability exists; most change-score variance is measurement error
Fox 1996 twins reared apart B Motor-skill heritability rose with practice; rate of learning itself heritable
Rankinen 2016 (GAMES) GWAS, 1,520 elite B Zero replicated variants for elite endurance status
Montero & Lundby 2017 dose-varying training study, n=78 C Non-response 69% → 40% → 29% → 0% → 0% as dose rises 60→300 min/wk; gone entirely on re-training
Churchward-Venne 2015 retrospective pooled pre-post, n=110 D "No nonresponders" — but the rule is "improved on ≥1 of 4 outcomes", no control arm
Hecksteden 2015 statistical methods paper D A control arm is not sufficient; stable trainability needs repeated interventions
Ross 2019 (BJSM consensus) expert consensus statement C Disagrees with this topic: response variability real, genetic, h²≈50%
Bouchard 2011 GWAS, HERITAGE n=473 C 21-SNP panel "49% of variance" — stepwise in-sample fit, panel never replicated

Added 2026-07-30: the "non-responder" construct, and the strongest published case against us

A citation-graph audit surfaced five well-cited papers that answer back to HERITAGE and were never read here. Together they sharpen the trainability half of this topic without overturning it. Verdict stays strong-support and confidence stays high — both are carried by the capacity claim, which none of these papers touch, and by four grade-B sources (Schutte, Renwick, Fox, Rankinen) that still point the same way. What changes is the description of the trainability dispute, which was previously "contested" and is now better stated as "the non-responder construct is failing three independent tests at once."

1. Non-response is dose-dependent, which is not what a trait does. In the only study that varied the dose, the share of people classified as non-responders fell monotonically from 69% at one session a week to 0% at four, and non-responders re-trained at a higher dose all became responders (Montero & Lundby 2017). If "non-responder" named a stable heritable type, two extra sessions a week should not abolish it. The honest discount: groups were self-selected, not randomized, there was no control arm, and re-testing people who were selected for a low change score is the textbook setup for regression to the mean — which the paper does not rule out. Graded C for those reasons. Even discounted, it establishes that the label is dose-relative.

2. The most-cited "no nonresponders" claim is mostly a classification rule. Churchward-Venne 2015 reports that every older adult responded — under a criterion of improving on at least one of four noisy outcomes with no control group. With four measures that is close to arithmetically guaranteed. We grade it D and cite it as an example of how the responder/non-responder literature manufactures its own answers, not as support. Its useful content is the ranges, which are enormous and two-tailed: lean mass −3.3 to +5.4 kg, leg press −36 to +87 kg. Plenty of participants got worse on individual measures.

3. A control arm is not enough — and this is a correction to our own standard. This database has been treating "control-arm design" as the bar that separates real trainability evidence from artifact (that is why Renwick 2024 is graded B and HERITAGE C). Hecksteden 2015 shows the bar is higher than we have been saying: within-subject variation in training efficacy — the same person responding differently to the same training on different occasions — "may not be disclosed by comparison to a control group but calls for repeated interventions." And at the individual level, "a low signal-to-noise ratio may not be compensated by increasing sample size." So the study that would actually settle trainability is not merely an RCT with a control arm; it is one that trains the same people twice. That study has not been run.

4. The strongest molecular claim is an overfit. Bouchard 2011 is the paper most often cited as proof that trainability is genetic: a 21-SNP panel "accounting for 49% of the variance in VO₂max trainability", with low-allele carriers gaining 221 ml/min and high-allele carriers 604 ml/min. Those 21 SNPs were chosen by stepwise regression from 39 that were themselves pre-selected out of 324,611 in the same 473 people. No SNP reached genome-wide significance and the panel was never validated as a panel in an independent cohort. That is an in-sample fit, not a predictor — and it lands suspiciously close to the 47% "maximal heritability" from the same cohort. It is consistent with Rankinen 2016, where 45 candidate markers replicated in zero of seven independent cohorts.

5. And the expert consensus disagrees with us — say so plainly. Ross 2019, a BJSM consensus statement with ~270 citations, concludes that exercise response variability is real and that "a genetic component contributes" to it, citing heritability of roughly 50% of CRF response variance. This database's position is the minority one and readers should know that. Two things blunt the disagreement rather than dissolve it: its human heritability figure is the same HERITAGE family-design upper bound already downgraded here, so it is not independent confirmation; and the statement concedes that observed response ranges "will tend to inflate the true interindividual response variability" while recommending precisely the designs — randomized control arm plus multiple pre- and post-tests — that produced nulls when Renwick 2024 and Bonafiglia 2022 applied them. The consensus and the sceptics agree about the method; they disagree about what the method has so far shown.

Hereditarian-lens assessment

Risk: low — these are the genetically-sensitive designs (twin, reared-apart, family, GWAS), so the usual worry that a finding is really selection does not apply. The interpretive work is in what heritability does and does not license.

Practice amplifies genetic differences. The most striking result here is Fox 1996: in twins reared apart, heritability of a motor skill was already high at the first trial and increased with practice, and the rate of learning was itself heritable. This is the identical signature found in reading, where good instruction raises heritability — removing an environmental bottleneck lets genetic potential express. Coaching does not equalise athletes; done well, it spreads them out.

Heritable ≠ predictable. The molecular record is null — 1,520 elite endurance athletes, zero replicated variants, and the one apparent exception (Bouchard 2011's 21-SNP panel "explaining 49%") is a stepwise in-sample fit that was never validated as a panel — and the BJSM consensus statement, co-signed by Bouchard himself, concludes that genetic tests have no role in talent identification or individualised prescription and that no child should be screened. Heritability is a population-variance statistic; it supports no inference about an individual child.

The correction to our own premise. This project's workflow methodology asserted "response to training is heritable (HERITAGE ~0.47)". That is downgraded here: HERITAGE established familial aggregation of training response, which is real and interesting, but the 47% is an upper bound from a design that cannot deliver a clean heritability, and the control-arm literature questions whether the underlying trait exists. HERITAGE is graded C for this reason.

Boundaries & what critics say

  • Capacity vs trainability must stay separate — the first is solid, the second is not.
  • Twin designs assume equal environments; that assumption is doing work here, though the complete absence of a shared-environment signal across independent cohorts is hard to explain away.
  • Mosing's outcome is auditory discrimination, a capacity measure, not instrument performance.
  • Range restriction contaminates elite-athlete samples in both directions.
  • The whole trainability literature is adult, and this topic is about children (ages 4–18). Every source in the trainability half is out of band: Montero & Lundby are men aged 18–35, Churchward-Venne are over-65s, HERITAGE spans 17–65, Hecksteden and Ross carry no sample at all. Nothing here has measured whether children respond to training the way adults do, or whether the response distribution is even the same shape in a growing body. Flagged 2026-07-30; it is a real gap, not a nitpick.
  • The published consensus disagrees with this verdict (Ross 2019). We think it is restating a family-design upper bound rather than adding evidence, but a reader should weigh that we are the minority position in the field.

Practical guidance

  • Plan for wide, durable differences in physical capacity within any class or squad — and expect good coaching to widen the spread of outcomes, not compress it.
  • Never screen children genetically for athletic potential; it is both invalid and, per expert consensus, inappropriate.
  • Judge programmes on whether they raise the floor and teach real skills, not on whether they close gaps — closing gaps is not what training does.
  • Treat "this child is a responder / non-responder" claims with scepticism; the trait may not exist as advertised.

Open questions

  • Whether true individual differences in trainability exist at all, once measurement error and within-subject variation are properly modelled, is genuinely open.
  • No twin design has estimated trainability of VO₂max with a control arm — the study that would settle it has not been run. Sharpened 2026-07-30: per Hecksteden 2015 a control arm is not even sufficient. The decisive study needs a randomized control arm, multiple pre- and post-tests, and repeated training periods in the same people, so that within-subject variation in training efficacy can be separated from a stable subject-by-training interaction.
  • Whether the dose-dependence of non-response (Montero & Lundby) is a genuine dose effect or regression to the mean in a selected subgroup — the design cannot tell them apart, and no one has replicated it with randomized dose assignment.
  • Whether any of this transfers to children at all; the trainability literature has no paediatric arm.

Evidence (14 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore