The Evidence on Teaching

Fadeout — why early gains disappear, and what actually persists

Early gains fade by default — halving every 12–18 months, faster when bigger. What persists is trajectory (placement, graduation), not ability.

strong supportconf: highgc: low

early-childhood · ages 418 · structure

Effect summary

Fadeout is the base rate, not the exception: intervention effects decay roughly geometrically, halving every 12-18 months after treatment ends (0.23 SD → 0.10 SD within a year). Bigger short-run effects fade MORE. The mechanism is outcome-specific — IQ fades because the treated group DECLINES, achievement fades because controls CATCH UP. What survives is not ability but institutional trajectory: less special-ed placement, less retention, more graduation.

Practical takeaway

Never buy an intervention on its end-of-treatment effect size — that number predicts almost nothing about durability, and larger ones fade faster. Judge programs on ≥1-year follow-ups measured with independent tests, and expect roughly half the effect to be gone within a year. If you want durable value, target institutional trajectory (placement, retention, graduation) rather than test scores.

Who this applies to

Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.

Verdict

Fadeout is the single most important fact in this database, and the one that has recurred in every tranche — reading, math, the methods layer, and now early childhood. Intervention effects on test scores decay roughly geometrically, halving every 12-18 months after treatment stops. This is not a defect of bad programs; it is the base rate that good programs are measured against, and it is established across three independent meta-analyses covering 67, 39, and 86 studies.

Two findings should change how a school builder evaluates any program:

  1. Larger short-run effects fade more (Hart et al. 2024, 86 RCTs). The impressive end-of-treatment number is actively anti-predictive of durability.
  2. Social-emotional skills are not stickier than cognitive ones — if anything cognitive skills persist slightly better at 1-2 years. The popular belief that "non-cognitive" gains last is unsupported.

What does survive is a different outcome class: institutional trajectory — special-education placement, grade retention, graduation — and behaviour, not measured ability.

What the evidence shows

Source Design Grade Key finding
Bailey 2017 67 ECE interventions C 0.23 SD → 0.10 SD at 12 months, halving again over 1-2 years
Protzko 2015 39 RCTs, N=7,584 B IQ fadeout is treatment-group decline (b=−1.09), not control catch-up (b=−0.21, ns)
Hart 2024 86 RCTs, N=56,662 B Larger posttest impacts fade more; cognitive ≥ social-emotional in persistence
McCoy 2017 22 studies, attainment outcomes C Special ed −0.33, retention −0.26, graduation +0.24 — these persist
Bitler 2014 HSIS quantiles B Effects >1 SD at the bottom, ~0.25 SD at the top — floor remediation
Watts 2018 conceptual replication, SECCYD n=918 C Marshmallow → age-15 achievement 0.236 → 0.081 → 0.050 (ns) as controls are added

The mechanism is outcome-class-specific, and conflating the two is a common error:

  • IQ fades by treatment-group decline. Protzko's growth-curve analysis is decisive: controls stay flat while the treated group loses its gain. The illustrative case is devastating — a nutritional supplement raised IQ 0.44 SD at age 4 with no environmental change at all, and it was fully gone by age 7. So fadeout is not simply "the child returns to a disadvantaged environment."
  • Achievement fades by control-group catch-up. Head Start's effect is >1 SD at the bottom of the distribution and ~0.25 SD at the top — it remediates a floor, and school remediates the same floor for everyone else within a couple of years.

Added 2026-07-30: the marshmallow test, and why it does not change this verdict

A citation-graph audit surfaced Watts, Duncan & Quan 2018, the conceptual replication of the marshmallow test, and it belongs here — but as a different fact from fadeout, so we state the distinction rather than blur it. The verdict and confidence are unchanged.

Fadeout is about a causal effect decaying after treatment stops. Watts is about a correlation evaporating once you control for what produced it. In a SECCYD sample roughly ten times the original 35–48 children, an extra minute of delay at age 4 predicted about 0.1 SD in age-15 achievement — half the original correlation — and that fell by two thirds with family background, home environment and early cognitive ability controlled (β 0.236 → 0.081), then to a non-significant 0.050 (p = .140) with concurrent 54-month measures added. Behavioural outcomes at 15 showed essentially nothing even without controls. Most of the signal came from waiting 20 seconds, not from heroic self-control.

Why it earns a row anyway: this topic's practical rule is "never buy an intervention on its end-of-treatment number." Watts extends the same discipline one step earlier in the pipeline — never buy a program on an early predictive correlation either, because the correlation and the durability are failing for related reasons. Both the fadeout literature and Watts say the early number is not the thing you think you are buying.

Honest limit: this is observational (grade C), it is contested in both directions — Falk, Kosse & Pinger (2020) argue the direction survives a like-for-like comparison, Doebel et al. (2020) argue the controls over-adjust — and it adds no causal weight to a verdict already carried by three RCT meta-analyses. It does not shift genetic_confound_risk, which stays low because the verdict still rests on randomized designs.

Hereditarian-lens assessment

Risk: low — the evidence base is RCT-heavy. And this is the topic where the lens supplies the mechanism rather than merely a caution. The behavioural-genetic architecture predicts exactly what the intervention literature observes:

  • The shared-environment component of cognitive ability — the channel interventions load onto — shrinks with age: c² falls from ~0.33 at age 9 to ~0.16 at 17 (Haworth 2010) and lands near 0.10 in adulthood (Bouchard 2013), while heritability rises toward ~0.80. An intervention that boosts a component which is itself contracting will see its effect contract too.
  • But note the honest ceiling: c² is ~0.10, not zero. This database does not use the common overstatement "shared environment goes to zero." That residual leaves room for roughly 4-5 IQ points — which is almost exactly what a whole childhood of improved rearing actually delivers.
  • Crucially, c² for attainment stays much larger than c² for ability. That single fact explains the signature pattern of this whole literature: early programs move attainment and behaviour durably while ability reverts.
  • The lens also reads Watts 2018 a particular way: the covariates that dissolve the marshmallow association are early cognitive ability and the home environment — both substantially heritable, both supplied by the same parents. "Controlling for confounds" and "controlling for heredity" are largely the same operation here, which is why the residual is so small.

What defeats fadeout (and what doesn't)

  • "Sustaining environments" is the field's favourite hypothesis and it is largely untested. Direct tests are sparse and mostly null — Abecedarian's supplementary school-age arm added nothing.
  • Targeting institutional trajectory works. Getting a child out of a special-education placement or a retention, or into a diploma, does not require sustained cognitive advantage — the "foot in the door" channel. This is the best-supported route to durable value.
  • The Boston case is the cleanest existence proof: a precisely estimated zero on achievement across grades 3-10 alongside real gains in graduation and college enrolment (Gray-Lobe 2023).

Practical guidance

  • Discount end-of-treatment effect sizes by default. Assume ~half is gone within a year, and be more suspicious of larger numbers, not less.
  • Demand ≥1-year follow-ups on independent, standardized measures before believing any program claim. Most published effects never report one.
  • Judge early programs on trajectory outcomes (placement, retention, graduation, behaviour) rather than on test-score bumps.
  • Don't pay a premium for "non-cognitive" or "social-emotional" durability — the evidence says those effects are no stickier.

Open questions

  • Whether "sustaining environments" is a real causal channel or a post-hoc label — few studies randomise the later environment, so the hypothesis is largely untested.
  • Whether attainment effects reflect genuine skill mediation or selection/measurement artifacts; the strongest evidence (Boston) is a lottery, but the sibling-design evidence has failed replication.
  • Whether Protzko's treatment-decline mechanism generalises beyond IQ to achievement and behaviour.

Evidence (10 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore

Fadeout — why early gains disappear, and what actually persists · The Evidence on Teaching