The Evidence on Teaching

What actually works in teaching — read honestly.

A graded evidence base for educating children 4–18, built on a premise most of the field won’t state: academic ability is substantially heritable. Most education research rests on correlations between what adults do and how children turn out — correlations that genes and selection produce for free. So no claim counts here until a design that randomizes, lotteries, or compares within families says it survives.

Two more things travel with every verdict. Practice history is evidence: a method teachers could have abandoned for centuries and didn’t carries weight a 2015 invention doesn’t — in proportion to how tight the survival filter was, and never as a substitute for trials. And every judgment is stored as data under declared weights — design rigor, evaluator independence, effect durability — published as defaults, not theorems. Disagree with the weights? Move them, and watch every verdict recompute.

88
decisions answered
one falsifiable conclusion each
1171
papers read
129 read and set aside — kept visible
51%
of recorded effects are end-of-treatment
the fadeout blindspot, measured
271
re-checked against the paper
the rest are marked unaudited, not hidden

What the archive is surest of

Start somewhere sharp

Subjects — how to teach each

Teaching character — executive function, self-control, mindset and social-emotional learning

Do not buy character as a curriculum. This is the literature with modern psychology's worst replication record and the largest recent school spending, and the pattern is almost perfectly consistent: interventions move the ratings and beliefs they are evaluated on, and do not move the attainment they are sold on. Growth mindset is exactly zero on national tests in two independent trials. Executive-function training decays to 0.008 at follow-up. Grit is conscientiousness renamed, adding 0.4% of attainment variance. SEL's famous +0.27 is now 0.10 and has never been reproduced independently on an externally marked test. The traits themselves are 37–100% heritable with shared-environment variance near zero, which is why the correlational case for character education is worthless and why the randomised case is so thin. What a school should actually run is a well-managed room — small, replicated, cheap — and teach the subjects.

4 decisions

Teaching children a foreign language

Foreign language is the subject where the single most confident popular belief — start as young as possible — has the least evidence behind it, and where the one structure with a lottery behind it has never measured the thing it is bought for. Three things are defensible. Hours beat starting age: at equal instruction, later starters learn faster, and two of the three large European early-start comparisons found the early cohorts level or behind years later. Immersion is the only structure with causal evidence, and its findings are that it does not cost English, mathematics or science and slightly helps English reading (lottery ITT +0.13 to +0.22 SD by grades 5-8) — while no lottery has ever measured how much of the second language students learn. And teach the forms explicitly: unlike the L1 grammar verdict in this archive, explicit form-focused L2 instruction beats implicit exposure repeatedly, including on spontaneous production — though the effect sizes come from treatment-aligned tests given to university students immediately after short treatments, so discount them hard. Refuse two claims outright: native-like attainment from school instruction at any starting age, and any cognitive or executive-function bonus from early bilingualism.

3 decisions

Teaching children history and civics

This is the subject people argue about most confidently on the least evidence, and the three best-identified trials in it are all nulls. What replicates: teaching history content richly raises history knowledge (g 0.19-0.46), teaching civics raises civic knowledge (d 0.49), and teaching with documents improves document work (g 0.42). What fails, repeatedly and under randomisation: transfer. Standardized reading moved in exactly one of eight content trials and did not replicate; a kindergarten curriculum moved aligned social-studies knowledge 0.93 and the standardized measure of the same construct 0.00 (p = .999); civic knowledge decayed to the control mean within two years and never touched attitudes or validated turnout; a 60-school RCT of document-based history found nothing on history knowledge, document analysis or interest while its teachers reported the opposite. Two things do work at a longer horizon: multi-year cumulative content (three-year spiral 0.11-0.12, six-year Core Knowledge lottery ITT 0.24) and, for turnout, school quality and childhood noncognitive skill — neither of which is a civics curriculum. Teach history for history, teach civics for civic knowledge, and buy the transfer only with years.

3 decisions

Teaching children mathematics

Teach math explicitly with worked examples and lots of practice; build fact fluency deliberately (timed, low-stakes — the anti-timed-test movement has no causal evidence); teach word-problem types explicitly as their own strand; avoid minimal-guidance constructivist curricula outright; and place students in algebra by measured readiness — accelerating the ready, double-dosing the not-yet-ready. The pieces of math are modular: teach each one, on purpose.

6 decisions

Teaching children science

Science is the subject where the strongest-sounding claims have the weakest designs behind them, and where nearly every programme deflates on contact with independent evaluation. Four things survive. Attach a teacher's conceptual explanation to every hands-on activity — guided inquiry is positive in 16 of 16 PISA regions and unguided investigation negative in 18 of 20. Name misconceptions out loud and refute them: free, worth ~0.3 SD, and it does not delete the intuition. Buy sustained teacher professional development rather than science kits (+0.36 against +0.02 on independent measures). Teach fewer topics for longer, and stop optimising the sequence — the one test of course order found nothing. Do practical work for enthusiasm and identity, the only outcome class where it survives a 205-school trial. And refuse the far-transfer promise outright: the one large randomised test of 'science teaches you to think' returned −0.15 in English and −0.11 in maths.

5 decisions

Teaching children to read

Teach the alphabetic code explicitly from day one, drive fluency with volume of reading, and build comprehension through a knowledge-rich cumulative curriculum rather than 'reading skills' lessons. The reading method is close to a solved problem; the comprehension engine (knowledge) is the part most schools get wrong.

6 decisions

Teaching children to write

Writing is the subject where the gap between popular practice and evidence is widest, and where the published effect sizes are least trustworthy. The single clearest finding is that teaching grammar does not improve writing at any age from 6 to 18 — three independent randomised trials and four meta-analyses agree, and the meta-analytic estimate against alternative instruction is negative. What does work is a short list: an explicit, modelled routine for planning and revising; sentence-level construction work from about age 9; explicit letter formation and spelling in the first two years; and volume of writing. Discount every number in this subject hard — researcher-scored writing quality gives d ≈ 0.8–1.1 for strategy instruction, and the same literature filtered for measures independent of the developers gives +0.18.

6 decisions

Teaching computer science

Teach coding because coding is worth knowing and pays — a twelve-week course moves a coding assessment about half a standard deviation, and a high-quality high-school CS course raises the odds of a CS degree by 5.5 points and earnings at 24 by about 8%. Never teach it as a route to general thinking. That claim was tested hard in the 1980s Logo era, failed, was forgotten, and came back verbatim as 'computational thinking' — and it has now failed again, more decisively, because this era finally built validated instruments and then returned d≈0.00 on them twice. Which teaching method you choose matters much less than the literature implies: the anchor meta-analysis says approaches differ only marginally, blocks-before-text is a meta-analytic null, and almost every effect size in the subject rests on an instrument its own developers built.

2 decisions

Teaching music

Teach music because music is worth teaching. Instruction moves musical skill by about half a standard deviation, which is a normal, respectable instructional return and the only claim about music that needs no hedge. Every spillover promise is weaker than advertised: the two largest randomised trials — 3,004 five-year-olds given daily music for a year, and a 2,914-child lottery for El Sistema — found nothing on reading, maths or cognition, and the disadvantaged-pupil subgroup came in at exactly 0.00. A small effect on laboratory executive function is genuinely contested and worth watching, but it is not a budget argument. Meanwhile the selection you are being sold as an effect is measurable: taking up an instrument is 78% heritable, self-selecting children are already 0.29 SD ahead before the first lesson, and the SAT advantage of music students goes from +37 points to +0.23 the moment prior achievement enters the model.

2 decisions

Teaching physical skills and athletic development

Teach specific movement skills because those skills have value in themselves — not because they will transfer to cognition, academic achievement, or lifelong activity, none of which survive rigorous testing. Structure practice for durable retention rather than for in-session performance, but be sceptical of importing adult laboratory findings: the most-taught motor-learning principle is null in children. Delay specialization, and never select young athletes on current performance, which mostly measures birth month and accumulated load.

4 decisions

Teaching visual art and drama

Arts education is cheap, does no academic harm, and cannot be defended on academic grounds. Two randomised programmes covering more than 440 schools — 42 in Houston, 400 across England — find nulls on reading and maths; the wider arts-integration field averages g = 0.11 with maths null, and one of forty-four catalogued interventions clears the top evidence tier. What does hold up is narrow and near: a single guided museum or theatre visit moves blind-rated analysis of art and tolerance by 0.08-0.18 SD weeks later, and a film of the same play moves nothing. The uncomfortable finding is about the subject's own core: there is no controlled trial anywhere in which art instruction is the input and measured artistic skill is the outcome. The field that most loudly claims spillover never tested the thing it does.

4 decisions

Teaching vocational and practical skills

Vocational and life-skills teaching are graded on the wrong outcome almost everywhere: test scores are flat and irrelevant, and the real currency is attainment, earnings and conduct. On that currency, selective CTE high schools deliver — male graduation up 8-10 points across three US states and early-career earnings up 17-35%, with the randomised trial moving earnings 17% while moving educational attainment by exactly zero. But nothing moves for girls under current enrolment patterns, nothing moves for test scores, nothing moves for degrees, and nobody has measured earnings past 33. Life-skills courses are the mirror image: they reliably teach the content and rarely change the conduct. The only behaviour that moves is behaviour close in time to the instruction — and driver education, three randomised trials deep, gets teenagers licensed sooner without making them safer.

2 decisions

The knowledge base of Millyard School. The colors are the Amoskeag Manufacturing Company’s own — indigo and unbleached cotton from its ACA ticking, granite and brick from the mile of mills where Manchester learned to make things.