▸ 10/10 strong support Early gains fade by default — halving every 12–18 months, faster when bigger. What persists is trajectory (placement, graduation), not ability. 48/48
▸ 7/7 strong support Timed retrieval practice builds arithmetic automaticity — beating identical untimed tutoring head-to-head — and the anti-timed-test harm claims have no causal evidence. 11/48
▸ 10/10 strong support Testing yourself beats rereading — robust in real classrooms; honest durable size ~0.1–0.3 SD, biggest after delay, thinnest on transfer. 18/48
▸ 7/7 strong support Spacing beats massing at equal total time — the most robust finding in learning science. Space repetitions at ~10–20% of how long you need to remember. 12/48
▸ 18/18 strong support Systematic phonics beats whole language for word reading — largest in K-1 and for at-risk readers; near-null for comprehension and past grade 3. 36/48
▸ 14/14 strong support Physical capacity is about as heritable as cognitive ability (~60%, no shared environment). Whether trainability is a stable trait remains unproven — the skeptics are winning. 36/48
▸ 11/11 strong support Teacher quality is the largest within-school lever — 1 SD of teacher ≈ 0.10–0.15 SD/yr, worth more than ten fewer students — and credentials predict none of it. Select; don't workshop. 36/48
▸ 18/18 strong support High-dosage tutoring is education's most reliable lever: ~0.29 SD in trials, ~0.2 well-scaled — not Bloom's 2σ. Groups of 3–4 work; 1:1 is unnecessary. 30/48
▸ 12/12 moderate support Group within classes or across grades by subject: modest, nearly free wins. Whole-school streaming does nothing, and early between-school tracking harms the bottom. 36/48
▸ 8/8 moderate support Accelerate ready kids: they keep pace with older classmates, bank a year, and show no social-emotional harm at 50. The gifted label itself does nothing; the content does. 30/48
▸ 10/10 moderate support Minimal-guidance math trailed every rival in the one multi-curriculum RCT; explicit instruction for strugglers is math's most replicated result. 21/48
▸ 7/7 moderate support Word-problem solving is its own skill: computation fluency doesn't produce it. Teaching problem schemas explicitly does (~0.25–0.45 SD on trained content). 8/48
▸ 9/9 mixed Neither acceleration mandates nor delay mandates work — readiness-matched placement does. Push everyone and the median falls; hold everyone back and the top falls. 36/48
▸ 20/20 mixed Grit is conscientiousness renamed and adds 0.4% to grade prediction; self-control genuinely predicts life outcomes but is 60% heritable with zero shared environment, and training moves ratings, not lives. 36/48
▸ 7/7 mixed Smaller classes buy real K-1 gains that fade on tests yet persist in attainment — at roughly triple tutoring's cost, and diluted to nothing when scaled fast. 48/48
▸ 9/9 mixed Curriculum is nearly free, so choosing beats not choosing — but the payoff is avoiding a demonstrated loser, not finding a magic winner. Content is the high-upside bet. 27/48
▸ 8/8 mixed The 10,000-hour rule is dead: practice explains ~14% of performance variance, is itself heritable, and identical twins 20,000 hours apart didn't differ. Necessary, insufficient. 36/48
▸ 18/18 mixed Growth-mindset interventions change beliefs almost everywhere and change achievement almost nowhere: the two independent national-scale trials measured standardized tests and found exactly zero. 29/48
▸ 18/18 mixed Feedback on the task helps modestly; feedback on the person backfires — a stable third of studies reverse. The famous 0.4–0.7 number has no computed source. 12/48
▸ 7/7 mixed The famous faded-feedback demo doesn't survive meta-analysis; mental practice is real at half the advertised size and never substitutes for physical practice. 9/48
▸ 6/6 mixed Mastery learning moves tests of what it taught (~0.25) and barely moves independent measures (~0.05) — and time-to-mastery gaps widen, converting ability differences into time. 9/48
▸ 11/11 mixed The textbook matters at the bottom, not the top: avoid the demonstrated losers; mainstream choices now differ by ~0.02 SD. 24/48
▸ 7/7 mixed Varied practice looks worse today and better at retention — in the lab. Applied settings shrink the edge to zero, and it reverses in under-18s. 6/48
▸ 17/17 mixed At scale, pre-K doesn't durably raise test scores — Tennessee went negative — yet Boston shows real attainment gains beside a test-score zero. It buys trajectory, not ability. 48/48
▸ 8/8 mixed Unallocated money is the worst buy in the database; the same dollars as No-Excuses charters buy +0.3–0.4 SD/yr — and the average charter is a precise null. 39/48
▸ 6/6 mixed Starting school older mostly manufactures an age-at-test artifact: the IQ effect collapses to ~−0.07 once identified. Real residues: less hyperactivity, unchanged attainment. 36/48
▸ 19/19 mixed SEL programmes move the ratings they are evaluated on far more than the attainment they are sold on: the famous +0.27 is now 0.10, and no large independent trial has reproduced it on an externally marked test. 48/48
▸ 12/12 mixed Most measured 'home environment' effects are parents' genes: three-quarters of parent-child transmission isn't rearing, and a whole better childhood buys ~4 IQ points. 36/48
▸ 21/21 no effect Executive function trains like a task, not like a capacity: gains are real, narrow, and gone at follow-up, EF curricula are null under independent trial, and the latent construct is ~100% heritable. 48/48
▸ 7/7 no effect Teaching movement teaches movement: the skills improve, but transfer to cognition, achievement, or lifelong activity is near zero in the best trials. 10/48
▸ 18/18 no effect Brain training improves the trained task and nothing else: far transfer to intelligence or achievement is 0.001 against active controls. 24/48
▸ 13/13 debunked Matching instruction to 'learning styles' does nothing: real matching experiments return d=.04. The correlational literature IS the myth. 24/48
▸ 18/18 strong support Later bells buy adolescents 40+ measured minutes of sleep — the best-identified positive effect in the archive; the achievement payoff is real but far smaller. 36/48
▸ 4/4 moderate support Yes, on their own terms: lottery-assigned museum and theatre trips move blind-rated analysis of art, plot knowledge and tolerance by 0.08-0.18 SD weeks later, and a film of the same play moves nothing. 14/48
▸ 5/5 moderate support Comprehension is knowledge — but vocabulary teaching moves standardized comprehension only d≈0.10. The big content-knowledge bet (Core Knowledge lottery, 0.24) is real and unreplicated. 18/48
▸ 19/19 moderate support Selective CTE high schools raise male graduation 8-10pp and early-career earnings 17-35% on lottery and cutoff designs; test scores, degrees, and every outcome for girls are flat. 48/48
▸ 8/8 moderate support Immersion costs nothing in English and buys a little — a lottery puts English reading 0.13-0.22 SD ahead by grades 5 and 8 — but no lottery has ever measured how much of the second language students actually learn. 24/48
▸ 8/8 moderate support Teaching handwriting works and transfers: freeing the hand frees composition, with gains still present at six months — even taught in groups of three. 18/48
▸ 8/8 moderate support Sentence combining improves writing where grammar teaching fails — the same meta-analyses score them +0.50 and −0.32. 9/48
▸ 8/8 moderate support Strategy instruction is writing's best-supported method — at roughly a fifth of its advertised size once measures are independent (d≈0.8 → ~0.15). 9/48
▸ 6/6 moderate support Yes — sight-reading, performance and aural skill move about half a standard deviation under instruction — but the evidence is quasi-experimental and thinner than the transfer literature built on top of it. 36/48
▸ 22/22 moderate support Teaching coding teaches coding: the one clean school RCT gives g=0.47-0.68. Which approach you pick barely matters, and almost every effect size rests on an instrument the developers built. 25/48
▸ 7/7 moderate support Formal spelling instruction works (ES 0.54) and more formal beats less formal — explicit wins again, and it transfers beyond spelling itself. 9/48
▸ 13/13 moderate support Unassisted discovery loses to explicit teaching; well-scaffolded guided discovery beats both. The operative variable is guidance, not ideology. 15/48
▸ 7/7 moderate support Mixing problem types helps where confusion is the enemy — discriminating similar categories, mixed math practice — and is useless or worse for facts and prose. 18/48
▸ 3/3 moderate support Manipulatives help when bland, guided, and aged ~7–11 — a guided-representation effect, not 'hands-on learning.' Rich, toy-like objects hurt transfer. 6/48
▸ 6/6 moderate support Phonemic awareness transfers to reading only with letters attached: speech-only training peaks near 10 hours and moves reading d=0.19; with print, 0.66. 7/48
▸ 16/16 moderate support Diagnose what a child already knows: teachers cut 40–50% of curriculum for high-ability children and achievement ROSE. The payoff is skipping, not monitoring. 25/48
▸ 4/4 moderate support Guided oral reading works, but the ingredient is volume, not repetition: at equal exposure, re-reading has no edge over wide reading. 6/48
▸ 17/17 moderate support Content-rich history teaching raises history knowledge (g 0.19-0.46) but not standardized reading in under three years; generic 'historical thinking' has no adequately controlled positive result. 30/48
▸ 17/17 moderate support Outdoor time prevents myopia from starting (not progressing); screening plus free glasses raises test scores in children who need them. Two claims, both real. 36/48
▸ 6/6 moderate support Novices learn faster studying solutions than solving problems — then the effect reverses with expertise. Use worked examples early; fade them. 7/48
▸ 20/20 mixed Correcting a real deficiency moves cognition (iron in anaemic children: 0.79 SD); supplementing already-fed children moves nothing (35 RCTs, 19,343 children). 18/48
▸ 28/28 mixed Capitalisation and punctuation move when directly taught and directly measured, and don't move when embedded in writing programmes. A teachable subject, not a writing lever. 21/48
▸ 14/14 mixed A missed day costs little (~0.005 SD); interventions reliably buy days back cheaply, but nobody has shown the recovered days move achievement. 39/48
▸ 11/11 mixed Explicit form-focused teaching beats pure exposure, so the archive's negative verdict on L1 grammar does NOT transfer — but the effects are measured on the taught forms, immediately, mostly on adults. 18/48
▸ 11/11 mixed Life-skills courses reliably teach the content and rarely change the conduct; behaviour moves only where instruction sits close in time to the decision, and driver education is worse than nothing. 48/48
▸ 19/19 mixed Civics teaching buys civic knowledge (d = 0.49) that fades to the control mean in two years and moves neither attitudes nor validated turnout; what moves voting is school quality and noncognitive skill. 36/48
▸ 7/7 mixed The headline effects are a measurement artefact: drama gives d≈0.89 on researcher-made tests and d≈0.29 on standardized ones, and the two randomised trials with standardized outcomes are null. 15/48
▸ 13/13 mixed Document-based instruction reliably improves the sourcing and argument tasks it teaches (g ≈ 0.42) and, in the two best-identified trials, moves neither standardized reading nor history knowledge. 10/48
▸ 53/53 mixed Sort by what the software replaces: adaptive drill inside the school day buys +0.05-0.20 SD; a device, a connection or the teacher buys zero to negative; LLM tutors have no usable evidence at all. 39/48
▸ 19/19 mixed Not at school: the largest randomised trials find nothing on reading, maths or cognition. A small effect on laboratory executive-function tasks is contested, may be real at ~0.2 SD, and is end-of-treatment only. 36/48
▸ 7/7 mixed Early specialization predicts junior success; later starts plus other sports predict adult world-class. Selecting on junior results selects the wrong athletes. 24/48
▸ 19/19 mixed The received wisdom that retention is harmful was two measurement errors, not a finding. Fix them and the achievement effect is zero — but the grade at which you hold a child back decides everything. 36/48
▸ 16/16 mixed Near-zero in primary and almost entirely correlational in secondary. The causal base is three small trials; what the homework IS beats how much of it there is. 29/48
▸ 25/25 mixed Classroom order is causally worth a great deal — one disruptive peer costs classmates 3% of adult earnings — yet no branded behaviour programme reliably delivers it and exclusion makes things worse. 36/48
▸ 19/19 mixed Guided inquiry is positive in 16 of 16 PISA regions and unguided inquiry negative in 18 of 20 — but the programmes built on that finding go to zero at scale: +0.22, then +0.01, then +0.02 on the same test. 27/48
▸ 14/14 mixed Attainment falls steadily with age of first exposure, but the sharp critical period is contested and an earlier classroom start buys nothing durable: start young for immersion, not for two lessons a week. 24/48
▸ 22/22 mixed Two reviews screened ~12,000 records and found one RCT and six quasi-experiments. What evidence exists favours interest (+0.12) over attainment (+0.01), and simulations match real apparatus. 21/48
▸ 23/23 mixed Taking time away hurts measurably; adding it back buys almost nothing. The return per hour is ~0.02-0.03 SD, concave, and near zero in disorderly classrooms. 38/48
▸ 8/8 mixed Perry and Abecedarian: n=123 and n=111, IQ gains gone by adolescence, attainment effects real but multiplicity-fragile. Too thin to carry the policy built on them. 36/48
▸ 5/5 mixed Reading Recovery buys large short-term gains at high cost; the only US long-term estimate is negative. Early 1:1 tutoring works — this model's durability is unproven. 27/48
▸ 4/4 mixed Strategy instruction is a one-time boost, not a trainable skill: six sessions buy what fifty do, and modern at-scale trials are ~null. 6/48
▸ 10/10 mixed Conceptual-change teaching reliably moves the test and nothing shows it removes the intuition. Corrected for design and publication bias, a taught unit is g≈0.64 and a refutation text g≈0.28. 17/48
▸ 20/20 mixed Coaching buys ~0.25 SD of score and has for 42 years — a fraction of what's advertised — and none of it is ability. The score moves; the construct doesn't. 36/48
▸ 33/33 mixed What parents are asked to DO is the whole effect: tutoring a skill d≈1.15, listening to reading 0.51, reading aloud ~0.18 — same parent, same child. All deflate under better designs. 36/48
▸ 21/21 mixed Trust the composite; distrust the breakdown. Subtest strengths-and-weaknesses replicate at chance on retest, and most of what people read off a score report is noise. 24/48
▸ 4/4 no effect No. The two randomised programmes that tested it — 42 Houston schools and five EEF trials across 400 English schools — return nulls on reading and maths; only writing, the outcome nearest the treatment, moves. 9/48
▸ 21/21 no effect Coding instruction teaches coding (d≈0.68) and not thinking: against active controls far transfer is g=0.15, two validated-instrument RCTs found d≈0.00, and a 46-school maths trial found −0.16 to −0.21. 18/48
▸ 24/24 no effect School exercise programs don't raise achievement: four of the five largest RCTs are null, and more dose doesn't rescue it. Exercise for health, not for grades. 29/48
▸ 14/14 no effect Narrow reasoning routines are teachable and stay taught for seven months. The far-transfer claim failed its one randomised test: science −0.01, English −0.15, maths −0.11. 12/48
▸ 15/15 insufficient No trial has ever varied the sequence of science content and measured the result. The one quasi-experimental test of course order found nothing, and depth-over-breadth is a retrospective survey worth 0.08-0.13 SD. 15/48
▸ 4/4 insufficient Never operationalised well enough to test: no MI-based trial meets even a loose evidential bar. Four decades in, a valid evaluation is still impossible. 12/48
▸ 12/12 negative Teaching grammar does not improve writing at any age tested — and it displaces the things that do. 27/48
▸ 11/11 debunked Brain Gym, Fast ForWord, coloured overlays, the Mozart effect: tested and failed, or never evidence-based at all. Mozart works exactly as well as any other music. 18/48
▸ 23/23 insufficient No study of homeschooling has a counterfactual — the famous percentile claims are selection all the way down. The honest answer: nobody knows. 36/48
▸ 6/6 insufficient Nobody has tested it. There is no controlled trial in which art instruction is the input and independently measured art skill is the outcome — the field's core claim is the one it never studied. 18/48
▸ 23/23 insufficient Homeschoolers get plenty of social contact but measurably fewer, closer peer ties; whether that helps or harms has never been causally tested. 24/48