Explore the evidence
85 of 88 decisions · 1 collapse at your settings · clear everything
- math6
- reading6
- writing6
- early childhood5
- science5
- character4
- health4
- arts3
- assessment3
- debunked3
- foreign language3
- history civics3
- practice3
- grouping2
- instruction style2
- motor skills2
- music2
- practical2
- time2
- transfer2
- behavior1
- class size1
- computer science1
- curriculum1
- deliberate practice1
- edtech1
- feedback1
- homeschool1
- homework1
- mastery1
- physical development1
- retention1
- school structure1
- sport1
- talent1
- teachers1
- tutoring1
Early gains fade by default — halving every 12–18 months, faster when bigger. What persists is trajectory (placement, graduation), not ability.
strong supportconf: highgc: lowFadeout — why early gains disappear, and what actually persists · early-childhood · ages 4–18 · structure
Timed retrieval practice builds arithmetic automaticity — beating identical untimed tutoring head-to-head — and the anti-timed-test harm claims have no causal evidence.
strong supportconf: highgc: lowMath-fact fluency and timed practice (and the anti-timed-test claims) · math · ages 5–12
Testing yourself beats rereading — robust in real classrooms; honest durable size ~0.1–0.3 SD, biggest after delay, thinnest on transfer.
strong supportconf: highgc: lowRetrieval practice (the testing effect) · practice · ages 8–18 · method
Spacing beats massing at equal total time — the most robust finding in learning science. Space repetitions at ~10–20% of how long you need to remember.
strong supportconf: highgc: lowSpaced (distributed) practice · practice · ages 6–18 · method
Systematic phonics beats whole language for word reading — largest in K-1 and for at-risk readers; near-null for comprehension and past grade 3.
strong supportconf: highgc: lowSystematic/explicit phonics vs whole language and balanced literacy · reading · ages 4–9
Physical capacity is about as heritable as cognitive ability (~60%, no shared environment). Whether trainability is a stable trait remains unproven — the skeptics are winning.
strong supportconf: highgc: lowTalent and trainability — what is heritable, and what that does not license · talent · ages 4–18 · input
Teacher quality is the largest within-school lever — 1 SD of teacher ≈ 0.10–0.15 SD/yr, worth more than ten fewer students — and credentials predict none of it. Select; don't workshop.
strong supportconf: highgc: lowTeacher quality — selection over credentials and workshops · teachers · ages 4–18 · structure
High-dosage tutoring is education's most reliable lever: ~0.29 SD in trials, ~0.2 well-scaled — not Bloom's 2σ. Groups of 3–4 work; 1:1 is unnecessary.
strong supportconf: highgc: lowTutoring — the honest effect, the Bloom 2-sigma myth, and what survives scale · tutoring · ages 5–18 · method
Group within classes or across grades by subject: modest, nearly free wins. Whole-school streaming does nothing, and early between-school tracking harms the bottom.
moderate supportconf: highgc: lowAbility grouping and tracking — four practices, four verdicts · grouping · ages 5–18 · structure
Accelerate ready kids: they keep pace with older classmates, bank a year, and show no social-emotional harm at 50. The gifted label itself does nothing; the content does.
moderate supportconf: highgc: mediumAcceleration and gifted programs — the label vs the content · grouping · ages 5–18 · structure
Minimal-guidance math trailed every rival in the one multi-curriculum RCT; explicit instruction for strugglers is math's most replicated result.
moderate supportconf: highgc: lowExplicit vs reform/constructivist math instruction · math · ages 5–18
Word-problem solving is its own skill: computation fluency doesn't produce it. Teaching problem schemas explicitly does (~0.25–0.45 SD on trained content).
moderate supportconf: highgc: lowWord problems, schema instruction, and the conceptual-vs-procedural question · math · ages 6–14
Neither acceleration mandates nor delay mandates work — readiness-matched placement does. Push everyone and the median falls; hold everyone back and the top falls.
mixedconf: highgc: lowAlgebra timing — acceleration mandates, delay mandates, and readiness-based placement · math · ages 11–18
Grit is conscientiousness renamed and adds 0.4% to grade prediction; self-control genuinely predicts life outcomes but is 60% heritable with zero shared environment, and training moves ratings, not lives.
mixedconf: highgc: mediumAre self-control and grit teachable levers on achievement? · character · ages 4–18
Smaller classes buy real K-1 gains that fade on tests yet persist in attainment — at roughly triple tutoring's cost, and diluted to nothing when scaled fast.
mixedconf: highgc: lowClass-size reduction · class-size · ages 4–18 · structure
Curriculum is nearly free, so choosing beats not choosing — but the payoff is avoiding a demonstrated loser, not finding a magic winner. Content is the high-upside bet.
mixedconf: highgc: lowCurriculum choice as a school-level lever · curriculum · ages 4–14 · structure
The 10,000-hour rule is dead: practice explains ~14% of performance variance, is itself heritable, and identical twins 20,000 hours apart didn't differ. Necessary, insufficient.
mixedconf: highgc: mediumDeliberate practice and the 10,000-hour rule · deliberate-practice · ages 6–18 · method
Growth-mindset interventions change beliefs almost everywhere and change achievement almost nowhere: the two independent national-scale trials measured standardized tests and found exactly zero.
mixedconf: highgc: lowDoes teaching a growth mindset raise achievement? · character · ages 4–18
Feedback on the task helps modestly; feedback on the person backfires — a stable third of studies reverse. The famous 0.4–0.7 number has no computed source.
mixedconf: highgc: lowFeedback and formative assessment · feedback · ages 5–18 · method
The famous faded-feedback demo doesn't survive meta-analysis; mental practice is real at half the advertised size and never substitutes for physical practice.
mixedconf: highgc: lowFeedback frequency and mental practice in motor learning · motor-skills · ages 4–18
Mastery learning moves tests of what it taught (~0.25) and barely moves independent measures (~0.05) — and time-to-mastery gaps widen, converting ability differences into time.
mixedconf: highgc: lowMastery learning (teach → test → reteach to criterion → advance) · mastery · ages 6–18 · method
The textbook matters at the bottom, not the top: avoid the demonstrated losers; mainstream choices now differ by ~0.02 SD.
mixedconf: highgc: lowMath curriculum choice — how much does the textbook matter? · math · ages 5–14
Varied practice looks worse today and better at retention — in the lab. Applied settings shrink the edge to zero, and it reverses in under-18s.
mixedconf: highgc: lowPractice scheduling for motor skills — varied vs blocked, spaced vs massed · motor-skills · ages 4–18
At scale, pre-K doesn't durably raise test scores — Tennessee went negative — yet Boston shows real attainment gains beside a test-score zero. It buys trajectory, not ability.
mixedconf: highgc: lowPreschool at scale — Head Start, state pre-K, and what universal provision delivers · early-childhood · ages 4–5 · structure
Unallocated money is the worst buy in the database; the same dollars as No-Excuses charters buy +0.3–0.4 SD/yr — and the average charter is a precise null.
mixedconf: highgc: lowSchool spending, charters, and what school-level choices actually move outcomes · school-structure · ages 4–18 · structure
Starting school older mostly manufactures an age-at-test artifact: the IQ effect collapses to ~−0.07 once identified. Real residues: less hyperactivity, unchanged attainment.
mixedconf: highgc: lowSchool starting age, relative age, and academic redshirting · early-childhood · ages 4–7 · structure
Should a school buy a social and emotional learning programme? · character · ages 4–18
Most measured 'home environment' effects are parents' genes: three-quarters of parent-child transmission isn't rearing, and a whole better childhood buys ~4 IQ points.
mixedconf: highgc: lowThe early home environment — what parents can and cannot causally move · early-childhood · ages 4–10 · input
Executive function trains like a task, not like a capacity: gains are real, narrow, and gone at follow-up, EF curricula are null under independent trial, and the latent construct is ~100% heritable.
no effectconf: highgc: lowCan executive function be trained, and does it transfer to learning? · character · ages 4–18
Teaching movement teaches movement: the skills improve, but transfer to cognition, achievement, or lifelong activity is near zero in the best trials.
no effectconf: highgc: mediumDoes teaching movement transfer to cognition, achievement, or lifelong activity? · physical-development · ages 4–18
Brain training improves the trained task and nothing else: far transfer to intelligence or achievement is 0.001 against active controls.
no effectconf: highgc: lowDoes working-memory or brain training transfer to intelligence and achievement? · transfer · ages 4–18 · debunked
Matching instruction to 'learning styles' does nothing: real matching experiments return d=.04. The correlational literature IS the myth.
debunkedconf: highgc: lowDoes matching instruction to a child's "learning style" improve learning? · debunked · ages 4–18 · debunked
Later bells buy adolescents 40+ measured minutes of sleep — the best-identified positive effect in the archive; the achievement payoff is real but far smaller.
strong supportconf: mediumgc: lowDo later school start times increase adolescent sleep, and does that raise achievement? · health · ages 11–18 · input
Yes, on their own terms: lottery-assigned museum and theatre trips move blind-rated analysis of art, plot knowledge and tolerance by 0.08-0.18 SD weeks later, and a film of the same play moves nothing.
moderate supportconf: mediumgc: lowAre museum and theatre trips worth the day out of school? · arts · ages 8–18
Comprehension is knowledge — but vocabulary teaching moves standardized comprehension only d≈0.10. The big content-knowledge bet (Core Knowledge lottery, 0.24) is real and unreplicated.
moderate supportconf: mediumgc: lowBackground knowledge and vocabulary as drivers of reading comprehension · reading · ages 5–14
Selective CTE high schools raise male graduation 8-10pp and early-career earnings 17-35% on lottery and cutoff designs; test scores, degrees, and every outcome for girls are flat.
moderate supportconf: mediumgc: lowDo career and technical education tracks help students — and which students? · practical · ages 14–18
Immersion costs nothing in English and buys a little — a lottery puts English reading 0.13-0.22 SD ahead by grades 5 and 8 — but no lottery has ever measured how much of the second language students actually learn.
moderate supportconf: mediumgc: lowDoes dual-language immersion work, and is the gain in the second language or in English? · foreign-language · ages 4–18
Teaching handwriting works and transfers: freeing the hand frees composition, with gains still present at six months — even taught in groups of three.
moderate supportconf: mediumgc: mediumDoes handwriting and transcription instruction matter, and should young children type instead? · writing · ages 5–12
Sentence combining improves writing where grammar teaching fails — the same meta-analyses score them +0.50 and −0.32.
moderate supportconf: mediumgc: lowDoes sentence combining improve writing? · writing · ages 5–18
Strategy instruction is writing's best-supported method — at roughly a fifth of its advertised size once measures are independent (d≈0.8 → ~0.15).
moderate supportconf: mediumgc: lowDoes strategy instruction (SRSD, the writing process) improve writing? · writing · ages 5–18
Yes — sight-reading, performance and aural skill move about half a standard deviation under instruction — but the evidence is quasi-experimental and thinner than the transfer literature built on top of it.
moderate supportconf: mediumgc: mediumDoes teaching music produce musical skill? · music · ages 4–18
Teaching coding teaches coding: the one clean school RCT gives g=0.47-0.68. Which approach you pick barely matters, and almost every effect size rests on an instrument the developers built.
moderate supportconf: mediumgc: lowDoes teaching programming actually teach programming — and does the method matter? · computer-science · ages 5–18
Formal spelling instruction works (ES 0.54) and more formal beats less formal — explicit wins again, and it transfers beyond spelling itself.
moderate supportconf: mediumgc: lowDoes teaching spelling work, and what does it transfer to? · writing · ages 5–18
Unassisted discovery loses to explicit teaching; well-scaffolded guided discovery beats both. The operative variable is guidance, not ideology.
moderate supportconf: mediumgc: lowExplicit instruction vs discovery/inquiry — how much guidance? · instruction-style · ages 4–18 · method
Mixing problem types helps where confusion is the enemy — discriminating similar categories, mixed math practice — and is useless or worse for facts and prose.
moderate supportconf: mediumgc: lowInterleaving (mixing problem types vs blocking) · practice · ages 10–18 · method
Manipulatives help when bland, guided, and aged ~7–11 — a guided-representation effect, not 'hands-on learning.' Rich, toy-like objects hurt transfer.
moderate supportconf: mediumgc: lowManipulatives and concrete-representational-abstract (CRA) sequences · math · ages 5–14
Phonemic awareness transfers to reading only with letters attached: speech-only training peaks near 10 hours and moves reading d=0.19; with print, 0.66.
moderate supportconf: mediumgc: lowPhonemic awareness training (and whether it needs letters) · reading · ages 4–7
Diagnose what a child already knows: teachers cut 40–50% of curriculum for high-ability children and achievement ROSE. The payoff is skipping, not monitoring.
moderate supportconf: mediumgc: mediumPlacement and mastery diagnosis — deciding what to teach next from evidence of current skill · assessment · ages 5–18 · structure
Guided oral reading works, but the ingredient is volume, not repetition: at equal exposure, re-reading has no edge over wide reading.
moderate supportconf: mediumgc: lowReading fluency — guided repeated oral reading, and whether repetition matters · reading · ages 6–12
Content-rich history teaching raises history knowledge (g 0.19-0.46) but not standardized reading in under three years; generic 'historical thinking' has no adequately controlled positive result.
moderate supportconf: mediumgc: lowShould history be built on content knowledge or on generic historical-thinking skills? · history-civics · ages 5–18
Outdoor time prevents myopia from starting (not progressing); screening plus free glasses raises test scores in children who need them. Two claims, both real.
moderate supportconf: mediumgc: lowVision — outdoor time against myopia, and screening and correction for achievement · health · ages 4–18 · input
Novices learn faster studying solutions than solving problems — then the effect reverses with expertise. Use worked examples early; fade them.
moderate supportconf: mediumgc: lowWorked examples (studying solutions vs solving problems) · instruction-style · ages 10–18 · method
Correcting a real deficiency moves cognition (iron in anaemic children: 0.79 SD); supplementing already-fed children moves nothing (35 RCTs, 19,343 children).
mixedconf: mediumgc: lowBreakfast, school meals, and micronutrients — what feeding children actually buys · health · ages 4–18 · input
Capitalisation and punctuation move when directly taught and directly measured, and don't move when embedded in writing programmes. A teachable subject, not a writing lever.
mixedconf: mediumgc: mediumCan capitalisation, punctuation and usage be taught directly as their own subject? · writing · ages 5–18
A missed day costs little (~0.005 SD); interventions reliably buy days back cheaply, but nobody has shown the recovered days move achievement.
mixedconf: mediumgc: mediumChronic absenteeism — does raising attendance raise achievement? · time · ages 4–18 · input
Explicit form-focused teaching beats pure exposure, so the archive's negative verdict on L1 grammar does NOT transfer — but the effects are measured on the taught forms, immediately, mostly on adults.
mixedconf: mediumgc: lowComprehensible input or explicit grammar teaching — which builds a second language? · foreign-language · ages 4–18
Life-skills courses reliably teach the content and rarely change the conduct; behaviour moves only where instruction sits close in time to the decision, and driver education is worse than nothing.
mixedconf: mediumgc: lowDo life-skills courses — money, cooking, driving, health — change what students actually do? · practical · ages 6–18
Civics teaching buys civic knowledge (d = 0.49) that fades to the control mean in two years and moves neither attitudes nor validated turnout; what moves voting is school quality and noncognitive skill.
mixedconf: mediumgc: lowDoes civic education produce civic knowledge, and does civic knowledge produce civic behaviour? · history-civics · ages 5–18
The headline effects are a measurement artefact: drama gives d≈0.89 on researcher-made tests and d≈0.29 on standardized ones, and the two randomised trials with standardized outcomes are null.
mixedconf: mediumgc: mediumDoes classroom drama build verbal and literacy skills? · arts · ages 5–16
Document-based instruction reliably improves the sourcing and argument tasks it teaches (g ≈ 0.42) and, in the two best-identified trials, moves neither standardized reading nor history knowledge.
mixedconf: mediumgc: lowDoes document-based source work move history knowledge, reading comprehension, or neither? · history-civics · ages 8–18
Sort by what the software replaces: adaptive drill inside the school day buys +0.05-0.20 SD; a device, a connection or the teacher buys zero to negative; LLM tutors have no usable evidence at all.
mixedconf: mediumgc: lowDoes educational technology raise learning — CAI, adaptive software, devices, screens, and AI tutors? · edtech · ages 5–18 · structure
Not at school: the largest randomised trials find nothing on reading, maths or cognition. A small effect on laboratory executive-function tasks is contested, may be real at ~0.2 SD, and is end-of-treatment only.
mixedconf: mediumgc: mediumDoes learning music make children smarter or better at school? · music · ages 4–16
Early specialization predicts junior success; later starts plus other sports predict adult world-class. Selecting on junior results selects the wrong athletes.
mixedconf: mediumgc: mediumEarly sport specialization, talent selection, and the relative age effect · sport · ages 6–18
The received wisdom that retention is harmful was two measurement errors, not a finding. Fix them and the achievement effect is zero — but the grade at which you hold a child back decides everything.
mixedconf: mediumgc: lowGrade retention — holding a child back versus promoting, and test-based promotion policies · retention · ages 5–18 · structure
Near-zero in primary and almost entirely correlational in secondary. The causal base is three small trials; what the homework IS beats how much of it there is.
mixedconf: mediumgc: mediumHomework — effects by age and by dosage · homework · ages 5–18 · structure
Classroom order is causally worth a great deal — one disruptive peer costs classmates 3% of adult earnings — yet no branded behaviour programme reliably delivers it and exclusion makes things worse.
mixedconf: mediumgc: lowHow should a school run its classrooms and its discipline system? · behavior · ages 4–18 · structure
Guided inquiry is positive in 16 of 16 PISA regions and unguided inquiry negative in 18 of 20 — but the programmes built on that finding go to zero at scale: +0.22, then +0.01, then +0.02 on the same test.
mixedconf: mediumgc: lowInquiry-based science teaching vs explicit and textbook science teaching · science · ages 8–18
Attainment falls steadily with age of first exposure, but the sharp critical period is contested and an earlier classroom start buys nothing durable: start young for immersion, not for two lessons a week.
mixedconf: mediumgc: mediumIs there a critical period for learning a second language, and when should a child start? · foreign-language · ages 4–18
Two reviews screened ~12,000 records and found one RCT and six quasi-experiments. What evidence exists favours interest (+0.12) over attainment (+0.01), and simulations match real apparatus.
mixedconf: mediumgc: lowLaboratory and practical work in school science · science · ages 5–18
Taking time away hurts measurably; adding it back buys almost nothing. The return per hour is ~0.02-0.03 SD, concave, and near zero in disorderly classrooms.
mixedconf: mediumgc: lowLength of the school day and year, time-on-task, extended time, and summer · time · ages 4–18 · structure
Perry and Abecedarian: n=123 and n=111, IQ gains gone by adolescence, attainment effects real but multiplicity-fragile. Too thin to carry the policy built on them.
mixedconf: mediumgc: lowPerry Preschool and Abecedarian — what the famous studies actually establish · early-childhood · ages 4–5 · structure
Reading Recovery buys large short-term gains at high cost; the only US long-term estimate is negative. Early 1:1 tutoring works — this model's durability is unproven.
mixedconf: mediumgc: lowReading Recovery and Tier-2 1:1 early-literacy tutoring · reading · ages 5–8
Strategy instruction is a one-time boost, not a trainable skill: six sessions buy what fifty do, and modern at-scale trials are ~null.
mixedconf: mediumgc: lowReading-comprehension strategy instruction (reciprocal teaching, summarizing, questioning) · reading · ages 8–14
Conceptual-change teaching reliably moves the test and nothing shows it removes the intuition. Corrected for design and publication bias, a taught unit is g≈0.64 and a refutation text g≈0.28.
mixedconf: mediumgc: lowScience misconceptions and conceptual change — can naive intuitions be taught away? · science · ages 6–18
Coaching buys ~0.25 SD of score and has for 42 years — a fraction of what's advertised — and none of it is ability. The score moves; the construct doesn't.
mixedconf: mediumgc: mediumTest preparation — does coaching raise scores, and does a raised score mean raised ability? · assessment · ages 5–18 · structure
What parents are asked to DO is the whole effect: tutoring a skill d≈1.15, listening to reading 0.51, reading aloud ~0.18 — same parent, same child. All deflate under better designs.
mixedconf: mediumgc: mediumThe parent as instructor — what survives when the person teaching is the person who raised the child · homeschool · ages 4–18 · method
Trust the composite; distrust the breakdown. Subtest strengths-and-weaknesses replicate at chance on retest, and most of what people read off a score report is noise.
mixedconf: mediumgc: lowWhat a standardized achievement score does and does not license · assessment · ages 5–18 · structure
No. The two randomised programmes that tested it — 42 Houston schools and five EEF trials across 400 English schools — return nulls on reading and maths; only writing, the outcome nearest the treatment, moves.
no effectconf: mediumgc: lowDoes arts education raise academic achievement? · arts · ages 5–14
Coding instruction teaches coding (d≈0.68) and not thinking: against active controls far transfer is g=0.15, two validated-instrument RCTs found d≈0.00, and a 46-school maths trial found −0.16 to −0.21.
no effectconf: mediumgc: lowDoes learning to code improve general thinking? · transfer · ages 5–16
School exercise programs don't raise achievement: four of the five largest RCTs are null, and more dose doesn't rescue it. Exercise for health, not for grades.
no effectconf: mediumgc: lowPhysical activity as an input — dose, fitness, and school outcomes · health · ages 4–18 · input
Narrow reasoning routines are teachable and stay taught for seven months. The far-transfer claim failed its one randomised test: science −0.01, English −0.15, maths −0.11.
no effectconf: mediumgc: lowTeaching students to think scientifically — does it transfer? · science · ages 7–18
No trial has ever varied the sequence of science content and measured the result. The one quasi-experimental test of course order found nothing, and depth-over-breadth is a retrospective survey worth 0.08-0.13 SD.
insufficientconf: mediumgc: mediumScience content sequencing — coherence, prerequisites, and course order · science · ages 5–18
Never operationalised well enough to test: no MI-based trial meets even a loose evidential bar. Four decades in, a valid evaluation is still impossible.
insufficientconf: mediumgc: mediuminsufficient at your settingsShould Gardner's multiple intelligences be used as a basis for instruction? · debunked · ages 4–18 · debunked
Teaching grammar does not improve writing at any age tested — and it displaces the things that do.
negativeconf: mediumgc: lowDoes teaching grammar improve children's writing? · writing · ages 5–18
Brain Gym, Fast ForWord, coloured overlays, the Mozart effect: tested and failed, or never evidence-based at all. Mozart works exactly as well as any other music.
debunkedconf: mediumgc: lowBrain Gym, Whole Brain Teaching, and the smaller classroom brain fads · debunked · ages 4–18 · debunked
The URL is the state — any view you build here is a permalink you can hand to someone mid-argument.