Explore the evidence
11 of 88 decisions · clear everything
Testing yourself beats rereading — robust in real classrooms; honest durable size ~0.1–0.3 SD, biggest after delay, thinnest on transfer.
strong supportconf: highgc: lowRetrieval practice (the testing effect) · practice · ages 8–18 · method
Spacing beats massing at equal total time — the most robust finding in learning science. Space repetitions at ~10–20% of how long you need to remember.
strong supportconf: highgc: lowSpaced (distributed) practice · practice · ages 6–18 · method
Teacher quality is the largest within-school lever — 1 SD of teacher ≈ 0.10–0.15 SD/yr, worth more than ten fewer students — and credentials predict none of it. Select; don't workshop.
strong supportconf: highgc: lowTeacher quality — selection over credentials and workshops · teachers · ages 4–18 · structure
Group within classes or across grades by subject: modest, nearly free wins. Whole-school streaming does nothing, and early between-school tracking harms the bottom.
moderate supportconf: highgc: lowAbility grouping and tracking — four practices, four verdicts · grouping · ages 5–18 · structure
Accelerate ready kids: they keep pace with older classmates, bank a year, and show no social-emotional harm at 50. The gifted label itself does nothing; the content does.
moderate supportconf: highgc: mediumAcceleration and gifted programs — the label vs the content · grouping · ages 5–18 · structure
Smaller classes buy real K-1 gains that fade on tests yet persist in attainment — at roughly triple tutoring's cost, and diluted to nothing when scaled fast.
mixedconf: highgc: lowClass-size reduction · class-size · ages 4–18 · structure
Curriculum is nearly free, so choosing beats not choosing — but the payoff is avoiding a demonstrated loser, not finding a magic winner. Content is the high-upside bet.
mixedconf: highgc: lowCurriculum choice as a school-level lever · curriculum · ages 4–14 · structure
Feedback on the task helps modestly; feedback on the person backfires — a stable third of measured effects reverse. The famous 0.4–0.7 number has no computed source.
mixedconf: highgc: lowFeedback and formative assessment · feedback · ages 5–18 · method
Mastery learning moves tests of what it taught (~0.25) and barely moves independent measures (~0.05) — and time-to-mastery gaps widen, converting ability differences into time.
mixedconf: highgc: lowMastery learning (teach → test → reteach to criterion → advance) · mastery · ages 6–18 · method
At scale, pre-K doesn't durably raise test scores — Tennessee went negative — yet Boston shows real attainment gains beside a test-score zero. It buys trajectory, not ability.
mixedconf: highgc: lowPreschool at scale — Head Start, state pre-K, and what universal provision delivers · early-childhood · ages 4–5 · structure
Starting school older mostly manufactures an age-at-test artifact: the IQ effect collapses to ~−0.07 once identified. Real residues: less hyperactivity, unchanged attainment.
mixedconf: highgc: lowSchool starting age, relative age, and academic redshirting · early-childhood · ages 4–7 · structure
The URL is the state — any view you build here is a permalink you can hand to someone mid-argument.