The School Builder's Playbook
If you were building a school for ages 4-18 and cared only about what causally works — judged under the premise that ability is substantially heritable, so only selection-resistant evidence counts — this is what the database says to do, ranked by evidence strength times cost. Twenty-six verified decisions, 165 graded sources, every verdict adversarially reviewed.
How to read this
Every claim below traces to a verified topic page (linked) with graded sources behind it. The grading constitution (METHODOLOGY) privileges RCTs, lotteries, and natural experiments; flags genetic confounding in observational work; discounts researcher-designed outcome measures ~2×; and tracks fadeout separately from end-of-treatment effects. Three database-wide regularities frame everything:
- Every famous number deflates. Hattie's 0.79 feedback, Black & Wiliam's 0.7 formative assessment, Bloom's 2-sigma tutoring, mastery learning's 1.0 — all shrink 3-20× on independent measures at realistic scale. The honest big effects in education are ~0.2-0.4 SD.
- Nothing moves g. Effective instruction moves domain skills, attainment, and life outcomes — sometimes dramatically — while general ability stays put (with one small Swedish exception noted below). Good instruction raises the heritability of skills by removing environmental bottlenecks: it is floor-raising, not gap-erasure.
- Test-score fadeout is the default; attainment persistence is the prize. Early gains revert toward genetically-anchored trajectories — but class size, tutoring, and double-dose algebra all show credential and life-outcome effects that outlive their score effects. Build for the long game and judge levers on the right ledger.
The ten commitments
1. Teach everything explicitly, and never use minimal guidance. The single most consistent finding across domains: unassisted discovery reliably loses (d≈−0.38); the constructivist curriculum lost the only multi-curriculum RCT by 0.22 SD; explicit instruction for strugglers is the most-replicated finding in intervention research. Guided inquiry and structured problem-first work are fine — minimal guidance is the poison. → explicit vs inquiry · explicit vs reform math · worked examples
2. Build retrieval and spacing into the curriculum's bones. The two strong-support methods that apply to every subject: low-stakes quizzing with feedback (durable-retention effects across 222 classroom studies) and distributed practice (the most robust finding in learning science). Students won't self-administer either — schedule them: cumulative reviews, spiral sequencing, cumulative exams, spaced homework. → retrieval practice · spaced practice · interleaving
3. Teach reading as code-then-knowledge. Explicit systematic phonics K-1 (the settled part), fluency through volume, and then the under-rated lever: a knowledge-rich cumulative curriculum as the actual comprehension program — not comprehension-strategy units (a one-time trick, ~null at scale). The Core Knowledge lottery (ITT 0.24 on state reading) is the best curriculum bet in the database. → teaching children to read · background knowledge
4. Teach math as explicit modular strands. Fluency (timed, low-stakes — the anti-timed-test movement has no causal evidence), word-problem schemas, and concepts↔procedures in alternation. Math sub-skills don't create each other; teach each on purpose. Place students in algebra by measured readiness; accelerate the ready; double-dose the rest (+3.3pp BA attainment, twelve years out). → teaching children mathematics · algebra timing
5. Hire for demonstrated performance; fire the credentials budget. Teacher quality is the largest within-school lever (1 SD ≈ 0.10-0.15 SD/yr — more than a 10-student class-size cut) and invisible on paper: certification null, master's null. Open the funnel, judge on 2-3 years of measured performance, gate retention on it. Replace workshop PD (~$18k/teacher/yr, causally dead) with content-specific coaching and evaluation-with-feedback — the two development levers that survive. → teacher quality
6. Tutor the students who fall behind — the expensive boring way. Education's most reliable intervention at its honest size (~0.2-0.3 SD; never Bloom's 2.0): groups of 2-4, paid consistent tutors, inside the school day, 3-5×/week, 30+ delivered hours, $1,400-4,800/student. Never opt-in, after-school, once-weekly, or rotating tutors. Assume decay; re-dose across years. → tutoring · Reading Recovery cautionary tale
7. Group by measured readiness, flexibly — and accelerate. Within-class and cross-grade grouping with genuinely differentiated instruction are ~free wins; the one tracking RCT raised the whole distribution via teach-to-level. Keep placements permeable and re-measured (locking 10-year-olds into tracks is where causal harm lives). Screen everyone objectively (universal screening tripled disadvantaged gifted placement); accelerate on mastery criteria (no harm + a saved year); run achievement-selected advanced classrooms (+0.5 SD for high-achieving minority students). Skip the gifted-label apparatus. → grouping & tracking · acceleration & gifted
8. Give feedback as task information; never as person evaluation. Feedback is the only method with a ~1/3 backfire rate. What works: specific task information, comments without grades on formative work (the grade cancels the comment), and slashed marking volume (no causal evidence for the hours). What backfires: praise, ego comparisons, discouragement. Treat "formative assessment" as a cheap ~0.1 bet, not a 0.7 transformation. → feedback & formative assessment
9. Choose curricula once, by exclusion list — then buy implementation. The whole cost is attention (~$36/student): exclude minimal-guidance programs (worth ~0.1-0.2 SD/yr of avoided harm), pick structured knowledge-rich materials, and spend on training and actual use — the field's default (3 days per career) guarantees the null. Don't chase rankings among mainstream texts (~0-0.05 SD at stake). → curriculum choice · math curricula
10. Buy the bundle, not the inputs. Five practices explain ~45% of school effectiveness — high-dosage tutoring, extended time, data-driven instruction, frequent teacher feedback, high expectations — while class size, spending level, and credentials explain ~zero. The bundle transplants into ordinary schools at ordinary budgets ($355-1,837/student, IRR 13-29%). Keep classes ~22-25 with excellent teachers; consider small classes only in K-1 and never by lowering the hiring bar. Match intensity to intake: the No-Excuses evidence is grade-A for disadvantaged populations and absent-to-negative for advantaged ones. → spending & school effects · class size
The refuse list
Money and hours the evidence says to withhold: whole-language/balanced-literacy and three-cueing · minimal-guidance discovery curricula in any subject · comprehension-strategy programs as the comprehension plan · standalone speech-only phonemic-awareness curricula · sustained-silent-reading as an intervention · dropping timed practice on anxiety grounds · concepts-first sequencing dogma · rich toy-like manipulatives and unguided hands-on time · algebra-timing mandates in either direction · strict mastery gates that tax fast learners (keep the diagnostic core, drop the gate) · master's-degree pay bumps · workshop PD · written-marking volume and grades on formative work · person praise · gifted-label gatekeeping by subjective referral · detracking on equity grounds · opt-in/after-school/once-weekly tutoring · commercial interim-assessment data platforms · Reading Recovery at its price · chasing textbook rankings and WWC curriculum labels · class-size reduction that dilutes teacher quality — and Bloom's 2-sigma, Hattie's 0.79, and Black & Wiliam's 0.7 as planning numbers.
The budget hierarchy
Per unit of learning, spend in roughly this order: (1) curriculum exclusion list + knowledge-
rich adoption (attention-priced) · (2) retrieval/spacing architecture (free, scheduling) ·
(3) grouping-by-readiness + acceleration + universal screening (free, saves money net) ·
(4) teacher selection & retention machinery (reallocates existing payroll) · (5) coaching +
evaluation-feedback ($3-7k/teacher, replaces dead PD spend) · (6) high-dosage tutoring for
whoever is behind ($1.4-4.8k/student served) · (7) extended time · (8) small K-1 classes
for disadvantaged intakes — the only defensible class-size buy · (9) unallocated spending
increases (the worst buy: ~$12.7k per 0.1 SD).
The honest limits
- Fadeout is undefeated. No intervention reliably prevents early score gains from reverting; the defensible responses are cumulative curriculum, maintained practice, re-dosed tutoring, and judging levers on attainment ledgers where they exist (STAR, double-dose, finance reforms).
- Reading comprehension is the hardest thing to move (math moves 2-5× more in every structural study) — which is exactly why the knowledge-curriculum bet and the early-code investments matter.
- The right tail is under-evidenced. Most causal evidence concerns struggling and average students; for profoundly gifted students the honest offer is acceleration (no harm + saved years) and advanced content, on thinner data.
- Implementation is the graveyard. The 2004 double-dose cohort, the Nashville tutoring program, the national tutoring schemes, the comparison-methods scale-up — every one died on delivered-dosage and fidelity, not on theory. Whatever you adopt, measure what was actually delivered.
- One Swedish exception to "nothing moves g": a class-size RD moved IQ-type measures durably (+0.02 SD/pupil, for low-income sons). Small, unreplicated, honestly recorded.