The Evidence on Teaching

Methodology: How Evidence Is Graded

This document is the constitution of the database. Every verdict, grade, and confidence rating in db/ obeys these rules. When a rule and a famous finding conflict, the rule wins.

The premise and what it implies

We operate from the premise that intelligence and academic ability are substantially heritable (twin/adoption/GWAS literature consistently puts heritability of cognitive ability at ~0.5–0.8 by adolescence, and heritability of school achievement around 0.6). This premise does specific work:

  1. Observational education research is presumptively confounded. Correlations between inputs (parenting practices, home literacy environment, school quality proxies, teacher credentials) and child outcomes are contaminated by selection and by passive gene–environment correlation: the same parents who supply books supply alleles. A correlation is not evidence of a teachable lever until a causal design says so.
  2. Claims of durably raising general intelligence (g) are near-certainly wrong. The record of interventions claiming lasting IQ gains is a graveyard (fadeout is the norm). Verdicts about "raising ability" require extraordinary evidence.
  3. The premise does NOT imply instruction is useless. Heritability of individual differences is compatible with large mean effects of instruction: literally nobody reads without being taught. The right questions are: which methods most efficiently build domain skills and knowledge, for whom, and do gains persist?
  4. Interactions matter. Good instruction can increase heritability of a skill by removing environmental bottlenecks (everyone gets the input, remaining variance is genetic). This is evidence instruction worked, not that it failed.

Evidence hierarchy and quality grades

Every source in db/sources/ gets a grade. Meta-analyses are graded by what they aggregate — a meta-analysis of confounded studies inherits the confounding.

Grade Designs
A Preregistered or independently replicated RCTs; large multi-site RCTs; lottery-based natural experiments (oversubscription randomization)
B Single well-powered RCT (adequately clustered); strong quasi-experiments with credible identification (regression discontinuity, difference-in-differences with clean shocks); within-family / discordant-twin / adoption designs
C Meta-analyses of mixed-quality designs; small or unregistered RCTs (n < ~100 or cluster-randomization errors); longitudinal studies with strong covariate + prior-achievement controls
D Cross-sectional correlational work; pre-post without control group; studies with researcher-developed outcome measures only; expert opinion; qualitative

Two further judgments are recorded on every source, because the methodology applies them in prose constantly and a discount applied in prose cannot be recomputed:

  • independence — independent | developer-involved | developer-led | unclear. The efficacy-to-effectiveness decay curve — this archive's most repeated finding — is a claim about this variable: developer-led +0.21 becomes independently evaluated 0.10 becomes −0.02 at scale. A discount that large cannot live in design_notes prose.
  • replication — replicated | mixed | failed | unreplicated | not-applicable | unclear. Makes the Replication rule below mechanical; prose commentary stays in replication_notes. unreplicated asserts the literature contains no independent replication; unclear records that we have not checked — the two must never be conflated.

These fields are required, with unclear as the honest unknown. The applies_to experience showed that optional blocks rot; requiredness lives in the schema and honesty lives in the vocabulary — a field may say "the record does not say," but it may not be silent.

Genetic-confound risk (topic-level field)

Every topic carries genetic_confound_risk, an assessment of whether the evidence base for the verdict could be explained by selection/heredity rather than the intervention:

  • low — verdict rests on randomized or lottery designs; genes can't differ between arms.
  • medium — verdict rests on a mix of experiments and observational work, or experiments with differential attrition / selection into treatment.
  • high — verdict rests mainly on correlational evidence where heritable parent/child traits plausibly drive both the input and the outcome. Such topics cannot exceed confidence: low for causal claims.

Survivorship and the Lindy prior (topic-level practice_record block)

Evidence can overturn practice — but long-term practice is itself evidence. A method that parents and teachers kept choosing across generations is not epistemically equal to a 2015 invention with the same RCT count: every generation was a fresh opportunity to abandon it. This section states how survival enters the grading, and how it does not. It guards against two failure modes at once: worshiping age (survivorship bias — practices survive for many reasons besides working) and ignoring it (treating the entire pre-RCT past as one undifferentiated insufficient).

  • Survival is a prior, never a source. Tradition never appears in an evidence table, never raises confidence, and never converts insufficient into support. It changes the burden on the evidence, not the reading of the evidence.
  • Grade the filter, not the age. Survival counts only insofar as a functioning selection filter could have killed the practice. The test: who felt failure, how fast, and could they exit? Practices also survive for reasons unrelated to working: inertia and lock-in, state mandate, signaling, serving the institution's adults rather than its children, cheapness, selection of students. A practice kept alive by statute has no survival record, only a sponsor. Whole language is the canonical filter failure — near-universal spread for decades under a mandate/monopoly filter (ed schools, basal publishers), and wrong.
  • The hereditarian discount, scoped. The premise weakens naive Lindy: if children of able parents learn under almost any method, then practices judged on cohort outcomes survive on their students' genes. Heritability voids outcome-selection filters (institutions judged on aggregate results with selected intakes); it does not void immediate-feedback filters (individual failure, visible fast, to someone with exit power). Tutoring survived under the tightest filter education has — a paying parent watching one child, free to fire the tutor. School grammar survived under a weak one — exams, mandates, inheritance from Latin teaching. Same age, different priors.
  • Function-match. Survival evidence attaches to the function that explains the survival, which the rationale must name. Physical education survived for fitness; a null on achievement transfer does not fight that prior. A favorable prior is only legal where the topic's verdict concerns the survived-for function; otherwise the prior is neutral and normal rules apply.
  • Anti-revival. Date the form as continuously practiced, not the branding. Retrieval practice as a 2006 lab framing is new; recitation is ancient. A "classical education" package assembled in the 1980s is a novel product in old clothes, and gets a novel practice's prior.
  • Scope. The block belongs on practice-shaped topics — methods, programmes, named curricula. Comparison topics attach the record to the incumbent arm. Structure parameters (class size, starting age) and phenomena topics (fadeout) take no block.

The record is a topic-frontmatter block, mirroring applies_to:

practice_record:
  filter: tight        # tight | loose | none — quality of the survival filter
  prior: favorable     # favorable | neutral | unfavorable — the resulting presumption
  rationale: "..."     # dating, spread, survived-for function, who felt failure, confounds
  burden_met: [...]    # for overturning verdicts: the source slugs satisfying the burden

favorable requires filter: tight: under the default weighting, only immediate-feedback survival earns a prior. unfavorable covers novel inventions and practices abandoned after a fair trial — abandonment is weak evidence against, subject to its own confounds (cost, fashion, litigation). The block is required on any topic with an overturning verdict (no-effect | negative | debunked), so the burden rule cannot be escaped by omission; elsewhere it is optional, and absence means "prior not yet assessed."

The burden asymmetry is the operative rule:

  1. Overturning a favorable practice — issuing no-effect, negative, or debunked — requires ≥ 2 independent grade-A/B causal sources, named in burden_met, testing the practice as practiced or in its best current form (a null on the steelman counts a fortiori; a null on a caricature counts for nothing), on the survived-for function. When the burden is unmet, the verdict still reads what the evidence says — verdicts never lie to express a prior — but confidence is capped at low and the practical layer says why: evidence leans against, but not enough to overturn a practice whose survival is evidence.
  2. Novel practices carry the burden the other way. The base rate for new interventions is this archive's through-line: fadeout, effect deflation, efficacy-to-effectiveness decay. One well-powered independent null suffices for no-effect (confidence at most medium, per the single-source ceiling). An unreplicated developer-led positive caps the verdict at mixed (see Replication rule).
  3. For insufficient topics the prior colors only the practical layer. An old practice with a tight filter and no causal evidence is "a reasonable default, unproven." A novel practice with no causal evidence is "unproven, and base rates say bet against."

Calibration. The archive's own record justifies the heuristic. Every topic in the debunked layer is a post-1970 invention (learning styles, multiple intelligences, brain training, Brain Gym). The one traditional practice this archive has overturned — grammar teaching as a route to better writing — took three independently evaluated randomised trials plus four meta-analyses to do it: exactly the mountain the burden rule codifies. Meanwhile phonics, spaced practice, and worked examples are old practices the trial literature keeps re-vindicating.

Effect-size discipline

  • Measure type matters more than the number. Researcher-designed, treatment-aligned outcome measures inflate effects roughly 2× vs independent standardized measures (Cheung & Slavin 2016; de Boer et al.). Evidence tables must say which type each effect used. When both exist, the standardized measure is the headline.
  • Baseline, horizon, and outcome class travel with every effect row. "d = 0.4" means nothing without against what, when measured, and on what kind of outcome — so every effects row carries three required enums beside measure_type: baseline (business-as-usual | active-alternative | none | unclear) because d against nothing and d against the next-best use of the same hour are different quantities; horizon (end-of-treatment | under-1yr | 1-2yr | over-2yr | adulthood | not-applicable | unclear), which is the fadeout rule below made recordable; and outcome_class, a slug from the outcome taxonomy below. unclear means the record does not say; none/not-applicable cover descriptive estimates (heritability, gradients) where control and follow-up have no meaning.
  • Naive meta-meta-analysis is rejected. Hattie-style rankings (the d = 0.40 "hinge") average across incommensurable ages, designs, outcome proximities, and study qualities; several Visible Learning inputs contain computational errors (see critiques by Wecker et al., Slavin, Bergeron). We never cite a Visible Learning average as evidence; we go to the underlying meta-analyses and grade them.
  • Small-study and publication bias. Large effects from small studies are discounted; where funnel-plot or PET-PEESE corrections exist, report the corrected estimate.
  • Fadeout is tracked separately. End-of-treatment effects and effects at ≥1 year are different fields. An intervention whose effect vanishes by the next grade is scored on the durable effect, with the immediate effect noted. Persistence on attainment (graduation, earnings) despite test-score fadeout is noted where documented.
  • Benchmarks for interpretation. Against business-as-usual controls, d ≈ 0.10 on a standardized measure from a scalable cheap intervention is meaningful; d > 0.60 from a small study with researcher-designed measures is a red flag, not a triumph. A year of schooling itself moves standardized achievement roughly d ≈ 0.2–0.4 depending on age.

Outcome taxonomy

Verdicts must state which outcome class they concern. These are not interchangeable:

  1. Domain skills/knowledge (decoding, arithmetic fluency, vocabulary) — teachable.
  2. Near transfer (trained-task-adjacent gains) — common.
  3. Far transfer (music → math, chess → IQ, working-memory training → reasoning) — presumptively absent; the burden of proof is on the claim.
  4. General intelligence (g) — no pedagogical intervention has credibly and durably raised it. Note the distinction, which is load-bearing: schooling DOES durably raise IQ test scores (~1–5 points per year of education, persisting across the lifespan; Ritchie & Tucker-Drob 2018), but that effect is not mediated by g — it consists of direct effects on specific taught skills (Ritchie, Bates & Deary 2015). Write "no durable gains in g", never "nothing durably raises IQ". Related ceiling: a whole childhood of improved rearing (full adoption) buys ~4–5 IQ points (Kendler 2015), matching the residual shared-environment variance in adulthood (c² ≈ 0.10, Bouchard 2013 — small but NOT zero).
  5. Attainment/credentials (graduation, college, earnings) — can move even when test scores fade; track separately.
  6. Non-cognitive/behavioral (attendance, discipline, conscientiousness proxies).

Replication rule

A finding with a failed direct replication or a collapsed-at-scale record (e.g., growth mindset large-scale RCTs, stereotype threat) is capped at verdict: mixed regardless of the original effect size. A finding that has never been independently replicated is capped at confidence: medium. For a novel practice the establishment bar is higher still: an unreplicated developer-led positive caps the verdict at mixed, not merely confidence (see Survivorship and the Lindy prior). Both halves of this rule are now mechanical rather than prose: each source's replication and independence enums are what the caps key on.

Verdict vocabulary (topic frontmatter)

Verdict Meaning
strong-support Multiple grade-A/B sources agree; effect survives adversarial review
moderate-support Causal evidence positive but thinner, smaller, or less consistent
mixed Credible evidence on both sides, or effects that appear only in some conditions
no-effect Well-powered causal evidence finds ~nothing
negative Causal evidence of harm relative to alternatives
debunked Core claim contradicted by replication failures or was never evidence-based
insufficient Not enough causal evidence to say (often: only confounded observational work exists)

Every topic also carries conclusion: — the verdict rendered as one falsifiable sentence (≤220 chars), with certainty carried inside the sentence: a mixed verdict gets an honestly hedged conclusion, and insufficient gets "nobody has causally tested this," never silence. The title names the decision; the conclusion answers it — identifier and claim are kept separate on purpose, and answer-facing surfaces lead with the conclusion.

Confidence rules

  • high requires ≥ 2 independent grade-A/B sources pointing the same direction AND a surviving adversarial-review pass.
  • medium is the ceiling for single-source verdicts and unreplicated findings.
  • low is mandatory when genetic_confound_risk: high.
  • low is mandatory when an overturning verdict (no-effect | negative | debunked) on a prior: favorable practice does not meet the Lindy burden (see Survivorship and the Lindy prior).

Defaults, not theorems (the fixed core and the tunable rim)

The rules above are of two kinds, and the distinction is architectural.

The fixed core is the honesty machinery, and it holds at every setting: tradition is never cited as a source; a collapsed verdict reads "insufficient causal evidence," never "refuted"; measure type, baseline, horizon, and outcome class travel with every effect; dropped studies stay visible; certainty travels with every estimate. These rules are what make every other setting trustworthy.

The tunable rim is the weights. Several rules in this document are declared defaults of a recomputable function, not theorems: the evidence-grade threshold (the /confound rigor bar already lets a reader drag it from grade D to grade A and watch verdicts recompute), the genetic-confound weighting (stored per topic as genetic_confound_risk), and the survival prior (this document's defaults: a burden of two grade-A/B sources, and favorable only from tight filters). Published verdicts use the defaults. The judgments underneath — source grades, independence and replication flags, effect rows tagged with baseline, horizon and outcome class, filter assessments, burden lists — are stored as data precisely so verdicts can be recomputed under different weights, by a reader or by an advisor reasoning over the archive. Disagree with the weights, not the facts.

Verification workflow

  1. Topic drafted from scout + deep-read agent output (status: surveyed or deep).
  2. Adversarial pass: skeptic agents attempt to refute the verdict (failed replications, critiques, confounds, effect-size errors). Only after surviving does a topic get status: verified.
  3. Source files get verified: true only after their extracted numbers are spot-checked against the underlying paper/abstract.

Node model (the school-builder decision tree)

The database is organized around a founder's decisions (see TAXONOMY.md), in three node types:

  • Subject (db/subjects/, layer: subject) — a founder-facing "how to teach X" program that synthesizes many choices. Not a single verdict; carries a bottom_line and a subtopics list.
  • Topic (db/topics/) — one evaluable decision with a single verdict. Either rolls_up_to: a subject (its evidence layer, e.g. phonicsreading) or layer: a standalone cross-cutting decision (method | structure | debunked | input).
  • Source (db/sources/) — one study/meta, graded A–D.

Controlled vocabularies

  • domain: see TAXONOMY.md slugs.
  • layer: subject | method | structure | debunked | input
  • verdict: strong-support | moderate-support | mixed | no-effect | negative | debunked | insufficient
  • confidence: high | medium | low
  • genetic_confound_risk: low | medium | high
  • practice_record.filter: tight | loose | none
  • practice_record.prior: favorable | neutral | unfavorable
  • status: stub | surveyed | deep | verified
  • source type: meta-analysis | rct | quasi-experiment | natural-experiment | longitudinal | twin-adoption | review | replication | critique
  • quality_grade: A | B | C | D
  • independence (source): independent | developer-involved | developer-led | unclear
  • replication (source): replicated | mixed | failed | unreplicated | not-applicable | unclear — unreplicated asserts no replication exists; unclear records that we do not know
  • effects[].baseline: business-as-usual | active-alternative | none | unclear
  • effects[].horizon: end-of-treatment | under-1yr | 1-2yr | over-2yr | adulthood | not-applicable | unclear
  • outcome (effects[].outcome_class and applies_to.outcome): domain-skill | near-transfer | far-transfer | g | attainment | non-cognitive | behaviour | health — outcome_class additionally allows unclear for sources whose content was never readable; applies_to.outcome must commit or be absent

Applicability (topic-level applies_to block)

Grading for truth is necessary but not sufficient. Evidence quality and applicability fail independently: a verdict can be well-established and still useless to the person asking ("works one-to-one, collapses in groups"), and a perfectly-fitted method can rest on junk evidence. A recommendation has to pass both tests, so they are kept in separate blocks and neither is allowed to stand in for the other.

The driving case is a real question — "I want to teach three kids between 5 and 7 to have better grammar". Group size, age band, target skill and who is delivering are all retrieval facets. None of them were recorded, so the question was unanswerable no matter how strong the underlying evidence was.

applies_to:
  delivery: [small-group, one-to-one]   # whole-class | small-group | one-to-one |
                                        # independent | software | home | school-wide
  who: [teacher, tutor]                 # teacher | tutor | parent | peer | software |
                                        # self | clinician | administrator
  dose: "3×/week, 30 min, 12+ weeks"    # the amount that produced the effect
  cost: medium                          # free | low | medium | high
  outcome: [domain-skill]               # the outcome CLASS moved (see Outcome taxonomy)
  age_evidence: [5, 9]                  # ages the EVIDENCE covers
  prerequisites: "..."                  # what must already be true
  not_for: "..."                        # where it is known to fail

Rules:

  1. dose is not optional in spirit. "d = 0.4" is not actionable; "d = 0.4 at 3×/week for 12 weeks with a trained tutor" is. An effect size divorced from the conditions that produced it invites exactly the over-application this project exists to correct.
  2. age_evidence is the honest age field. Most topics carry ages: [4, 18], which over-claims and makes age useless as a filter. age_evidence records the ages actually studied and must fall inside the declared range; the validator enforces it.
  3. not_for carries as much weight as the positive case. A recorded boundary is what stops a bad recommendation, and boundaries are usually better evidenced than the general claim.
  4. The block is optional and absence means "not yet assessed" — which reads as not recommendable, the honest default. It is not required retrospectively because forcing a guess would put fabricated applicability into a database whose entire value is that it does not guess.
  5. outcome uses the outcome taxonomy above. Collapsing domain skills, far transfer, g, and attainment into "improves learning" is the most common way education claims are oversold; the recommendation layer must not reintroduce it.