Methodology: How Evidence Is Graded
This document is the constitution of the database. Every verdict, grade, and confidence
rating in db/ obeys these rules. When a rule and a famous finding conflict, the rule wins.
The premise and what it implies
We operate from the premise that intelligence and academic ability are substantially heritable (twin/adoption/GWAS literature consistently puts heritability of cognitive ability at ~0.5–0.8 by adolescence, and heritability of school achievement around 0.6). This premise does specific work:
- Observational education research is presumptively confounded. Correlations between inputs (parenting practices, home literacy environment, school quality proxies, teacher credentials) and child outcomes are contaminated by selection and by passive gene–environment correlation: the same parents who supply books supply alleles. A correlation is not evidence of a teachable lever until a causal design says so.
- Claims of durably raising general intelligence (g) are near-certainly wrong. The record of interventions claiming lasting IQ gains is a graveyard (fadeout is the norm). Verdicts about "raising ability" require extraordinary evidence.
- The premise does NOT imply instruction is useless. Heritability of individual differences is compatible with large mean effects of instruction: literally nobody reads without being taught. The right questions are: which methods most efficiently build domain skills and knowledge, for whom, and do gains persist?
- Interactions matter. Good instruction can increase heritability of a skill by removing environmental bottlenecks (everyone gets the input, remaining variance is genetic). This is evidence instruction worked, not that it failed.
Evidence hierarchy and quality grades
Every source in db/sources/ gets a grade. Meta-analyses are graded by what they
aggregate — a meta-analysis of confounded studies inherits the confounding.
| Grade | Designs |
|---|---|
| A | Preregistered or independently replicated RCTs; large multi-site RCTs; lottery-based natural experiments (oversubscription randomization) |
| B | Single well-powered RCT (adequately clustered); strong quasi-experiments with credible identification (regression discontinuity, difference-in-differences with clean shocks); within-family / discordant-twin / adoption designs |
| C | Meta-analyses of mixed-quality designs; small or unregistered RCTs (n < ~100 or cluster-randomization errors); longitudinal studies with strong covariate + prior-achievement controls |
| D | Cross-sectional correlational work; pre-post without control group; studies with researcher-developed outcome measures only; expert opinion; qualitative |
Two further judgments are recorded on every source, because the methodology applies them in prose constantly and a discount applied in prose cannot be recomputed:
independence— independent | developer-involved | developer-led | unclear. The efficacy-to-effectiveness decay curve — this archive's most repeated finding — is a claim about this variable: developer-led +0.21 becomes independently evaluated 0.10 becomes −0.02 at scale. A discount that large cannot live indesign_notesprose.replication— replicated | mixed | failed | unreplicated | not-applicable | unclear. Makes the Replication rule below mechanical; prose commentary stays inreplication_notes.unreplicatedasserts the literature contains no independent replication;unclearrecords that we have not checked — the two must never be conflated.
These fields are required, with unclear as the honest unknown. The applies_to
experience showed that optional blocks rot; requiredness lives in the schema and honesty
lives in the vocabulary — a field may say "the record does not say," but it may not be
silent.
Genetic-confound risk (topic-level field)
Every topic carries genetic_confound_risk, an assessment of whether the evidence base
for the verdict could be explained by selection/heredity rather than the intervention:
- low — verdict rests on randomized or lottery designs; genes can't differ between arms.
- medium — verdict rests on a mix of experiments and observational work, or experiments with differential attrition / selection into treatment.
- high — verdict rests mainly on correlational evidence where heritable parent/child
traits plausibly drive both the input and the outcome. Such topics cannot exceed
confidence: lowfor causal claims.
Survivorship and the Lindy prior (topic-level practice_record block)
Evidence can overturn practice — but long-term practice is itself evidence. A method that
parents and teachers kept choosing across generations is not epistemically equal to a
2015 invention with the same RCT count: every generation was a fresh opportunity to abandon
it. This section states how survival enters the grading, and how it does not. It guards
against two failure modes at once: worshiping age (survivorship bias — practices survive
for many reasons besides working) and ignoring it (treating the entire pre-RCT past as one
undifferentiated insufficient).
- Survival is a prior, never a source. Tradition never appears in an evidence table,
never raises
confidence, and never convertsinsufficientinto support. It changes the burden on the evidence, not the reading of the evidence. - Grade the filter, not the age. Survival counts only insofar as a functioning selection filter could have killed the practice. The test: who felt failure, how fast, and could they exit? Practices also survive for reasons unrelated to working: inertia and lock-in, state mandate, signaling, serving the institution's adults rather than its children, cheapness, selection of students. A practice kept alive by statute has no survival record, only a sponsor. Whole language is the canonical filter failure — near-universal spread for decades under a mandate/monopoly filter (ed schools, basal publishers), and wrong.
- The hereditarian discount, scoped. The premise weakens naive Lindy: if children of able parents learn under almost any method, then practices judged on cohort outcomes survive on their students' genes. Heritability voids outcome-selection filters (institutions judged on aggregate results with selected intakes); it does not void immediate-feedback filters (individual failure, visible fast, to someone with exit power). Tutoring survived under the tightest filter education has — a paying parent watching one child, free to fire the tutor. School grammar survived under a weak one — exams, mandates, inheritance from Latin teaching. Same age, different priors.
- Function-match. Survival evidence attaches to the function that explains the
survival, which the
rationalemust name. Physical education survived for fitness; a null on achievement transfer does not fight that prior. Afavorableprior is only legal where the topic's verdict concerns the survived-for function; otherwise the prior isneutraland normal rules apply. - Anti-revival. Date the form as continuously practiced, not the branding. Retrieval practice as a 2006 lab framing is new; recitation is ancient. A "classical education" package assembled in the 1980s is a novel product in old clothes, and gets a novel practice's prior.
- Scope. The block belongs on practice-shaped topics — methods, programmes, named curricula. Comparison topics attach the record to the incumbent arm. Structure parameters (class size, starting age) and phenomena topics (fadeout) take no block.
The record is a topic-frontmatter block, mirroring applies_to:
practice_record:
filter: tight # tight | loose | none — quality of the survival filter
prior: favorable # favorable | neutral | unfavorable — the resulting presumption
rationale: "..." # dating, spread, survived-for function, who felt failure, confounds
burden_met: [...] # for overturning verdicts: the source slugs satisfying the burden
favorable requires filter: tight: under the default weighting, only immediate-feedback
survival earns a prior. unfavorable covers novel inventions and practices abandoned
after a fair trial — abandonment is weak evidence against, subject to its own confounds
(cost, fashion, litigation). The block is required on any topic with an overturning
verdict (no-effect | negative | debunked), so the burden rule cannot be escaped by
omission; elsewhere it is optional, and absence means "prior not yet assessed."
The burden asymmetry is the operative rule:
- Overturning a
favorablepractice — issuingno-effect,negative, ordebunked— requires ≥ 2 independent grade-A/B causal sources, named inburden_met, testing the practice as practiced or in its best current form (a null on the steelman counts a fortiori; a null on a caricature counts for nothing), on the survived-for function. When the burden is unmet, the verdict still reads what the evidence says — verdicts never lie to express a prior — butconfidenceis capped atlowand the practical layer says why: evidence leans against, but not enough to overturn a practice whose survival is evidence. - Novel practices carry the burden the other way. The base rate for new
interventions is this archive's through-line: fadeout, effect deflation,
efficacy-to-effectiveness decay. One well-powered independent null suffices for
no-effect(confidence at mostmedium, per the single-source ceiling). An unreplicated developer-led positive caps the verdict atmixed(see Replication rule). - For
insufficienttopics the prior colors only the practical layer. An old practice with a tight filter and no causal evidence is "a reasonable default, unproven." A novel practice with no causal evidence is "unproven, and base rates say bet against."
Calibration. The archive's own record justifies the heuristic. Every topic in the
debunked layer is a post-1970 invention (learning styles, multiple intelligences, brain
training, Brain Gym). The one traditional practice this archive has overturned — grammar
teaching as a route to better writing — took three independently evaluated randomised
trials plus four meta-analyses to do it: exactly the mountain the burden rule codifies.
Meanwhile phonics, spaced practice, and worked examples are old practices the trial
literature keeps re-vindicating.
Effect-size discipline
- Measure type matters more than the number. Researcher-designed, treatment-aligned outcome measures inflate effects roughly 2× vs independent standardized measures (Cheung & Slavin 2016; de Boer et al.). Evidence tables must say which type each effect used. When both exist, the standardized measure is the headline.
- Baseline, horizon, and outcome class travel with every effect row. "d = 0.4" means
nothing without against what, when measured, and on what kind of outcome — so
every
effectsrow carries three required enums besidemeasure_type:baseline(business-as-usual | active-alternative | none | unclear) because d against nothing and d against the next-best use of the same hour are different quantities;horizon(end-of-treatment | under-1yr | 1-2yr | over-2yr | adulthood | not-applicable | unclear), which is the fadeout rule below made recordable; andoutcome_class, a slug from the outcome taxonomy below.unclearmeans the record does not say;none/not-applicablecover descriptive estimates (heritability, gradients) where control and follow-up have no meaning. - Naive meta-meta-analysis is rejected. Hattie-style rankings (the d = 0.40 "hinge") average across incommensurable ages, designs, outcome proximities, and study qualities; several Visible Learning inputs contain computational errors (see critiques by Wecker et al., Slavin, Bergeron). We never cite a Visible Learning average as evidence; we go to the underlying meta-analyses and grade them.
- Small-study and publication bias. Large effects from small studies are discounted; where funnel-plot or PET-PEESE corrections exist, report the corrected estimate.
- Fadeout is tracked separately. End-of-treatment effects and effects at ≥1 year are different fields. An intervention whose effect vanishes by the next grade is scored on the durable effect, with the immediate effect noted. Persistence on attainment (graduation, earnings) despite test-score fadeout is noted where documented.
- Benchmarks for interpretation. Against business-as-usual controls, d ≈ 0.10 on a standardized measure from a scalable cheap intervention is meaningful; d > 0.60 from a small study with researcher-designed measures is a red flag, not a triumph. A year of schooling itself moves standardized achievement roughly d ≈ 0.2–0.4 depending on age.
Outcome taxonomy
Verdicts must state which outcome class they concern. These are not interchangeable:
- Domain skills/knowledge (decoding, arithmetic fluency, vocabulary) — teachable.
- Near transfer (trained-task-adjacent gains) — common.
- Far transfer (music → math, chess → IQ, working-memory training → reasoning) — presumptively absent; the burden of proof is on the claim.
- General intelligence (g) — no pedagogical intervention has credibly and durably raised it. Note the distinction, which is load-bearing: schooling DOES durably raise IQ test scores (~1–5 points per year of education, persisting across the lifespan; Ritchie & Tucker-Drob 2018), but that effect is not mediated by g — it consists of direct effects on specific taught skills (Ritchie, Bates & Deary 2015). Write "no durable gains in g", never "nothing durably raises IQ". Related ceiling: a whole childhood of improved rearing (full adoption) buys ~4–5 IQ points (Kendler 2015), matching the residual shared-environment variance in adulthood (c² ≈ 0.10, Bouchard 2013 — small but NOT zero).
- Attainment/credentials (graduation, college, earnings) — can move even when test scores fade; track separately.
- Non-cognitive/behavioral (attendance, discipline, conscientiousness proxies).
Replication rule
A finding with a failed direct replication or a collapsed-at-scale record (e.g., growth
mindset large-scale RCTs, stereotype threat) is capped at verdict: mixed regardless of
the original effect size. A finding that has never been independently replicated is capped
at confidence: medium. For a novel practice the establishment bar is higher still: an
unreplicated developer-led positive caps the verdict at mixed, not merely
confidence (see Survivorship and the Lindy prior). Both halves of this rule are now
mechanical rather than prose: each source's replication and independence enums are
what the caps key on.
Verdict vocabulary (topic frontmatter)
| Verdict | Meaning |
|---|---|
strong-support |
Multiple grade-A/B sources agree; effect survives adversarial review |
moderate-support |
Causal evidence positive but thinner, smaller, or less consistent |
mixed |
Credible evidence on both sides, or effects that appear only in some conditions |
no-effect |
Well-powered causal evidence finds ~nothing |
negative |
Causal evidence of harm relative to alternatives |
debunked |
Core claim contradicted by replication failures or was never evidence-based |
insufficient |
Not enough causal evidence to say (often: only confounded observational work exists) |
Every topic also carries conclusion: — the verdict rendered as one falsifiable
sentence (≤220 chars), with certainty carried inside the sentence: a mixed verdict
gets an honestly hedged conclusion, and insufficient gets "nobody has causally
tested this," never silence. The title names the decision; the conclusion answers
it — identifier and claim are kept separate on purpose, and answer-facing surfaces
lead with the conclusion.
Confidence rules
highrequires ≥ 2 independent grade-A/B sources pointing the same direction AND a surviving adversarial-review pass.mediumis the ceiling for single-source verdicts and unreplicated findings.lowis mandatory whengenetic_confound_risk: high.lowis mandatory when an overturning verdict (no-effect|negative|debunked) on aprior: favorablepractice does not meet the Lindy burden (see Survivorship and the Lindy prior).
Defaults, not theorems (the fixed core and the tunable rim)
The rules above are of two kinds, and the distinction is architectural.
The fixed core is the honesty machinery, and it holds at every setting: tradition is never cited as a source; a collapsed verdict reads "insufficient causal evidence," never "refuted"; measure type, baseline, horizon, and outcome class travel with every effect; dropped studies stay visible; certainty travels with every estimate. These rules are what make every other setting trustworthy.
The tunable rim is the weights. Several rules in this document are declared defaults
of a recomputable function, not theorems: the evidence-grade threshold (the /confound
rigor bar already lets a reader drag it from grade D to grade A and watch verdicts
recompute), the genetic-confound weighting (stored per topic as genetic_confound_risk),
and the survival prior (this document's defaults: a burden of two grade-A/B sources, and
favorable only from tight filters). Published verdicts use the defaults. The judgments
underneath — source grades, independence and replication flags, effect rows tagged with
baseline, horizon and outcome class, filter assessments, burden lists — are stored as
data precisely so verdicts can be recomputed under different weights, by a reader or by
an advisor reasoning over the archive. Disagree with the weights, not the facts.
Verification workflow
- Topic drafted from scout + deep-read agent output (
status: surveyedordeep). - Adversarial pass: skeptic agents attempt to refute the verdict (failed replications,
critiques, confounds, effect-size errors). Only after surviving does a topic get
status: verified. - Source files get
verified: trueonly after their extracted numbers are spot-checked against the underlying paper/abstract.
Node model (the school-builder decision tree)
The database is organized around a founder's decisions (see TAXONOMY.md), in three node types:
- Subject (
db/subjects/,layer: subject) — a founder-facing "how to teach X" program that synthesizes many choices. Not a single verdict; carries abottom_lineand asubtopicslist. - Topic (
db/topics/) — one evaluable decision with a single verdict. Eitherrolls_up_to:a subject (its evidence layer, e.g.phonics→reading) orlayer:a standalone cross-cutting decision (method|structure|debunked|input). - Source (
db/sources/) — one study/meta, graded A–D.
Controlled vocabularies
- domain: see TAXONOMY.md slugs.
- layer: subject | method | structure | debunked | input
- verdict: strong-support | moderate-support | mixed | no-effect | negative | debunked | insufficient
- confidence: high | medium | low
- genetic_confound_risk: low | medium | high
- practice_record.filter: tight | loose | none
- practice_record.prior: favorable | neutral | unfavorable
- status: stub | surveyed | deep | verified
- source type: meta-analysis | rct | quasi-experiment | natural-experiment | longitudinal | twin-adoption | review | replication | critique
- quality_grade: A | B | C | D
- independence (source): independent | developer-involved | developer-led | unclear
- replication (source): replicated | mixed | failed | unreplicated | not-applicable | unclear —
unreplicatedasserts no replication exists;unclearrecords that we do not know - effects[].baseline: business-as-usual | active-alternative | none | unclear
- effects[].horizon: end-of-treatment | under-1yr | 1-2yr | over-2yr | adulthood | not-applicable | unclear
- outcome (effects[].outcome_class and applies_to.outcome): domain-skill | near-transfer | far-transfer | g | attainment | non-cognitive | behaviour | health —
outcome_classadditionally allowsunclearfor sources whose content was never readable;applies_to.outcomemust commit or be absent
Applicability (topic-level applies_to block)
Grading for truth is necessary but not sufficient. Evidence quality and applicability fail independently: a verdict can be well-established and still useless to the person asking ("works one-to-one, collapses in groups"), and a perfectly-fitted method can rest on junk evidence. A recommendation has to pass both tests, so they are kept in separate blocks and neither is allowed to stand in for the other.
The driving case is a real question — "I want to teach three kids between 5 and 7 to have better grammar". Group size, age band, target skill and who is delivering are all retrieval facets. None of them were recorded, so the question was unanswerable no matter how strong the underlying evidence was.
applies_to:
delivery: [small-group, one-to-one] # whole-class | small-group | one-to-one |
# independent | software | home | school-wide
who: [teacher, tutor] # teacher | tutor | parent | peer | software |
# self | clinician | administrator
dose: "3×/week, 30 min, 12+ weeks" # the amount that produced the effect
cost: medium # free | low | medium | high
outcome: [domain-skill] # the outcome CLASS moved (see Outcome taxonomy)
age_evidence: [5, 9] # ages the EVIDENCE covers
prerequisites: "..." # what must already be true
not_for: "..." # where it is known to fail
Rules:
doseis not optional in spirit. "d = 0.4" is not actionable; "d = 0.4 at 3×/week for 12 weeks with a trained tutor" is. An effect size divorced from the conditions that produced it invites exactly the over-application this project exists to correct.age_evidenceis the honest age field. Most topics carryages: [4, 18], which over-claims and makes age useless as a filter.age_evidencerecords the ages actually studied and must fall inside the declared range; the validator enforces it.not_forcarries as much weight as the positive case. A recorded boundary is what stops a bad recommendation, and boundaries are usually better evidenced than the general claim.- The block is optional and absence means "not yet assessed" — which reads as not recommendable, the honest default. It is not required retrospectively because forcing a guess would put fabricated applicability into a database whose entire value is that it does not guess.
outcomeuses the outcome taxonomy above. Collapsing domain skills, far transfer, g, and attainment into "improves learning" is the most common way education claims are oversold; the recommendation layer must not reintroduce it.