The Evidence on Teaching

Subject

Teaching children science

Bottom line. Science is the subject where the strongest-sounding claims have the weakest designs behind them, and where nearly every programme deflates on contact with independent evaluation. Four things survive. Attach a teacher's conceptual explanation to every hands-on activity — guided inquiry is positive in 16 of 16 PISA regions and unguided investigation negative in 18 of 20. Name misconceptions out loud and refute them: free, worth ~0.3 SD, and it does not delete the intuition. Buy sustained teacher professional development rather than science kits (+0.36 against +0.02 on independent measures). Teach fewer topics for longer, and stop optimising the sequence — the one test of course order found nothing. Do practical work for enthusiasm and identity, the only outcome class where it survives a 205-school trial. And refuse the far-transfer promise outright: the one large randomised test of 'science teaches you to think' returned −0.15 in English and −0.11 in maths.

The program (what a school should actually do)

Reading is close to solved at the level of method, and maths has a clear modular structure. Science has neither. It has a large, confident, well-funded literature in which the effect size falls every time somebody independent looks at it — and a set of national curriculum documents that fixed the sequence for a generation with no trial behind them.

So the programme is short, and much of it is about what to stop paying for.

Everywhere, all ages — the one instructional rule that replicates.

  • Attach the teacher's conceptual explanation to the activity. This is the whole finding. When PISA's inquiry items are split into guided inquiry (activity plus teacher explanation) and independent inquiry (students design their own investigations), the sign flips: guided positive in 16 of 16 regions, independent negative in 18 of 20, across 151,721 students. The classroom literature says the same thing from inside — teacher-led inquiry beats student-led by ~0.40, and "emphasis on active thinking" predicts learning where "amount of inquiry" does not. (evidence: inquiry vs explicit)
  • Never use student-designed investigation as a primary method. It is the one practice in this subject that is negatively associated with achievement almost everywhere it has been measured — and with enjoyment too.

Ages 5–11 — buy training, not boxes, and build knowledge inside the literacy block.

  • Buy the professional development, not the science kits. In the only synthesis that requires a control group, four weeks' minimum duration, and measures independent of the treatment, kit-based inquiry programmes return +0.02 across seven studies; the same inquiry built on sustained teacher PD returns +0.36 across ten. (evidence: content sequencing)
  • Then check who delivers that training. The same English CPD programme was worth +0.22 when its authors trained the teachers and +0.01 across 205 schools when a trainer chain did — on the same test. Repairing the training model and trying again returned +0.02.
  • Run science as a reading-and-writing-heavy content block. Integrated literacy-and-content instruction moves standardized reading comprehension +0.25, with science knowledge as the by-product. Note the direction: science content helps reading more reliably than reading helps science. This is the same bet as curriculum choice's highest-upside recommendation and background knowledge's.

Ages 8–18 — teach the misconceptions on purpose.

  • Say the wrong idea out loud, label it wrong, then give the right one. A refutation text beats an ordinary text stating the same true things by ~0.28–0.41 SD, and it costs nothing. Aim it at confident wrong beliefs; low-confidence ones gain nothing. (evidence: misconceptions)
  • Budget hours for the hard ones: about +0.06 per additional hour of instruction from an intercept of 0.59. A serious misconception is a ten-to-twelve-hour teaching job, not a paragraph.
  • Do not believe you have deleted the intuition. It is still measurable in chemistry professors with fourteen years' teaching experience, who get 31% of conflict items wrong in their own subject. Plan maintenance retrieval, and expect resurfacing.

Ages 11–18 — sequence with a light hand.

  • Teach fewer topics for longer. Weakly evidenced (d ≈ 0.1, correlational) but unanimous in direction — and the one experimental curriculum that worked did exactly this, reducing teacher-reported coverage while raising attainment.
  • Stop arguing about course order. Physics-first has been tested once against biology-first, concurrently in the same district, on independent state end-of-course exams: "little or no impact."
  • Do not build a prerequisite chain between the sciences. Out-of-discipline high-school science does not predict college performance in biology, chemistry or physics — and per the authors' own reasoning, that null is the trustworthy half of their data. (evidence: content sequencing)

Practical work — keep it, and change what you claim for it.

  • Do it for enthusiasm, subject identity, and apparatus competence. Those are the outcomes with randomised support that survives independent scale-up: +0.12 interest and +0.09 self-efficacy in the same 205-school trial where attainment moved +0.01. (evidence: practical work)
  • Teach the concept first, then use the practical to anchor it. In 55 observed English science lessons, not one contained an explicit plan for linking the observation back to the idea.
  • Use the simulation where the phenomenon simulates. Three K-12 randomised trials from kindergarten to grade 8 find real apparatus and virtual materials exactly equivalent — on knowledge, transfer, and interest.

What to refuse to spend money or time on

  • Thinking-skills programmes that replace science lessons. The flagship one lost 0.11–0.15 SD in English and maths and gained nothing in science, across 53 randomised schools.
  • Any promise of far transfer — "science teaches students to think", "the gains show up at GCSE three years later", "raises general intellectual ability across the board." Every positive estimate of this kind in the literature is developer-led and non-randomised.
  • Science kits as the intervention. +0.02 across seven studies with independent measures.
  • Out-of-school science-centre labs bought on learning grounds. Under randomisation the ordinary classroom beat them on achievement; they won only on situational enjoyment.
  • Branded conceptual-change programmes. Which strategy you buy does not moderate the effect (Q=0.95, p=0.62), and bolting peer argumentation onto a refutation text adds nothing.
  • Reordering the science curriculum in expectation of a return.
  • Any science effect size above ~0.5 that does not say who trained the teachers and who administered the test.

The measurement problem, and the bigger one behind it

Every subject in this archive has a measure-type problem. Science's is close to METHODOLOGY's nominal — notably not the ~5× the writing tranche found. Five estimates, four of them computed inside a single analysis:

Same intervention, different measurement Aligned Independent Ratio
Activity-based science, outcome tests coded for bias (Bredderman) 0.54 0.24 2.25×
Integrated literacy + content, reading comprehension (Hwang) 0.54 0.25 2.2×
Critical thinking, content-specific vs generic (Abrami) 0.57 0.30 1.9×
Conceptual change, researcher-built vs pre-existing instrument (Pacacı) 1.17 0.88 1.33×
One inquiry programme, inside its own end-of-year test (Slavin) taught units 0.58 remaining items 0.09 ~6×

The last row is the one to remember: same children, same test, same day, and the entire programme effect lived in the items covering the topics the programme taught. Slavin's cross-field benchmark says it more brutally still — treatment-inherent measures average +0.45 in maths and +0.51 in reading against −0.03 and +0.06 on treatment-independent ones.

But in science, measure type is not the biggest deflator. Three others are larger, and a founder who corrects only for measurement will still be badly wrong:

  • Efficacy to effectiveness — up to 20×. Thinking, Doing, Talking Science: +0.22 with the developers training teachers and teachers administering the test; +0.01 across 205 schools under independent evaluators; +0.02 across 180 schools after the training model was repaired. The instrument never changed. EEF's own scale-up synthesis puts seven programmes' mean effect falling from 0.25 to 0.01.
  • Design and selection — ~1.9×, and up to 0.8 SD on its own. Randomised conceptual-change studies return 0.64 against quasi-experiments' 1.23. And cognitive acceleration's famous far-transfer result came from genuinely independent GCSE examinations — so the inflation there was carried entirely by developer-led, non-randomised value-added analysis of volunteer schools, not by measurement at all. That is a distinct channel and this archive's 2× rule does not catch it.
  • Publication bias — up to 4×. The only significant moderator in the refutation-text meta-analysis is publication type: journals 0.45, dissertations 0.11 and not significant.

One counter-example, so the correction is not applied mechanically: project-based science in Detroit scored higher on the state test (0.37–0.44, non-randomised) than the same programme family scored on researcher-built tests under randomisation (0.21–0.25). That is selection running the other way.

The hereditarian bottom line for a founder

Science achievement is a heritable trait like every other academic outcome, and the lens does three specific jobs here.

  1. It explains why the sequencing literature is so confident and so weak. Every positive claim about course order, depth versus breadth, and prerequisites comes from retrospective surveys of students who chose their own high-school courses and then chose to enrol in college science, with a professor's grade as the outcome and a quarter of the strongest students missing because they placed out with AP credit. A robust correlation between demanding preparation and later success is exactly what the premise predicts, with or without any causal effect of the sequence. The authors' own calibration makes the point: students' self-estimate of how much high-school physics helped them was about three times the measured association.
  2. Heredity is not needed to explain the instructional nulls, which is what makes them credible. The scale-up and cognitive-acceleration nulls rest on school-level randomisation with matched-pair stratification. Genes cannot differ systematically between arms. When an effect disappears under randomisation and the disappearance cannot be attributed to measurement, the honest reading is that the effect was not there.
  3. The one explicit g claim in the subject failed. Cognitive acceleration's developer escalated to "improves students' general intellectual ability across the board"; the independent 53-school randomised trial returned science −0.01, English −0.15, maths −0.11. METHODOLOGY requires extraordinary evidence for claims about g. What was offered was a value-added residual against a regression line drawn through seventeen other schools.

Nothing in this subject moves g. Everything that works is domain-skill, with non-cognitive (interest, self-efficacy) as the reliable secondary product.

Confidence

Decision Verdict Confidence
Guided inquiry (activity + teacher explanation) over unguided investigation mixed — mechanism supported, programmes deflate medium
Naming and refuting misconceptions explicitly mixed — moves the test, not the intuition medium
Practical work for science content learning (don't) — null at scale medium
Practical work for interest and subject identity mixed — positive but situational medium
Sequencing, course order, learning progressions insufficient medium
Teaching general scientific thinking as a transferable skill no-effect medium

Why this subject has five topics, not four

The taxonomy scoped science as "inquiry vs explicit, misconceptions, labs, content sequencing." The evidence carved one extra joint, and it is the most important one on the page: "learns the science content" and "learns to think scientifically" have opposite answers. Guided inquiry has a real (if inflated) domain-skill effect; the far-transfer claim has a well-powered independent null with negative point estimates. Their evidence bases barely overlap — none of the inquiry meta-analyses carries a far-transfer outcome — so merging them would let the domain-skill positive launder the far-transfer null. Hence a separate reasoning and transfer topic.

Open questions a founder should watch

  • The trial that would settle practical work has never been run at K-12: practical work against a content-matched, time-matched non-practical lesson, on an independent measure. Two systematic reviews screening ~12,000 records between them found one randomised trial and six quasi-experiments.
  • Neither has the trial that would settle sequencing: same content, same time, same teachers, two orders, independent outcome. One quasi-experimental district study is the entire causal literature on course order, and spiral-versus-mastery in science is completely untested — the debate is imported wholesale from mathematics.
  • Nothing in this subject measures durability past about five months, and the one delayed external measure available halved and lost significance.
  • The GCSE follow-up of the randomised cognitive-acceleration cohort was flagged as future work and never published. It is the single piece of evidence that could rehabilitate the far-transfer claim on its own terms.
  • Why did one full-year project-based programme survive? It is the only positive here on a fully external, externally scored, high-stakes exam (+0.19 on AP totals, +0.30 in environmental science, 74 schools, five districts). Either full-year PBL with older students is genuinely different, or it is the next result to collapse. Nobody has tried to replicate it.
  • No K-12 science study in this tranche reports the same construct on an aligned and an independent measure at the same time. The 2× figure above is assembled across studies plus one within-test decomposition; the within-study head-to-head has not been run.
  • The conceptual-change literature has never used a standardized science achievement test at all, and has essentially no data on children past two months.

Evidence topics

← Home

Teaching children science · The Evidence on Teaching