The Evidence on Teaching

Teaching students to think scientifically — does it transfer?

Narrow reasoning routines are teachable and stay taught for seven months. The far-transfer claim failed its one randomised test: science −0.01, English −0.15, maths −0.11.

no effectconf: mediumgc: low

science · ages 718

Effect summary

Separate the two claims. TEACHABLE ROUTINES: yes. One explicit-instruction session on the control-of-variables strategy raised correct use from 34% to 65% in 7-10 year-olds and held at seven months for grade 4 (79% vs 15%) — but never outside a CVS-shaped task. GENERAL SCIENTIFIC THINKING: no. Cognitive Acceleration through Science Education claimed GCSE gains of 0.67 in science, 0.72 in maths and 0.69 in English three years after a two-year programme, from developer-led non-randomised value-added analyses of volunteer schools. The independent 53-school EEF randomised trial found science g=-0.01, English -0.15 (CI -0.24 to -0.06) and maths -0.11. Reasoning-intervention meta-analyses report g=0.80 and g=0.893 on researcher-designed instruments; the one randomised independent test on a standardized instrument returns zero.

Practical takeaway

Teach specific reasoning moves — control of variables, what counts as evidence, how to structure an argument — explicitly and directly, and expect them to work on tasks shaped like the ones you taught. That is worth doing on its own terms. Do not buy a thinking-skills programme that replaces science lessons, and do not accept a promise of gains in other subjects years later: the one large randomised test of exactly that promise returned negative point estimates on independently marked English and maths tests.

Who this applies to

Group size
whole-classsmall-groupone-to-one
Delivered by
teacher
Ages studied
716(narrower than the 718 this topic is filed under — outside it is extrapolation)
Dose
Control of variables: a single session of explicit instruction with probe questions raised correct strategy use from 34% to 65% and held seven months later in 9-10 year-olds. Cognitive acceleration: 19-30 one-hour lessons over two years, replacing normal science lessons, plus one training day and three support sessions per year — and produced nothing on independent measures.
Cost
low
Moves
domain-skillnear-transfer
Needs first
Explicit instruction in the strategy itself. Probe questions alone do nothing and exploration alone does nothing — the randomised comparison is decisive on this. Younger children need more: the seven-month retention held for grade 4 and not for grade 3.
Not for
Attainment in any other subject, exam results generally, critical thinking as a disposition, or general ability. Displacing science curriculum time for thinking lessons cost 0.11-0.15 SD in English and maths in the one trial that measured it on independent standardized tests.

Verdict

This topic exists to keep two questions apart that science education routinely welds together: does school science teach the content, and does it teach students to think scientifically in a way that travels. The first is the inquiry topic. This is the second, and it is a far-transfer claim, which this archive treats as presumptively absent with the burden of proof on the claimant.

The claimant had a strong case and lost it under randomisation.

  • Narrow routines are teachable, durable, and worth teaching. One session of explicit instruction on the control-of-variables strategy took 7–10 year-olds from 34% to 65% correct use, and grade-4 children were still at 79% versus 15% seven months later on tasks about plants, baking, aeroplanes and running. That is a real result and the archive should say so.
  • Nothing has ever transferred beyond a task of the same shape. Not to a school subject, not to a general reasoning test, not to an achievement measure. Chen & Klahr's own "remote transfer" is the same strategy under a new cover story — near-transfer by this archive's rules, and tagged that way.
  • The one big far-transfer claim failed a well-powered independent randomised test, with the transfer outcomes pointing the wrong way.

Hence no-effect on transfer, medium confidence — the ceiling for a verdict resting on a single decisive trial.

What the evidence shows

Source Design Grade Key effect
EEF Let's Think Secondary Science 2016 cluster RCT, 53 schools / 8,000+ pupils, independent B Science g=−0.01; English −0.15 (CI −0.24 to −0.06); maths −0.11 — all standardized
Shayer & Adey 1993 quasi-experiment, developer-led D GCSE science 0.67, maths 0.72, English 0.69 — standardized, three years post
Adey & Shayer 1990 quasi-experiment, developer-led D Piagetian reasoning tasks 0.89; science achievement: null
Shayer 1999 value-added residuals, developer-led, no control group D Piagetian 0.29–1.26; KS3/GCSE residuals against a regression line through 17 other schools
Chen & Klahr 1999 lab RCT, 87 children aged 7–10 C CVS use 34% → 65%; 7-month test grade 4 79% vs 15%; grade 3 not significant
Klahr & Nigam 2004 RCT, elementary science B Direct instruction → 77% mastery vs 23% for discovery; transfer tracks mastery, not path
Kuhn & Dean 2005 small quasi-experiment, grade 6 D A one-sentence hint moves variable control
Engelmann 2016 meta, ~30 interventions C g≈0.80 on researcher-designed instruments; no transfer outcome pooled
Antonio 2024 meta, 20 studies / 1,349 pupils D g=0.893, I²=85%, asymmetric funnel, largest study g=3.20
Abrami 2015 meta, 341 effects C Generic CT 0.30 (standardized) vs content-specific 0.57; STEM no better than non-STEM
Higgins 2005 meta, 29 studies D Overall 0.74; science 0.78, maths 0.89, reading 0.40 — mixed measures
Zohar & Nemet 2002 quasi-experiment, developer-led D Genetics knowledge and argument quality both up, on author-built instruments
Osborne 2013 quasi-experiment, 4 departments, 2 years C "Few significant changes" on standard instruments
Lederman 2002 instrument paper D Nature-of-science understanding has no independent standardized measure, by design

The cognitive-acceleration story, told straight

CASE is the largest far-transfer claim in science education and its evidence trail is worth reading in order, because the effect moves each time it is looked at:

  1. 1990, end of treatment. Piagetian reasoning tasks — the programme's own theory turned into a test — d = 0.89, in the boys-only 12+ cell. Science achievement: null.
  2. 1992. A science effect appears, one year after the programme ended.
  3. 1993. Far transfer to GCSE maths and English appears, three years after the programme ended — now in the girls-only 11+ cell. The authors report that the Piagetian gains did not mediate the achievement gains for a substantial subset.
  4. 1999. The result is scaled up, and the control group is replaced by a regression line drawn through 17 other schools.
  5. 2006. The developer's own summary escalates to "improves students' general intellectual ability across the board" — an explicit claim about g.

Then the independent test. EEF-funded, evaluated at York by a team including Robert Slavin, 53 schools randomised in matched pairs, 8,000+ pupils, intention-to-treat. The primary outcome was a Key Stage 3-based science test; the transfer outcomes were commercially published, nationally standardised GL Assessment English and maths tests, scored by GL. Science g = −0.01. English −0.15. Maths −0.11.

The most important methodological fact here is that the inflation is not carried by measure type. Adey & Shayer's headline far-transfer numbers came from GCSE — genuinely independent, nationally marked examinations. The independent trial used equally independent measures and got the opposite sign. So this is a design-and-selection collapse, not a measure-alignment one: a developer-led, non-randomised, value-added analysis of volunteer schools inflated a standardized outcome by roughly 0.8 SD on its own. That is a distinct inflation channel from the one METHODOLOGY's 2× rule covers, and it is bigger.

The developers' live defences are real and should be carried with the null: Let's Think Secondary Science cut CASE from 30 lessons to 19; fidelity was poor, with a quarter of teachers teaching six or fewer of the 19; the transfer test was pre-specified one-sided and each secondary outcome rested on only 17–20 schools; and the original claim was about GCSE three years later, which this trial never measured. The GCSE follow-up the evaluators flagged as future work has not been published. That is the single outstanding empirical question in this topic.

What Chen & Klahr actually reached

Because this is the positive result, it is worth being exact about its edges.

Reached: correct control-of-variables use inside the trained mechanics domain; new mechanics domains a week later (61–64% of trials); and a seven-month paper-and-pencil test across five non-mechanics cover stories — plants, baking, aeroplanes, drink sales, running — where grade-4 children who had been explicitly taught scored 79% good reasoners versus 15%.

Did not reach: grade 3 at seven months (40% vs 22%, not significant); grade 2 beyond very near within-domain transfer; and anything that is not a CVS-shaped task, ever. No school subject, no general reasoning test, no standardized instrument, no achievement measure. The authors call the cover-story change "remote transfer"; by this archive's outcome taxonomy it is near-transfer, because the structure of the task is identical and only its surface changed.

The companion finding matters as much: explicit instruction produced the strategy, and exploration or probe questions alone did not. This is the same guidance-dose result the rest of the subject keeps returning.

Has anything moved an independent standardized measure of something other than the trained task?

No. That sentence is the topic.

The only affirmative claim in the literature is Shayer & Adey's GCSE result, which is developer-led, non-randomised, flips subgroups between cohorts, has its own proposed mediator failing to mediate, and drops non-implementing schools from the analysis. Every well-powered independent randomised test of the mechanism returns zero or slightly negative — Let's Think Secondary Science inside science, and a 198-school, 7,677-pupil randomised trial of a sibling thinking-skills programme outside it (reading 0.01–0.02, maths 0.04–0.05), recorded here as off-scope because it is not science teaching.

Meanwhile the meta-analyses that are quoted in support run on researcher-designed instruments: 0.80 for scientific-reasoning interventions, 0.893 for inquiry-and-higher-order-thinking with an asymmetric funnel and a largest study at g=3.20. Abrami's much more careful critical-thinking meta gives the cleanest single confirmation of the archive's 2× rule found anywhere in this tranche — content-specific critical thinking 0.57 versus generic standardized critical thinking 0.30, a 1.9× gradient inside one analysis — and reports that STEM subjects are no better a vehicle than any other. Whatever generic thinking instruction achieves, science is not privileged in achieving it.

Hereditarian-lens assessment

Risk: low, and the asymmetry is instructive. The null rests on school-level randomisation with matched-pair stratification, where genotype cannot differ systematically between arms; the residual exposure is 26% missing outcome data, not selection on family background.

The positive claim rests on exactly the design this archive's premise predicts will inflate: volunteer schools, self-selected into a research programme, compared against a regression line through other schools, with non-implementing sites dropped. Where the far-transfer effect appears, it appears in a subgroup that changes between cohorts — boys at one point, girls at another — which is what multiple testing on selected samples produces.

Nothing here credibly bears on g, and the developer's own escalation to "general intellectual ability across the board" is precisely the claim METHODOLOGY requires extraordinary evidence for. What was offered was a value-added residual.

Boundaries & what critics say

  • This is not an argument against teaching reasoning. Explicit instruction in control of variables works, holds for seven months, and is cheap. Teach it. Just do not sell it as general thinking.
  • Poor implementation is a genuine alternative explanation for the null, and the archive should not pretend otherwise. A quarter of teachers taught six or fewer of nineteen lessons.
  • The transfer outcomes in the decisive trial were underpowered and one-sided. Only positive transfer could formally be detected, and each transfer outcome rested on 17–20 schools. The evaluators explicitly decline to call the result negative transfer. Neither should this topic — the finding is a null with negative point estimates, not evidence of harm.
  • Nature-of-science instruction has no measurable independent outcome at all, by design: the field measures it with open-ended questionnaires interpreted by the researchers. It is not graded here because there is nothing to grade.
  • Argumentation shrinks the same way everything else does. A developer-led quasi-experiment on one genetics unit found gains on author-built instruments; four whole science departments over two years, measured on standard instruments, produced "few significant changes."
  • Preece's published critique of the CASE evidence base could not be obtained and is recorded as unresolved rather than dropped.

Practical guidance

  • Teach specific reasoning moves explicitly: control of variables, what counts as evidence, how an argument is structured. One session moves the strategy; explicit instruction beats probe questions and beats exploration.
  • Expect near transfer and plan for it deliberately. Retrain the strategy in each new domain rather than assuming it will arrive there on its own — and note that the seven-month retention held for 9–10 year-olds and not for 8-year-olds.
  • Do not replace science lessons with thinking lessons. The one trial that did it lost 0.11–0.15 SD in English and maths and gained nothing in science.
  • Refuse delayed-transfer promises. "The gains show up at GCSE three years later" is the exact claim that failed. Ask for the randomised trial.
  • Discount any scientific-reasoning effect size above ~0.5 on sight. The metas that produce them pool researcher-designed instruments with funnel asymmetry and no transfer outcome.
  • If general critical thinking is the goal, science is not a privileged route to it — and the honest number is 0.30 on a standardized measure, not 0.80.

Open questions

  • The GCSE follow-up of the randomised cohort has never been published. It is the one piece of evidence that could rehabilitate the far-transfer claim on its own terms, the evaluators flagged it as future work, and it does not exist.
  • Would a faithful, full-dose CASE at 30 lessons with high fidelity return a different answer? The developers say yes. Nobody has run it under randomisation.
  • How far does explicit CVS instruction reach if transfer is trained on purpose — across many domains, over years, rather than in one session? Every study here trains once and tests the edge.
  • Is there any age at which broad reasoning instruction pays? Abrami finds middle school 0.37 and high school 0.25 on standardized measures, a difference that is not significant. If the effect is real anywhere it is small and young, and nobody has targeted it.

Evidence (14 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore