Teaching students to think scientifically — does it transfer?
Narrow reasoning routines are teachable and stay taught for seven months. The far-transfer claim failed its one randomised test: science −0.01, English −0.15, maths −0.11.
no effectconf: mediumgc: lowscience · ages 7–18
Separate the two claims. TEACHABLE ROUTINES: yes. One explicit-instruction session on the control-of-variables strategy raised correct use from 34% to 65% in 7-10 year-olds and held at seven months for grade 4 (79% vs 15%) — but never outside a CVS-shaped task. GENERAL SCIENTIFIC THINKING: no. Cognitive Acceleration through Science Education claimed GCSE gains of 0.67 in science, 0.72 in maths and 0.69 in English three years after a two-year programme, from developer-led non-randomised value-added analyses of volunteer schools. The independent 53-school EEF randomised trial found science g=-0.01, English -0.15 (CI -0.24 to -0.06) and maths -0.11. Reasoning-intervention meta-analyses report g=0.80 and g=0.893 on researcher-designed instruments; the one randomised independent test on a standardized instrument returns zero.
Teach specific reasoning moves — control of variables, what counts as evidence, how to structure an argument — explicitly and directly, and expect them to work on tasks shaped like the ones you taught. That is worth doing on its own terms. Do not buy a thinking-skills programme that replaces science lessons, and do not accept a promise of gains in other subjects years later: the one large randomised test of exactly that promise returned negative point estimates on independently marked English and maths tests.
Who this applies to
Verdict
This topic exists to keep two questions apart that science education routinely welds together: does school science teach the content, and does it teach students to think scientifically in a way that travels. The first is the inquiry topic. This is the second, and it is a far-transfer claim, which this archive treats as presumptively absent with the burden of proof on the claimant.
The claimant had a strong case and lost it under randomisation.
- Narrow routines are teachable, durable, and worth teaching. One session of explicit instruction on the control-of-variables strategy took 7–10 year-olds from 34% to 65% correct use, and grade-4 children were still at 79% versus 15% seven months later on tasks about plants, baking, aeroplanes and running. That is a real result and the archive should say so.
- Nothing has ever transferred beyond a task of the same shape. Not to a school subject, not to a
general reasoning test, not to an achievement measure. Chen & Klahr's own "remote transfer" is the
same strategy under a new cover story —
near-transferby this archive's rules, and tagged that way. - The one big far-transfer claim failed a well-powered independent randomised test, with the transfer outcomes pointing the wrong way.
Hence no-effect on transfer, medium confidence — the ceiling for a verdict resting on a single
decisive trial.
What the evidence shows
| Source | Design | Grade | Key effect |
|---|---|---|---|
| EEF Let's Think Secondary Science 2016 | cluster RCT, 53 schools / 8,000+ pupils, independent | B | Science g=−0.01; English −0.15 (CI −0.24 to −0.06); maths −0.11 — all standardized |
| Shayer & Adey 1993 | quasi-experiment, developer-led | D | GCSE science 0.67, maths 0.72, English 0.69 — standardized, three years post |
| Adey & Shayer 1990 | quasi-experiment, developer-led | D | Piagetian reasoning tasks 0.89; science achievement: null |
| Shayer 1999 | value-added residuals, developer-led, no control group | D | Piagetian 0.29–1.26; KS3/GCSE residuals against a regression line through 17 other schools |
| Chen & Klahr 1999 | lab RCT, 87 children aged 7–10 | C | CVS use 34% → 65%; 7-month test grade 4 79% vs 15%; grade 3 not significant |
| Klahr & Nigam 2004 | RCT, elementary science | B | Direct instruction → 77% mastery vs 23% for discovery; transfer tracks mastery, not path |
| Kuhn & Dean 2005 | small quasi-experiment, grade 6 | D | A one-sentence hint moves variable control |
| Engelmann 2016 | meta, ~30 interventions | C | g≈0.80 on researcher-designed instruments; no transfer outcome pooled |
| Antonio 2024 | meta, 20 studies / 1,349 pupils | D | g=0.893, I²=85%, asymmetric funnel, largest study g=3.20 |
| Abrami 2015 | meta, 341 effects | C | Generic CT 0.30 (standardized) vs content-specific 0.57; STEM no better than non-STEM |
| Higgins 2005 | meta, 29 studies | D | Overall 0.74; science 0.78, maths 0.89, reading 0.40 — mixed measures |
| Zohar & Nemet 2002 | quasi-experiment, developer-led | D | Genetics knowledge and argument quality both up, on author-built instruments |
| Osborne 2013 | quasi-experiment, 4 departments, 2 years | C | "Few significant changes" on standard instruments |
| Lederman 2002 | instrument paper | D | Nature-of-science understanding has no independent standardized measure, by design |
The cognitive-acceleration story, told straight
CASE is the largest far-transfer claim in science education and its evidence trail is worth reading in order, because the effect moves each time it is looked at:
- 1990, end of treatment. Piagetian reasoning tasks — the programme's own theory turned into a test — d = 0.89, in the boys-only 12+ cell. Science achievement: null.
- 1992. A science effect appears, one year after the programme ended.
- 1993. Far transfer to GCSE maths and English appears, three years after the programme ended — now in the girls-only 11+ cell. The authors report that the Piagetian gains did not mediate the achievement gains for a substantial subset.
- 1999. The result is scaled up, and the control group is replaced by a regression line drawn through 17 other schools.
- 2006. The developer's own summary escalates to "improves students' general intellectual ability across the board" — an explicit claim about g.
Then the independent test. EEF-funded, evaluated at York by a team including Robert Slavin, 53 schools randomised in matched pairs, 8,000+ pupils, intention-to-treat. The primary outcome was a Key Stage 3-based science test; the transfer outcomes were commercially published, nationally standardised GL Assessment English and maths tests, scored by GL. Science g = −0.01. English −0.15. Maths −0.11.
The most important methodological fact here is that the inflation is not carried by measure type. Adey & Shayer's headline far-transfer numbers came from GCSE — genuinely independent, nationally marked examinations. The independent trial used equally independent measures and got the opposite sign. So this is a design-and-selection collapse, not a measure-alignment one: a developer-led, non-randomised, value-added analysis of volunteer schools inflated a standardized outcome by roughly 0.8 SD on its own. That is a distinct inflation channel from the one METHODOLOGY's 2× rule covers, and it is bigger.
The developers' live defences are real and should be carried with the null: Let's Think Secondary Science cut CASE from 30 lessons to 19; fidelity was poor, with a quarter of teachers teaching six or fewer of the 19; the transfer test was pre-specified one-sided and each secondary outcome rested on only 17–20 schools; and the original claim was about GCSE three years later, which this trial never measured. The GCSE follow-up the evaluators flagged as future work has not been published. That is the single outstanding empirical question in this topic.
What Chen & Klahr actually reached
Because this is the positive result, it is worth being exact about its edges.
Reached: correct control-of-variables use inside the trained mechanics domain; new mechanics domains a week later (61–64% of trials); and a seven-month paper-and-pencil test across five non-mechanics cover stories — plants, baking, aeroplanes, drink sales, running — where grade-4 children who had been explicitly taught scored 79% good reasoners versus 15%.
Did not reach: grade 3 at seven months (40% vs 22%, not significant); grade 2 beyond very near
within-domain transfer; and anything that is not a CVS-shaped task, ever. No school subject, no
general reasoning test, no standardized instrument, no achievement measure. The authors call the
cover-story change "remote transfer"; by this archive's outcome taxonomy it is near-transfer,
because the structure of the task is identical and only its surface changed.
The companion finding matters as much: explicit instruction produced the strategy, and exploration or probe questions alone did not. This is the same guidance-dose result the rest of the subject keeps returning.
Has anything moved an independent standardized measure of something other than the trained task?
No. That sentence is the topic.
The only affirmative claim in the literature is Shayer & Adey's GCSE result, which is developer-led, non-randomised, flips subgroups between cohorts, has its own proposed mediator failing to mediate, and drops non-implementing schools from the analysis. Every well-powered independent randomised test of the mechanism returns zero or slightly negative — Let's Think Secondary Science inside science, and a 198-school, 7,677-pupil randomised trial of a sibling thinking-skills programme outside it (reading 0.01–0.02, maths 0.04–0.05), recorded here as off-scope because it is not science teaching.
Meanwhile the meta-analyses that are quoted in support run on researcher-designed instruments: 0.80 for scientific-reasoning interventions, 0.893 for inquiry-and-higher-order-thinking with an asymmetric funnel and a largest study at g=3.20. Abrami's much more careful critical-thinking meta gives the cleanest single confirmation of the archive's 2× rule found anywhere in this tranche — content-specific critical thinking 0.57 versus generic standardized critical thinking 0.30, a 1.9× gradient inside one analysis — and reports that STEM subjects are no better a vehicle than any other. Whatever generic thinking instruction achieves, science is not privileged in achieving it.
Hereditarian-lens assessment
Risk: low, and the asymmetry is instructive. The null rests on school-level randomisation with matched-pair stratification, where genotype cannot differ systematically between arms; the residual exposure is 26% missing outcome data, not selection on family background.
The positive claim rests on exactly the design this archive's premise predicts will inflate: volunteer schools, self-selected into a research programme, compared against a regression line through other schools, with non-implementing sites dropped. Where the far-transfer effect appears, it appears in a subgroup that changes between cohorts — boys at one point, girls at another — which is what multiple testing on selected samples produces.
Nothing here credibly bears on g, and the developer's own escalation to "general intellectual ability across the board" is precisely the claim METHODOLOGY requires extraordinary evidence for. What was offered was a value-added residual.
Boundaries & what critics say
- This is not an argument against teaching reasoning. Explicit instruction in control of variables works, holds for seven months, and is cheap. Teach it. Just do not sell it as general thinking.
- Poor implementation is a genuine alternative explanation for the null, and the archive should not pretend otherwise. A quarter of teachers taught six or fewer of nineteen lessons.
- The transfer outcomes in the decisive trial were underpowered and one-sided. Only positive transfer could formally be detected, and each transfer outcome rested on 17–20 schools. The evaluators explicitly decline to call the result negative transfer. Neither should this topic — the finding is a null with negative point estimates, not evidence of harm.
- Nature-of-science instruction has no measurable independent outcome at all, by design: the field measures it with open-ended questionnaires interpreted by the researchers. It is not graded here because there is nothing to grade.
- Argumentation shrinks the same way everything else does. A developer-led quasi-experiment on one genetics unit found gains on author-built instruments; four whole science departments over two years, measured on standard instruments, produced "few significant changes."
- Preece's published critique of the CASE evidence base could not be obtained and is recorded as unresolved rather than dropped.
Practical guidance
- Teach specific reasoning moves explicitly: control of variables, what counts as evidence, how an argument is structured. One session moves the strategy; explicit instruction beats probe questions and beats exploration.
- Expect near transfer and plan for it deliberately. Retrain the strategy in each new domain rather than assuming it will arrive there on its own — and note that the seven-month retention held for 9–10 year-olds and not for 8-year-olds.
- Do not replace science lessons with thinking lessons. The one trial that did it lost 0.11–0.15 SD in English and maths and gained nothing in science.
- Refuse delayed-transfer promises. "The gains show up at GCSE three years later" is the exact claim that failed. Ask for the randomised trial.
- Discount any scientific-reasoning effect size above ~0.5 on sight. The metas that produce them pool researcher-designed instruments with funnel asymmetry and no transfer outcome.
- If general critical thinking is the goal, science is not a privileged route to it — and the honest number is 0.30 on a standardized measure, not 0.80.
Open questions
- The GCSE follow-up of the randomised cohort has never been published. It is the one piece of evidence that could rehabilitate the far-transfer claim on its own terms, the evaluators flagged it as future work, and it does not exist.
- Would a faithful, full-dose CASE at 30 lessons with high fidelity return a different answer? The developers say yes. Nobody has run it under randomisation.
- How far does explicit CVS instruction reach if transfer is trained on purpose — across many domains, over years, rather than in one session? Every study here trains once and tests the edge.
- Is there any age at which broad reasoning instruction pays? Abrami finds middle school 0.37 and high school 0.25 on standardized measures, a difference that is not significant. If the effect is real anywhere it is small and young, and nobody has targeted it.
- grade BLet's Think Secondary Science: Evaluation report and executive summaryHanley, P., Böhnke, J. R., Slavin, B., Elliott, L., & Croudace, T. · 2016 · rct
- grade DAccelerating the development of formal thinking in middle and high school students IV: Three years after a two-year interventionShayer, M., & Adey, P. · 1993 · quasi-experiment
- grade DAccelerating the development of formal thinking in middle and high school studentsAdey, P., & Shayer, M. · 1990 · quasi-experiment
- grade DCognitive acceleration through science education II: its effects and scopeShayer, M. · 1999 · quasi-experiment
- grade CAll Other Things Being Equal: Acquisition and Transfer of the Control of Variables StrategyChen, Z., & Klahr, D. · 1999 · rct
- grade DIs Developing Scientific Thinking All About Learning to Control Variables?Kuhn, D., & Dean, D. · 2005 · quasi-experiment
- grade CFostering scientific reasoning in education – meta-analytic evidence from intervention studiesEngelmann, K., Neuhaus, B. J., & Fischer, F. · 2016 · meta-analysis
- grade DEffects of Inquiry-Based Approaches on Students' Higher-Order Thinking Skills in Science: A Meta-AnalysisAntonio, R. P., & Prudente, M. S. · 2024 · meta-analysis
- grade CStrategies for Teaching Students to Think CriticallyAbrami, P. C., Bernard, R. M., Borokhovski, E., Waddington, D. I., Wade, C. A., & Persson, T. · 2015 · meta-analysis
- grade DA meta-analysis of the impact of the implementation of thinking skills approaches on pupilsHiggins, S., Hall, E., Baumfield, V., & Moseley, D. · 2005 · meta-analysis
- grade DFostering students' knowledge and argumentation skills through dilemmas in human geneticsZohar, A., & Nemet, F. · 2002 · quasi-experiment
- grade CLearning to argue: A study of four schools and their attempt to develop the use of argumentation as a common instructional practice and its impact on studentsOsborne, J., Simon, S., Christodoulou, A., Howell-Richardson, C., & Richardson, K. · 2013 · quasi-experiment
- grade DViews of nature of science questionnaire: Toward valid and meaningful assessment of learners' conceptions of nature of scienceLederman, N. G., Abd-El-Khalick, F., Bell, R. L., & Schwartz, R. S. · 2002 · review
- grade BThe Equivalence of Learning Paths in Early Science Instruction: Direct Instruction and Discovery LearningKlahr, D., & Nigam, M. · 2004 · rct
Related decisions
- Inquiry-based science teaching vs explicit and textbook science teachingmixedconf: mediumgc: low