The Evidence on Teaching

Inquiry-based science teaching vs explicit and textbook science teaching

Guided inquiry is positive in 16 of 16 PISA regions and unguided inquiry negative in 18 of 20 — but the programmes built on that finding go to zero at scale: +0.22, then +0.01, then +0.02 on the same test.

mixedconf: mediumgc: low

science · ages 818

Effect summary

Science replicates the archive's guidance-dose rule and then adds a harder lesson. On the guidance side the science-specific evidence is unusually clean: splitting PISA's inquiry items into GUIDED (inquiry plus teacher conceptual explanation) and INDEPENDENT (students design their own investigations) flips the sign — guided positive in 16/16 regions (Finland +0.51, Singapore +0.50), independent negative in 18/20 (Finland -0.63, Singapore -0.55). Pooled 'inquiry saturation' predicts nothing (Minner, 101 studies); what predicts is emphasis on active thinking. On the magnitude side the news is bad: the flagship English CPD programme returned +0.22 with developers delivering the training and teachers administering the test, then +0.01 (205 schools) and +0.02 (180 schools) under independent evaluation ON THE SAME INSTRUMENT. The surviving positives are scripted, heavily-scaffolded, PD-intensive curricula measured on aligned tests (0.21-0.32) and one full-year PBL programme measured on College Board AP exams (+0.19 overall, +0.30 in environmental science).

Practical takeaway

Buy guided inquiry, not inquiry. The operative ingredient is teacher conceptual explanation attached to the practical activity; without it the same activities predict LOWER achievement in almost every country measured. Then budget honestly: a scripted, coached, well-implemented unit is worth roughly 0.2 SD on an aligned test and about half that on an external one, and a CPD programme handed to a trainer chain is worth zero. If a science programme quotes an effect above ~0.3, ask who trained the teachers and who administered the test.

Who this applies to

Group size
whole-class
Delivered by
teacher
Ages studied
818
Dose
Either a scripted unit (7-10 weeks; e.g. 24 x 60-min sessions with a coach) or 4-5 days of teacher CPD across a year with funded in-school planning time. Effects appear only where the developers or intensively-coached trainers delivered the training; a train-the-trainer chain removed the whole effect.
Cost
low
Moves
domain-skillnon-cognitive
Needs first
Teacher subject knowledge, plus sustained coaching through the first year of delivery. The treatment must include explicit teacher conceptual explanation — that is the ingredient the PISA split isolates, and removing it reverses the sign.
Not for
Minimal-guidance or student-designed investigation as a primary method: negatively associated with achievement in 18 of 20 PISA regions, on enjoyment as well as attainment. Also not for transfer — the efficacy trial's own one-year follow-up found KS2 reading -0.01 and maths -0.06, and a sibling reasoning programme found English -0.13 and maths -0.13.

Verdict

The archive already grades the general question at moderate-support / medium, and its finding is about guidance dose, not about "explicit beats inquiry." The science-specific evidence agrees with that verdict and does not disturb it. What it adds is a second variable the cross-cutting topic does not carry: in science, the guidance finding is robust and the programmes built on it are not.

Two claims, graded separately:

  1. Guidance is the operative variable, and science gives the cleanest demonstration of it anywhere in the archive. Aditomo & Klieme split PISA 2015's nine inquiry items into guided inquiry (hands-on activity plus teacher conceptual explanation) and independent inquiry (students design their own investigations) and ran them separately in the twenty highest- and lowest-scoring science systems. Guided inquiry is positive in 16 of 16 regions; independent inquiry is negative in 18 of 20 — with coefficients of similar magnitude in both directions, on the same students, in the same models. That is the guidance-dose claim replicated on 151,721 students.
  2. The instructional programmes that package guided inquiry deflate hard. The best-documented chain in this archive runs +0.22 → +0.01 → +0.02, on the same science assessment each time.

Hence mixed: the mechanism is well supported, the marketed effect sizes are not.

What the evidence shows

Source Design Grade Key effect
Slavin 2014 best-evidence synthesis, 23 studies, independent measures required B Inquiry via science kits +0.02 (7 studies); inquiry emphasising teacher PD +0.36 (10 studies)
Aditomo & Klieme 2020 PISA 2015 SEM, 20 regions, 151,721 students D Guided inquiry positive 16/16 regions; independent inquiry negative 18/20 (standardized)
Minner 2010 research synthesis, 138 studies C No significant association between inquiry saturation and conceptual learning; emphasis on active thinking does predict. Only 11% used an established instrument
Furtak 2012 inquiry meta, 37 studies C Inquiry d=0.50; teacher-led beats student-led by ~0.40 (researcher-designed)
Schroeder 2007 meta, 61 US studies C Inquiry d=0.65; enhanced context 1.48 (researcher-designed)
Granger 2012 cluster RCT, 140 teachers / 2,594 pupils B d=0.17 immediate → 0.07 ns at 5 months, on the externally developed MOSART
Saavedra 2022 cluster RCT, 74 schools / 3,645 students B AP exam total +0.192; environmental science +0.304 — College Board scored, fully external
Krajcik 2023 cluster RCT, 46 schools, developer-led B 0.277 SD; WWC classifies the outcome as researcher-developed although the paper calls it standardized
Harris 2015 RCT, 42 schools C +0.21 / +0.25 (researcher-designed); WWC: does not meet standards
Lynch 2005 10-cluster RCT, 4,176 students C 0.32 / 0.25 — WWC clustering correction → 0.29, not significant
Geier 2008 non-equivalent comparison D 0.44 / 0.37 on the state MEAP; WWC: does not meet standards
Wilson 2010 RCT, same teacher both arms, n=58 C 5E advantage at 0 and 4 weeks; BSCS evaluating BSCS; effect size not extractable
EEF TDTS efficacy 2015 cluster RCT, 41 schools / 655 pupils, developer-delivered CPD B +0.22; FSM +0.38; KS2 reading −0.01, maths −0.06 a year later
EEF TDTS effectiveness 2018 cluster RCT, 205 schools / 7,806 pupils, independent (AIR) A +0.01 (−0.08 to 0.10), p=0.79; interest +0.12, self-efficacy +0.09
EEF TDTS effectiveness 2025 cluster RCT, 180 schools / 6,107 pupils, independent (York) A g=0.02 (−0.07 to 0.12); FSM 0.00 — the null reproduced after the training model was repaired
Teig 2018 TIMSS 2015, 4,382 students / 211 classes D Quadratic term −0.264 → −0.174, not significant once SES is controlled
OECD PISA 2015 / Jerrim 2019 correlational D The famous raw negative attenuates to null under within-country controls
Zhang 2016 critique D Pooled "inquiry" estimates aggregate incommensurable treatments

The scale-up chain, which is the most important thing on this page

Thinking, Doing, Talking Science is the cleanest efficacy-to-effectiveness record in the archive, and it is clean precisely because nothing about the measurement changed:

Trial Schools / pupils Who trained the teachers Who administered the test Effect
Hanley, Slavin & Elliott 2015 (efficacy) 41 / 655 the programme developers, 5 days + 2 funded planning days the participating teachers +0.22
Kitmitto (AIR) 2018 (effectiveness) 205 / 7,806 trainers taught by the authors, 4 days, planning days removed NatCen field staff, independent scorers +0.01
Robinson-Smith (York) 2025 (effectiveness) 180 / 6,107 strengthened accredited trainer programme independent +0.02

AIR obtained permission to use the same science assessment, from the same 2001 Russell & McGuigan item bank, that the efficacy trial used. So the collapse from +0.22 to +0.01 cannot be charged to measure type. What changed was who delivered the training, who handled the test, and how many schools were involved. Two independent trials, five years apart, with the developers' own diagnosis of the first null acted upon in between, and the effect did not come back. EEF's own scale-up synthesis puts TDTS among seven programmes whose mean effect fell from 0.25 to 0.01.

The programme did not fail to change practice — teachers in treatment schools were 19–30 percentage points more likely to report confidence adapting teaching and challenging high achievers, and pupil interest in science rose +0.12 (p=0.015). It changed practice and attitudes and did not change learning.

The word "inquiry" is doing too much work

Zhang's construct-validity complaint is fully vindicated by this file set. These are not the same treatment and pooling them produces a number about nothing:

  • Scripted commercial curriculum, teacher-led, heavy PD — Granger (GEMS: 24 × 60 min, 4 PD days, an assigned coach), Lynch (Chemistry That Applies: 18 prescribed lessons), Harris and Krajcik (scripted big-question project units).
  • A structured 5E instructional sequence — Wilson, same teacher teaching both arms.
  • Teacher CPD in a discussion pedagogy with no scheme of work — TDTS, all three trials.
  • Full-year project-based learning with the teacher as facilitator — Saavedra.
  • Self-reported activity frequency with guidance unknown — Teig, and the PISA headline.
  • Guidance explicitly separated from activity — Aditomo & Klieme, and only there does the sign become interpretable.

Every randomised win in this table is a scripted, heavily scaffolded, PD-intensive programme. Nothing here is evidence for open discovery, and the one design that isolates unguided investigation finds it negative almost everywhere.

And the active ingredient is the teacher training, not the materials. Slavin's best-evidence synthesis is the only review here that requires a control group, four weeks' duration, and outcome measures independent of the treatment. Under those rules, inquiry programmes built around science kits return +0.02 across seven studies, while inquiry-oriented programmes built around sustained teacher professional development return +0.36 across ten. Buying the box does nothing; buying the training does something — which is exactly what the TDTS chain shows from the other direction, since what broke on scale-up was the training model.

Measure type, and why it is not the main villain here

The paired estimates that exist point at roughly , in line with METHODOLOGY's nominal correction and nowhere near the ~5× the writing tranche found:

  • The PBIS developers' own evaluators state it outright: past studies gave d = 0.37–0.44 where standardized test scores were used and d = 1.0–1.5 in smaller studies using researcher-built performance tasks closely aligned to the learning goals — a ≈3× gradient asserted inside the literature, not imposed from outside.
  • Across randomised trials here: aligned researcher-designed outcomes run 0.21, 0.25, 0.25, 0.277, 0.32; independent/external outcomes run 0.17 (MOSART), 0.19–0.30 (College Board AP), 0.01–0.02 (TDTS at scale), −0.02 (the reasoning sibling). Aligned-to-independent is roughly 1.5×; meta-to-well-identified-RCT is roughly 3–4×.
  • Efficacy-to-effectiveness is ~20× and dwarfs both. In science, who ran it and at what scale deflates an estimate far harder than who wrote the test.

One counter-example is worth keeping so the correction is not applied mechanically: the project-based Detroit work scores higher on the state test (0.37–0.44, non-randomised) than the same programme family scores on researcher-built tests under randomisation (0.21–0.25). That is selection running the other way, not measure inflation.

Hereditarian-lens assessment

Risk: low. The verdict rests on cluster-randomised trials — TDTS ×3, Granger, Saavedra, Krajcik, Lynch — where allocation is at school level and genotype cannot differ systematically between arms. The correlational leg (PISA, TIMSS) is explicitly firewalled: it is used only for the contrast between guided and independent inquiry within the same students and models, never for the absolute size of an inquiry effect. That firewall matters, because the raw PISA correlation is exactly what selection predicts — Teig's curvilinear result loses significance once SES is controlled, and Jerrim showed the famous negative attenuates to null under within-country controls (weaker classes get assigned more hands-on work).

Nothing here touches g. The outcome class throughout is domain-skill, with non-cognitive (interest, self-efficacy) as the one secondary effect that survives independent evaluation.

Boundaries & what critics say

  • Only 11% of the 138 studies in the largest synthesis used an established instrument. That is the field's own summary of its measurement, and it is why the meta-analytic 0.50–0.65 should never be the planning number.
  • The developers' defences of the TDTS nulls are real and were tested. The 2018 null was blamed on a diluted train-the-trainer chain; the 2025 trial strengthened exactly that and returned 0.02.
  • Granger's effect halves and loses significance five months later — the only delayed follow-up on an external measure in this set. Assume fadeout; nobody has measured past a year.
  • Krajcik 2023's outcome label is disputed. The paper calls it standardized; the project's own technical report says it was designed by the Michigan Department of Education; WWC classifies it as researcher-developed. Cite it as 0.277 on an aligned measure, developer-led.
  • The undergraduate active-learning literature does not transfer here and is recorded as off-scope: Freeman 2014 (225 studies, +0.47, failure rates 33.8% → 21.8%) and Deslauriers 2019 (students learn more from active instruction and feel they learn less) are about self-selected adults in lecture courses, not about children in schools.

Practical guidance

  • Require teacher conceptual explanation as part of the activity. This is the single actionable finding: the same practical activities predict higher achievement when the teacher explains the concepts and lower achievement when students are left to design and conclude alone.
  • Never adopt student-designed investigation as a primary method. Negative in 18 of 20 systems, on enjoyment as well as attainment.
  • Buy the scripted unit, not the pedagogy workshop. Every randomised positive here came with prescribed lessons, materials and coaching. The one intervention that was pure CPD is the one that went to zero twice.
  • Budget 0.1–0.2 SD on an external measure, not 0.5. And expect roughly half of it to be gone five months later.
  • Do not claim transfer. Both in-scope follow-ups measured it and both found nothing: KS2 reading −0.01 and maths −0.06 a year after the efficacy trial.
  • Treat interest as the reliable secondary product. +0.12 on interest and +0.09 on self-efficacy survived a 205-school independent trial in which attainment did not move at all.

Open questions

  • Why did Saavedra's AP result survive when everything else deflated? It is the only positive here on a fully external, high-stakes, externally scored exam — 74 schools, five districts, +0.19 overall and +0.30 in environmental science. Either full-year PBL with older students is genuinely different, or this is the next result to collapse on replication. Nobody has tried.
  • No K-12 science trial in this set reports the same construct on both an aligned and an independent measure at the same time. The 2× estimate above is assembled across studies and from the PBIS evaluators' own statement; the within-study test has not been run.
  • Nothing measures beyond five months except the efficacy trial's transfer test.
  • Whether the guided/independent split survives outside PISA self-report is untested. It is the most important finding on this page and it rests on a cross-sectional questionnaire.

Evidence (19 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore

Inquiry-based science teaching vs explicit and textbook science teaching · The Evidence on Teaching