The Evidence on Teaching

Laboratory and practical work in school science

Two reviews screened ~12,000 records and found one RCT and six quasi-experiments. What evidence exists favours interest (+0.12) over attainment (+0.01), and simulations match real apparatus.

mixedconf: mediumgc: low

science · ages 518

Effect summary

Practical work reliably teaches students to do the practical (activity-based science: process skills 0.52) and barely moves what they know (content 0.16). The measure-type split inside that same 1983 meta is the cleanest in the tranche: outcome tests judged biased toward the hands-on programme 0.54, unbiased 0.24, biased toward the control 0.13 — a 2.25x inflation. The best-identified modern evidence is a 205-school independent trial of a practical-and-discussion-heavy science pedagogy: attainment +0.01 (CI -0.08 to 0.10) but interest +0.12 and self-efficacy +0.09, replicated at 180 schools (g=0.02). Physical apparatus and simulation are equivalent in three K-12 RCTs from kindergarten to grade 8. Cost is not the issue people think: English state secondaries spend GBP 0.75-31.25 per student per year on consumables, and the dominant cost is lesson time.

Practical takeaway

Do practical work for enthusiasm, apparatus competence and subject identity — those are what the causal evidence supports and what it survived for. Do not do it to teach content, and do not let it displace explanation: the same activities predict LOWER achievement where the teacher does not attach a conceptual explanation. Teach the idea first, then use the practical to anchor it, and plan out loud how the observation connects back. Where a phenomenon simulates well, a simulation is as good as the apparatus and costs less.

Who this applies to

Group size
whole-classsmall-group
Delivered by
teacher
Ages studied
515(narrower than the 518 this topic is filed under — outside it is extrapolation)
Dose
Trialled doses run from a single 4-6 week unit to a full school year with 4-5 CPD days. The commonly recommended benchmark (practical activity in at least half of science lessons) has never been tested against any alternative frequency.
Cost
low
Moves
non-cognitive
Needs first
The target concept must already have been taught. Practical work functions as a bridge from an idea the student already holds to an observation, not as a route to discovering the idea — and it needs an explicit planned link from the observation back to the concept. In the observational study of 55 English science lessons, not one contained such a plan.
Not for
Teaching science content. The best-identified K-12 evidence — a 205-school independent trial replicated at 180 schools — finds nothing on attainment. Also not for paying a premium for physical apparatus over simulation where the phenomenon is simulable, and not for out-of-school science-centre labs, which lose to the ordinary classroom on achievement and win only on situational enjoyment.

Verdict

Practical work is the single most expensive assumption in school science — in lesson time, in buildings, in technicians, in teacher preparation — and it is close to the least evidenced thing in this archive.

The number that matters most is not an effect size. Gericke's systematic review screened 11,771 records covering 1996–2019 and retained 39 studies. Oliveira & Bonito screened 163 and retained 53, of which one was a randomised trial and six were quasi-experiments. After a century of universal practice, the causal literature on secondary-school laboratory work is roughly seven studies.

mixed, because the pieces point different ways and a single verdict would misrepresent it:

  • Practical work → science content learning: nothing, at the only scale anyone has tested. Activity-based science moves content by 0.16 while moving process skills by 0.52. The best-identified modern trial of a practical-and-discussion-heavy pedagogy returns +0.01 on attainment across 205 schools and +0.02 on replication across 180.
  • Practical work → interest and self-efficacy: real, randomised, and it survives scale-up. +0.12 and +0.09 in the trial where attainment did not move at all.
  • Physical apparatus versus simulation: equivalent. Three independent K-12 RCTs, kindergarten to grade 8, all null; 50 of 56 studies in a review equal-or-better for virtual and remote labs.
  • Practical work versus a content-matched, time-matched non-practical lesson: never tested. This is the contrast the entire debate presupposes, and it does not exist at K-12.

That last point is the reason this is mixed and not no-effect. On content learning the evidence genuinely is a null and it is a well-powered independent one — but it is a null on a programme that includes practical work, not on practical work against its real alternative.

What the evidence shows

Source Design Grade Key effect
EEF TDTS effectiveness 2018 cluster RCT, 205 schools / 7,806 pupils, independent A Attainment +0.01 (−0.08 to 0.10); interest +0.12 (p=0.015); self-efficacy +0.09
EEF TDTS effectiveness 2025 cluster RCT, 180 schools / 6,107 pupils, independent A g=0.02; the null reproduced with a repaired training model
Slavin 2014 best-evidence synthesis, 23 studies, independent measures required B Inquiry programmes built on science kits: +0.02 across 7 studies; the same inquiry built on teacher PD: +0.36
Bredderman 1983 meta, 57 studies / ~13,000 pupils C Overall 0.35; process 0.52 vs content 0.16; biased tests 0.54, unbiased 0.24, control-biased 0.13
Itzek-Greulich 2015 / 2017 cluster RCT, 67–68 classes / ~1,800 students B All lab arms beat an untaught control; school lab beats science-centre lab on achievement, science centre wins on motivation
Itzek-Greulich & Vollmer 2017 same trial, motivational outcomes B State affect rises; trait motivation roughly flat
Klahr 2007 RCT, grades 7–8, n=56 B Physical vs virtual materials: no difference on design performance, causal knowledge, transfer, confidence or interest
Triona & Klahr 2003 RCT ×2, grades 4–5 B Physical vs virtual: no difference on CVS or domain knowledge
Zacharia 2012 RCT, kindergarten, n=80 C Physical vs virtual: no difference
Brinson 2015 review, 56 studies C 50/56 (89%) equal-or-higher for virtual/remote, including 7 of 9 on practical skill itself
de Jong 2013 review C Physical and virtual labs have complementary, largely substitutable affordances
Gericke 2022 systematic review C 39 studies retained from 11,771 screened, 1996–2019
Oliveira & Bonito 2023 systematic review D 53 studies retained; 1 RCT and 6 quasi-experiments
Abrahams & Millar 2008 observation of 25 lessons D Effective at getting students to do what was intended; ineffective at getting them to think what was intended
Abrahams & Reiss 2012 observation, 55 lessons, primary + secondary D Same pattern in both phases; no observed lesson contained an explicit plan linking observation to concept
Sharpe & Abrahams 2020 survey D Under 40% interested in the epistemic parts (hypothesising, measuring); more enjoy setting up apparatus
Gatsby 2017 / SCORE 2013 grey literature, costing D Staff time dominates cost; consumables GBP 0.75–31.25 per student per year in state secondaries

Two clean paired measures, and they agree

  1. Inside Bredderman's meta, on the same studies: outcome tests coded as biased toward the activity-based programme d = 0.54; unbiased d = 0.24; biased toward the conventional control d = 0.13. Inflation factor 2.25× — the tidiest confirmation of METHODOLOGY's nominal 2× anywhere in this tranche.
  2. Same programme, two trials: the bespoke, teacher-administered test in the efficacy trial gave +0.22; the same instrument administered and scored independently at five times the scale gave +0.01. That is not 2×; the effect goes to zero. It confounds measure administration with scale-up, so treat it as an upper bound on the measurement component alone.

Bredderman's second split is worth as much as the first: process skills 0.52 versus content knowledge 0.16. Activity-based science reliably teaches students to do the activity. It barely moves what they know.

And the hands-on materials themselves are the part with nothing behind them. Slavin's synthesis required a control group, four weeks' duration, and outcome measures independent of the treatment. Under those rules, elementary inquiry programmes built around science kits return +0.02 across seven studies, while inquiry programmes built around sustained teacher professional development return +0.36 across ten. The apparatus is not the ingredient.

The cost argument inverts

The intuitive case against practical work is that it is expensive. The costing evidence says the expensive part is not what people assume.

  • Consumables and equipment are cheap. English state secondaries spent GBP 0.75 to 31.25 per student per year on practical science (independent schools 7.18–83.21; primaries 0.04–19.08). Secondaries held about 70% of the "essential" equipment list, primaries 46%, and around 70% of secondary science staff bought materials with their own money.
  • Staff time dominates — overwhelmingly teachers' time, which the school is paying anyway. Capital cost of laboratories and equipment is small by comparison.
  • Three of six comparison countries in the Gatsby review — the USA, Germany and Finland — employ no school laboratory technicians at all.

So the real price of practical work is lesson time, which means every effect estimate in this literature uses the wrong denominator. All of them compare practical work with business-as-usual or with nothing; none compares it with the next-best use of the same hour. That is the missing trial.

Hereditarian-lens assessment

Risk: low. The load-bearing sources are cluster-randomised (two EEF effectiveness trials, the Itzek-Greulich trial, three physical-versus-virtual RCTs), so allocation is exogenous to family background. The observational leg — Abrahams & Millar, Abrahams & Reiss, Sharpe & Abrahams — is used only descriptively, for what happens in lessons and what students say about them, never for a causal estimate.

The one selection lesson worth importing comes from the excluded undergraduate evidence: a multi-institution study of introductory physics found a raw t=4.86 advantage for lab-takers that vanished entirely under a within-student design (+0.24 ± 0.67 percentage points). The apparent benefit of labs was who took them, not what they did. Nothing here bears on g.

Boundaries & what critics say

  • The strongest null in the world literature is out of scope. Holmes and colleagues' within-student difference-in-differences across 2,617 undergraduates and nine courses put the educational benefit of the introductory physics lab at +0.24 ± 0.67 percentage points. It is recorded as off-scope because it is undergraduate, and because it is out of scope the archive's Lindy burden for overturning a long-standing practice is not met — only the EEF trials are in-scope grade A/B evidence, and they test a programme rather than practical work itself.
  • The trial that would settle it does not exist. Itzek-Greulich comes closest and its control arm received no teaching on the topic at all, so "lab beats control" there means "teaching beats not teaching."
  • TDTS is not a pure test of practical work. It is a CPD programme in a practical-and-discussion pedagogy. It is the best-identified relevant evidence and it is not the same thing.
  • Physical-versus-virtual equivalence has real limits. The samples are small, the outcomes immediate and researcher-designed, and every domain tested (springs, beam balances, mousetrap cars) is one that simulates well. Nobody has tested whether physicality matters for smell, for hazard, or for apparatus that fails.
  • The misconceptions topic disagrees. A meta-analysis of floating-and-sinking instruction finds hands-on experiments (0.96, 53 studies) beating virtual ones (0.52, 17 studies), F(1,189)=8.02, p<0.001. That is a between-study moderator across different researchers and decades; the three head-to-head randomised trials are within-study. The randomised evidence should win, but the disagreement is real and unresolved.
  • Students do not enjoy the part the advocates value. Fewer than 40% report interest in hypothesising and measuring; more enjoy setting up the equipment. In an older survey, experiments ranked third for enjoyment (71%) but only 38% for usefulness — behind class discussion (48%) and note-taking (45%).

Practical guidance

  • Keep practical work, and change what you claim for it. Enthusiasm, subject identity and apparatus competence are the outcomes with randomised support that survives independent scale-up. Attainment is not one of them.
  • Teach the concept first, then do the practical. The observational work is unanimous that practical work functions as a bridge from an idea already held to an observation — and that in 55 observed lessons, not one contained an explicit plan for making that link.
  • Attach an explanation to every activity. This is the same finding as guided vs independent inquiry: the activity without the teacher's conceptual explanation is the version that predicts lower achievement.
  • Buy the simulation where the phenomenon simulates. Three K-12 RCTs from kindergarten to grade 8 find no advantage for real apparatus, and 89% of a 56-study review favours or ties virtual.
  • Do not pay for the science-centre trip on learning grounds. Under randomisation the ordinary classroom beat the out-of-school lab on achievement; the out-of-school lab won on situational enjoyment, and on state rather than trait motivation.
  • Count the hour, not the consumables. At under GBP 31 per student per year for materials, the budget question is a rounding error and the timetable question is the whole decision.

Open questions

  • The decisive experiment has never been run at K-12: practical work versus a content-matched, time-matched non-practical lesson on the same concept, with an independent outcome measure. Every estimate in this topic answers a different question.
  • Nobody has measured downstream subject uptake. If the survived-for function of practical work is enthusiasm and identity, the outcome that matters is whether students go on taking science — and no study in this set follows anyone that far.
  • Whether the motivational effect is worth anything. It is +0.12, it is state rather than trait, and its downstream consequences are entirely unmeasured.
  • Physicality for non-simulable phenomena — hazard, smell, failure of apparatus, the smell of a gas prepared badly — is untested and is exactly where a real advantage would be expected.

Evidence (22 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore