The Evidence on Teaching

Learning to argue: A study of four schools and their attempt to develop the use of argumentation as a common instructional practice and its impact on students

Osborne, J., Simon, S., Christodoulou, A., Howell-Richardson, C., & Richardson, K. · 2013

grade Cquasi-experimentdeveloper-lednot-applicable
Sample
Four secondary school science departments over two years, two lead teachers per school, students aged 11-16; plus a comparison sample
Population
England; whole science departments in four secondary schools, students aged 11-16.
Design
The scale-up test of the argumentation programme, run by the researchers who built the case for argumentation — so developer-led, and if it had worked that would be the ceiling, not the floor. Two lead teachers per school embedded argumentation activities in the science curriculum and spread the practice to colleagues, with deliberately MINIMAL support and professional development, over two years. Not randomised; a non-equivalent comparison sample. The redeeming methodological feature, and the reason this source matters more than the enthusiastic small studies, is that outcomes were measured with "a set of standard instruments" covering conceptual understanding, reasoning and attitudes to science — i.e. the authors did not build the test around their own treatment.
Key findings
Argumentation as a whole-department instructional practice, tested with standard instruments against a comparison sample, produced almost nothing: "few significant changes were found in students compared to the comparison sample" on conceptual understanding, reasoning or attitudes. The authors turn the paper toward implications for teacher professional development. This is the efficacy-to-effectiveness decay pattern arriving early: the same research programme that produced encouraging small-unit results (Zohar & Nemet 2002; Osborne et al. 2004) could not move standard measures when argumentation was spread across four real departments with realistic support.
Genetic confound
Medium: four volunteer schools with a non-equivalent comparison sample; school selection and departmental capability differ between arms by construction.
Replication notes
No independent replication of the whole-department model. Its null is consistent with the independent randomised nulls for reasoning-focused science programmes in England (EEF Let's Think Secondary Science 2016) and for general thinking-skills programmes (EEF Philosophy for Children 2021).
DOI / URL
10.1002/tea.21073

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Students' conceptual understanding, reasoning and attitudes to science, argumentation departments vs comparison samplesignificance onlyfew significant changesstandardizedafter two years of departmental implementationbusiness-as-usualend-of-treatmentnear-transfer
Science conceptual understanding specificallysignificance onlyno consistent significant advantagestandardizedafter two yearsbusiness-as-usualend-of-treatmentdomain-skill

Cited by