Randomised Controlled Trial and Process Evaluation of Code Clubs
Straw, S., Bamford, S., & Styles, B. · 2017
grade Brctindependentreplicated
Sample
688 recruited, 537 entered the trial, 317 pupils from 21 schools completed the primary outcome at baseline and endpoint; powered to detect d = 0.17
Population
Year 5 pupils aged 9-11 in England.
Design
A PUPIL-randomised trial (stratified by school) of Code Club, a year-long after-school volunteer-led programming club, registered on ISRCTN and conducted by the National Foundation for Educational Research. Two features make it the most decision-relevant trial in this topic. First, the primary outcome is the Bebras Computational Thinking Assessment — a validated, independently developed instrument, not a researcher-made computational-thinking test — and NFER explicitly checked and reported that Bebras was sensitive enough to detect a difference if one existed. Second, it was commissioned by Code Club UK and the Raspberry Pi Foundation, and the funder published the null openly. Control pupils continued the normal computing curriculum, which itself includes Scratch, so the baseline is business-as-usual with acknowledged contamination — a conservative test for the intervention and a hard test for transfer.
Key findings
After a full academic year of Code Club the effect on computational thinking was +0.93 Bebras points (95% CI -2.65 to 4.51) on a 117-point scale with an SD of about 17 — roughly d = 0.05 (-0.15 to 0.26), t = 0.51, not significant. Both arms gained about 16 points, so children got better at Bebras anyway. Coding skill did move: +1.49 points (95% CI 0.92-2.05) on the Coding Quiz, roughly d = 0.67, t = 5.17, p < .001. This is the cleanest available demonstration that coding instruction teaches coding and does not teach computational thinking as independently measured.
Genetic confound
None material. Pupil-level randomisation stratified by school.
Replication notes
Independently reproduced in signature by the Yang and Bers second-grade cluster trial — a different curriculum, a different validated computational-thinking instrument, a different country, and school rather than pupil randomisation — which found the same pattern: domain skill up around 0.68, computational thinking flat.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Computational thinking (Bebras Computational Thinking Assessment) | point difference and derived d | +0.93 points (95% CI -2.65 to 4.51); d about 0.05 (-0.15 to 0.26); t = 0.51, not significant | validated-instrument | endpoint, after one academic year | business-as-usual | end-of-treatment | far-transfer |
| Coding skill (Scratch, HTML/CSS and Python Coding Quiz) | point difference and derived d | +1.49 points (95% CI 0.92-2.05); d about 0.67; t = 5.17, p < .001 | researcher-designed | endpoint, after one academic year | business-as-usual | end-of-treatment | domain-skill |
| Self-rated ability at making things with code | direction | positive, small | researcher-designed | endpoint | business-as-usual | end-of-treatment | non-cognitive |
Cited by
- Does learning to code improve general thinking?no effectconf: mediumgc: low