Coding in Primary Grades Boosts Children’s Executive Functions
Arfé, B., Vardanega, T., Montuori, C., & Lavanga, M. · 2019
grade Crctdeveloper-involvedunreplicated
Sample
Study 1: 76 analysed (42 experimental, 34 control) across 4 classrooms. Study 2: 38 second graders (17 experimental, 19 control)
Population
Italian first graders aged 5-6 and second graders with a mean age of 6.89.
Design
This is the steelman for the transfer claim and the red flag at the same time. Study 1 is a cluster-randomised stepped-wedge design with a WAITING-LIST control — the weakest possible baseline. Study 2 is a mixed randomised/longitudinal design with a business-as-usual control and n = 17 in the longitudinal arm. The dose was 8 hours of Code.org across four weeks. The transfer outcomes were standardized neuropsychological instruments (Elithorn Maze, Tower of London, NEPSY-II Inhibition, numerical Stroop), which is a genuine strength and simultaneously makes the size of the reported effects implausible. Read at secondary depth.
Key findings
Study 1: coding accuracy d = 1.62; Tower of London d = 0.95; Elithorn d = 0.80; NEPSY-II inhibition errors d = -0.65; Stroop errors d = -0.90. Study 2: coding d = 1.91; Elithorn d = 0.96; Tower of London d = 0.93; NEPSY-II inhibition d = -1.05. The longitudinal comparison claims one month of coding — eight lessons — produced planning and inhibition improvement equivalent to or greater than seven months of standard schooling. Gains were retained at a one-month delayed post-test. Against the archive's own benchmark that a year of schooling moves standardized achievement about 0.2-0.4 SD, an eight-hour intervention producing d ≈ 0.9 on the Tower of London against a waiting-list control in 76 children is a red flag, not a triumph.
Genetic confound
Low. Cluster randomisation. The problem is baseline quality and effect implausibility, not heredity.
Replication notes
Extended by the same laboratory (Arfé, Vardanega & Ronconi 2020, n = 179) and pooled by the same laboratory (Montuori et al. 2023). No independent replication located. The two independent randomised trials using validated computational-thinking instruments found nothing on the transfer side.
DOI / URL
10.3389/fpsyg.2019.02713
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Coding accuracy | Cohen d | 1.62 (Study 1); 1.91 (Study 2) | researcher-designed | post-test | none | end-of-treatment | domain-skill |
| Planning (Tower of London accuracy) | Cohen d | 0.95 (Study 1); 0.93 (Study 2) | standardized | post-test | none | end-of-treatment | far-transfer |
| Planning (Elithorn Maze accuracy) | Cohen d | 0.80 (Study 1); 0.96 (Study 2) | standardized | post-test | none | end-of-treatment | far-transfer |
| Response inhibition (NEPSY-II errors; negative d means fewer errors) | Cohen d | -0.65 (Study 1); -1.05 (Study 2) | standardized | post-test | none | end-of-treatment | far-transfer |
| Maintenance of gains | significance | gains retained at a one-month delayed post-test | standardized | one month post-intervention | none | under-1yr | far-transfer |
Cited by
- Does learning to code improve general thinking?no effectconf: mediumgc: low