The Evidence on Teaching

Does learning to code improve general thinking?

Coding instruction teaches coding (d≈0.68) and not thinking: against active controls far transfer is g=0.15, two validated-instrument RCTs found d≈0.00, and a 46-school maths trial found −0.16 to −0.21.

no effectconf: mediumgc: low

transfer · ages 516

Effect summary

This claim was tested hard in the 1980s, failed, was forgotten, and returned verbatim as 'computational thinking'. The best meta-analysis contains its own refutation: overall transfer g=0.49, far transfer g=0.47 — but far transfer against ACTIVE controls is g=0.15 (vs 0.64 against untreated controls, Q_M=40.12, p<.001), Egger t(537)=4.10 p<.001, trim-and-fill 0.43, standardized tests 0.33 vs unstandardized 0.49, school achievement 0.28 falling to a non-significant 0.22, literacy −0.02. Every correction moves the same way. Then the trials arrived. A pupil-randomised English trial of a year of Code Club: Bebras +0.93 points on a 117-point scale, d≈0.05, CI −0.15 to 0.26, while coding skill moved d≈0.67. A school-randomised US trial of a 24-lesson curriculum: TechCheck d=−0.02 (p=0.81) while coding moved d=0.68 — and the developer built both instruments. A 46-school class-randomised French trial pitting Scratch against content-matched maths teaching: −0.16, −0.19, −0.21 on three mathematical notions. Structural facts about the corpus: 87.6% of 105 studies measured far transfer only, never checking whether anyone learned to program; zero of 105 reported a delayed follow-up.

Practical takeaway

Teach coding as a subject. Never buy it as a thinking programme. The claim has now been tested twice across forty years — once as Logo, once as computational thinking — and the modern test is the more decisive one because the modern era finally built validated instruments and then failed on them. If a vendor shows you a computational-thinking gain, ask three questions: was the control group doing something, was the instrument built by the people selling you the curriculum, and what did the students lose the hours from. The one trial that answered the third question found the hours came out of mathematics and cost 0.16 to 0.21 SD.

Who this applies to

Group size
whole-classsmall-group
Delivered by
teacher
Ages studied
516
Dose
The corpus runs 2 to 120 hours with a median of about 20; the modern trials are 8 hours (Code.org), 24 lessons, 12 weeks, or a full academic year. None of these doses moves an independently measured thinking outcome. The dose that matters most for a school builder is the one that came out of mathematics: a class-randomised trial across 46 schools substituted Scratch activities for content-matched mathematics teaching on three notions and lost 0.16 to 0.21 SD on each.
Cost
medium
Moves
domain-skill
Needs first
None are identified. The anchor meta-analysis explicitly found no moderation by programming language, educational level or intervention length — which is itself informative: a real transfer mechanism should have conditions.
Not for
Do not adopt programming as a route to mathematics, literacy, reasoning, creativity, executive function or general problem solving. Against an active alternative far transfer is g = 0.15; against a matched mathematics curriculum it was NEGATIVE; two randomised trials using validated computational-thinking instruments (Bebras, TechCheck) found d ≈ 0.00 while the same trials showed d ≈ 0.68 on coding skill. Do not accept a computational-thinking test score as evidence of transfer: CT test scores correlate r = .44-.67 with reasoning, spatial and problem-solving ability and r = .56 with fluid intelligence in programming-naive adults, so a large share of that score is ordinary ability.

Verdict

no-effect on far transfer, general thinking and academic achievement — and deliberately not debunked, for the same reason brain training is not. Near transfer and domain skill are real and well evidenced. The phenomenon is bounded, not fake. What fails is the spillover, and it is the spillover that policy arguments are made of.

This is the third time this archive has graded the same structure. Working-memory training improves the trained task and nothing else. Fundamental movement skills improve on the trained battery and do not cascade. Programming teaches programming and does not teach thinking. The shape is identical: what transfers is knowledge and skill, not capacity.

What the evidence shows

Source Design Grade Key effect
Scherer, Siddiq & Sánchez Viveros 2019 meta, 105 studies / 539 ES / N = 9,139 C Far transfer 0.47 → 0.15 against ACTIVE controls (Q_M = 40.12, p < .001); literacy −0.02
Straw 2017 (Code Club) pupil-randomised RCT, 21 schools, 1 year B Bebras d ≈ 0.05 (CI −0.15 to 0.26); coding skill d ≈ 0.67
Yang, Blake-West, Yang & Bers 2025 cluster RCT, 11 schools, 331 children B TechCheck d = −0.02 (p = 0.81); coding d = 0.68. Developer-led null
Laurent 2022 class-randomised, 46 schools, N ≈ 1,880 B Mathematics −0.16, −0.19, −0.21 against content-matched teaching
Pea & Kurland 1984 (TR16) two year-long quasi-experiments C No planning benefit, on product or process measures, twice
Pea & Kurland 1984 (TR9) critique D Every Logo transfer claim traced to non-blind ratings or unevidenced assertion
WWC 2007 federal evidence review D Zero Logo studies met design standards for elementary mathematics
Liao & Bright 1991 meta, 65 studies / 432 comparisons D 0.41 — and publication type is a significant moderator, in 1991
Montuori 2023 meta, 11 studies / 862 children C Problem solving 0.890 (labelled near transfer); flexibility 0.118, ns
Arfé 2019 RCT, n = 76 / 38, waiting-list control C The steelman: Tower of London d = 0.95 from 8 hours of Code.org
Robledo-Castro 2025 registered RCT, 111 children, 2 clusters/arm C Working memory η² = 0.167; cognitive flexibility: nothing
Zhang 2025 3-arm RCT, 198 preschoolers C Robots beat unplugged on EF — but unplugged beat conventional on CT and on NO executive function
Li & Oon 2024 meta, 37 studies C Far transfer 0.444 with Egger p < .001; n < 50 → 0.826, n > 150 → 0.233
Román-González 2017 criterion validity, n = 1,251 D CT test correlates .67 / .44 / .44 with problem solving, spatial, reasoning
Aranyi 2024 correlational, programming-naive adults D CT test vs fluid intelligence r = .56; crystallised r = .13 (ns)

The meta-analysis everyone cites for the modern claim contains its own refutation, and it is in the moderator table. Far transfer against treated control groups is g = 0.15; against untreated controls it is 0.64. Add Egger t(537) = 4.10 (p < .001), trim-and-fill 0.43, published 0.58 against grey literature 0.43, standardized measures 0.33 against unstandardized 0.49, school achievement falling from 0.28 to a non-significant 0.22 once ten influential effect sizes are dropped, and literacy at −0.02. Every correction moves the same direction. This is precisely the design-quality gradient the archive has already documented for chess, music and working-memory training, reproduced a fourth time in a fourth domain.

Then the trials arrived, and the modern era's one real advantage turned against it. The Logo era had no validated, independent measure of thinking. The computational-thinking era built two — Bebras and TechCheck — and both have now been the primary outcome of a properly randomised multi-school trial. Both returned approximately zero, in two countries, with two curricula, while the same trials moved coding skill by d ≈ 0.67 and d = 0.68. The English evaluators pre-emptively checked and reported that Bebras was sensitive enough to detect a difference if one existed. One of the two nulls is developer-led, which under this archive's efficacy-decay logic is the strongest kind of null there is.

And the modern era ran the academic test the Logo era never ran, and lost it. Forty-six schools, class-randomised, Scratch activities against traditional teaching on the same three mathematical notions: −0.16, −0.19, −0.21, all significant. The meta-analytic far-transfer-to-mathematics estimate is +0.57. Under a properly active baseline at scale it is negative. That is the same shape as the −3.0 mathematics points at two years in the archive's brain training file, and the same mechanism: the hours come out of something.

The pro-transfer case is one laboratory. The modern positive results — planning d = 0.93-1.27, inhibition 0.65-1.05, problem solving 1.31-1.91 — come from a single group in Padua, from doses of 8 hours of Code.org, largely against waiting-list controls, and the meta-analysis of that literature was written by the same group, used a fixed-effects model, excluded grey literature, and states that publication bias could not be robustly assessed. Eight hours producing d ≈ 1.0 on the Tower of London is not a finding; against the archive's own benchmark that a year of schooling moves standardized achievement 0.2-0.4 SD, it is a warning light. The older pro-transfer classic, Clements & Gullo 1984, has nine children per arm — and, to its credit, an active control, which most modern studies still do not manage.

Two structural facts about the corpus deserve to be said out loud. In 105 studies asking whether learning to program improves thinking, only 13 checked whether anyone learned to program — 87.6% measured far transfer only. And not one of the 105 reported a delayed follow-up. In a topic whose entire promise is a durable change in how a child thinks, the maximum measured horizon in the corpus is the end of treatment.

Is the computational-thinking claim better evidenced than the Logo claim it replaced?

No. On the dimension that decides it, worse — and where it is better evidenced, the better evidence is negative.

  • The corpus is largely the same corpus. The 2019 meta-analysis notes that most included publications date from the 1980s and 1990s. Mean study n = 87, median 66.
  • The artefact moderators are identical across the generation gap. Liao & Bright found publication type and treatment duration significant in 1991; Scherer found publication status and control-group treatment significant in 2019. Twenty-eight years, the same two variables.
  • The new positive evidence is more fragile than the old positive evidence, not less. One lab, 8-hour doses, waiting-list controls, and a self-authored fixed-effect meta-analysis that cannot assess bias.
  • The construct degraded. Logo denoted something concrete: a program a child wrote. "Computational thinking" still has no consensus definition — the 2017 complaint is conceded verbatim by a 2024 meta-analysis. And the tests built to measure it are substantially ability tests (r = .67 with a problem-solving battery, r = .56 with fluid intelligence in programming-naive adults, crystallised intelligence unrelated).
  • What actually improved is the other claim. In 1984 the field could not reliably show children learned to program at all — concept grasp after 30 hours was "highly context-specific." Today three randomised trials converge on d ≈ 0.5-0.7 for coding skill. The claim that survived is smaller than the one made in 1980, and smaller than the one being made now.

Hereditarian-lens assessment

Risk: low for the verdict. It rests on trials randomised at pupil, school and class level. Genes cannot differ between arms.

The lens is decisive on the other half of the file, though — the correlational half, which is high risk and is where most public confidence in this claim actually comes from. Computational-thinking test scores correlate r = .44-.67 with reasoning, spatial ability and problem solving in 1,251 schoolchildren, and r = .56 with fluid intelligence in 97 programming-naive adults, with crystallised intelligence unrelated (r = .13, ns). In other words a computational-thinking test, administered to people who have never programmed, is already measuring reasoning ability. So:

  • "Programmers think better" is ability selection, not evidence of transfer.
  • A gain on a researcher-made CT test is partly a reasoning score, which is exactly the class of estimate that shrinks when a validated independent instrument is substituted — as it did, to zero, twice.

This is also the archive's cleanest available restatement of premise (2). "No durable gains in g" is not "nothing raises test scores": schooling raises IQ scores through directly taught content. Programming is another attempt at the capacity route rather than the content route, and it produces the same nothing.

Boundaries & what critics say

  • Near transfer is real and should not be denied. Children who learn Scratch get better at Scratch, and at tasks that look like Scratch. The 55-study Scratch review documents this well. It is routinely cited in policy documents as though it evidenced cognitive transfer; it does not address transfer at all.
  • The steelman is the Padua line plus the preschool three-arm trial. The latter is the best-designed pro-transfer study here — three arms, a genuine alternative activity, twelve weeks — and its own internal structure argues against it: the marginal R² for the fixed factors is 0.29 for computational thinking and only 0.04-0.08 for the executive functions, and unplugged programming beat conventional kindergarten on computational thinking and on NO executive function.
  • The most-cited defence of the field is not what it appears. The 2019 authors respond to the definitional critique by asserting that the claim of no transfer evidence "cannot be substantiated" — on the strength of a pooled estimate their own moderator table puts at 0.15 against active controls, and their 2021 summary for a general audience omits that moderator entirely.
  • The construct-validity objection cuts both ways. If CT tests are substantially ability tests, then the nulls on Bebras and TechCheck are also nulls on instruments that may be insensitive to whatever coding actually builds. The answer to that is the mathematics trial, which used ordinary curricular outcomes and found harm.
  • This is not an argument against teaching computer science. It is an argument about which justification to use. See programming instruction for the one that survives.
  • Two of the flagship nulls are small. Straw analysed 317 pupils; the ScratchJr trial had 11 clusters and its authors flag underpowering for the computational-thinking outcome. Confidence is capped at medium accordingly, and because the adversarial pass METHODOLOGY requires has not been run.

Practical guidance

  • Never buy coding as a thinking programme, at any age, at any dose, in any packaging — Logo, Scratch, robotics, unplugged, computational thinking, twenty-first-century skills.
  • Interrogate any positive claim on three axes: was the control group doing something, was the instrument built by the people selling the curriculum, and where did the hours come from. Those three questions account for essentially the entire positive literature.
  • Protect mathematics time specifically. The only trial that substituted programming for matched mathematics teaching at scale lost 0.16-0.21 SD on every notion tested.
  • Do not treat a computational-thinking score as a learning outcome. It correlates with fluid intelligence at .56 in people who have never programmed.
  • If the goal is mathematics, teach mathematics; if it is reasoning, there is nothing here to buy. The archive's fadeout and persistence finding applies: what holds is content, not capacity.

Open questions

  • Nobody has measured anything after the course ended. Zero of 105 studies in the anchor meta-analysis report a delayed follow-up; the only maintenance data anywhere in this file is a one-month post-test.
  • England's 2014 compulsory computing curriculum is an unexploited natural experiment on mathematics and science attainment, with clean pre- and post-administrative data. It is the single highest-value study that could be run tomorrow and has not been.
  • Nobody has independently replicated the Padua results, which are the entire modern pro-transfer case.
  • The one transfer-specific meta-analysis of the CT era that could not be obtained reports, in its abstract, a generally significant effect with no number attached. That is worth chasing.
  • Whether the two validated CT instruments are sensitive to anything coding builds is genuinely open — they are the best measures available and they are also the reason the modern claim fails, which is a position worth being uncomfortable about.

Evidence (21 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore