Subject
Teaching computer science
The program (what a school should actually do)
Computer science is the subject where the gap between the marketing and the evidence is at its widest, and where the archive's transfer rule earns its keep for the third time. The programme is therefore short, and half of it is a refusal.
Teach programming, and justify it on programming and on work.
- Explicit programming instruction works on programming. A 13-school cluster randomised trial of a twelve-week K-2 curriculum moved coding skill g = 0.47; a second cluster trial moved it d = 0.68; a pupil-randomised English trial of a year-long after-school club moved it d ≈ 0.67. Four meta-analyses converge on roughly g = 0.5-0.8.
- The economic case is the strongest attainment evidence in the subject. In 635,771 Maryland students, taking a high-quality high-school CS course raised the probability of declaring a CS major 10.2 points, of earning a CS degree 5.5 points, and annual earnings at 24 by about 8%. Note the honest caveat: much of that is reallocation within STEM, pulling students out of other STEM (10-13 points) and engineering (5-9 points), not out of non-STEM. (evidence: programming instruction)
Stop choosing between methods and start choosing between measures.
- The anchor meta-analysis of 139 interventions reports that effect sizes "differed only marginally between the instructional approaches and conditions." Blocks versus text — the single design principle the whole K-12 computing movement rests on — is a pooled null (g = 0.245, CI −0.078 to 0.567, p = .137) with detected publication bias.
- Two levers are worth having. Parsons problems deliver the same post-test and the same one-week retention in roughly half the practice time (84 seconds per problem against 172). Subgoal-labelled worked examples halve the rate of failing or missing exams and cut score variance — a lever on the tail, not on the mean, and their own founding paper is a documented non-replication in CS specifically.
- Demand an assessment the vendor did not write. That one question is worth more than the entire curriculum-comparison debate in this subject.
Refuse the transfer claim, explicitly and in writing.
- Far transfer against active control groups is g = 0.15. Two randomised trials using validated, independently developed computational-thinking instruments returned d ≈ 0.05 and d = −0.02 — while the same trials moved coding skill by 0.67 and 0.68. A 46-school class-randomised French trial that substituted Scratch for content-matched mathematics teaching lost 0.16 to 0.21 SD on every notion tested. (evidence: coding & cognitive transfer)
What to refuse to spend money or time on
- Any coding product sold as a thinking, reasoning, creativity or executive-function programme. Logo, Scratch, robotics, unplugged, computational thinking, twenty-first-century skills — same claim, same failure, four decades apart.
- Programming time taken out of mathematics. The one trial that tested exactly that substitution at scale found significant negative effects on all three mathematical notions it measured.
- Any curriculum whose evidence is a computational-thinking score. CT tests correlate r = .44-.67 with reasoning, spatial ability and problem solving in schoolchildren, and r = .56 with fluid intelligence in adults who have never programmed. A large share of that score is ability, not learning.
- Effect sizes above about 0.5 from this literature, and anything from the early-childhood robotics meta-analyses (k = 10, I² = 81%, prediction interval touching zero).
- The geek-gene story. Only 5.8% of 778 CS grade distributions are multimodal, worldwide CS1 pass rates are 67.7%, and failure rates (28%) are better than US college algebra's (42-50%).
The measurement problem, which is this subject's defining feature
Computing education has one serious language-independent assessment line — the FCS1, then the SCS1, then the SCS1Rv2 — and it has needed psychometric repair at every generation. The version the community actually adopted reports an internal consistency around 0.59, below the 0.65 threshold its own authors set. Fifteen years after the first instrument, its authors were still writing retrospectives about why nobody uses it.
The consequence is structural: the overwhelming majority of published effect sizes in this subject rest on instruments the intervention's own designers wrote. Apply the archive's standing 2× correction and "g = 0.81 for programming instruction" translates to something near 0.4 — which is approximately the one randomised number the field has.
And the pattern holds wherever it can be checked. Every time the measuring instrument and the intervention stop sharing an author, the effect shrinks or vanishes:
| Same claim, different measurement | Effect |
|---|---|
| Programming interventions, meta-analytic, mostly developer-built measures | g = 0.81 |
| Coding skill, cluster RCT, developer-built instrument | g = 0.47-0.68 |
| Computational thinking, cluster RCT, developer-built instrument (same trial) | d = −0.02 |
| Computational thinking, pupil-randomised RCT, independent validated instrument | d ≈ 0.05 |
| Far transfer, meta-analytic, untreated controls | g = 0.64 |
| Far transfer, meta-analytic, active controls | g = 0.15 |
| Mathematics, class-randomised, content-matched comparison | −0.16 to −0.21 |
The hereditarian bottom line for a founder
Two things, and they point in opposite directions from what a founder might expect.
- The strongest innatist claim in this field is false, and not for the reason people think. The "geek gene" — that programming aptitude splits students into two populations — fails on its own terms: 5.8% of 778 grade distributions are multimodal, and a randomised priming experiment shows the belief predicts the perception of bimodality. Be precise about what that establishes. Individual differences in a heritable trait produce a normal distribution, not a bimodal one, so a unimodal distribution is not evidence against heritable variation in programming aptitude. What fails is the two-populations story, which was never what heritability predicted. The archive's premise is untouched; the folk belief that borrows its prestige is not.
- The correlational case for coding-as-thinking is the archive's premise in action. Computational-thinking test scores correlate .67 with a problem-solving battery, .44 each with spatial and reasoning ability, and .56 with fluid intelligence in adults who have never programmed — while crystallised intelligence is unrelated. "Programmers think better" is ability selection into programming, and a gain on a researcher-made CT test is partly a reasoning score. Substitute a validated independent instrument in a randomised trial and the gain goes to zero, twice.
No genetically informative study exists anywhere in this subject — no twin, adoption or sibling design on programming skill or on response to programming instruction. That is a clean gap.
Confidence
| Decision | Verdict | Confidence |
|---|---|---|
| Teaching programming to produce programming skill | moderate-support | medium |
| Choosing one instructional approach over another | (barely matters) | medium |
| Block-based before text-based environments | (no pooled support) | medium |
| Programming as a route to general thinking or achievement | no-effect | medium |
Open questions a founder should watch
- There is no preregistered, multi-site, independently evaluated RCT of a school computing curriculum anywhere in the world. The subject's best evidence is one developer-led cluster trial.
- Nothing measures retention. Zero of 105 studies in the transfer meta-analysis report a delayed follow-up; the instructional trials measure at end of treatment. Every headline in this subject is acquisition-phase.
- England's 2014 compulsory computing curriculum has never been used as a natural experiment, despite being a clean national shock with administrative data on both sides. It would answer the transfer question and the attainment question at once.
- The block-to-text transition — the design decision every school CS programme actually faces — has one recent systematic review, and it could not be obtained.
- Coding bootcamps have no credible causal evaluation at all; the searchable literature is alumni surveys.
Evidence topics
- Does teaching programming actually teach programming — and does the method matter?moderate supportconf: mediumgc: low
- Does learning to code improve general thinking?no effectconf: mediumgc: low