Does teaching programming actually teach programming — and does the method matter?
Teaching coding teaches coding: the one clean school RCT gives g=0.47-0.68. Which approach you pick barely matters, and almost every effect size rests on an instrument the developers built.
moderate supportconf: mediumgc: lowcomputer-science · ages 5–18
Two claims, different strengths. THAT IT WORKS: a 13-school cluster RCT of a 12-week K-2 curriculum moved coding skill g=0.47 against business-as-usual; a pupil-randomised English trial moved coding skill d=0.67 over a year; a second cluster RCT moved it d=0.68. Four meta-analyses converge on g=0.5-0.8 (programming per se g=0.81, CI 0.42-1.21; robotics 0.535-0.539; early childhood d=0.73). WHICH METHOD: much weaker. The anchor meta reports effect sizes 'differed only marginally between the instructional approaches and conditions'. Block-based versus text-based is a null with detected publication bias (g=0.245, CI -0.078 to 0.567, p=.137). Pair programming gives 0.41-0.64, but the assignment estimate is jointly produced work. Subgoal labels give d=0.44 on formative quizzes, a non-significant d=0.20-0.22 on exam components, and their founding paper is a documented non-replication in CS specifically. Parsons problems are an efficiency win only — same post-test and same one-week retention in roughly half the time (84s vs 172s per problem). Base rates: worldwide CS1 pass rate 67.7%, failure 28% against 42-50% for college algebra, and only 5.8% of 778 grade distributions are multimodal — there is no geek gene in the data, only in instructors' perception of it.
Teach programming because programming is worth knowing, budget about half a standard deviation on a coding assessment for a twelve-week course, and stop agonising over the method — the anchor meta-analysis says approaches differ only marginally, and the one design principle the whole K-12 movement rests on (blocks before text) is a meta-analytic null. Spend the effort you would have spent choosing a curriculum on choosing an assessment you did not build: in this literature the difference between a developer's instrument and an independent one is the difference between the headline and the finding. And drop the geek-gene story — it is not in the grade distributions, only in the instructors.
Who this applies to
Verdict
moderate-support, and it is the honest justification for computer science in school — the one this
subject should be sold on instead of the transfer claim it usually leans on.
Three things are true at once and the verdict has to hold all of them.
- The domain-skill claim survives. One genuine cluster RCT moved coding skill g = 0.47 against business-as-usual; a second moved it d = 0.68; a pupil-randomised English trial moved it d ≈ 0.67 over a full academic year. Four meta-analyses converge on roughly g = 0.5-0.8. Under this archive's premise that is exactly what should happen: nobody programs without being taught, so mean instructional effects on a taught skill ought to be real and large.
- Nothing reaches grade A. There is no preregistered multi-site RCT, no independently replicated RCT and no lottery design anywhere in this literature. The single best trial is developer-led, uses developer-built instruments, and was downgraded by the What Works Clearinghouse for attrition.
- The instructional-approach question comes back much weaker than the does-it-work question. If someone asks which method to use, the honest answer from this literature is: it does not much matter — teach it seriously and measure it with something you did not build.
What the evidence shows
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Yang, Yang & Bers 2023 | cluster RCT, 13 schools, K-2 | B | Coding g = 0.47; computational thinking 0.04; literacy −0.01 |
| Scherer 2020 | meta, 139 interventions / 375 ES | C | Programming per se g = 0.81 (0.42-1.21); approaches "differed only marginally" |
| Xu 2019 | meta, blocks vs text, k = 13 | C | g = 0.245 (−0.078 to 0.567), p = .137, publication bias detected |
| Umapathy & Ritzhaupt 2017 | meta, pair programming, N = 3,308 | C | 0.41-0.64 on assignments, exams, pass rates; nothing affective |
| Margulieux 2020 (semester) | quasi-experiment, CS1, n = 265 | C | Quizzes d = 0.44; exams 0.26 (components ns); failing/missing exams halved |
| Morrison 2015 | lab RCT | D | The non-replication — given-label gains failed in CS while replicating in maths and science |
| Ericson 2017 | 3-arm RCT, n = 135, 1-wk retention | C | Same learning, half the time; no difference at post-test or one week |
| Sun & Zhou 2022 | meta, robotics, 36 studies | C | 0.539 / 0.535 — largest at kindergarten, at 1-5 weeks, at n < 50 |
| Alonso-García 2024 | meta, k = 10, open access only | D | SMD 0.93, I² = 81%, prediction interval −0.045 to 1.908 |
| Liu, Conrad & Blazar 2024 | natural experiment, 635,771 students | B | CS major +10.2pp, CS degree +5.5pp, earnings +8% at 24 |
| Patitsas 2016 | 778 grade distributions + priming experiment | C | 5.8% multimodal; primed professors see bimodality that is not there |
| Watson & Li 2014 | review, 161 courses / 15 countries | C | CS1 pass rate 67.7%; no change over time, no effect of language |
| Parker, Guzdial & Engleman 2016 | psychometric replication | D | The field's flagship instrument: α ≈ 0.59, below the authors' own threshold |
The measurement problem is not a caveat; it is the headline. Computing education's only serious language-independent instrument line — FCS1, then SCS1, then SCS1Rv2 — has needed psychometric repair at every generation, and the version the community actually adopted reported an alpha around 0.59, below the 0.65 threshold its own authors set. That means the overwhelming majority of the effect sizes in this file rest on instruments the intervention's own designers wrote. Apply the archive's standing correction — researcher-designed measures inflate roughly 2× against independent standardized ones — and the honest translation of "g = 0.81 for programming instruction" is something near 0.4, which is approximately the one randomised number this field has.
Every time the measuring instrument and the intervention stop sharing an author, the effect shrinks. Blocks versus text, the one contrast metaed by people who built neither environment: null with publication bias. The CAL trial's independently scored outcomes: null. The federal What Works Clearinghouse reviewed that trial and reviewed only the literacy and computational-thinking domains — the coding result, which carries the entire positive case, sits outside the independent review.
The subgoal literature contains its own replication failure, in its founding paper. Morrison, Margulieux & Guzdial set out to test given against generated subgoal labels and report that the given-label gains did not replicate as expected in the introductory CS task, despite replicating in mathematics and science. The honest reading of the whole line is that subgoal labelling is a real but small and fragile effect on near-term formative performance, whose best-documented benefit is reducing catastrophic failure rather than raising the mean — exam variance fell and the rate of failing or missing exams roughly halved.
Parsons problems are the one clean, useful, internally replicated lever, and it is a lever on time. Same post-test, same one-week retention, roughly half the practice time. The authors flag a ceiling effect on their post-tests, which is exactly the condition under which a no-difference result is least informative, so treat "same learning" as "no detectable loss" rather than as equivalence.
Two base rates deflate the field's own framing. Worldwide CS1 pass rates are 67.7% and failure rates 28%, against 42-50% for US college algebra — programming is not uniquely brutal. And CS grades are not bimodal in 94% of the 778 distributions where anyone tested it. The second finding comes with an experiment: professors randomly primed with the bimodality folklore, and professors who believed in innate CS aptitude, were more likely to label ambiguous histograms bimodal. The field's most durable belief about who can learn to program is a confirmation loop.
The strongest attainment evidence is not about instruction at all. A staggered-rollout instrumental-variable design on 635,771 Maryland students finds that taking a high-quality high-school CS course raises the probability of declaring a CS major by 10.2 points, of earning a CS degree by 5.5 points, and annual earnings at 24 by about 8%. Two honest caveats: the first stage is modest (F 13.0-16.6) and the authors concede it may bias the estimate upward, and the gains come partly from reallocation within STEM — CS pulled students out of other STEM (10-13 points) and engineering (5-9 points), not out of non-STEM.
Hereditarian-lens assessment
Risk: low. The anchor evidence is randomised at pupil, class or school level, and the Maryland attainment result is identified off the timing of a school's course offering — none of which any family chose.
The lens does real work in one place, though, and it cuts against a hereditarian folk belief rather than for it. The "geek gene" — the claim that programming aptitude splits people into two kinds — is the strongest innatist claim in this field, and it fails on its own terms: only 5.8% of 778 CS grade distributions are multimodal, and the belief predicts the perception of bimodality under experimental priming. Be precise about what that does and does not establish. A unimodal grade distribution is not evidence against heritable variation in programming aptitude; individual differences in a heritable trait produce a normal distribution, not a bimodal one. What fails is the specific two-populations claim, which was never what heritability predicts. The archive's premise is untouched, and the folk belief that borrows its prestige is not.
No genetically informative study exists anywhere in this literature — no twin, adoption or sibling design on programming skill or on response to programming instruction. That is a clean gap, and it is the same gap the archive flags for trainability in motor skill.
Boundaries & what critics say
- The design distribution is bad. One real cluster RCT (developer-led, WWC-with-reservations), one credible natural experiment (which measures earnings, not skill), one randomised-block school study that discarded its randomisation to report dose-response associations, and otherwise small unregistered lab experiments and single-institution section comparisons. The meta-analyses inherit exactly this.
- The robotics meta-analyses should be read as hypothesis-generating. Sun & Zhou's moderator table — biggest effects at the shortest durations, the smallest samples, and dependent on the choice of measurement scale — is a portrait of small-study bias with the labels rearranged. Alonso-García's k = 10 open-access-only search returns SMD 0.93 with a prediction interval touching zero.
- The most awkward result in the subject is that unplugged activities beat screen-based ones in the early-childhood meta-analysis. If the mechanism is learning to program a computer, the arm that does not involve programming a computer should not win.
- Vendor involvement is real and should be named. Code.org's own research lead co-authors an evaluation of a Code.org unit; the CAL curriculum, its app, its coding assessment and its computational-thinking assessment were all built by one group; the TIPP&SEE strategy, curriculum and assessment likewise. None of this makes the findings false. It means every positive estimate here carries the archive's standard developer discount, and the nulls — including the developer-led ones — carry none.
- Half the instructional evidence is undergraduate. Pair programming, Parsons problems and subgoal labelling are all CS1 findings. Whether they transfer to 11-year-olds is untested.
- A whole question is unanswered. The block-to-text transition — the design decision every school CS programme actually faces — has one recent systematic review, and it could not be obtained.
Practical guidance
- Justify computer science on programming, employment and earnings, not on thinking. That is what the evidence supports, and the Maryland natural experiment gives it a defensible economic case.
- Budget about 0.4-0.5 SD on a coding assessment from a twelve-week course, and treat any figure above about 0.5 from this literature as a measurement artefact until shown otherwise.
- Do not agonise over blocks versus text. The pooled contrast is a null with detected publication bias. Pick one and teach it.
- Use Parsons problems for practice volume. Same measured learning in roughly half the time is a real lever, and it is the one instructional finding here with internal replication.
- Use subgoal-labelled worked examples if your problem is the failing tail, not if your problem is the mean. That is what the variance and withdrawal results actually show.
- Demand an assessment the vendor did not write. In this literature that single question is worth more than the entire curriculum-comparison debate.
Open questions
- There is no preregistered, multi-site, independently evaluated RCT of a school computing curriculum anywhere. The field's best evidence is one developer-led cluster trial.
- No credible design exists for coding bootcamps — the searchable literature is alumni surveys and interviews.
- England's 2014 compulsory computing curriculum has never been exploited as a natural experiment, despite being an unusually clean national shock with pre- and post-administrative data.
- Nobody has measured retention of programming skill beyond one week except at end of treatment. Every headline here is acquisition-phase.
- No genetically informative design exists on programming aptitude or on response to programming instruction.
- grade CA meta-analysis of teaching and learning computer programming: Effective instructional approaches and conditionsScherer, R., Siddiq, F., & Sánchez Viveros, B. · 2020 · meta-analysis
- grade BThe efficacy of a computer science curriculum for early childhood: evidence from a randomized controlled trial in K-2 classroomsYang, D., Yang, Z., & Bers, M. U. · 2023 · rct
- grade CA Meta-Analysis of Pair-Programming in Computer Programming CoursesUmapathy, K., & Ritzhaupt, A. D. · 2017 · meta-analysis
- grade CBlock-based versus text-based programming environments on novice student learning outcomes: a meta-analysis studyXu, Z., Ritzhaupt, A. D., Tian, F., & Umapathy, K. · 2019 · meta-analysis
- grade CEvidence That Computer Science Grades Are Not BimodalPatitsas, E., Berlin, J., Craig, M., & Easterbrook, S. · 2016 · critique
- grade CReducing withdrawal and failure rates in introductory programming with subgoal labeled worked examplesMargulieux, L. E., Morrison, B. B., & Decker, A. · 2020 · quasi-experiment
- grade DSubgoals, Context, and Worked Examples in Learning Computing Problem SolvingMorrison, B. B., Margulieux, L. E., & Guzdial, M. · 2015 · rct
- grade DEffect of Implementing Subgoals in Code.org's Intro to Programming Unit in Computer Science PrinciplesMargulieux, L. E., Morrison, B. B., Franke, B., & Ramilison, H. · 2020 · quasi-experiment
- grade CSolving parsons problems versus fixing and writing codeEricson, B. J., Margulieux, L. E., & Rick, J. · 2017 · rct
- grade CEvaluating the Efficiency and Effectiveness of Adaptive Parsons ProblemsEricson, B. J., Foley, J. D., & Rick, J. · 2018 · rct
- grade DA Review of Research on Parsons ProblemsDu, Y., Luxton-Reilly, A., & Denny, P. · 2020 · review
- grade DTIPP&SEESalac, J., Thomas, C., Butler, C., Sanchez, A., & Franklin, D. · 2020 · quasi-experiment
- grade CEffective instruction conditions for educational robotics to develop programming ability of K-12 students: A meta-analysisSun, L., & Zhou, D. · 2022 · meta-analysis
- grade DEnhancing computational thinking in early childhood education with educational robotics: A meta-analysisAlonso-García, S., Rodríguez Fuentes, A. V., Ramos Navas-Parejo, M., & Victoria-Maldonado, J. J. · 2024 · meta-analysis
- grade CThe effects of programming interventions in early childhood: A meta-analysisSimonsmeier, B. A., Kampmann, K., Staub, J., & Scherer, R. · 2025 · meta-analysis
- grade DA systematic review of learning computational thinking through Scratch in K-9Zhang, L., & Nouri, J. · 2019 · review
- grade DA theory of instruction for introductory programming skillsXie, B., Loksa, D., Nelson, G. L., Davidson, M. J., Dong, D., Kwik, H., Tan, A. H., Hwa, L., Li, M., & Ko, A. J. · 2019 · quasi-experiment
- grade BComputer Science for All? The Impact of High School Computer Science Courses on College Majors and EarningsLiu, J., Conrad, C., & Blazar, D. · 2024 · natural-experiment
- grade CFailure rates in introductory programming revisitedWatson, C., & Li, F. W. B. · 2014 · review
- grade DReplication, Validation, and Use of a Language Independent CS1 Knowledge AssessmentParker, M. C., Guzdial, M., & Engleman, S. · 2016 · replication
Related decisions
- Does learning to code improve general thinking?no effectconf: mediumgc: low
- Talent and trainability — what is heritable, and what that does not licensestrong supportconf: highgc: low