The Evidence on Teaching

Does learning music make children smarter or better at school?

Not at school: the largest randomised trials find nothing on reading, maths or cognition. A small effect on laboratory executive-function tasks is contested, may be real at ~0.2 SD, and is end-of-treatment only.

mixedconf: mediumgc: medium

music · ages 416

Effect summary

Split the claim in three. ACADEMIC ACHIEVEMENT: null in every large randomised test. The EEF's First Thing Music trial randomised 3,004 five-year-olds to daily music for a school year and got reading g = 0.07 [−0.02, 0.16], p = .13, with the free-school-meal subgroup at exactly 0.00. A 2,914-child lottery evaluation of El Sistema found no cognitive or academic effects. Elpus's national US analysis takes the raw +37-point SAT advantage of music students to +0.23 once prior achievement enters. COGNITION: contested. Sala & Gobet's 54-study meta gives g = 0.056 [−0.069, 0.182] under active controls, falling to −0.021 with zero heterogeneity after bias correction; Cooper's independent meta falls from 0.28 to 0.08 [−0.04, 0.20] on quality adjustment. Against that, Román-Caballero's instrumental-only meta gives gΔ = 0.26 [0.13, 0.39] that survives five bias corrections, and Jamey's 8 active-control RCTs give g = 0.60 [0.39, 0.82] on inhibition control with I² = 0. MATHS is null in every synthesis. SELECTION is measured, not assumed: instrument engagement is 78% heritable, self-selecting children are already 0.29 SD ahead before training starts, and the music–IQ link survives controlling family income but dies controlling music aptitude (training's unique variance: 0.63%).

Practical takeaway

Do not fund music on the promise that it raises test scores. Two randomised trials with more than 5,900 children between them, on the exact outcome a school board cares about, found nothing — and the free-school-meal subgroup where the benefit is always promised came in at exactly zero. The live scientific question is narrower and less useful to a budget: whether instrumental training moves laboratory executive function by ~0.2–0.6 SD at end of treatment. That question is genuinely open, two independent teams locate the surviving effect in the same place (executive function), and it deserves watching — but it is not a school-board argument, because nobody has shown it persists, generalises, or shows up on a test.

Who this applies to

Group size
small-groupone-to-onewhole-class
Delivered by
teachertutor
Ages studied
414(narrower than the 416 this topic is filed under — outside it is extrapolation)
Dose
Where a positive is claimed, programmes averaged 17 months and ~53 hours of instrumental training; the executive-function RCTs are shorter. Training duration did NOT moderate the effect in the two largest meta-analyses (hours p = .266, weeks p = .952; intensity p = .69), so there is no dose at which the effect is known to become reliable.
Cost
high
Moves
far-transfer
Needs first
If any effect exists it is for learning to play an INSTRUMENT. Classroom music-appreciation, Kodály/Orff-style whole-class programmes and computerised music training are where the randomised nulls concentrate.
Not for
Mathematics — null in every meta-analysis (0.23 [−0.41, 0.87]; 0.17 [−0.02, 0.36]; 0.06 with active controls) and in the Houston school RCT. Reading and literacy — null in the two largest randomised trials. Disadvantaged pupils specifically: the free-school-meal subgroup of the largest music RCT is 0.00 [−0.20, 0.19]. And nothing at all is known about persistence: no study in the cluster measures an effect after training stops.

Verdict

Mixed — and the mixedness is not a hedge, it is the actual state of a literature where two methodologically serious camps run the same numbers and disagree. But the mixedness lives in a much smaller place than the claim does, so state the parts separately.

  1. Listening to Mozart. Dead, and filed under misc-fads. Mozart d = 0.37 against silence, any other music d = 0.38, Mozart versus other music d = 0.15, with lab affiliation as the moderator that finishes it. There is an arousal effect of music. There is no Mozart effect.
  2. Music training and school achievement. Null, and this is the part a school board is actually being asked about. Every large randomised test comes back empty, and every observational association collapses when prior achievement enters.
  3. Music training and laboratory cognition. Genuinely contested. Depending on whose coding decisions you accept, the answer is g ≈ 0.00 or g ≈ 0.26, and on the narrow construct of inhibition control it may be g ≈ 0.60. This is the live question, and it is not the question the advocacy is asking.

What the evidence shows

The randomised trials on achievement

Source Design Grade Key effect
EEF First Thing Music 2021 cluster RCT, 3,004 pupils / 121 classes / 64 schools, daily music for a year B Reading 0.07 [−0.02, 0.16], p = .13. FSM subgroup 0.00 [−0.20, 0.19]
Alemán 2017 (El Sistema) lottery, 2,914 children, 16 centres A No cognitive or academic effects; self-control +0.095, behavioural difficulties −0.081
Schellenberg 2004 individual RCT, 4 arms, 36 weeks, n = 132 B IQ +2.7 points, d = 0.35; no reliable difference on standardized achievement; drama arm won on social behaviour, d = 0.57
Mehr 2013 two RCTs, n = 29 and 45 C Exp 1 significant in both directions; Exp 2 replication attempt null on everything

First Thing Music is the trial that should have settled it and did not settle it favourably. Three thousand five-year-olds, fifteen minutes of specialist-designed music every school day for a year, independent evaluators, a published standardised reading test — and one month of progress that the evaluators decline to call real. The free-school-meal subgroup, where advocacy always locates the benefit, came in at exactly 0.00.

Schellenberg 2004 deserves better than either side gives it. It is a genuine, well-run, active-controlled randomised trial, its 2.7-point IQ difference is real, and it has never been directly replicated. Two details are almost never quoted. First, it found no reliable difference on the standardized achievement test — the IQ effect did not show up in school learning. Second, the drama arm produced a larger effect on adaptive social behaviour (d = 0.57) than music produced on IQ. And Schellenberg's own proposed mechanism is not music at all: music lessons "are like school but still enjoyable," and he expects chess, science or reading programmes to do the same.

The meta-analytic war

Source Design Grade Key effect
Sala & Gobet 2020 multilevel meta, 54 studies / 254 effects / N = 6,984 B Naive 0.184; active controls 0.056 [−0.069, 0.182]; bias-corrected −0.021, τ² = 0
Sala & Gobet 2017 meta, 38 studies / 118 effects B Randomised + active control: d = −0.12 [−0.27, 0.03]. Literacy −0.07
Román-Caballero 2022 meta of instrumental training only, 34 samples / N = 5,998 C gΔ = 0.26 [0.13, 0.39], survives 4 of 5 bias corrections; randomised+active 0.23 [−0.13, 0.58]
Bigand & Tillmann 2022 reanalysis of Sala & Gobet's own OSF script C g = 0.234 [0.141, 0.327] after equalising near/far coding
Jamey 2024 meta, 22 studies, 8 active-control RCTs C Inhibition control g = 0.60 [0.39, 0.82], I² = 0.0%; design rigour increases the effect
Cooper 2020 independent meta, 21 studies / 100 effects C 0.28 naive → 0.08 [−0.04, 0.20] after quality adjustment
Gordon 2015 meta, 13 samples / n = 901 C Phonological awareness 0.20 [0.04, 0.36]; reading fluency null

The disagreement is not about the data. Bigand and Tillmann downloaded Sala and Gobet's OSF script and re-ran it. Three coding decisions do all the work: whether the randomisation moderator survives a consistently specified model (they say p = .08, not .042); whether active-control groups were held to the same far-transfer standard as music groups (they say no — a phonologically trained control tested on phonological awareness is being scored on near transfer); and whether pooling instrumental tuition with Kodály/Orff classroom programmes masks a real instrumental effect (Román-Caballero says yes, and reports that 73% of Sala and Gobet's high-quality randomised studies were non-instrumental, giving gΔ = 0.11, p = .197). Sala and Gobet have published no reply to either critique. That is worth recording as a fact about the field.

Two things nonetheless survive on the sceptical side. Even inside Román-Caballero's own paper, the cell that is both randomised and active-controlled gives 0.23 with a confidence interval crossing zero, and the selection-model correction restricted to randomised studies gives 0.15, p = .207. And Cooper, working independently of both camps, reproduces the collapse: 0.28 to 0.08.

One thing survives on the positive side and the archive should not bury it. Jamey et al. pooled eight randomised trials with non-musical active controls on a single theory-motivated construct and found g = 0.60 [0.39, 0.82] with zero between-study heterogeneity — and their design moderators run the opposite way to Sala and Gobet's, with RCT design adding +0.19 and active controls adding +0.13. Román-Caballero independently found executive function to be the strongest domain cell (gΔ = 0.41). Two independent teams locating the surviving effect in the same place is the best argument that something is there. Caveats: eight studies, N in the hundreds, a laboratory construct with no demonstrated school consequence, no follow-up, and the authors benchmark themselves against Hattie's d = 0.40 hinge, which this archive rejects outright.

The selection story, measured rather than assumed

Source Design Grade Key effect
Gustavson 2021 twin + adoption, n = 1,684 B Instrument engagement a² = 0.78, c² ≈ 0; r_g with age-12 IQ = 0.44–0.80
Elpus 2013 national transcript + ELS:2002, school FE C SAT gap +37.3 → +25.6 → +0.23 as covariates enter; choral students −10 points
Swaminathan 2017 cross-sectional, n = 133 (+ n = 91 child replication) C Survives SES control (pr = 0.26); dies controlling music aptitude; training's unique variance 0.63%
Guhn 2020 population admin data, n = 112,916 C Adjusted d = 0.22–0.34 (41–53% attenuation); no parental education in the model
Foster 2017 propensity weighting, PSID C "Selection into arts education is at least as strong as any direct effect"
Mosing 2014 MZ co-twin control, 10,539 twins B Practice gaps to 20,228 hours, no within-pair ability difference

Elpus is the cleanest demonstration in education research of a claim dying one covariate at a time. Music students really do score 37 SAT points higher. Add demographics: 25.6. Add prior academic achievement: 0.23. Add time use and attitudes: negative. Four years of music, at a looser threshold, buys 0.04 SD — less than the College Board's own measurement band. His conclusion is that the advocacy argument should be "seriously questioned, or even retired."

Swaminathan identifies which confound it is, and it is not the one everyone controls for. The music–IQ association survives controlling maternal education. It dies controlling music aptitude: training's unique contribution is 0.63% of variance against rhythm's 7.14%. Combined with Gustavson's a² = 0.78 for instrument engagement and its genetic correlation with verbal IQ, the mechanism is not "rich families buy lessons." It is that musical aptitude and general ability share causes, and children who have both are the children who take lessons.

Guhn is the honest counter-case and should not be dismissed. With 112,916 students, provincial exams and Grade 7 prior achievement controlled, the adjusted advantage is still d = 0.22–0.34, concentrated in instrumental rather than vocal music, with a dose gradient. It is the single best observational result in the subject. It also has no parental education in the model, no household SES, no conscientiousness, and no openness — the authors say all of this themselves — and Schellenberg and Lima's 2024 review reads it straightforwardly as selection.

Hereditarian-lens assessment

Risk: medium, and this topic is where the lens earns its keep, because the confound is measured rather than posited.

  • Instrument engagement is 78% heritable [0.55, 0.94] with shared environment indistinguishable from zero — adoptive siblings in the same musical household do not converge. Whether a child takes up an instrument is among the most heritable "choices" in this archive.
  • It is genetically correlated with IQ (r_g = 0.44 in twins, 0.80 in siblings). The same genes push a child toward the instrument and toward the test score.
  • The self-selection gap is 0.29 SD before the first lesson. Román-Caballero measured it directly: in studies where children chose music, the music group was already ahead at baseline by g_pre = 0.29 [0.12, 0.47], while randomised studies showed g_pre = 0.00. That pre-existing gap is roughly the size of the entire claimed training effect.
  • Practice does not cause the ability. In MZ pairs discordant by up to 20,228 practice hours, the twin who practised more was no better at music discrimination.

Two scoping corrections the evidence forces, and both cut against lazy hereditarianism:

  1. The confound is aptitude, not income. Schellenberg's own data show family income and parental education do not attenuate the music–IQ link. Repeating "it's just SES" is wrong.
  2. Gustavson's residual is real. Instrument engagement at 12 still predicted verbal ability at 16 after controlling age-12 full-scale IQ (β = 0.09 [0.04, 0.15]). Small, but the authors call it evidence of weak direct benefit, and honesty requires reporting it.

Boundaries & what critics say

  • This is not an argument against music. Teaching music teaches music at d ≈ 0.5 (music instruction). That is the case, and it is sufficient.
  • The randomised literature is underpowered, and the sceptics concede it. Román-Caballero compute that 188 per group would be needed for 80% power at the effect they report; typical arms hold ~25. Sala and Gobet's own null cell has df = 4.2. "Well-powered causal evidence finds nothing" is true of achievement and not true of cognition, which is why the verdict is mixed and not no-effect.
  • Instrumental versus classroom music may be the real fault line. Román-Caballero's reanalysis and Guhn's instrumental-versus-vocal contrast point the same way. If that distinction holds, the EEF's whole-class daily singing null and a small instrumental-tuition positive are both true.
  • Randomised children do not practise. Schellenberg's randomly assigned children practised 10–15 minutes a week — nothing like a child who chose the instrument. Randomisation buys internal validity by destroying the treatment.
  • Persistence is completely unmeasured. Not one estimate in this entire topic has a follow-up after training stops. Given fadeout, that is not a small gap.
  • The non-cognitive results are the ones that keep showing up. El Sistema moved self-control and behavioural difficulties, concentrated in boys exposed to violence; the Houston arts RCT moved discipline; Schellenberg's drama arm moved social behaviour. Whether structured group arts activity does something behavioural, for vulnerable children specifically, is a better-supported question than the cognitive one — and it is a question about structured group activity, not about music.

Practical guidance

  • Never take the transfer argument into a budget meeting. The two largest randomised trials in the field are null on exactly the outcome that gets measured, and the disadvantaged-pupil subgroup is 0.00. If you win the argument on transfer, you lose the programme the year someone checks.
  • Discount any music-and-achievement statistic that lacks prior achievement as a control. Elpus shows that single covariate takes the SAT gap from +37 points to zero.
  • If you want executive function, that is not yet a purchase decision. The evidence is eight small trials on one laboratory construct, with no persistence data and no demonstrated school consequence. Watch it; do not budget on it.
  • Instrumental tuition is where any effect lives. Whole-class music appreciation and branded classroom-music systems are where the randomised nulls concentrate.
  • If the goal is behaviour and self-control in vulnerable children, El Sistema-type provision has the best evidence in this topic (+0.20 SD self-control, −0.24 aggression for boys exposed to violence) — and the authors' own reading is that it works partly because better-off families can already buy alternatives.

Open questions

  • Whether the instrumental/non-instrumental distinction really explains the meta-analytic disagreement is untested by anyone outside the disputing parties.
  • Whether the inhibition-control effect replicates in a large preregistered trial with an active control is the single most valuable experiment left in this literature. One such trial is registered (NCT05912270) and has not reported.
  • Nothing is known about persistence after training stops — anywhere in the cluster.
  • Whether any cognitive effect, if real, has any consequence a school could observe is untested; every positive lives on a laboratory task.

Evidence (19 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore