Please don't stop the music: A meta-analysis of the cognitive and academic benefits of instrumental musical training in childhood and adolescence
Román-Caballero R, Vadillo MA, Trainor LJ, Lupiáñez J · 2022
grade Cmeta-analysisindependentmixednumbers spot-checked
Sample
32 studies / 34 independent samples, 176 effect sizes, 5,998 participants (1,664 trained; 4,334 controls, of whom 3,670 passive and only 664 active)
Population
Children and adolescents, mean age 8.0 years (SD 2.2, range 3.9-14.7). Mean programme duration 17 months (SD 16.3, range 0.75-60). 8 samples randomised, 12 non-randomised without self-selection, 14 self-selection.
Design
This is the steelman for music-makes-you-smarter and it has to be taken seriously: it is the one meta-analysis in the cluster that restricts to LEARNING TO PLAY AN INSTRUMENT (excluding Kodály/Orff/Kindermusik classroom music, computerised music training and listening programmes), requires pre-post designs with a control group, searches grey literature, and applies five separate publication-bias procedures. Second author Vadillo is a bias-correction methodologist, which is a strong independence signal in a field full of advocacy. Grade C rather than B because of what it aggregates: only 8 of 34 samples randomised, typical arm ~25 participants, and the authors compute that 188 per group would be needed for 80% power at the effect they report. The paper is explicitly a rebuttal to Sala & Gobet (2020) and includes a reanalysis of their corpus. Third author Trainor directs a music-and-the-mind institute, so the field-level stake sits on the positive side here.
Key findings
Overall gDelta = 0.26 [0.13, 0.39] after removing three implausible outliers, with high residual heterogeneity (I2 = 70.4%). Crucially, the effect did NOT shrink with design quality: randomised samples gave 0.26 and non-randomised 0.25, and blinding did not moderate. Publication-bias procedures mostly left the estimate intact (selection model 0.30; trim-and-fill 0.22-0.26), the one exception being PET (0.01), which the authors argue is known to underestimate under their conditions. The two results that most constrain the claim are inside the paper itself. First, the subset that is BOTH randomised and active-controlled gives gDelta = 0.23 with a confidence interval crossing zero [-0.13, 0.58], and the selection-model correction restricted to randomised studies gives 0.15, p = .207 - so the cleanest design cell cannot exclude nil. Second, in the self-selection studies music groups were already ahead BEFORE the intervention, gpre = 0.29 [0.12, 0.47], while randomised studies showed gpre = 0.00 [-0.14, 0.15]: the selection confound this archive assumes is measured here directly, and it is roughly the same size as the entire claimed training effect. By outcome domain the pattern is not the one advocacy uses: executive functions 0.41 and short-term memory 0.28 are the reliable cells, literacy is 0.20, and MATHEMATICS is null (0.23, CI -0.41 to 0.87), as is intelligence (0.29, CI -0.00 to 0.58) and phonological processing (-0.10). Every effect is end-of-treatment; the meta-analysis contains no follow-up estimate, so persistence is untested.
Genetic confound
Partly addressed and partly the finding. Randomised samples remove genetic confounding by construction and still show 0.26; but the self-selection baseline gap (gpre = 0.29) is a direct measurement of the confound in the non-randomised two thirds of the literature. The authors cite Gustavson et al. (2021) approvingly for the "nature and nurture" reading: instrument engagement is highly heritable and genetically correlated with verbal ability.
Replication notes
Directly contradicts Sala & Gobet (2017, 2020) on an overlapping corpus, which is why replication is recorded as `mixed` rather than `replicated`. The disagreement is not about the raw numbers but about scope: Román-Caballero et al. reanalyse Sala & Gobet's set and report that 73% of their high-quality randomised studies used NON-instrumental music programmes, which yield gDelta = 0.11 (p = .197) in randomised designs and 0.01 with active controls, versus 0.26 and 0.23 for instrumental programmes. If that decomposition holds, the two meta-analyses are compatible: the null belongs to classroom music-appreciation programmes and the small positive to learning an instrument. Independent adjudication of this claim does not yet exist.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Cognitive and academic outcomes pooled, randomised + non-randomised (outliers removed) | Hedges gDelta (standardised difference of mean pre-post change) | 0.26 (95% CI 0.13-0.39), p < .001, I2 = 70.35%; 0.19 (0.10-0.28) when self-selection studies are added | standardized | end of programme (mean 17 months of training) | business-as-usual | end-of-treatment | far-transfer |
| Randomised samples only (m = 8) | Hedges gDelta | 0.26, p = .013 - numerically identical to non-randomised (0.25) and larger than self-selection (0.11), i.e. no design-quality gradient | standardized | end of programme | unclear | end-of-treatment | far-transfer |
| Randomised samples with an ACTIVE control group - the cleanest design cell | Hedges gDelta | 0.23 (95% CI -0.13 to 0.58) - the confidence interval crosses zero; passive-control randomised studies give 0.32 (0.18-0.46) | standardized | end of programme | active-alternative | end-of-treatment | far-transfer |
| Pre-existing advantage of children who CHOSE music, before any training (self-selection studies) | Hedges gpre (baseline difference) | 0.29 (95% CI 0.12-0.47), Bayes factor 7.20 for a real difference; randomised studies gpre = 0.00 (-0.14 to 0.15) and non-randomised-without-selection 0.03 (-0.08 to 0.14) | standardized | pre-test, before the intervention began | none | not-applicable | far-transfer |
| Effect by outcome domain | Hedges gDelta [95% CI], m studies | executive functions 0.41 [0.12, 0.70] (m=7); short-term memory 0.28 [0.15, 0.41] (m=7); literacy 0.20 [0.06, 0.34] (m=10); intelligence 0.29 [-0.00, 0.58] ns (m=9); processing speed 0.26 [-0.06, 0.57] ns (m=5); MATHEMATICS 0.23 [-0.41, 0.87] ns (m=4); phonological processing -0.10 [-0.51, 0.31] (m=1); long-term memory 0.61 [-0.28, 1.5] ns (m=3); visuospatial 0.48 [0.13, 0.83] from a single study | standardized | end of programme | business-as-usual | end-of-treatment | far-transfer |
| Publication-bias corrected estimates | Hedges gDelta after correction | all studies: selection model 0.30 (p = .005), trim-and-fill 0.22-0.26, PEESE 0.19 (ns), PET 0.01 (ns), Mathur-VanderWeele sensitivity 0.23 at eta = 1.5 and 0.13 at eta = 5. RANDOMISED ONLY: selection model 0.15 (p = .207, non-significant), trim-and-fill L0 0.19, PET -0.04 (ns). No bias test was itself significant. | standardized | not-applicable | unclear | not-applicable | far-transfer |
| Reanalysis of Sala & Gobet's corpus, instrumental vs non-instrumental programmes | Hedges gDelta | randomised designs: instrumental 0.26 (p = .013) vs non-instrumental 0.11 (p = .197); active controls: instrumental 0.23 vs non-instrumental 0.01. 73% of Sala & Gobet's randomised studies were non-instrumental. | standardized | end of programme | unclear | end-of-treatment | far-transfer |
Cited by
- Does learning music make children smarter or better at school?mixedconf: mediumgc: medium