Cognitive and academic benefits of music training with children: A multilevel meta-analysis
Sala G, Gobet F · 2020
grade Bmeta-analysisindependentmixed
Sample
m = 54 studies (1986-2019), k = 254 effect sizes, N = 6,984; 41 studies / 144 effect sizes with non-active controls, 23 studies / 110 effect sizes with active controls
Population
Children aged 3-16, sample mean age 6.45 years (median 5.90). Mean training 53.4 hours (range 2-507) over 29.3 weeks.
Design
The deflationary pole of the music-transfer debate and the most methodologically aggressive synthesis in it: robust variance estimation for effect-size dependence, plus Bayesian analysis with informative priors, plus trim-and-fill and selection models applied separately to each design subgroup. The design principle is the one this archive uses everywhere - grade the estimate by the comparison that produced it, not by the pooled average. Grade B rather than A because it aggregates a mixed-quality literature, and because the headline null rests on a thin cell: 20 studies / 96 effect sizes with df = 4.2, reached after a two-step exclusion sequence. The Bayesian priors are imported from the authors' own earlier cognitive-training work, which builds part of the null into the analysis.
Key findings
Naive pooled effect g = 0.184, SE 0.041, 95% CI [0.101, 0.268], I2 = 43.2%. The whole result is the moderator: with NON-ACTIVE controls g = 0.228 [0.137, 0.320], with ACTIVE controls g = 0.056 [-0.069, 0.182], p = .350. After removing three studies with large baseline imbalance the active-control estimate is g = -0.021 [-0.109, 0.068] with tau-squared = 0 and I2 = 0 - a null with no residual heterogeneity, which is the signature of an absent phenomenon rather than a moderated one. Publication-bias correction pushes the active-control estimate to -0.020 [-0.183, 0.142] and the randomised-studies estimate to -0.034 [-0.131, 0.063]. Two results cut against the advocacy framing in a way worth recording: outcome domain did NOT moderate (all ps >= .362; cognitive vs academic p = .981), and neither did age (p = .403) or training duration (hours p = .266, weeks p = .952) - so there is no "start earlier" or "train longer" dose story inside this data. Note that the paper reports no per-domain pooled effect sizes, only the null moderator tests. Randomisation was NOT a significant moderator in the main model (p = .518); it reached p = .042 only after the second sensitivity step. The authors' verdict: "Once the quality of study design is controlled for, the overall effect of music training programs is null (g ~ 0)... researchers' optimism about the benefits of music training is empirically unjustified and stems from misinterpretation of the empirical data and, possibly, confirmation bias."
Genetic confound
Handled indirectly through design controls: the active-control and randomised subgroups remove selection into music, and the effect vanishes there. The authors read the residual positive in the non-randomised studies as exactly the confound this archive assumes.
Replication notes
Recorded as `mixed` because two independent reanalyses of overlapping data reach the opposite conclusion. Bigand & Tillmann (2022) re-ran Sala & Gobet's own OSF script and argue the randomisation moderator is p = .08 rather than .042 under a consistently specified model, and that active-control groups were not held to the same far-transfer standard as music groups. Román-Caballero et al. (2022) argue the null is driven by non-instrumental classroom-music programmes making up 73% of the high-quality studies. Sala and Gobet have published no reply to either; their 2023 restatement in Perspectives on Psychological Science does not mention either critique. The absence of a reply is itself part of the record.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| All cognitive and academic outcomes, naive pooled | Hedges g (robust variance estimation) | 0.184 (SE 0.041, 95% CI 0.101-0.268), p < .001, tau-squared 0.041, I2 = 43.16% | standardized | end of training | unclear | end-of-treatment | far-transfer |
| Studies with NON-ACTIVE (do-nothing) control groups | Hedges g | 0.228 (95% CI 0.137-0.320), m = 41, k = 144, p < .001 | standardized | end of training | business-as-usual | end-of-treatment | far-transfer |
| Studies with ACTIVE control groups - the load-bearing comparison | Hedges g | 0.056 (95% CI -0.069 to 0.182), m = 23, k = 110, p = .350; after removing three baseline-imbalanced studies, -0.021 (-0.109 to 0.068) with tau-squared = 0 and I2 = 0 | standardized | end of training | active-alternative | end-of-treatment | far-transfer |
| Publication-bias-corrected estimates by design subgroup | Hedges g after trim-and-fill / selection model | active controls -0.020 (-0.183 to 0.142), selection model 0.039; randomised studies 0.009 (L0) and -0.034 (R0), selection model -0.002; non-active controls 0.170 (0.064-0.276), selection model 0.119; non-randomised 0.189-0.211, selection model 0.126 | standardized | not-applicable | unclear | not-applicable | far-transfer |
| Moderators that did NOT explain variance | moderator p-values | outcome domain all ps >= .362 (cognitive vs academic p = .981); participant age p = .403; training duration in hours p = .266, weeks p = .952, sessions p = .662; randomisation p = .518 in the main model. Only baseline difference (p = .031) and control type (p = .035) mattered. | standardized | not-applicable | unclear | not-applicable | far-transfer |
Cited by
- Does learning music make children smarter or better at school?mixedconf: mediumgc: medium