Near and far transfer: Is music special?
Bigand E, Tillmann B · 2022
grade Ccritiqueindependentmixed
Sample
Reanalysis of Sala & Gobet's (2020) own dataset - 54 studies, 254 effect sizes, N = 6,984 - using their published OSF R script
Population
Same as the reanalysed meta-analysis; children aged 3-16
Design
The strongest adversarial check available on the archive's preferred deflationary source, and it deserves weight because it is the good kind of critique: the authors downloaded Sala & Gobet's data and script from OSF and re-ran them, so the disagreement cannot be about the numbers. Graded C - it is a reanalysis-commentary, not new data, and the authors are music-cognition researchers arguing for their own field, which is the direction of bias to expect.
Key findings
Three specific charges. (1) The randomisation moderator is fragile: it was non-significant in Sala & Gobet's main model (p = .518) and became significant (p = .042) only at a sensitivity step where - per lines 737-739 of their own script - the analysis switched from one multi-moderator model to three separate single-moderator models. Re-specified consistently it gives p = .08. The influential-case procedure also flags nine effect sizes rather than the five removed; dropping all nine leaves g = 0.203, p < .0001, with tau-squared = 0 and randomisation non-significant (p = .194). (2) The active-control comparison is not equidistant: music-related outcomes were stripped from the music groups but the equivalent was not done for control groups, so a phonologically trained control tested on phonological awareness is being scored on NEAR transfer while the music group is scored on FAR transfer using the same test - which, as the authors put it, "leads to underestimation of the effect of musical training." (3) Post-test-only studies were treated as having zero baseline difference. Re-running with those 21 mismatched effect sizes removed and post-test-only studies excluded gives g = 0.234 [0.141, 0.327] with tau-squared = 0.017, and neither randomisation nor control type is a significant moderator. Their most pointed sub-analysis: on the 21 excluded effect sizes, music's FAR transfer versus language training's NEAR transfer on the same tests is g = -0.126 [-0.350, 0.099], p = .17 - statistically indistinguishable.
Genetic confound
Not addressed. The reanalysis inherits whatever selection remains in the underlying studies; the disagreement is about effect-size coding, not about confounding.
Replication notes
Sala and Gobet have published no reply. A search of all 75 papers citing this one found none authored by either, and their 2023 restatement (Gobet & Sala, Perspectives on Psychological Science) does not mention Bigand & Tillmann or Román-Caballero. That restatement gives second-order figures for music in typically developing children of g = 0.19 naive and g = -0.02 under active controls with bias correction, i.e. the position is unchanged. The near/far coding charge therefore stands unanswered rather than refuted, and readers should treat it as an open methodological dispute.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Sala & Gobet's dataset reanalysed with mismatched near-transfer control outcomes and post-test-only studies removed | Hedges g (robust variance estimation) | 0.234 (SE 0.046, 95% CI 0.141-0.327), p < .0001, m = 41, k = 190, tau-squared = 0.017, I2 = 13.03%; intermediate steps 0.208 [0.121, 0.295] and 0.243 [0.136, 0.345] | standardized | end of training | unclear | end-of-treatment | far-transfer |
| Music training's far transfer versus language training's near transfer on the same tests | Hedges g | -0.126 (SE 0.071, 95% CI -0.350 to 0.099), p = .17, I2 = 0, Bayes factor 0.447 - the two are indistinguishable, which is the authors' central point about the active-control comparison | standardized | end of training | active-alternative | end-of-treatment | far-transfer |
| Re-specification of the randomisation moderator | moderator p-value | p = .08 under a consistently specified three-moderator model versus the p = .042 reported by Sala & Gobet after switching to three single-moderator models; p = .194 after removing all nine flagged influential cases | standardized | not-applicable | unclear | not-applicable | far-transfer |
Cited by
- Does learning music make children smarter or better at school?mixedconf: mediumgc: medium