The Evidence on Teaching

Unfounded authority, underpowered studies, and non-transparent reporting perpetuate the Mozart effect myth: a multiverse meta-analysis

Oberleiter S, Pietschnig J · 2023

grade Bmeta-analysisindependentfailednumbers spot-checked
Sample
k = 8 studies, N = 207 (26 studies met inclusion criteria; the rest were dropped for insufficient reporting and author non-response)
Population
Patients with epilepsy (k = 6), stroke (k = 1) and other-reported premature-infant pain (k = 1); children and adults; primary-study samples of 11 to 70. This is the modern revival of the Mozart claim, not the original spatial-reasoning version.
Design
Preregistered (OSF) multiverse meta-analysis of the second-generation Mozart claim (Mozart's sonata KV448 for epilepsy) that has dominated recent popular coverage, with PRISMA screening of 1,573 titles and Newcastle-Ottawa quality assessment. Three separate random-effects syntheses were run rather than one pooled model, because the designs are not commensurable: independent MO (KV448 vs silence, independent-groups pre-post, k = 3), dependent MO (KV448 vs silence, one-group pre-post, k = 6), and OM (any other music vs no stimulus, one-group pre-post, k = 5). Bias handled ten ways, plus leave-one-out, specification-curve and exhaustive combinatorial (GOSH) analyses. Grade caveat for anyone recomputing: what it aggregates is weak - only 3 of 8 studies were RCTs, 4 were mirror or counterbalanced designs run WITHOUT washout periods (the authors flag those effects as possible overestimates from carry-over), and 1 was a bare one-group pre-post. Maximum single-study power was 15% in the dependent MO-condition and 6.5% in the OM-condition. Under METHODOLOGY's grid this reads closer to a grade-C meta of mixed-quality designs than to grade B; the grade is left as recorded, but the underlying design mix is now on the record so the judgement can be recomputed. The reporting failure is itself a finding: eligible studies had to be dropped because authors would not supply data, and one author reported the data were no longer accessible for six of his own team's published studies.
Key findings
Non-significant trivial-to-small summary effects across all three independent analyses (g = 0.431, 0.158 and 0.088), none of which reaches significance, and the one non-trivial estimate is a single-study artefact: dropping the premature-infant-pain study from the independent MO-condition flips the sign to g = -0.171. The two RCTs with objectively operationalised outcomes showed a trivial positive effect on epileptic discharges (g = 0.096) and a NEGATIVE effect on stroke patients' blood pressure (g = -0.610). Listening to any other music produced essentially the same (non-)effect as KV448, which is the direct test of a Mozart-specific mechanism and it fails. Bias analyses flag the literature: trim-and-fill and selection models indicate bias, the excess-significance test finds more published significant effects than the observed power can support, and p-curve indicates the dependent MO-condition has no evidential value at all. The multiverse is honest about the exceptions - 2 of 48 specifications across the dependent MO and OM conditions were nominally significant, and larger effects tracked higher heterogeneity, meaning single uncharacteristic studies drive every spectacular result. The named drivers of the myth are unfounded authority, underpowered studies and non-transparent reporting.
Genetic confound
Low; experimental exposure contrasts.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
KV448 vs silence, independent-groups pre-post designs (independent MO-condition)Hedges g0.431 (SE 0.676, p = 0.524, 95% CI -0.89 to 1.76, k = 3, I-squared = 89.4%)standardizedpost-exposurebusiness-as-usualend-of-treatmenthealth
KV448 vs silence, same analysis with the leverage point removed (leave-one-out)Hedges g-0.171 (SE 0.342, p = 0.617, 95% CI -0.84 to 0.50) after removing the premature-infant-pain study - the summary effect reverses sign, so the only non-trivial estimate in the paper rests on one non-epilepsy studystandardizedpost-exposurebusiness-as-usualend-of-treatmenthealth
KV448 vs silence, one-group pre-post designs (dependent MO-condition)Hedges g0.158 (SE 0.103, p = 0.127, 95% CI -0.04 to 0.36, k = 6, I-squared < 0.001%); trim-and-fill and selection models both indicate bias, and p-curve indicates the data have no evidential value; max single-study power 15%standardizedpost-exposurebusiness-as-usualend-of-treatmenthealth
Any OTHER music vs no stimulus, one-group pre-post (OM-condition) - the Mozart-specificity testHedges g0.088 (SE 0.160, p = 0.581, 95% CI -0.23 to 0.40, k = 5); statistically indistinguishable from the KV448 conditions, so there is no evidence the effect is specific to Mozart; max single-study power 6.5%standardizedpost-exposureactive-alternativeend-of-treatmenthealth
RCT-only evidence (3 of the 8 studies)Hedges gepileptic discharges 0.096 (SE 0.38, trivial positive); stroke systolic blood pressure -0.610 (SE 0.52, negative); premature-infant pain 1.655 (SE 0.27, the outlier driving every non-trivial summary)standardizedpost-exposurebusiness-as-usualend-of-treatmenthealth
Multiverse spread - specifications where an effect survivesHedges g rangespecification curve g = 0.40 to 1.30 (independent MO, mostly non-significant with wide CIs), 0.08 to 0.62 (dependent MO) and -0.10 to 0.48 (OM), with only 2 of 48 specifications across the latter two nominally significant. Exhaustive combinatorial analysis: -0.61 to 1.65 (independent MO; -0.60 to 0.09 once the outlier study is excluded), 0.04 to 1.04 (dependent MO), -0.21 to 0.74 (OM). Larger effects always tracked higher heterogeneity.standardizedpost-exposurebusiness-as-usualend-of-treatmenthealth

Cited by