Is it really a neuromyth? A meta-analysis of the learning styles matching hypothesis
Clinton-Lisell V, Litzinger C · 2024
grade Cmeta-analysisindependentmixednumbers spot-checked
Sample
21 studies, 101 effect sizes, 1,712 participants (plus 12 studies inspected but not poolable, in Table 4). Crossover tallies are over 42 learning-outcome measures. Only 5 of the 21 met What Works Clearinghouse standards; 2 were quasi-experiments without random assignment and without demonstrated baseline equivalence (Moussa-Inaty et al. 2019; Mujtaba et al. 2022). 6,299 citations screened, 1,810 after de-duplication, 40 full texts.
Population
Mixed ages; predominantly higher-education and adult samples with some school-age (fifth-grade - Chen & Sun 2012, Rogowsky et al. 2020; 15-16 - Riding & Douglas 1993) studies, one sample of trainee pilots and one of medical students. Every included study was a SINGLE SESSION of instruction, so nothing here speaks to durability; a handful of studies did include delayed post-tests (Rassaei 2018, 2019; Mujtaba et al. 2022) but these are pooled in with the immediate ones rather than analysed as a separate horizon.
Design
The strongest published challenge to the debunked verdict, and it must be answered rather than ignored. Robust variance estimation over matched-vs-unmatched contrasts, assumed within-study dependency rho = 0.8 (results stable across rho = 0 to 1: g moves only 0.33 to 0.32). Its critical weakness is that a positive pooled main effect of 'matching' is NOT evidence for meshing - meshing requires a crossover, and the authors report crossover separately as a COUNT of outcome measures, not as a pooled interaction effect size. There is no pooled interaction estimate in this paper; anyone quoting one is inventing it. Study quality is described by the authors themselves as low, and outcome measures are overwhelmingly researcher-designed immediate post-tests (multiple-choice recall of a history lesson, verbatim fill-in-the-blank, vocabulary production, flight-simulator performance, first-attempt IV placement) - the measure type METHODOLOGY.md says inflates effects roughly 2x. The one clear exception is Rogowsky et al. (2020), which used items from a standardized comprehension test. Heterogeneity is enormous and unexplained: tau-squared = 0.77, I-squared = 91.17, and no moderator reached significance - design (within/between) b = 0.49, p = .28; style b = -0.25, p = .35; assessment-modality match b = -0.18, p = .46; study quality b = 0.22, p = .56. That last null matters in both directions: the effect is not demonstrably smaller in the WWC-compliant studies, so "it is only the bad studies" is asserted rather than shown - though with 21 studies the meta-regression is badly underpowered to detect it. PUBLICATION BIAS: the funnel plot is visually symmetrical and Egger's test of the intercept is null (b = -0.058, 95% CI [-0.52, 0.41], p = 0.11). NO trim-and-fill, PET-PEESE or other bias-corrected estimate is reported anywhere in the paper, so there is no corrected number for this archive to prefer; the uncorrected g is all that exists. The authors nonetheless flag a residual concern: all but two of the included reports were journal articles, and only two came from the grey literature where nulls are likelier to surface. Funded by a University of North Dakota professorship (Rose Isabella Kelly Fischer); no conflicts declared; both authors give positionality statements describing training that treated learning styles as a myth, and their framing is hostile to the "completely lacks empirical evidence" characterisation - which makes their negative crossover finding the more telling.
Key findings
Overall benefit of matched instruction g = 0.32 (SE 0.12, 95% CI 0.07 to 0.57, p = .01) in the Results section; the abstract states the same model as g = 0.31, 95% CI 0.05 to 0.57, p = .02. Sensitivity: removing the sole kinesthetic study gives g = 0.34, SE 0.13 [0.08, 0.60], p = .01 (20 studies, 98 effects); removing the two quasi-experiments that failed baseline equivalence gives g = 0.33, SE 0.13 [0.05, 0.61], p = .02 (19 studies, 91 effects) - i.e. the positive main effect survives restriction to the randomized studies. Egger's test and the funnel plot show no publication bias, and no bias-corrected estimate is reported. But the main effect is not the hypothesis. Only 11 of 42 learning-outcome measures (26.19%) showed the crossover interaction the matching hypothesis actually predicts. Among the 12 unpoolable studies the picture is muddled in the source itself: the abstract and the Results text both state 25.00% (= 3 of 12), but Table 4 bolds FOUR studies as showing a crossover (Fajari et al. 2020; Hoffler & Schwartz 2011; Huang 2019; Riding & Ashmore 1980), which is 33%. The count in the Results sentence cannot be read because the typesetting swallowed it into a mangled "Tables 3, 4" cross-reference. Note also that one of the four bolded studies, Huang (2019), is described in the same table as showing a difference that "was not significant and the interaction between styles and condition was not reported" - so the Table 4 crossover tally is generous. Most of the studies that DID show a crossover failed WWC quality standards. Three further findings cut against the main effect and belong in any citation of it: (1) Moser & Zumbach (2015) told students a FAKE learning-style result, and matching to the fake style produced higher scores while matching to their actual measured style produced none - i.e. the effect behaves like expectancy, not aptitude; (2) 85% of the studies dropped participants whose style scores did not permit confident categorisation, so the pooled estimate is over a screened subsample; (3) the authors themselves benchmark g = 0.32 against content-side alternatives that beat it and cost less - the modality effect g = 0.70, decluttering lessons g = 0.33, segmenting g = 0.32, general multimedia principles g = 0.28 - all of which apply one change to every student instead of building two versions of the lesson. The authors' own conclusion is that benefits are "too small and too infrequent to warrant widespread adoption" and that it is "far from conclusive that there is actually a benefit", with the recommendation being multimodal instruction for everyone rather than matching.
Genetic confound
Medium at the level of the pooled main effect: many included studies are small and quasi-experimental, so a matched arm can differ from an unmatched arm in ability composition. The crossover statistic, which is the one that matters, is not vulnerable this way - and it is the one that fails.
Replication notes
No independent replication of the +0.32 pooled estimate exists - this is the first meta-analysis of the matched-vs-unmatched contrast. But it is not unopposed: the only prior pooled estimate of the same contrast, Aslaksen & Loras (2018), was NEGATIVE (g = -0.09 visual, -0.27 auditory) over a partly overlapping literature. The two papers AGREE on the statistic that actually tests the matching hypothesis - the crossover interaction - which fails in both. The disagreement is confined to the pooled main effect, which the authors of this paper themselves say is not the hypothesis. Recorded as `mixed` rather than `unreplicated` because a prior pooled estimate with the opposite sign is a fact about the literature that `unreplicated` would hide.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Matched vs unmatched instruction (main effect) | Hedges g (RVE, rho = 0.8) | g = 0.32 (SE 0.12, 95% CI 0.07 to 0.57), p = .01 in Results; abstract reports the same model as g = 0.31 (95% CI 0.05 to 0.57), p = .02 | researcher-designed | post-instruction | active-alternative | end-of-treatment | domain-skill |
| Matched vs unmatched, two non-randomized quasi-experiments removed (sensitivity) | Hedges g (RVE) | g = 0.33 (SE 0.13, 95% CI 0.05 to 0.61), p = .02, 19 studies / 91 effect sizes - the positive main effect survives restriction to randomized studies | researcher-designed | post-instruction | active-alternative | end-of-treatment | domain-skill |
| Crossover interaction supportive of matching, pooled studies | proportion of learning-outcome measures | 11 of 42 measures (26.19%); most of the studies showing one failed WWC quality standards | researcher-designed | post-instruction | none | end-of-treatment | domain-skill |
| Crossover interaction supportive of matching, 12 unpoolable studies | proportion of studies | source is internally inconsistent - abstract and Results state 25.00% (3 of 12), Table 4 bolds 4 of 12 (33%); one of the four (Huang 2019) had a non-significant, unreported interaction | researcher-designed | post-instruction | none | end-of-treatment | domain-skill |
| Publication bias (funnel asymmetry) | Egger's test of the intercept | b = -0.058 (95% CI -0.52 to 0.41), p = 0.11, funnel plot visually symmetrical; NO trim-and-fill or PET-PEESE correction is reported, so no bias-corrected estimate exists | researcher-designed | not-applicable | none | not-applicable | domain-skill |
Cited by
- Does matching instruction to a child's "learning style" improve learning?debunkedconf: highgc: low