The Evidence on Teaching

Meta-Analytic Findings of the Self-Controlled Motor Learning Literature: Underpowered, Biased, and Lacking Evidential Value

McKay, B., Yantha, Z. D., Hussien, J., Carter, M. J., & Ste-Marie, D. M. · 2022

grade Bmeta-analysisindependentmixednumbers spot-checked
Sample
k=52 experiments, N=2,061 (delayed-retention analysis)
Population
k=52 studies, N=2,061; predominantly healthy young adults in laboratory settings
Design
Preregistered meta-analysis of self-controlled (learner chooses a practice/feedback parameter) vs yoked-control experiments with formal selection-bias correction (Vevea-Hedges weight-function model, one-tailed p-cutpoint .025) plus p-curve, PEESE and z-curve evidential-value analysis. The PRIMARY analysis is DELAYED RETENTION only (k=52, N=2,061); transfer was a separate, pre-planned exploratory analysis the authors themselves declare unreliable. Two extreme outliers (g=3.7, g=3.95) removed before all analyses. Naive-model heterogeneity Q(51)=103.45, p<.0001, I2=47.9%. The weight-function estimate is a FIXED-effect estimate: weightr failed to fit the random-effects version. Published in Meta-Psychology (open, methods-focused).
Key findings
The popular OPTIMAL-theory claim that giving learners autonomy over feedback improves learning shrinks ~75% under selection-bias correction: retention g=.44 becomes g=.107, "small and not currently distinguishable from zero". Three independent corrections converge on trivial effects (weight-function .107, PEESE .05, p-curve .04; the latter two cannot reject the null). The single significant moderator was PUBLICATION STATUS (R2=48%): published experiments g=.54, unpublished g=.003. Transfer behaves the same way — naive g=.52 falls to g=.17 (p=.24) under correction. Advice to 'let learners choose when they get feedback' rests on a near-null effect. Note the authors' own caution: their model-performance bounds leave a benefit anywhere from g=-.11 to .26 (plausible upper 95% limit .33) on the table, so this is "insufficient evidence that the effect exists", not a demonstrated zero.
Genetic confound
Not applicable (randomized allocation); no modeling of who benefits, so aptitude-treatment interaction is invisible.
Replication notes
CONFLICT: Wang, Tao, Yuan & Guo 2025 (Behav Sci, k=29, N=1147, doi 10.3390/bs15091291) report SMD .63 retention / .68 transfer / .20 ns acquisition — but Egger's test was significant and uncorrected, reproducing McKay's naive g.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Self-controlled vs yoked, DELAYED RETENTION, naive pooled estimateHedges g (naive random-effects)g = 0.44 [0.31, 0.56]; k=52, N=2,061; Q(51)=103.45, p<.0001, I2=47.9%researcher-designeddelayed retention; lab motor tasksactive-alternativeend-of-treatmentdomain-skill
Same retention effect after selection-bias correction (weight-function model)bias-corrected Hedges gg = 0.107 [0.047, 0.18] (results text [.05, .17]) — 'not currently distinguishable from zero'; adjusted model fits better, chi2(1)=21.18, p<.0001; non-significant results only 6% as likely to survive selectionresearcher-designeddelayed retention; lab motor tasksactive-alternativeend-of-treatmentdomain-skill
Retention effect by PUBLICATION STATUS (the only significant moderator, R2=48%, p<.0001)Hedges g by subgrouppublished g = 0.54 [0.28, 0.81]; unpublished g = 0.003 [-0.23, 0.24] — the entire effect lives in the published literatureresearcher-designeddelayed retention; lab motor tasksactive-alternativeend-of-treatmentdomain-skill
Self-controlled vs yoked, DELAYED TRANSFER (separate exploratory analysis)Hedges g, naive then bias-correctednaive g = 0.52; bias-corrected g = 0.17, p = 0.24 (non-significant); selection model again fit better than naive, p=.008researcher-designeddelayed transfer; lab motor tasksactive-alternativeend-of-treatmentnear-transfer
Retention effect under alternative bias corrections (sensitivity)PEESE and p-curve effect estimatesPEESE g = 0.05; p-curve g = 0.035 (reported as ~.04); neither can reject the null — all three correction methods converge on trivially small effectsresearcher-designeddelayed retention; lab motor tasksactive-alternativeend-of-treatmentdomain-skill
Evidential value / power of the literaturez-curve and p-curve power estimatesz-curve ERR 12% [3%, 34%], EDR 6% [5%, 13%] vs observed discovery rate 48% (both CIs exclude it — significant publication bias); p-curve of published significant results significantly FLATTER than 33% power (p=.0035), estimated power 5% [5%, 17%]; median n per experiment = 36unknownn/a; meta-levelnonenot-applicabledomain-skill
Authors' own bound on what the data can rule out (CUTS AGAINST a flat null reading)plausible range under model-performance assumptionsresults consistent with a true benefit anywhere from g = -0.11 to 0.26, plausible upper 95% limit g = 0.33 — 'this analysis does not rule out the possibility that self-controlled practice provides meaningful motor learning benefits on average'researcher-designeddelayed retention; lab motor tasksactive-alternativeend-of-treatmentdomain-skill

Cited by