The Impact of Test Preparation on Performance of Large-Scale Educational Tests: A Meta-analysis of Experimental Studies
Hao Z, Baird J, El Masri Y, Double K · 2025
grade Cmeta-analysisindependentreplicated
Sample
28 experimental and quasi-experimental studies
Population
Test-takers on large-scale educational tests; school-age and above, international
Design
The modern successor to the 1983-84 Kulik-group metas, restricted to experimental and quasi-experimental designs rather than pooling uncontrolled pre-post studies. Still a meta-analysis of mixed-quality primaries, so C. Its main value to this archive is that it re-estimates a forty-year-old parameter with contemporary standards and lands in the same place, which is a replication in the sense that matters.
Key findings
Pooled effect of test preparation on the coached test is g = 0.26 (95% CI 0.10 to 0.42) - statistically distinguishable from zero, and almost exactly the 0.25 SD that Bangert-Drowns, Kulik & Kulik reported in 1983. Forty-two years of research have not moved the number. The moderator pattern is the practically useful part and it is not what the test-prep industry sells: workbooks, socio-affective strategy training and explicit test-taking-skills teaching all moderated the effect, while there was LITTLE EVIDENCE that drilling sample items and practice tests alone helps. The authors attribute the effect to gains in domain knowledge or in test-specific cognitive skill, which is precisely the ambiguity this topic exists to keep open. But the meta-analysis contains its own answer to that ambiguity and it is the most decision-relevant number in the paper: across the studies that measured whether preparing for one test improved performance on a DIFFERENT test in the same domain, the pooled effect was g = 0.06 (SE 0.08, p = .502) - indistinguishable from zero. Prepared performance does not travel. The numeric moderator table says the same thing from the other side and reverses the intuitive story: teaching TEST-TAKING SKILLS gave g = 0.35 against 0.05 where it was absent, commercial workbooks 0.40 against 0.15, and NARROW (0.32) and SPECIFIC (0.33) preparation beat BROAD preparation, which came in at g = -0.04 (n.s.); programmes developing broad knowledge or skill gave 0.17 (p = .069, n.s.) against 0.31 for those that did not. College admission tests moved least of all test types (g = 0.14, p = .156, n.s.), and mathematics least of the subject domains (g = 0.10). The narrower and more test-shaped the preparation, the larger the score gain - which is the signature of score inflation, not of learning.
Genetic confound
Low to medium. Restriction to experimental and quasi-experimental designs removes most of the self-selection that contaminates the observational coaching literature.
Replication notes
Replicates the 0.25 SD achievement-test coaching effect established in 1983 on a largely non-overlapping and much more recent study pool, with modern inclusion criteria.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Test preparation effect on the prepared-for test | Hedges g | 0.26 (95% CI 0.10-0.42) | standardized | post-preparation administration | business-as-usual | end-of-treatment | domain-skill |
| Effect of drilling sample items and practice tests alone | moderator | little evidence of benefit | standardized | post-preparation administration | business-as-usual | end-of-treatment | domain-skill |
| Moderators associated with larger effects | moderator set | workbooks; socio-affective strategy training; explicit test-taking-skills teaching | standardized | post-preparation administration | business-as-usual | end-of-treatment | domain-skill |
| Transfer of the preparation effect to a DIFFERENT test in the same domain | Hedges g | 0.06 (SE 0.08, p = .502) - not significant; 5 effect sizes from 4 studies | standardized | second, unprepared-for test | business-as-usual | end-of-treatment | near-transfer |
| Preparation teaching test-taking skills vs preparation that did not | Hedges g by subgroup | 0.35 (k=62) with test-taking-skills teaching vs 0.05 (k=25, n.s.) without | standardized | post-preparation administration | business-as-usual | end-of-treatment | domain-skill |
| Breadth of the preparation | Hedges g by subgroup | narrow 0.32, specific 0.33, broad -0.04 (n.s.); broad knowledge/skill development 0.17 (p = .069, n.s.) vs 0.31 without | standardized | post-preparation administration | business-as-usual | end-of-treatment | domain-skill |
| College admission tests specifically | Hedges g by subgroup | 0.14 (p = .156, not significant), against 0.41 for other test types | standardized | post-preparation administration | business-as-usual | end-of-treatment | domain-skill |
Cited by
- Test preparation — does coaching raise scores, and does a raised score mean raised ability?mixedconf: mediumgc: medium