The Evidence on Teaching

The Impact of Test Preparation on Performance of Large-Scale Educational Tests: A Meta-analysis of Experimental Studies

Hao Z, Baird J, El Masri Y, Double K · 2025

grade Cmeta-analysisindependentreplicated
Sample
28 experimental and quasi-experimental studies
Population
Test-takers on large-scale educational tests; school-age and above, international
Design
The modern successor to the 1983-84 Kulik-group metas, restricted to experimental and quasi-experimental designs rather than pooling uncontrolled pre-post studies. Still a meta-analysis of mixed-quality primaries, so C. Its main value to this archive is that it re-estimates a forty-year-old parameter with contemporary standards and lands in the same place, which is a replication in the sense that matters.
Key findings
Pooled effect of test preparation on the coached test is g = 0.26 (95% CI 0.10 to 0.42) - statistically distinguishable from zero, and almost exactly the 0.25 SD that Bangert-Drowns, Kulik & Kulik reported in 1983. Forty-two years of research have not moved the number. The moderator pattern is the practically useful part and it is not what the test-prep industry sells: workbooks, socio-affective strategy training and explicit test-taking-skills teaching all moderated the effect, while there was LITTLE EVIDENCE that drilling sample items and practice tests alone helps. The authors attribute the effect to gains in domain knowledge or in test-specific cognitive skill, which is precisely the ambiguity this topic exists to keep open. But the meta-analysis contains its own answer to that ambiguity and it is the most decision-relevant number in the paper: across the studies that measured whether preparing for one test improved performance on a DIFFERENT test in the same domain, the pooled effect was g = 0.06 (SE 0.08, p = .502) - indistinguishable from zero. Prepared performance does not travel. The numeric moderator table says the same thing from the other side and reverses the intuitive story: teaching TEST-TAKING SKILLS gave g = 0.35 against 0.05 where it was absent, commercial workbooks 0.40 against 0.15, and NARROW (0.32) and SPECIFIC (0.33) preparation beat BROAD preparation, which came in at g = -0.04 (n.s.); programmes developing broad knowledge or skill gave 0.17 (p = .069, n.s.) against 0.31 for those that did not. College admission tests moved least of all test types (g = 0.14, p = .156, n.s.), and mathematics least of the subject domains (g = 0.10). The narrower and more test-shaped the preparation, the larger the score gain - which is the signature of score inflation, not of learning.
Genetic confound
Low to medium. Restriction to experimental and quasi-experimental designs removes most of the self-selection that contaminates the observational coaching literature.
Replication notes
Replicates the 0.25 SD achievement-test coaching effect established in 1983 on a largely non-overlapping and much more recent study pool, with modern inclusion criteria.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Test preparation effect on the prepared-for testHedges g0.26 (95% CI 0.10-0.42)standardizedpost-preparation administrationbusiness-as-usualend-of-treatmentdomain-skill
Effect of drilling sample items and practice tests alonemoderatorlittle evidence of benefitstandardizedpost-preparation administrationbusiness-as-usualend-of-treatmentdomain-skill
Moderators associated with larger effectsmoderator setworkbooks; socio-affective strategy training; explicit test-taking-skills teachingstandardizedpost-preparation administrationbusiness-as-usualend-of-treatmentdomain-skill
Transfer of the preparation effect to a DIFFERENT test in the same domainHedges g0.06 (SE 0.08, p = .502) - not significant; 5 effect sizes from 4 studiesstandardizedsecond, unprepared-for testbusiness-as-usualend-of-treatmentnear-transfer
Preparation teaching test-taking skills vs preparation that did notHedges g by subgroup0.35 (k=62) with test-taking-skills teaching vs 0.05 (k=25, n.s.) withoutstandardizedpost-preparation administrationbusiness-as-usualend-of-treatmentdomain-skill
Breadth of the preparationHedges g by subgroupnarrow 0.32, specific 0.33, broad -0.04 (n.s.); broad knowledge/skill development 0.17 (p = .069, n.s.) vs 0.31 withoutstandardizedpost-preparation administrationbusiness-as-usualend-of-treatmentdomain-skill
College admission tests specificallyHedges g by subgroup0.14 (p = .156, not significant), against 0.41 for other test typesstandardizedpost-preparation administrationbusiness-as-usualend-of-treatmentdomain-skill

Cited by