Test preparation — does coaching raise scores, and does a raised score mean raised ability?
Coaching buys ~0.25 SD of score and has for 42 years — a fraction of what's advertised — and none of it is ability. The score moves; the construct doesn't.
mixedconf: mediumgc: mediumassessment · ages 5–18 · structure
Two different questions with two different answers. SCORES: yes, a little, and the number has not moved in 42 years — 0.25 SD for achievement-test coaching in 1983, g=0.26 (CI 0.10–0.42) in 2025. On the SAT that is ~8 verbal and ~18 maths points net of selection, against advertised 100–140. Pure retest with no instruction is d=0.26, and 0.42 on an identical form vs 0.23 on a parallel one. ABILITY: no. The same 2025 meta-analysis that finds g=0.26 on the prepared-for test finds g=0.06 (p=.50) on a DIFFERENT test in the same domain, and the effect is LARGEST for narrow, test-shaped preparation (0.32–0.33) and absent for broad preparation (−0.04, n.s.) — the signature of score inflation, not learning. The correlation between a subtest's g-loading and its score gain is ≈ −1.00; a real high-stakes retest gain ELIMINATED the test's criterion validity and did not predict later GPA; Chicago's +0.30 SD on the high-stakes ITBS produced no gain on the low-stakes state test; Kentucky's gains ran 3–4× NAEP's; re-administering an old form dropped scores back to baseline. Even 8 sessions of pure test-taking technique — no reading taught — moved second graders' standardized reading scores ~0.25 SD, still there 4 months later.
Buy a couple of hours of format familiarisation and stop. One or two practice tests under real timing capture most of the available gain; beyond roughly 10–20 hours the returns go log-linear and the money is better spent teaching the subject. Never read a post-preparation score as a measure of the child — that is the one thing the evidence rules out. And if you are the one INTERPRETING a score, ask what preparation preceded it: an unprepared child's first sitting probably understates them by up to a quarter of a standard deviation, and a heavily prepared child's score is measuring partly the preparation.
Who this applies to
Verdict
The topic asks two questions and they have opposite answers, which is why the verdict is mixed rather
than moderate-support or debunked.
Does coaching raise scores? Yes, by a small and remarkably stable amount. Bangert-Drowns, Kulik & Kulik put achievement-test coaching at 0.25 SD in 1983. Hao et al. re-estimated it on a different, modern, experimental-only study pool in 2025 and got g = 0.26 (95% CI 0.10–0.42). Forty-two years of research have not moved the number. On the SAT, net of self-selection, it is roughly 8 verbal and 18 maths points — against the 100–140 combined that commercial companies advertised.
Does the raised score mean raised ability? No, and this is the better-evidenced half. Every design that adds an independent check on the same construct finds the gain fails to arrive there: a different test of the same subject, an audit test nobody prepared for, the old form re-administered, the test's own g-loadings, or the criterion the test was built to predict.
This maps directly onto the archive's outcome taxonomy, which is what the taxonomy is for. A coached
gain is a domain-skill gain on the coached instrument. It is not near-transfer, and it is certainly
not g. The archive already holds the general form of this distinction: schooling durably raises IQ
test scores without raising g (Ritchie, Bates & Deary 2015). Test preparation is the same phenomenon
with a shorter time constant and a much smaller effect.
What the evidence shows
Half one — the score moves, a little
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Hao 2025 | meta of 28 experimental/quasi studies | C | g = 0.26 (0.10–0.42) on the coached test — but g = 0.06 (p = .50) on a DIFFERENT test in the same domain; narrow prep 0.32, specific 0.33, broad −0.04 (n.s.); test-taking-skills teaching 0.35 vs 0.05 without |
| Bangert-Drowns 1983 | meta, 30 controlled studies, achievement tests | C | 0.25 SD (~4 IQ points, ~2.5 GE months). Tiers: orientation 0.17 (22/30 studies), drill 0.43, broad skills 0.66 (one study) |
| Kulik 1984 | meta, 40 studies, practice only, no instruction | C | Identical form +0.42 SD, parallel form +0.23 SD; rises with number of practice tests; larger for high-ability students |
| Hausknecht 2007 | meta, 50 studies / 107 samples / N = 134,436 | C | Retest d = 0.26; with coaching δ = 0.70 vs 0.24 without; identical form 0.46 vs alternate 0.24 |
| Callenbach 1973 | small RCT, 24 second graders | C | 8×30 min of pure test-taking technique, no reading taught → ~0.25 SD on the Stanford Reading Test, still significant at 4 months |
| Powers & Rock 1999 | national sample, n≈4,200, five estimation models | C | ~8 V / ~18 M; uncoached retesters gain 12 V / 16 M on their own; advertised 100–140 |
| Becker 1990 | meta, 48 studies | C | Published controlled studies: 0.09 SD V, 0.16 SD M |
| DerSimonian & Laird 1983 | meta | C | Uncontrolled studies ~40 V / ~50 M; controlled ~15 each — 4–5× inflation from the missing control group |
| Messick & Jungeblut 1981 | ETS meta-analysis | C | Hours↔effect ρ = .77 V / .71 M, log-linear; time for >20–30 points/section "rapidly approaches that of full-time schooling" |
| Briggs 2001 | NELS:88, n = 3,492 | C | 14–15 M / 6–8 V with full controls; ACT effects ≈ nil |
| Briggs 2002 | dissertation: NELS + 32-study census | C | 3–20 V / 10–28 M; median programme 10–12 h; richer covariates → smaller estimates; Heckman estimates unstable to tiny specification changes |
| Montgomery & Lilly 2011 | systematic review, 10 controlled studies | C | The dissenting high estimate: 23.5 V / 32.7 M (~0.21 / 0.30 SD) |
Half two — the ability does not
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Lievens 2007 | 941 real high-stakes admissions candidates, real criterion | B | Retest scores load more on memory, less on g; the gain eliminated the test's criterion validity and did not predict later GPA |
| te Nijenhuis 2007 | meta, 64 retest studies, N = 26,990 | C | Correlation between subtest g-loading and score gain ≈ −1.00; practice and coaching reduce a test's g-loadedness |
| Jacob 2005 | natural experiment, student-level panel, built-in audit test | B | High-stakes ITBS +0.30 SD maths / +0.20 reading; low-stakes IGAP no jump, slightly negative; gains on coachable item types and end-of-test items |
| Koretz 2002 | synthesis of audit-test studies | C | Re-administering the old, unfamiliar form dropped scores back to first-year baseline; Kentucky KIRIS maths gains 3–4× NAEP; HS gains not reflected in ACT at all |
| Klein 2000 | TAAS vs NAEP trend comparison + primary study | C | Grade-4 reading: NAEP 0.13–0.15 SD, TAAS 0.31–0.49 SD; racial gap increasing on NAEP, shrinking fast on TAAS |
| Shepard 1990 | analytic essay | D | Inflation is stale norms and teaching the test — and teaching the test contaminates the norms themselves |
| Ritchie, Bates & Deary 2015 | existing archive source | — | The general case: education raises test scores via specific taught skills, not via g |
| Appelrouth 2015 | excluded — company's own clients, no control arm | D | +22.6 pts/sitting, but the whole uncoached retest effect is inside that number; author owns the company |
| Buchmann 2010 | excluded — same NELS data as Briggs, weaker selection controls | C | +30 course / +37 tutor; books, software and school courses ≈ nil |
The decisive comparison
The cleanest version of the answer is inside the meta-analysis that produced the headline number. Hao et al. report g = 0.26 for preparation measured on the test that was prepared for. Across the studies that also measured a different test in the same domain, the pooled effect is g = 0.06 (SE 0.08, p = .50) — indistinguishable from zero. Same meta-analysis, same studies, same students; move the measurement one test to the left and the effect disappears. That single contrast is the topic's verdict in two numbers, and it is worth more than any amount of argument about selection bias, because it is a within-study comparison.
Lievens et al. is the study that settles the second question, and it is worth stating why it is graded B when nearly everything else here is C. It is not a laboratory manipulation and not a self-report survey. Nine hundred and forty-one real candidates sat a real medical-school entrance examination with real consequences, retook it, and were then followed to an external criterion — subsequent academic performance. Every other study in this literature has to compare one test score to another test score. This one asks the only question that matters: does the improvement predict anything? It does not. The retest gain wiped out the test's criterion-related validity.
Jacob 2005 is the school-age analogue and reaches the same place by a completely different route. Chicago's accountability policy moved the high-stakes ITBS by 0.30 SD in maths. The same children, in the same years, showed no corresponding movement on the low-stakes state exam — slightly negative overall, and in the top schools it went down. The item-level analysis is what turns a correlation into a mechanism: the gains concentrated on item types that are easy to coach and disproportionately common on the ITBS, and on end-of-test items conditional on difficulty, which is a pacing-and-effort signature rather than a knowledge signature.
And Callenbach 1973 is the demonstration in miniature, on the youngest children in the topic. Twenty- four second graders got eight half-hour sessions of content-independent test-taking technique. No reading was taught. Their standardized reading scores went up about a quarter of a standard deviation and were still up four months later. That single small trial is the whole verdict in one design: a standardized score is a performance, and roughly 0.25 SD of a young child's performance is available without touching the skill.
Hereditarian-lens assessment
Risk: medium, concentrated entirely in the observational coaching literature and essentially absent from the transfer literature.
Where the confound bites. Families who buy coaching differ from families who do not on income, parental education, and academic push — every one of which correlates with heritable ability. Briggs's dissertation is the archive's cleanest worked example of this in action: as he adds covariate blocks (grades, course-taking, socioeconomic status, motivation proxies), the estimated coaching effect gets smaller at every step. That monotone shrinkage is the signature of selection bias being progressively removed, not of a real effect being uncovered, and it means the surviving estimate is an upper bound. Powers & Rock attacked the same problem with five different estimators and got a range rather than a point, which is the honest outcome. DerSimonian & Laird quantified the cost of not attacking it at all: studies without a control group report effects four to five times larger.
Where it does not bite, and this is the important half. The transfer failures are not vulnerable to selection. Jacob compares the same children on two tests. Koretz re-administers an old form to the same population. te Nijenhuis correlates a property of subtests with gains on those subtests. Lievens compares first and second attempts by the same candidates against an external criterion. Whatever selects students into coaching cannot explain why the gain shows up on one test and not another taken by the same people — and it certainly cannot explain why a gain fails to predict a criterion. The half of the verdict that matters most for the intake tool is the half the confound cannot touch.
Boundaries & what critics say
- Montgomery & Lilly is the live dissent and is recorded as such, not buried. Their systematic review pools ten controlled studies and lands at 23.5 verbal and 32.7 maths points — roughly double Becker's 0.09/0.16 SD on an overlapping literature. The archive does not get to pick the number it prefers. The honest reading is that the upper bound is contested and the disagreement is about pooling weak primaries, not about direction; and that even at the dissenting estimate the effect is a fifth to a third of a standard deviation, not the transformation being sold.
- The dose-response relation is real but shakier than it is usually quoted. Messick & Jungeblut's ρ ≈ .7–.8 between contact hours and effect is the most cited number in the field, and DerSimonian & Laird criticised the regression while the Kulik group failed to replicate it on a different study pool. It is partly rehabilitated by Hausknecht, who found contact time a significant positive predictor across 134,436 people. Treat the shape — log-linear, sharp early returns, rapid flattening — as established and the coefficients as soft.
- The top of the intensity gradient rests on one study. Bangert-Drowns et al.'s 0.66 for "broad cognitive skills" training is a single study, and 22 of their 30 studies sit in the low-yield 0.17 orientation tier. Quoting 0.66 as what coaching can achieve inverts the actual distribution of the evidence.
- Hao 2025's moderator table is the whole verdict in one place, and it inverts the intuitive story. Preparation that taught test-taking skills returned g = 0.35 against 0.05 where it was absent. Narrow preparation returned 0.32 and specific preparation 0.33, while broad preparation returned −0.04 (n.s.) and programmes aimed at developing broad knowledge or skill returned 0.17 (n.s.) against 0.31 for those that did not. The narrower and more test-shaped the preparation, the larger the score gain. If preparation were teaching the subject, the gradient would run the other way. Note also that college admission tests moved least of all test types (g = 0.14, n.s.) — the tests the entire commercial industry is built around are the ones with the weakest experimental evidence of coachability. The one tension the archive cannot resolve: Hao found little evidence that drilling sample items and practice tests alone helps, which is the opposite of what Kulik's 1984 pure-practice meta-analysis found.
- No post-2005 randomised trial of any name-brand commercial course was located. The modern experimental pool is school- and researcher-delivered preparation. For an industry this large, that absence is itself a finding.
- Appelrouth 2015 is excluded and is a useful specimen. It reports 22.6 points per official sitting from a commercial tutoring company's own client records, analysed by the company's founder, with no control arm — and Powers & Rock already measured the uncoached retest gain at 12–16 points per section. Two of its incidental findings (individual tutoring beats group per hour; spaced sessions beat massed) are consistent with this archive's tutoring and spaced-practice topics, which is why it is recorded rather than ignored.
- This is not an argument that preparation is pointless. Some of it is a genuine equity issue — Buchmann et al. document how sharply access is stratified by income — and a child who has never seen a bubble sheet is being measured partly on that unfamiliarity. The argument is that the ceiling is low, the returns flatten fast, and the resulting score is worth less as a measurement than it was before.
Practical guidance
- Do the cheap part and skip the expensive part. One or two full-length practice tests under real timing conditions capture most of the available gain (Kulik: the first practice trial is worth 0.23–0.42 SD depending on form similarity). The marginal hour after roughly 10–20 is where the log-linear curve flattens.
- If the goal is a higher score, teach the subject. Messick & Jungeblut's conclusion is the correct cost comparison and it has never been overturned: the contact time needed for gains much beyond 20–30 points per section "rapidly approaches that of full-time schooling." At that price, schooling is the better product.
- Never interpret a post-preparation score as a measure of the child. This is the one hard rule. The gain is concentrated in exactly the parts of the test that measure the construct least (te Nijenhuis), and it destroys the score's predictive meaning (Lievens).
- When reading someone else's score, ask what came before it. A test-naive child's first sitting probably understates them by up to a quarter of a standard deviation (Callenbach). A child on their third sitting of the same form is carrying a 0.42 SD practice effect. Both are common in home testing.
- Watch for the high-ability practice effect specifically. Kulik found practice gains were larger for high-ability students, which is the opposite of the usual intuition and matters for any child being screened at the top of a distribution.
- If you are running the test, use a parallel form, not the same one. Identical-form gains run roughly double parallel-form gains in both the 1984 and 2007 meta-analyses. That gap is the part of the "gain" that is purely test-specific, and choosing a parallel form removes it.
- Treat any programme evaluated without a control group as reporting nothing. DerSimonian & Laird's 4–5× inflation factor is the single most useful consumer-protection number in this topic.
Open questions
- Confidence is
medium, capped by the pending adversarial pass rather than by the evidence. Two grade-B sources (Lievens, Jacob) point the same way on the transfer question and would supporthighafter a skeptic review; the score-gain half rests entirely on grade-C meta-analyses of old primaries. - Kulik's pure-practice result and Hao's null on sample-item drilling appear to contradict each other and nobody has reconciled them. One is a 1984 meta of practice-form studies, the other a 2025 meta of experimental preparation studies. If drilling practice items does nothing, the retest effect must be coming from something else — familiarity, anxiety reduction, pacing — and that distinction is practically important.
- The elementary-age evidence is one 1973 trial with 24 children. Callenbach is doing an enormous amount of work in this topic for its size. The obvious modern replication — teach young children test-taking technique, measure a standardized achievement test, add an audit measure — does not exist.
- Nothing measures preparation effects in home-schooled or home-administered testing, which is the delivery context this archive is being built for and the one where practice effects, motivation, and administration conditions are least controlled.
- How long does a coached gain last? Callenbach's held for four months; Koretz's sawtooth implies it evaporates when the form changes. Nobody has followed a coached cohort with an audit measure at one and two years.
- grade CThe Impact of Test Preparation on Performance of Large-Scale Educational Tests: A Meta-analysis of Experimental StudiesHao Z, Baird J, El Masri Y, Double K · 2025 · meta-analysis
- grade CEffects of Coaching Programs on Achievement Test PerformanceBangert-Drowns RL, Kulik JA, Kulik CC · 1983 · meta-analysis
- grade CEffects of Practice on Aptitude and Achievement Test ScoresKulik JA, Kulik CC, Bangert RL · 1984 · meta-analysis
- grade CRetesting in selection: A meta-analysis of coaching and practice effects for tests of cognitive ability.Hausknecht JP, Halpert JA, Di Paolo NT, Moriarty Gerrard MO · 2007 · meta-analysis
- grade CEffects of Coaching on SAT I: Reasoning Test ScoresPowers DE, Rock DA · 1999 · quasi-experiment
- grade CCoaching for the Scholastic Aptitude Test: Further Synthesis and AppraisalBecker BJ · 1990 · meta-analysis
- grade CEvaluating the Effect of Coaching on SAT Scores: A Meta-analysisDerSimonian R, Laird NM · 1983 · meta-analysis
- grade CThe Effect of Admissions Test Preparation: Evidence from NELS:88Briggs DC · 2001 · quasi-experiment
- grade CSystematic reviews of the effects of preparatory courses on university entrance examinations in high school-age studentsMontgomery P, Lilly J · 2011 · meta-analysis
- grade CScore gains on g-loaded tests: No gte Nijenhuis J, van Vianen AEM, van der Flier H · 2007 · meta-analysis
- grade BAn examination of psychometric bias due to retesting on cognitive ability tests in selection settings.Lievens F, Reeve CL, Heggestad ED · 2007 · quasi-experiment
- grade BAccountability, incentives and behavior: the impact of high-stakes testing in the Chicago Public SchoolsJacob BA · 2005 · natural-experiment
- grade CWhat Do Test Scores in Texas Tell Us?Klein SP, Hamilton LS, McCaffrey DF, Stecher BM · 2000 · quasi-experiment
- grade CLimitations in the Use of Achievement Tests as Measures of Educators' ProductivityKoretz DM · 2002 · review
- grade DInflated Test Score Gains: Is the Problem Old Norms or Teaching the Test?Shepard LA · 1990 · critique
- grade BIs Education Associated with Improvements in General Cognitive Ability, or in Specific Skills?Ritchie, S. J., Bates, T. C., & Deary, I. J. · 2015 · longitudinal
- grade BHow Much Does Education Improve Intelligence? A Meta-AnalysisRitchie, S. J., & Tucker-Drob, E. M. · 2018 · meta-analysis