The Evidence on Teaching

Effects of Coaching Programs on Achievement Test Performance

Bangert-Drowns RL, Kulik JA, Kulik CC · 1983

grade Cmeta-analysisindependentreplicated
Sample
30 controlled studies of coaching programmes
Population
School-age students taking ACHIEVEMENT tests (not admissions tests)
Design
Meta-analysis of controlled studies, so every included estimate has a comparison group - better than the SAT coaching literature on that dimension. Graded C because it pools primary studies of mixed quality from before modern reporting standards, and because the moderator analysis is meta-regression on study features rather than randomised comparison of programme types.
Key findings
The anchor source for this archive, because it is about achievement tests and school-age children rather than about college admissions. Coaching raised achievement test scores by 0.25 SD in the typical study. The dose-response structure is what makes it actionable: effects were SMALLEST for short test-taking orientation sessions, LARGER for extensive programmes of drill and practice, and LARGEST in a single lengthy programme designed to improve broad cognitive skills - and effects were directly related to the number of contact hours. Read together with the outcome measure (a standardized achievement test, not a researcher-made one), this says a quarter of a standard deviation of a child's achievement score is purchasable with enough hours of practice on the test format. That is roughly the size of a year of schooling at some ages, obtained without teaching the subject. The authors' own restatement of the same data puts it in units a parent reads: about 0.27 SD, roughly 4 IQ-scale points, roughly 2.5 grade-equivalent MONTHS. Two cautions the archive keeps: 22 of the 30 studies sit in the low-yield short-orientation tier, and the 0.66 top tier rests on a single study, so the gradient should be read as intervention intensity rather than as a clean hours dose-response - the same authors' companion analysis of aptitude tests found no significant duration effect and failed to replicate Messick & Jungeblut's hours relation.
Genetic confound
Low to medium. Controlled studies with comparison groups; the residual concern is differential attrition and volunteer effects in the primary trials rather than genetic selection.
Replication notes
The 0.25 SD figure has held up across the adjacent practice-effect and test-taking-skills literatures (Kulik, Kulik & Bangert 1984; Callenbach 1973) and is consistent in magnitude with the SAT coaching estimates once scale differences are accounted for.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Achievement test score after coaching, typical studySD0.25 SDstandardizedpost-coaching administrationbusiness-as-usualend-of-treatmentdomain-skill
Effect by intervention intensitySD by tiershort narrow test-taking orientation 0.17 (22 of the 30 studies); intensive drill/cramming 0.43; broad cognitive skills 0.66 (SINGLE study - the top of the gradient is fragile)standardizedpost-coaching administrationbusiness-as-usualend-of-treatmentdomain-skill
Effect by study designSDpretest-posttest designs 0.32; posttest-only designs lowerstandardizedpost-coaching administrationbusiness-as-usualend-of-treatmentdomain-skill
Relation of effect to programme contact hoursassociationdirectly related to number of contact hoursstandardizedpost-coaching administrationbusiness-as-usualend-of-treatmentdomain-skill

Cited by