SAT Coaching, Bias and Causal Inference
Briggs DC · 2002
grade Cquasi-experimentindependentreplicated
Sample
NELS:88 F1-F2 panel; coaching analyses on 2,554 test-takers of whom 379 were commercially coached. Literature review covers 32 SAT coaching studies, 1953-2001.
Population
US 11th and 12th graders in the nationally representative National Education Longitudinal Study of 1988, tested 1991-92
Design
A doctoral dissertation (UC Berkeley, 2002; committee included David Freedman and Paul Holland), which the archive treats as a primary source in the same way it treats the Anania and Burke tutoring dissertations. Two designs in one document: an observational estimate of commercial coaching effects from NELS:88 under both linear regression and the Heckman selection model, and a full census of the coaching literature. The methodological finding is as valuable as the substantive one - "small changes in the selection function are shown to have a big impact on estimated coaching effects", i.e. the standard fix for self-selection is itself unstable, which is why this literature's range is so wide.
Key findings
Commercial coaching moves the SAT by about 3 to 20 points on verbal and about 10 to 28 points on maths - ranges, not point estimates, because the estimate depends on modelling choices that the data cannot adjudicate. The literature census supplies the context an intake tool needs: across 32 studies in 48 years the MEDIAN coaching programme was only about 10 hours per section for verbal and 12 for maths, the range ran from 3.5 to 100 hours, and the largest school-based effects came from the longest programmes (30, 52, 68 hours) on the smallest samples. Where observational studies gathered richer covariates - grades, course-taking, socioeconomic status, motivation proxies - the estimated coaching effects got SMALLER, which is the signature of selection bias being progressively removed rather than of a real effect being uncovered.
Genetic confound
Medium-to-high and explicitly analysed. This is the archive's cleanest worked example of the confound in action: coached and uncoached students differ on academic background, family resources and both intrinsic and extrinsic motivation, and each additional covariate block shrinks the estimate. The remaining estimate is an upper bound on the causal effect, not the causal effect.
Replication notes
The NELS:88 estimates are consistent with Powers & Rock's independent national-sample estimates, and the literature census reproduces Becker's moderator findings on a larger study set.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| SAT-Verbal effect of commercial coaching | scaled score points | about 3 to 20 points | standardized | post-coaching administration | business-as-usual | end-of-treatment | domain-skill |
| SAT-Math effect of commercial coaching | scaled score points | about 10 to 28 points | standardized | post-coaching administration | business-as-usual | end-of-treatment | domain-skill |
| Median duration of a coaching programme across 32 studies | contact hours per test section | 10.2 h verbal, 12 h maths (range 3.5-100 h) | standardized | not applicable | none | not-applicable | domain-skill |
| Effect of adding covariates to observational coaching estimates | direction | richer covariate sets (grades, course-taking, SES, motivation) produce SMALLER estimated effects | standardized | not applicable | business-as-usual | not-applicable | domain-skill |
| Stability of Heckman selection-model estimates | sensitivity | small changes in the selection function produce large changes in the estimated effect | standardized | not applicable | unclear | not-applicable | domain-skill |
Cited by
- Test preparation — does coaching raise scores, and does a raised score mean raised ability?mixedconf: mediumgc: medium