The Evidence on Teaching

Accountability, incentives and behavior: the impact of high-stakes testing in the Chicago Public Schools

Jacob BA · 2005

grade Bnatural-experimentindependentreplicated
Sample
student-level panel data for Chicago Public Schools grades 3, 6 and 8 around the 1996-97 accountability policy, with comparison series from other midwestern cities
Population
US urban public elementary and middle school students
Design
Interrupted time series against the pre-policy trend and against other midwestern cities, on student-level administrative data, with a built-in audit measure - the high-stakes ITBS on one side and the low-stakes state IGAP on the other. That audit structure is what earns grade B: the design does not merely observe a score rise, it observes whether the same children rose on a test nobody was preparing for. Limitation, stated by the author: IGAP is not perfectly parallel to ITBS in content or format, so a modest share of the divergence could be format rather than inflation.
Key findings
By 2000 the high-stakes ITBS was up about 0.30 SD in maths and 0.20 SD in reading relative to the pre-policy trend, with the largest gains in the lowest-achieving schools (up to ~0.35 maths and 0.25 reading). On the low-stakes IGAP there was NO comparable jump - the overall change was slightly negative, and in the highest-performing schools IGAP fell about 0.14 SD in reading and 0.13 in maths. The item-level analysis names the mechanism rather than inferring it: the biggest ITBS gains appeared on item TYPES that are easy to coach and relatively more common on the ITBS than on the IGAP, and on END-OF-TEST items conditional on difficulty, which is a pacing-and-effort signature rather than a knowledge signature. Teachers also responded strategically - more special-education placements, preemptive retention, and time shifted away from science and social studies.
Genetic confound
Low. A policy shock applied to a whole district's cohorts, compared against its own prior trend and against other cities; the treated and untreated cohorts do not differ genetically.
Replication notes
The pattern - large gains on the incentivised test, little or none on an audit measure of the same domain - has been reproduced in Texas (Klein et al. 2000), Kentucky (Koretz & Barron 1998) and in Koretz's re-administration studies. This is one of the better-replicated findings in education policy.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
High-stakes ITBS maths, relative to pre-policy trendSD+0.30 SD by 2000 (up to +0.35 in lowest-achieving schools)standardized4 years post-policybusiness-as-usualover-2yrdomain-skill
High-stakes ITBS reading, relative to pre-policy trendSD+0.20 SD by 2000 (up to +0.25 in lowest-achieving schools)standardized4 years post-policybusiness-as-usualover-2yrdomain-skill
Low-stakes state IGAP, same students and periodSDno comparable jump; slightly negative overall; top schools fell ~0.14 SD reading and 0.13 mathsstandardized4 years post-policybusiness-as-usualover-2yrnear-transfer
Where the ITBS gains were concentrated (item-level)mechanismon easily coached item types more common on ITBS than IGAP, and on end-of-test items conditional on difficulty (pacing and effort)standardizeditem levelbusiness-as-usualover-2yrdomain-skill
Strategic responses by schoolsbehaviourincreased special-education placement, preemptive grade retention, substitution away from science and social studiesstandardizedpost-policybusiness-as-usualover-2yrbehaviour

Cited by