The Evidence on Teaching

Improving the Targeting of Treatment

Scott-Clayton J, Crosta PM, Belfield CR · 2014

grade Bquasi-experimentindependentunreplicated
Sample
two systems: a large urban community-college system (~70,000 first-time degree-seekers, 2004-07 cohorts; analysis n = 37,813 maths and 34,697 English) and a statewide system of 50+ colleges (~49,000 students, 2 cohorts)
Population
US community-college entrants placed by standardized placement examination (COMPASS and similar), late adolescence into adulthood
Design
A predictive-validity simulation on very large administrative datasets, not an experiment in placement policy - it asks what would have happened under alternative screening rules given observed outcomes, and the "severe misplacement" definition depends on a predicted-grades model (predicted B-or-better versus predicted to fail). Graded B for the scale, the two independent systems, and the fact that the outcome is real subsequent course performance rather than another test score. The age range sits above this archive's K-12 core, so the numbers transfer as a mechanism rather than as a parameter.
Key findings
A single cut-score-based placement test misplaces roughly a quarter to a third of the students it sorts. About one in four maths test-takers and one in three English test-takers are SEVERELY mis-assigned. The errors are strongly asymmetric: UNDER-placement (held back into remediation despite a predicted B or better) is two to six times more common than over-placement. In the urban system, 18.5% of maths test-takers and 28.9% of English test-takers were placed into remediation despite predicted B-or-better performance. Substituting high-school transcript information for the test, holding the remediation rate constant, removes four to eight severe misplacements per hundred students - up to a 30% reduction - and raises the success rate among those placed into college-level work by about 10 percentage points (76% to 89% C-or-better in one system). Adding the test on top of transcript data adds little. The general lesson for an intake tool: a single sitting is a weak placement instrument, an accumulated performance record beats it, and the errors it makes are systematically the errors of holding capable students back.
Genetic confound
Low relevance for the misplacement estimate, which is about the instrument's accuracy against an observed criterion rather than about differences between students.
Replication notes
The two systems in the paper agree with each other, but no independent replication on a different national context was identified. The finding drove real policy change (the discontinuation of COMPASS), which is corroboration of a different kind.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Severe misplacement rate under test-score-cutoff placement%~25% of maths test-takers, ~33% of English test-takersstandardizedfirst-term placementnoneunder-1yrattainment
Direction of placement errorratiounder-placement 2-6x more common than over-placementstandardizedfirst-term placementnoneunder-1yrattainment
Students remediated despite predicted B-or-better performance%18.5% of maths test-takers, 28.9% of English test-takers (urban system)standardizedfirst-term placementnoneunder-1yrattainment
Using high-school transcript instead of the placement testreduction in severe errors4-8 fewer severe misplacements per 100 tested (up to ~30% reduction); success rate among college-level placements up ~10pp (76% to 89% C-or-better)standardizedfirst-term placement and subsequent course outcomeactive-alternativeunder-1yrattainment
Marginal value of the test on top of transcript dataincremental benefitlittlestandardizedfirst-term placementactive-alternativeunder-1yrattainment

Cited by