The Evidence on Teaching

The Problem With "Proficiency": Limitations of Statistics and Policy Under No Child Left Behind

Ho AD · 2008

grade Ccritiqueindependentnot-applicable
Sample
large-scale state and national test score distributions
Population
US K-12 state assessment and NAEP score distributions
Design
A statistical demonstration on real large-scale score distributions rather than a conceptual essay, which is why it sits at C rather than D. It is not a causal study and makes no causal claim; what it establishes is that a particular way of summarising scores destroys information in ways that are not recoverable.
Key findings
Collapsing a score distribution to "percent proficient" - or to any single threshold - gives "limited and unrepresentative depictions" of trends, gaps and gap trends, and the distortions are "unpredictable, dramatic, and difficult to correct in the absence of other data." The mechanism generalises well past NCLB and straight onto an individual score report: a threshold statistic responds only to movement near the cut, so it is maximally sensitive to children sitting near the line and blind to everyone else. Ho's prescription is the one an intake tool should adopt - a distribution-wide view is required for any serious analysis of test data, including growth.
Genetic confound
Not applicable - a claim about score summarisation.
Replication notes
A methodological demonstration rather than an effect; the arithmetic is not disputed.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Fidelity of percent-proficient statistics to underlying score distributionsdemonstrationlimited and unrepresentative; distortions unpredictable, dramatic, hard to correctstandardizedcross-sectional and trendnonenot-applicabledomain-skill
Sensitivity of threshold statistics across the distributionmechanismresponds only to movement near the cut score; blind elsewherestandardizednot applicablenonenot-applicabledomain-skill

Cited by