The Evidence on Teaching

Measuring Test Measurement Error

Boyd D, Lankford H, Loeb S, Wyckoff J · 2013

grade Cquasi-experimentindependentunreplicated
Sample
New York City student-level longitudinal administrative data, math and ELA state assessments across consecutive grades
Population
US urban public-school students, elementary and middle grades
Design
A method plus an application. The method generalises the test-retest framework so it can be used on consecutive-grade state assessments - allowing for real growth between tests, tests that are neither parallel nor vertically scaled, and error that differs across tests - and requires only descriptive statistics. The application is New York City. The estimate is of a measurement parameter rather than a treatment effect, so the causal-design rubric does not fit; C reflects a well-identified estimate on large administrative data. Limitation: one district, and the decomposition rests on testable but unverified assumptions about the structure of achievement growth.
Key findings
The single most useful number for anyone reading a child's score: the overall extent of test measurement error is MORE THAN TWICE as large as the figure the test vendor reports. Vendor reliability is split-test reliability, which captures only inconsistency within one sitting; it cannot see the day-to-day variation in how a child performs, which is a large part of the real error. Every confidence band printed on a score report, and every rule of thumb about what counts as a meaningful difference between two subtests, is therefore built on an error estimate that is less than half the true one.
Genetic confound
Not applicable. Estimating measurement error in an instrument; no trait attribution and no causal claim about children.
Replication notes
The method has been taken up in the value-added literature but the archive found no independent re-estimation of the "more than twice the vendor figure" result on another testing programme, which is why it is recorded as unreplicated rather than replicated.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Total test measurement error vs vendor-reported reliabilityratiomore than 2x the measurement error reported by the test vendorstandardizedconsecutive-grade state assessmentsnonenot-applicabledomain-skill
Source of the understatementmechanismvendor split-test reliability omits day-to-day variation in student performancestandardizednot applicablenonenot-applicabledomain-skill

Cited by