Measuring Test Measurement Error
Boyd D, Lankford H, Loeb S, Wyckoff J · 2013
grade Cquasi-experimentindependentunreplicated
Sample
New York City student-level longitudinal administrative data, math and ELA state assessments across consecutive grades
Population
US urban public-school students, elementary and middle grades
Design
A method plus an application. The method generalises the test-retest framework so it can be used on consecutive-grade state assessments - allowing for real growth between tests, tests that are neither parallel nor vertically scaled, and error that differs across tests - and requires only descriptive statistics. The application is New York City. The estimate is of a measurement parameter rather than a treatment effect, so the causal-design rubric does not fit; C reflects a well-identified estimate on large administrative data. Limitation: one district, and the decomposition rests on testable but unverified assumptions about the structure of achievement growth.
Key findings
The single most useful number for anyone reading a child's score: the overall extent of test measurement error is MORE THAN TWICE as large as the figure the test vendor reports. Vendor reliability is split-test reliability, which captures only inconsistency within one sitting; it cannot see the day-to-day variation in how a child performs, which is a large part of the real error. Every confidence band printed on a score report, and every rule of thumb about what counts as a meaningful difference between two subtests, is therefore built on an error estimate that is less than half the true one.
Genetic confound
Not applicable. Estimating measurement error in an instrument; no trait attribution and no causal claim about children.
Replication notes
The method has been taken up in the value-added literature but the archive found no independent re-estimation of the "more than twice the vendor figure" result on another testing programme, which is why it is recorded as unreplicated rather than replicated.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Total test measurement error vs vendor-reported reliability | ratio | more than 2x the measurement error reported by the test vendor | standardized | consecutive-grade state assessments | none | not-applicable | domain-skill |
| Source of the understatement | mechanism | vendor split-test reliability omits day-to-day variation in student performance | standardized | not applicable | none | not-applicable | domain-skill |
Cited by
- What a standardized achievement score does and does not licensemixedconf: mediumgc: low