Examining the Reliability of Running Records: Attaining Generalizable Results
Fawson PC, Ludlow BC, Reutzel DR, Sudweeks R, Smith JA · 2006
grade Cquasi-experimentindependentreplicated
Sample
generalizability study - 10 teachers x 10 first-grade children x 2 leveled passages
Population
US first-grade children assessed with running records by their own teachers
Design
A generalizability study, which decomposes score variance into its sources rather than estimating an effect - the right design for asking how many observations a reliable placement decision needs. Small (10 teachers, 10 children), single grade, and the passages are commercially leveled texts. Graded C.
Key findings
The problem with a running record is not the scorer, it is the sample. RATERS contributed less than 1% of score variance - teachers agree with each other almost perfectly about what they heard. But passage-to-passage and occasion-to-occasion variance is large, and the authors conclude that a MINIMUM OF THREE PASSAGES is needed for a reliable estimate. A single running record, which is how reading level is assigned in practice everywhere, does not reliably place a child. A follow-up generalizability study on naturalistic lessons put the requirement higher still, at roughly eight to ten reads for reliable accuracy and self-correction measurement.
Genetic confound
Not applicable - a variance decomposition of an instrument.
Replication notes
Replicated in direction and strengthened in magnitude by the 2021 Reading Psychology generalizability study on naturalistic lessons, which found that even more observations are needed.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Share of score variance attributable to the rater | variance component | less than 1% | standardized | single administration | none | not-applicable | domain-skill |
| Number of passages needed for a reliable running-record estimate | count | minimum of 3 passages (a later naturalistic generalizability study implies ~8-10 reads) | standardized | across administrations | none | not-applicable | domain-skill |
Cited by
- Placement and mastery diagnosis — deciding what to teach next from evidence of current skillmoderate supportconf: mediumgc: medium