The Evidence on Teaching

Effects of Systematic Formative Evaluation: A Meta-Analysis

Fuchs LS, Fuchs D · 1986

grade Cmeta-analysisdeveloper-involvedmixed
Sample
21 controlled studies, 96 effect sizes
Population
Preschool through grade 12, predominantly students with mild to moderate disabilities; progress monitored 2-5 times per week
Design
The founding number of the progress-monitoring literature, and the archive should treat it as an efficacy ceiling rather than an expected value. Three reasons: the primary studies are 1980s special-education trials with small samples, the authors are central developers of curriculum-based measurement, and publication type was itself a significant moderator, which is the field admitting its own selection problem. The moderator values below come from a secondary account (Wiliam's 2011 summary), not from the paper itself - the full text was not obtained, and that is recorded debt.
Key findings
Weighted mean effect size 0.70 across 96 effect sizes (0.63 across the 22 effect sizes for non-disabled learners) - one of the largest effects in this archive, and one of the least likely to survive modern replication. The moderators are what make it useful. Where teachers followed EXPLICIT DATA-DECISION RULES the effect was about 0.92; where they used their own judgement about what the data meant it was about 0.42. Where the data were GRAPHED the effect was about 0.70; where they were not, about 0.26. The measurement is not the intervention. Collecting the data and leaving the teacher to interpret it recovers only a third of the effect; the structure imposed on the response is most of the effect. Read against the modern estimates (Filderman 2018: g = 0.24), the 1986 figure looks like an efficacy-to-effectiveness decay case as much as a finding.
Genetic confound
Low for the internal comparisons - controlled studies within special-education samples. The concern is measurement inflation and publication bias, not heredity.
Replication notes
Directionally replicated but massively deflated: modern syntheses of the same practice return g = 0.24 (Filderman 2018), 0.31-0.37 (Jung 2018) and about 0.19 for formative assessment on reading at scale. The rules-versus-judgement moderator is the part that has held up, being independently echoed by Stecker, Fuchs & Fuchs (2005).

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Systematic formative evaluation, overallweighted mean effect size0.70 across 96 effect sizes (0.63 for non-disabled learners)standardizedend of intervention periodbusiness-as-usualend-of-treatmentdomain-skill
Explicit data-decision rules vs teacher judgementeffect size by moderatorrules ~0.92 vs judgement ~0.42standardizedend of intervention periodbusiness-as-usualend-of-treatmentdomain-skill
Data graphed vs not graphedeffect size by moderatorgraphed ~0.70 vs not graphed ~0.26standardizedend of intervention periodbusiness-as-usualend-of-treatmentdomain-skill
Publication type as moderatorsignificancesignificant - the field's own indicator of publication bias in this corpusstandardizednot applicableunclearnot-applicabledomain-skill

Cited by