Effects of Systematic Formative Evaluation: A Meta-Analysis
Fuchs LS, Fuchs D · 1986
grade Cmeta-analysisdeveloper-involvedmixed
Sample
21 controlled studies, 96 effect sizes
Population
Preschool through grade 12, predominantly students with mild to moderate disabilities; progress monitored 2-5 times per week
Design
The founding number of the progress-monitoring literature, and the archive should treat it as an efficacy ceiling rather than an expected value. Three reasons: the primary studies are 1980s special-education trials with small samples, the authors are central developers of curriculum-based measurement, and publication type was itself a significant moderator, which is the field admitting its own selection problem. The moderator values below come from a secondary account (Wiliam's 2011 summary), not from the paper itself - the full text was not obtained, and that is recorded debt.
Key findings
Weighted mean effect size 0.70 across 96 effect sizes (0.63 across the 22 effect sizes for non-disabled learners) - one of the largest effects in this archive, and one of the least likely to survive modern replication. The moderators are what make it useful. Where teachers followed EXPLICIT DATA-DECISION RULES the effect was about 0.92; where they used their own judgement about what the data meant it was about 0.42. Where the data were GRAPHED the effect was about 0.70; where they were not, about 0.26. The measurement is not the intervention. Collecting the data and leaving the teacher to interpret it recovers only a third of the effect; the structure imposed on the response is most of the effect. Read against the modern estimates (Filderman 2018: g = 0.24), the 1986 figure looks like an efficacy-to-effectiveness decay case as much as a finding.
Genetic confound
Low for the internal comparisons - controlled studies within special-education samples. The concern is measurement inflation and publication bias, not heredity.
Replication notes
Directionally replicated but massively deflated: modern syntheses of the same practice return g = 0.24 (Filderman 2018), 0.31-0.37 (Jung 2018) and about 0.19 for formative assessment on reading at scale. The rules-versus-judgement moderator is the part that has held up, being independently echoed by Stecker, Fuchs & Fuchs (2005).
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Systematic formative evaluation, overall | weighted mean effect size | 0.70 across 96 effect sizes (0.63 for non-disabled learners) | standardized | end of intervention period | business-as-usual | end-of-treatment | domain-skill |
| Explicit data-decision rules vs teacher judgement | effect size by moderator | rules ~0.92 vs judgement ~0.42 | standardized | end of intervention period | business-as-usual | end-of-treatment | domain-skill |
| Data graphed vs not graphed | effect size by moderator | graphed ~0.70 vs not graphed ~0.26 | standardized | end of intervention period | business-as-usual | end-of-treatment | domain-skill |
| Publication type as moderator | significance | significant - the field's own indicator of publication bias in this corpus | standardized | not applicable | unclear | not-applicable | domain-skill |
Cited by
- Placement and mastery diagnosis — deciding what to teach next from evidence of current skillmoderate supportconf: mediumgc: medium