The effect of a formative assessment practice on student achievement in mathematics
Boström E, Palm T · 2023
grade Breplicationdeveloper-ledfailednumbers spot-checked
Sample
29 teachers / 566 pupils, Year 7 (14 PD teachers / 291 pupils vs 15 control teachers / 275 pupils)
Population
Swedish Year 7 (age ~13) mathematics classrooms in one mid-sized municipality; PD delivered spring 2011, outcomes measured May 2012.
Design
Stratified random selection of 20 of the 35 year-7 mathematics teachers in the municipality into a 96-hour formative-assessment PD programme (6 subsequently lost); the remaining 15 teachers are the control. The programme was "organised and led by the second author" (Palm), so this is a developer-led evaluation. The paper calls Andersson & Palm 2017 "a parallel study to the one presented here" — same programme and developer, school year 4 instead of 7 — rather than a formal replication. Outcome is NOT an independent standardized test: the pre- and post-tests were built for the municipality "by a group consisting of the authors, experienced secondary school teachers, and national test developers" against the national curriculum (alpha 0.88 / 0.92). Analysis is single-level ANCOVA; the authors show omitting the class level inflates type-1 not type-2 error, so the null is not an artefact of that choice.
Key findings
Same developer, same programme, randomised: no effect on achievement (partial eta-squared = 0.000, p = 0.85), flat across all three prior-attainment terciles, and no correlation between how much formative assessment a teacher actually implemented and how her pupils scored (r = 0.15, p = 0.62). Read against the developer's own parallel year-4 study (Andersson & Palm 2017, reported in this paper as an effect of 0.7 measured at the teacher level), this is a textbook hothouse-to-null collapse. The authors do not conclude formative assessment fails; they conclude the way it was implemented was insufficient.
Genetic confound
Randomised selection into the PD programme; the null is not a selection artefact.
Replication notes
The cleanest developer-constant efficacy-to-null collapse in the formative-assessment literature — the same developer, running the same programme, reports zero. Note that the outcome here is curriculum-aligned and researcher-co-authored, which should bias TOWARD an effect, and there is still none.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Year-7 mathematics achievement | d | ~0.00 — F(1,563) = 0.037, p = 0.85, partial eta-squared = 0.000; adjusted post-test means 28.55 (PD) vs 28.66 (control), implying d ~ -0.01. Null in every prior-attainment tercile (p = 0.44, 0.44, 0.24) | researcher-designed | end of school year | business-as-usual | end-of-treatment | domain-skill |
| Dose-response — number of new formative-assessment activities a teacher implemented vs pupil achievement | partial r (controlling pre-test) | r = 0.15, n = 14 teachers, p = 0.62 (zero-order r = 0.26); implementing more formative assessment did not predict higher achievement | researcher-designed | end of school year | none | end-of-treatment | domain-skill |
Cited by
- Feedback and formative assessmentmixedconf: highgc: low