Effects of Interim Assessments on Student Achievement: Evidence From a Large-Scale Experiment
Konstantopoulos S, Miller SR, van der Ploeg A, Li W · 2016
grade Brctindependentreplicated
Sample
~57 Indiana schools / ~25,000 K-8 students
Population
Indiana K-8 students in the 2009-2010 state benchmark-assessment experiment.
Design
Large cluster RCT in which Indiana schools were randomly assigned to use commercial interim/benchmark assessment systems (mCLASS and Acuity) or not. Well-powered by the standards of this literature, but the null is against a minimum detectable effect of ~0.25, so effects of ~0.1 cannot be excluded.
Key findings
Benchmark assessment systems produced no significant overall effect across 25,000 Indiana students, with only scattered grade-specific positives. Note the power limit: 'null at scale' here is compatible with a true effect around 0.1.
Genetic confound
Minimal — randomized, with state-test outcomes.
Replication notes
Third of the three independent US at-scale trials of commercial formative/interim-assessment products, all broadly null.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Overall K-8 achievement (intent-to-treat) | ITT | not significant overall | standardized | end of year | business-as-usual | end-of-treatment | domain-skill |
| Grade-specific results | pattern | scattered positives — reading in grades 3-4, math in grades 5-6 | standardized | end of year | business-as-usual | end-of-treatment | domain-skill |
Cited by
- Feedback and formative assessmentmixedconf: highgc: low