Classroom Assessment for Student Learning: Impact on Elementary School Mathematics in the Central Region
Randel B, Beesley AD, Apthorp H, Clark TF, Wang X, Cicchinelli LF, Williams JM · 2011
grade Brctindependentreplicatednumbers spot-checked
Sample
67 elementary schools (33 intervention / 34 control) in 32 Colorado districts; 9,596 students; 409 grade 4-5 teachers (178 intervention / 231 control)
Population
US grade 4 and 5 students and their mathematics teachers in Colorado; Colorado Student Assessment Program (CSAP) mathematics outcomes, 2007/08 training year plus 2008/09 implementation year.
Design
Cluster RCT: volunteer schools randomly assigned (blocked by district) to receive Classroom Assessment for Student Learning materials and form teacher learning teams, or to continue with regular professional development. Implemented under real-world conditions with no researcher involvement — teachers spent an average of 31 hours on CASL against the developer's recommended 60, and only 40% read every chapter fully, which is itself a finding about a self-executing product. Powered (>.80) to detect 0.25 SD, so it cannot rule out effects smaller than that. Volunteer sample, not representative of Colorado schools. Attrition on the teacher and motivation outcomes exceeded WWC-acceptable levels and was handled by imputation. IES report (NCEE 2011-4005), no DOI; ERIC ED517969.
Key findings
The best-known formative-assessment PD product raised teacher knowledge by 0.42 SD and left everything downstream untouched: assessment practice quality 0.03 (p = .85), student involvement 0.26 (p = .10, ns), student motivation null, and mathematics achievement 0.01 (p = .87). The mediator is where it fails — giving teachers assessment knowledge did not change what they did in class. Two honest limits the report itself insists on: power only to 0.25 SD, and a non-significant impact is not evidence of no impact.
Genetic confound
Minimal — school-level randomization with a state-test outcome.
Replication notes
One of a set of independent US cluster RCTs that break at the same mediator: teacher knowledge moves, teacher practice does not, achievement does not. The `replicated` enum is the archive's prior judgement about the pattern across trials, not something this report establishes.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| State mathematics achievement (CSAP scale score) | effect size | +0.01 — impact 0.58 scale points (SE 3.47, p = .87); adjusted means 502.49 vs 501.91. Robust to covariate specification, estimation method and missing-data treatment. Cohort splits also null (two-year exposure -4.53, p = .24; one-year exposure +5.59, p = .19) | standardized | end of implementation year (spring 2009) | business-as-usual | end-of-treatment | domain-skill |
| Teacher knowledge of classroom assessment (60-item test) | effect size | +0.42 — 41.36 vs 38.58 items correct, difference 2.78, p = .01. The ONLY significant impact in the trial. | researcher-designed | end of implementation year (spring 2009) | business-as-usual | end-of-treatment | domain-skill |
| Quality of teachers' classroom assessment practice (blind-scored work samples) | effect size | +0.03 — 1.61 vs 1.60 on a 1-4 rubric, p = .85. Null. | researcher-designed | end of implementation year (spring 2009) | business-as-usual | end-of-treatment | behaviour |
| Teacher-reported frequency of involving students in assessment | effect size | +0.26 — 0.39 vs 0.34, p = .10. NOT statistically significant. | self-report | end of implementation year (spring 2009) | business-as-usual | end-of-treatment | behaviour |
| Student motivation to learn mathematics | mean rating (4-point scale) | null at both waves — 3.29 vs 3.28 (May 2008) and 3.33 vs 3.32 (posttest) | self-report | May 2008 and end of implementation year | business-as-usual | end-of-treatment | non-cognitive |
Cited by
- Feedback and formative assessmentmixedconf: highgc: low