The Evidence on Teaching

Effectiveness of Reading and Mathematics Software Products: Findings From Two Student Cohorts

Campuzano, L., Dynarski, M., Agodini, R., & Rall, K. · 2009

grade Brctindependentmixednumbers spot-checked
Sample
Second-year data collection: 10 products, 23 districts, 77 schools, 176 teachers, 3,280 students. Experience-effect analysis restricted to the 115 teachers (27 percent of the 428 randomised in year 1) who stayed in the study, covering 5,345 student-year observations pooled across the two years - grade 1 n = 801 treatment / 610 control, grade 4 n = 322 / 282, grade 6 n = 1,239 / 1,040, algebra n = 547 / 504. Individual-product analysis used all students from either year.
Population
US public-school students in high-poverty districts, school years 2004-05 and 2005-06. Grade 1 and grade 4 reading, grade 6 pre-algebra, and Algebra I (9 percent grade 8, 87 percent grade 9, 4 percent grades 10-12).
Design
Same study, same funder and same contract as Dynarski et al. 2007 - U.S. Department of Education, IES/NCEE, contract ED-01CO-0039/0007 - conducted and authored by Mathematica Policy Research; vendors supplied product, training and usage logs but did not run or write the evaluation, hence `independence: independent`. THIS IS THE DIRECT TEST OF THE "IT JUST NEEDED MORE TIME" DEFENCE. Identification of the experience effect is a treatment-by-year interaction estimated on the 115 teachers who continued into 2005-06 with a new cohort of students: because the treatment/control randomisation is preserved within that retained sample, each year's effect is still experimental, but the CONTRAST BETWEEN YEARS is not - the retained 27 percent of teachers are a selected group (teachers who moved school, changed grade or left teaching dropped out), so the year-2 estimates cannot be read as the year-1 estimates re-run. That is the main reason this is graded B rather than A despite resting on the same randomisation. Five further design compromises were made to fit the second year inside budget: products implemented in only a few schools were dropped (16 products became 10), CLASSROOMS WERE NOT OBSERVED AT ALL in year 2, only one treatment and one control classroom were sampled per school where there were more, some districts supplied their own test scores instead of the study administering tests (Iowa Tests of Basic Skills, New Mexico Standards Based Assessment and Stanford scores mixed in, all converted to NCE units), and some items came from school records. Because there were no year-2 observations or teacher interviews, the report states plainly that it has NO information about how teachers changed their use of products, and NO information about whether control teachers changed their use of other software - a real hole in the counterfactual. WHAT THE CONTROL CLASSROOMS DID WITH THE SAME TIME: unchanged from year 1 - control teachers taught the same subject in the same period as usual and could use whatever technology they already had; eight of the ten products were supplements to the core curriculum, while Plato Focus (grade 1) and Cognitive Tutor (Algebra I) were the core curriculum, so for those two the software displaced the teacher-led curriculum rather than sitting alongside it. Products were deployed to whole classes, not as remediation for lagging students. IMPLEMENTATION AND ACTUAL USAGE, YEAR 1 vs YEAR 2 (software-logged, restricted to continuing teachers; all four year-on-year changes statistically significant): grade 1 fell from 2,556 to 1,182 minutes per student (43 h to 20 h), grade 4 rose from 720 to 936 minutes (12 h to 16 h), grade 6 fell from 852 to 678 minutes (14 h to 11 h), algebra rose from 1,309 to 1,450 minutes (22 h to 24 h) over 14.4 then 16.8 weeks of use. So experienced teachers did not converge on more use - in the two elementary reading and math cases they used the software substantially LESS. The report tested the obvious mediator directly and found the relationship between change in effects and change in usage was NOT statistically significant, and no significant usage-by-effect interaction in either year in grades 4 or algebra. Usage-data coverage was itself patchy (grade 1: 84 percent of treatment students in year 1, only 51 percent in year 2). MULTIPLICITY: the individual-product analysis ran ten separate models; one product came out significant at the 5 percent level, which is close to what chance alone delivers across ten tests, and the report itself warns that products were implemented in different self-selected districts and schools and that the estimates are not adjusted for unmeasured district/school differences, so they should not be read as a product ranking.
Key findings
The direct counterpoint to "implementation just needed more time": a second year of teacher experience with the same product did not produce the gains the first year had failed to produce. Two of the four experience effects were not significant (grade 1 reading -2.14 NCE, p = 0.08; grade 4 reading +2.02 NCE, p = 0.29), and the two that WERE significant point in opposite directions - grade-6 math got significantly WORSE with experience (-2.80 NCE, p = 0.02, taking the average student from the 49th to the 44th percentile) while Algebra I got better (+2.90 percent correct, p = 0.05). Software-logged usage did not rise with experience either; in grade 1 it nearly halved, and the change in usage was unrelated to the change in effects. Reported separately for the first time, 9 of the 10 individual products had no statistically significant effect (five of those nine negative, four positive); the single significant positive was LeapTrack in grade 4 at +1.97 NCE (95% CI 0.48 to 3.46), roughly 0.09 SD, one hit in ten simultaneous tests.
Genetic confound
Minimal within each year (teachers randomised within schools); students were not randomised to teachers but were equivalent on fall pretest, age and gender. The between-year contrast is not protected by randomisation because only 27 percent of teachers were retained.
Replication notes
This report IS the replication of Dynarski et al. (2007) - the same within-school teacher-randomised design, a new student cohort, and teachers who now had a year of practice with the same software. The overall null replicated. What did NOT replicate cleanly are the two statistically significant second-year effects, which point in OPPOSITE directions (grade-4 reading +0.22 SD; grade-6 math -0.15 SD, with a significantly more negative experience effect) and neither of which has an independent replication. The one significantly positive individual product (LeapTrack, +1.97 NCE) is a single hit out of ten simultaneous product tests and has not been independently reproduced. The Cognitive Tutor estimate here (-1.28 NCE, ns) sits against the fragile +0.21 year-2 result in the RAND trial (Pane et al. 2014), so that product's record across the two largest trials is mixed.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Grade 1 reading, SAT-9 - EXPERIENCE EFFECT (year 2 minus year 1)NCE units / ESyear 1 +0.86 NCE (ES 0.04, p > 0.50); year 2 -1.28 NCE (ES -0.06, p > 0.50); difference -2.14 NCE, p = 0.08 - not significant, and the point estimate goes the WRONG way for the experience storystandardizedspring 2005 and spring 2006business-as-usualend-of-treatmentdomain-skill
Grade 4 reading, SAT-10 - EXPERIENCE EFFECT (year 2 minus year 1)NCE units / ESyear 1 +2.65 NCE (ES 0.13, p = 0.18, ns); year 2 +4.67 NCE (ES 0.22, p = 0.01, significant); difference +2.02 NCE, p = 0.29 - the year-2 gain is significant on its own but the year-on-year improvement is not, on n = 322/282 studentsstandardizedspring 2005 and spring 2006business-as-usualend-of-treatmentdomain-skill
Grade 6 mathematics, SAT-10 - EXPERIENCE EFFECT (year 2 minus year 1)NCE units / ESyear 1 -0.44 NCE (ES -0.02, p > 0.50); year 2 -3.24 NCE (ES -0.15, p = 0.11); difference -2.80 NCE, p = 0.02 - SIGNIFICANTLY NEGATIVE. Experienced teachers using the same math software produced worse results, the average student falling from the 49th to the 44th percentile.standardizedspring 2005 and spring 2006business-as-usualend-of-treatmentdomain-skill
Algebra I, ETS End-of-Course Assessment - EXPERIENCE EFFECT (year 2 minus year 1)percent correct / ESyear 1 -0.34 percent correct (ES -0.02, p > 0.50); year 2 +2.56 (ES 0.15, p = 0.03, significant); difference +2.90, p = 0.05 - the one substudy where experience significantly helped, and the smallest of the four samplesstandardizedspring 2005 and spring 2006business-as-usualend-of-treatmentdomain-skill
Individual product effects, pooled across both years, all 10 products (reading, NCE units)NCE units (95% CI)LeapTrack (grade 4) +1.97 (0.48 to 3.46), the ONLY significant effect in the study; Destination Reading +1.91 (-1.56 to 5.38); Plato Focus +0.50 (-2.42 to 3.42); Waterford Early Reading +0.42 (-2.46 to 3.30); Headsprout +0.29 (-1.90 to 2.48); Academy of Reading -0.16 (-2.25 to 1.93). SD of the NCE norm sample is 21.06, so +1.97 NCE is about 0.09 SD.standardizedspring 2005 and/or spring 2006business-as-usualend-of-treatmentdomain-skill
Individual product effects, pooled across both years, math productsNCE units / percent correct (95% CI)Larson Pre-Algebra (grade 6) +2.37 (-0.86 to 5.60); Plato Achieve Now (grade 6) -0.58 (-3.58 to 2.42); Larson Algebra I -0.10 (-2.31 to 2.11); Cognitive Tutor Algebra I -1.28 (-3.62 to 1.06). None significant; three of four negative point estimates.standardizedspring 2005 and/or spring 2006business-as-usualend-of-treatmentdomain-skill
Actual software usage, year 1 vs year 2 (does experience raise the dose?)minutes per student per yeargrade 1 2,556 -> 1,182 (43 h -> 20 h); grade 4 720 -> 936 (12 h -> 16 h); grade 6 852 -> 678 (14 h -> 11 h); Algebra I 1,309 -> 1,450 (22 h -> 24 h) over 14.4 -> 16.8 weeks of use. All four changes statistically significant; two went DOWN. The relationship between change in usage and change in effect was not statistically significant.administrativeacross 2004-05 and 2005-06noneend-of-treatmentbehaviour
Grade 6 math, share of students below the 33rd national percentileppyear 1 treatment 32.9 vs control 31.8 (+1.1 pp, significant at 5 percent in year 1 per the report's text); year 2 treatment 38.2 vs control 30.0 (+8.2 pp, ns); year-on-year difference +7.1 pp, p = 0.28. Direction is unfavourable to the software in both years.standardizedspring 2005 and spring 2006business-as-usualend-of-treatmentdomain-skill

Cited by