The Evidence on Teaching

Effectiveness of Reading and Mathematics Software Products: Findings from the First Student Cohort

Dynarski, M., Agodini, R., Heaviside, S., Novak, T., Carey, N., Campuzano, L., Means, B., Murphy, R., Penuel, W., Javitz, H., Emery, D., & Sussex, W. · 2007

grade Arctindependentreplicatednumbers spot-checked
Sample
33 unduplicated districts, 132 unduplicated schools, 439 randomly assigned teachers, 9,424 students in the analysis sample (10,659 tested in fall 2004, 9,792 in spring 2005); 16 software products in four independent substudies - grade 1 reading (5 products, 2,619 students), grade 4 reading (4 products, 2,265 students), grade 6 math (3 products, 3,136 students), algebra (3 products, 1,404 students)
Population
US public-school students in 33 mostly high-poverty, high-minority districts, school year 2004-05. Grades 1 and 4 (reading), grade 6 (mathematics), and algebra classes (typically grade 9, mean age 14.8). Districts were recruited on the criterion that they were NOT already using products similar to the study products.
Design
Congressionally mandated under NCLB section 2421; funded by the U.S. Department of Education, Institute of Education Sciences / NCEE under contract ED-01CO-0039/0007, and conducted and authored by Mathematica Policy Research with SRI International. Vendors supplied the software, the teacher training and the technical assistance; they did not run or write the evaluation, hence `independence: independent`. IDENTIFICATION: within each participating school, volunteering teachers were randomly assigned to be allowed to use the study product or not; students were not randomised to teachers, and effects were estimated with three-level hierarchical linear models (students in classrooms in schools) adjusting for pretest, student age/gender, teacher gender/experience/degree, and school composition. Robustness checked against district-administered tests where available, with the same pattern. PRODUCT SELECTION MAKES THE NULL HARDER TO DISMISS, NOT EASIER: 160 products were submitted in response to a public invitation; a team rated submissions on prior evidence of effectiveness, ability to operate at national scale, and capacity to train teachers; two external expert panels reviewed; ED chose 16. Only products that could show at least some prior evidence of effectiveness were selected, and 12 of the 16 had won or been nominated for industry/media/teacher awards. The report states outright that ED recognised "selecting ostensibly more effective products could tilt the study toward finding higher levels of effectiveness" and accepted that tilt. This is therefore a null on a vendor-nominated, evidence-screened, award-winning sample of the market, not a null on a random draw of software. WHAT THE CONTROL CLASSROOMS DID WITH THE SAME TIME: control teachers taught the same subject in the same period as they normally would, and were permitted to keep using whatever technology they already had. Products functioned overwhelmingly as SUPPLEMENTS displacing part of the existing reading/math block rather than as added time - 96 percent of grade-1 treatment teachers used the product as a supplement. Grade-1 treatment teachers reported 8.7 hours a week of reading instruction vs 7.9 for controls (p = 0.13, ns); grade-4 treatment teachers reported one hour a week MORE reading instruction than controls (significant), so the grade-4 null absorbs a small time advantage for treatment. Control-classroom technology use was non-zero: grade-1 control teachers reported using other reading software about a fifth as much as treatment teachers used study products; grade-6 control teachers reported ~3 hours of other technology use against 51 teacher- reported hours of study-product use in treatment classrooms; grade-4 treatment teachers scheduled 6 hours of OTHER products and control teachers 7 hours. IMPLEMENTATION AND ACTUAL USAGE (recorded so that the standard "it just was not implemented" rebuttal can be checked rather than asserted): 94-98 percent of treatment teachers attended vendor training and most reported feeling prepared; 92 percent said they would use the product again; observers logged technical problems in about 20 percent of observed time segments, nearly all minor and quickly resolved. Software-logged student usage: grade 1 averaged 23 minutes a day when used, 76 days used, 29 hours a year, about 11 percent of reading instructional time (product-by-product range 8 to 61 hours a year, 11 to 34 minutes a day); grade 4 was 7 hours for one product and 20 for the other, under 10 percent of reading time; grade 6 was about 17 hours a year, about 11 percent of math time; algebra was 15 hours a year, about 10 percent of math time, with a 5-to-28-hour range across products. THE DOSE GAP IS LARGE AND EXPLICIT: grade-6 vendors recommended their products be used 120 to 225 minutes a week, i.e. roughly 70 to 135 hours over a school year, against 17 hours actually logged - students received on the order of a fifth of the prescribed dose. The report is candid that these are first-time users inside a study that bought the hardware and paid teachers honoraria to attend training, so usage could be higher OR lower than typical district practice. LIMITS: teachers were randomised WITHIN schools, so treatment-to-control spillover between colleagues would bias effects toward zero and cannot be ruled out; district and school participation was voluntary; year-1 results were reported only for GROUPS of products, because developers agreed to participate on that condition, so no product is individually identified here. The report's own sub-study sample counts disagree slightly between the design box, Table 1 and the chapter text (e.g. grade 1 is variously 11/13/14 districts and 42/43/46 schools); the unduplicated totals above are the ones the report treats as authoritative. GRADE A is justified on the evidence hierarchy's "large multi-site RCT" row: 132 schools across 33 districts and four separately powered substudies, prespecified in a published design report, correctly clustered analysis, and a replicated second cohort. The within-school randomisation and voluntary site recruitment are the reasons it is not stronger than A.
Key findings
The canonical large null in ed-tech. Across all four substudies, test scores in classrooms randomly assigned to use the software did not differ significantly from control classrooms: grade-1 SAT-9 reading ES 0.03 (p = 0.34), grade-4 SAT-10 reading ES 0.02 (p = 0.48), grade-6 SAT-10 math ES 0.07 (p = 0.15), and algebra on the ETS End-of-Course exam ES -0.06 (p = 0.33), with the algebra concepts subtest at -0.10 (p = 0.07). The products were not a random draw: they were volunteered by vendors, screened for prior evidence of effectiveness, and mostly award-winning, which is why this null bites - it is a null on the industry's own nominees. School- level effects varied enormously (within a single algebra district, school effect sizes ran from about -0.75 to +0.75; 62-63 percent of the variance in school effects sat between districts), so the aggregate zero is an average over real dispersion that the measured school and classroom characteristics almost entirely failed to explain. What the software reliably DID change was the classroom: teachers moved from leading (58 percent of observation intervals in control classrooms to 29 percent in treatment) to facilitating, students moved from lecture and question-and-answer to individual work (40 percent to 84 percent of intervals) - and on-task behaviour and test scores both stayed put.
Genetic confound
Minimal - teachers randomised within schools; students were not randomised to teachers, but treatment and control students were equivalent on fall pretest, age and gender.
Replication notes
The study's own second cohort (Campuzano, Dynarski, Agodini & Rall 2009, NCEE 2009-4041) re-ran the design with a new student cohort and teachers who now had a year of experience with the same products. The overall null held: 9 of the 10 individually reported products showed no significant effect, and the four year-on-year "experience effects" were either null or split in opposite directions (grade-6 math got significantly WORSE, Algebra I significantly better). That is a same-team replication rather than an independent one, but the pattern is corroborated externally by the RAND Cognitive Tutor trial (Pane et al. 2014, ~0 in year 1) and by Cheung & Slavin's meta-analytic finding that effects shrink toward zero as trials get larger and measures get more independent. No large independent trial has overturned it.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Grade 1 reading, SAT-9 overall scoreES0.03 (NCE difference 0.73; p = 0.34, ns)standardizedspring 2005, after one school year of usebusiness-as-usualend-of-treatmentdomain-skill
Grade 1 reading, Test of Word Reading Efficiency overall (individually administered fluency measure)ES0.04 (standard-score difference 0.51; p = 0.57, ns); subtests phonemic decoding 0.03, sight word 0.02standardizedspring 2005, after one school year of usebusiness-as-usualend-of-treatmentdomain-skill
Grade 4 reading, SAT-10 overall scoreES0.02 (NCE difference 0.41; p = 0.48, ns); subtests vocabulary 0.02, word study 0.00, comprehension 0.03 - all nsstandardizedspring 2005, after one school year of usebusiness-as-usualend-of-treatmentdomain-skill
Grade 6 mathematics, SAT-10 overall scoreES0.07 (NCE difference 1.43; p = 0.15, ns) - the largest point estimate in the study and still not significantstandardizedspring 2005, after one school year of usebusiness-as-usualend-of-treatmentdomain-skill
Algebra, ETS End-of-Course Algebra Assessment overall scoreES-0.06 (difference -0.86 percentage points correct; p = 0.33, ns). Concepts subtest -0.10 (p = 0.07), processes -0.06, skills +0.02. Point estimate is negative.standardizedspring 2005, end of the algebra coursebusiness-as-usualend-of-treatmentdomain-skill
Between-school dispersion of effects (the "it works somewhere" question)school-level ES rangevery large and largely unexplained - e.g. algebra school effects within one district ran from about -0.75 to +0.75, grade-4 reading within one district from -0.25 to +0.25; 63 percent (grade 4) and 62 percent (grade 6) of school-effect variance was between districts. Only two moderators correlated with effects at all (student-teacher ratio in grade 1; amount of product use in grade 4), and the report states these are not causal because districts and schools self-selected.standardizedspring 2005business-as-usualend-of-treatmentdomain-skill
DISPLACEMENT - what the software replaced in the classroom (percentage of observation intervals, grade 1)ppteacher as leader 29 treatment vs 58 control (p < .01); teacher as facilitator 56 vs 33 (p < .01); individual student work 84 vs 40 (p < .01); lecture 16 vs 26 (p = .01); question-and-answer 16 vs 33 (p < .01). The software substituted individual seat-work at a screen for whole-class teacher-led instruction, and achievement did not move.administrativethree one-hour structured classroom observations across 2004-05business-as-usualend-of-treatmentbehaviour
Student on-task behaviour (the engagement claim usually made for ed-tech)pp85 percent of observation intervals with >90 percent of students on task in treatment classrooms vs 85 percent in control (p > .50) - no differenceadministrativethree structured classroom observations across 2004-05business-as-usualend-of-treatmentbehaviour
Actual usage versus prescribed dose (implementation, recorded so the "implementation failure" rebuttal is checkable)hours per student per yearsoftware-logged usage: grade 1 = 29 h/yr (~11% of reading instruction; product range 8-61 h/yr), grade 4 = 7 and 20 h/yr for the two products that logged it (<10% of reading time), grade 6 = 17 h/yr (~11% of math time), algebra = 15 h/yr (~10% of math time, product range 5-28). Grade-6 vendors PRESCRIBED 120-225 minutes a week (~70-135 h/yr), so delivered dose was roughly a fifth of prescribed - while 94-98% of teachers attended vendor training and 92% said they would use the product again.administrativeacross the 2004-05 school yearnoneend-of-treatmentbehaviour

Cited by