The Evidence on Teaching

Technology's Edge: The Educational Benefits of Computer-Aided Instruction

Barrow, L., Markman, L., & Rouse, C. E. · 2009

grade Brctindependentunreplicatednumbers spot-checked
Sample
2,278 students randomly assigned (1,145 to CAI, 1,133 to traditional instruction) in 147 classes taught by 61 teachers across 17 schools and 60 randomisation pools; analysis sample 1,872 students with follow-up scores (141 classes, 59 teachers), and 1,585 with both pre- and post-tests. Statewide- test samples are much smaller (District 1 n = 237, District 2 n = 341, District 3 n = 199). NOTE - the "~17,000 students" figure sometimes attached to this study is not the sample; the three DISTRICTS enrol about 187,000 students between them, but under 1,900 students are in the analysis.
Population
Pre-algebra and algebra students, overwhelmingly grade 9 with some grade 8, in three large high-poverty urban US school districts (one northeast, one midwest, one south); study schools were 80-97 percent African American with substantial Hispanic enrolment in District 2. District 1 studied in 2004-05 (8 high schools, 2 middle schools), Districts 2 and 3 in 2003-04 (4 and 3 high schools). Average class sizes 24-29.
Design
THE INTERVENTION: I Can Learn (Interactive Computer Aided Natural Learning), a hardware-plus-software lab product from JRL Enterprises. Every student sits at a machine; each lesson runs pretest, prerequisite review, lesson, cumulative review and comprehensive test, and students who fail the pretest or review REPEAT until they reach mastery, so pacing is fully individual. The software also handles lesson planning, grading and homework, and the teacher's role is reduced to giving targeted help to whoever is struggling. FUNDING AND INDEPENDENCE: the project was funded by the Education Research Section at Princeton University; the authors are at the Federal Reserve Bank of Chicago and Princeton. The vendor supplied the product and lab support and nothing else - `independence: independent`. IDENTIFICATION: classroom-level random assignment. Schools handed over their pre-algebra/algebra timetable at the start of the year; the researchers formed 60 "randomisation pools" (usually a single class period) and randomly picked which classes in each pool would go to the computer lab. Crucially, schools were NOT told the outcome of the randomisation until after they had finished assigning students to classes, which blocks selection of students into or out of the lab; the authors verify that within-class dispersion of baseline scores matches what random student allocation would produce, so classes were not tracked. Standard errors are clustered at the classroom level and every model includes randomisation-pool fixed effects. WHAT THE CONTROL CLASSES DID WITH THE SAME TIME: they took the same pre-algebra or algebra course, in the same period, from a teacher using traditional whole-class instruction. This is a SUBSTITUTION design, not an add-on - treatment students received their mathematics instruction in the lab INSTEAD of from a teacher, so the estimate is software-versus-teacher, not software-plus-teacher versus teacher. Just over half the teachers taught both a lab class and a traditional class, which is what makes the teacher-fixed-effect models possible. THE OUTCOME MEASURE PROBLEM, which governs how much the headline number is worth: the primary outcome is a 30-item multiple-choice test the authors COMMISSIONED from NWEA and had built "to target specific pre-algebra and algebra skills outlined in the district's course objectives AND THE CAI CURRICULUM" - a treatment-aligned, study-designed instrument, recorded here as researcher-designed. The independent measures are the states' own high-stakes maths tests, which the authors warn contain as little as 10 percent pre-algebra/algebra content and are therefore low-powered for this intervention. Effects on the aligned test are significant overall; on the independent state tests only District 1 is significant. SCALING CAVEAT, and it halves the headline: all effect sizes are expressed in the WITHIN-STUDY baseline standard deviation of 9.20, not a national one. The authors state that using national standard deviations (16.7 at grade 8, 17.4 at grade 10+) "cuts the estimated effect sizes by roughly one-half", and in their own cost section they quote the average-class District 1 gain as 11 percent of a NATIONAL standard deviation. So the famous 0.17 is roughly 0.09 on the scale the rest of this archive uses. IMPLEMENTATION AND CONTAMINATION: treatment students completed on average 33 lessons, about 64 percent of what the CAI course expected of them; 84 percent completed at least 10. Control students completed 5.6 lessons on average (10 percent of expectation) and 15 percent of them completed at least 10 lessons in the lab, i.e. real but limited contamination, which the authors handle with an IV using random assignment as the instrument for having completed at least one lesson. ATTRITION: 2,278 randomised, 1,872 post-tested - about 18 percent lost to district mobility. Baseline algebra scores are identical across arms in both the full and the analysis sample, but in the analysis sample the share African American (p = 0.060) and Hispanic (p = 0.061) differ marginally, entirely from District 2; models control for sex/race/ethnicity in response. GRADE B: a well-executed, correctly clustered single RCT with randomisation-pool fixed effects and honest attrition analysis - the standard "single well-powered RCT" row. Not A because it is unreplicated, the headline measure is treatment-aligned and study-commissioned, the site-level results do not agree with each other, and the outcome was taken immediately at the end of the course with no follow-up (the authors say plainly they do not know how long the gains would last). COST for the cost-effectiveness comparison: a 30-seat lab costs $100,000 for hardware plus $150,000 for software plus about $17,000 a year for training, support and maintenance, roughly $53,000 a year amortised over a 7-year life; the authors compute that cutting class size to 13 in District 1 would cost $241 per pupil per year, and that the benefit of CAI falls to zero at a class size of 13. VERSION READ: all numbers above were checked against the full text of the October 2007 working paper (Federal Reserve Bank of Chicago WP 2007-17 / ERIC ED505645), which is the openly available version of the paper published as AEJ: Economic Policy 1(1), 52-74, February 2009; the abstract and headline estimate are identical across the two, but table numbering and any minor post-referee revisions were not re-checked against the typeset AEJ article.
Key findings
Students randomly assigned to computer-aided pre-algebra/algebra scored 0.17 SD higher (SE 0.076) than students randomly assigned to traditional instruction on a study-commissioned, curriculum-aligned test, rising to 0.25 SD for treatment-on-the-treated and to about 0.30/0.40 SD in teacher-fixed-effect models - though those SDs are within-study, and on national norms the same effect is roughly 0.09-0.11 SD. WHY THIS PAPER MATTERS TO THE TOPIC IS THE MECHANISM THE AUTHORS THEMSELVES PROPOSE, WHICH IS NOT ABOUT TECHNOLOGY: they attribute the gain to individualised pacing substituting for teacher attention that was scarce, and the moderator pattern is the evidence. The CAI advantage grows with class size - 0.21 SD (p < 0.001) in a class of 25 and 0.01 SD (p = 0.89) in a class of 15 - and grows where classmates are frequently absent - under 0.06 SD (ns) at average class attendance versus 0.35 SD (p = 0.08) one SD below it - and the biggest, clearly significant advantage appears in classes that are both LARGE and HETEROGENEOUS in prior attainment. In other words, the software wins exactly where a human teacher is most stretched, and ties where a teacher can already individualise. The authors conclude the gains are "comparable to those achieved with drastic class size reduction". One moderator cuts against the pacing story and is recorded honestly: CAI was NOT differentially effective for low- versus high-prior-attainment students (p > 0.60), which is what a pure "students move at their own pace" account would predict. The effect is also concentrated in pre-algebra (+0.48 SD) rather than algebra (~0.00 SD, p of difference = 0.001), and on independent statewide tests it is significant in only one of three districts (+0.26 in District 1; under 0.10 ns in District 2; negative and ns in District 3).
Genetic confound
Minimal - classrooms randomised within period-level randomisation pools, and schools did not know the assignment when they placed students in classes. Baseline test scores are identical across arms in both the full and analysis samples; small residual race/ethnicity imbalance after 18 percent attrition is controlled for.
Replication notes
No independent randomised replication of I Can Learn is on record. The prior evidence the authors themselves cite (Brooks 2000, Kerstyn 2001, Kirby 1995, Kirby 2004) is quasi-experimental with mixed results, and at least the Kirby reports were evaluations submitted to JRL Enterprises, the vendor - i.e. developer-commissioned. This trial is therefore a single independent positive standing alone, and by the archive's replication rule it caps any verdict it supports at `confidence: medium`. Its own internal replication across sites is imperfect: the effect is large in District 1, positive but insignificant in District 2, and NEGATIVE (insignificant) in District 3, where it is driven by a single randomisation pool.
DOI / URL
10.1257/pol.1.1.52

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Pre-algebra/algebra achievement, NWEA test commissioned for the study and aligned to the CAI curriculum - intent to treatSD (within-study baseline SD = 9.20)+0.173 (SE 0.076, p < .05); +0.172 (SE 0.074) adding demographics; +0.212 (SE 0.077) on the subsample with a pretest; +0.172 (SE 0.060) controlling for baseline score. Equivalent to 26 percent of a school year of growth. In NATIONAL standard-deviation units this is roughly half as large, about 0.09.researcher-designedend of the school year / end of coursebusiness-as-usualend-of-treatmentdomain-skill
Pre-algebra/algebra achievement, aligned NWEA test - treatment on the treated (IV)SD (within-study)+0.249 (SE 0.086), instrumenting completion of at least one lab lesson with random assignment; robust to defining treatment at 5 or 10 lessonsresearcher-designedend of coursebusiness-as-usualend-of-treatmentdomain-skill
Pre-algebra/algebra achievement with TEACHER fixed effects (identified off the ~half of teachers who taught both a lab class and a traditional class)SD (within-study)ITT about +0.30 and IV about +0.40, both significant - i.e. holding the teacher constant makes the effect larger, so it is not selection of better teachers into the lab. But on that same subsample WITHOUT fixed effects the ITT is already 0.27 and the IV 0.44, so most of the increase comes from which teachers are in the subsample, not from the fixed effects.researcher-designedend of coursebusiness-as-usualend-of-treatmentdomain-skill
INDEPENDENT STANDARDIZED MEASURE - statewide mathematics tests, by districtSD (within-district baseline)District 1 +0.26 (significant at 5 percent); District 2 under +0.10 (ns); District 3 negative and not significant. The authors caution that as little as 10 percent of these state tests covers pre-algebra/algebra, so power is low - but the headline 0.17 is on the aligned test and only one of three districts moves an independent one. District 1 also showed +0.4 and +0.6 SD on district benchmark tests, which are also aligned measures.standardizedend of school year (grade 8 test in District 1; grade 10 test in Districts 2 and 3)business-as-usualend-of-treatmentdomain-skill
MECHANISM - CAI effect by class size (the authors' central evidence that the gain is individualised pacing, not technology)SD (within-study)pooling Districts 1 and 2, the CAI-by-class-size interaction is significant at 10 percent (p = 0.067): the effect is +0.21 SD (p < 0.001) in a class of 25 and +0.01 SD (p = 0.89) in a class of 15. Pooling all three districts the interaction is the same sign but not significant (p = 0.19). Positive in Districts 1 (p = 0.09) and 2 (p = 0.80); small and negative with a large SE in District 3. The authors "cautiously conclude there is some evidence CAI is more effective in larger classes, consistent with the idea that the main benefit of CAI is the individualization of the instruction."researcher-designedend of coursebusiness-as-usualend-of-treatmentdomain-skill
MECHANISM - CAI effect by classmates' attendance rate (Districts 2 and 3)SD (within-study)in a classroom with average attendance the CAI effect is under +0.06 SD and not significant; in a classroom one SD below average attendance it is +0.35 SD (p = 0.08). CAI helps most where disrupted classes make ordinary teaching least effective. Individual-student attendance moderation pointed the same way but was not significant.researcher-designedend of coursebusiness-as-usualend-of-treatmentdomain-skill
MECHANISM - CAI effect by class heterogeneity in prior attainment, interacted with class sizeSD (within-study)the two-way interaction of CAI with the classroom's baseline test-score SD is not significant in any sample; but the THREE-way interaction of CAI, baseline heterogeneity and "large class" (more than 24 students) is large and statistically significant. Individualisation buys nothing in a small mixed class, where a teacher can cope; it buys a lot in a large mixed one.researcher-designedend of coursebusiness-as-usualend-of-treatmentdomain-skill
MECHANISM CHECK THAT FAILED - CAI effect by student's own prior maths attainmentSD (within-study)no differential effect across baseline test-score quartiles (p > 0.60) pooling all three districts or Districts 1 and 2. Recorded because a naive "students move at their own pace" story predicts differential gains at one or other end of the distribution, and there are none. The individualisation the data support is about relieving the TEACHER's attention constraint, not about matching each student's speed.researcher-designedend of coursebusiness-as-usualend-of-treatmentdomain-skill
CAI effect by course level - pre-algebra vs algebraSD (within-study)pre-algebra +0.48 SD vs algebra under +0.01 SD, difference p = 0.001 pooling all districts; within District 1 alone, pre-algebra +0.44 vs algebra +0.13 (p of difference < 0.07 in every district). The effect is essentially a pre-algebra effect.researcher-designedend of coursebusiness-as-usualend-of-treatmentdomain-skill
IMPLEMENTATION - lessons completed versus lessons prescribed, and contamination of the control armlessons / percent of expectationtreatment students completed 33 lessons on average, 64 percent of what the CAI course expected; 84 percent completed at least 10. Control students completed 5.6 lessons, 10 percent of expectation, and 15 percent of them completed at least 10 lessons in the lab. So the treatment arm received about two-thirds of the prescribed dose and the control arm was not clean.administrativeacross the school yearnoneend-of-treatmentbehaviour

Cited by