The Evidence on Teaching

New Evidence on Classroom Computers and Pupil Learning

Angrist, J. D., & Lavy, V. · 2002

grade Bquasi-experimentindependentmixednumbers spot-checked
Sample
Roughly 200 Jewish schools sampled nationally, of which 122 had applied for Tomorrow-98 money. Analysis (applicant) samples: 4th-grade Maths 3,271 pupils in 122 schools (2,891 in 107 schools with 1991 lagged scores); 4th-grade Hebrew 2,464; 8th-grade Maths 2,621; 8th-grade Hebrew 2,593. Full sampled frame before restricting to applicants was 4,779 4th-grade Maths pupils in 181 schools.
Population
Israeli elementary and middle school pupils, grades 4 and 8, Jewish sector (religious and secular), tested nationally in June 1996 in Maths and Hebrew. Programme schools had held the computers for an average of about 9 months (grade 4) to 13 months (grade 8) at the test date.
Design
THE INTERVENTION: Tomorrow-98 ("Mahar"), Israel's national school computerisation programme, funded mainly by the Israeli State Lottery with Ministry of Education and municipal money. 35,000 computers were installed in 905 schools between 1994 and 1996, targeting a 10:1 pupil-computer ratio, with a substantial teacher-training component; most machines went into a dedicated lab used on a schedule for both computer-skills training and computer-aided instruction. By June 1996 about 10 percent of elementary and 45 percent of middle-school pupils had received programme computers. Much of the software came from the Center for Educational Technology, which held most of the Israeli educational-software market. No study funder is named; the Ministry of Education's Chief Scientist's Office, Evaluation Division and Information Systems Division supplied the data and commissioned the implementation survey, and the paper was written by two academic economists - `independence: independent`. IDENTIFICATION, PRECISELY: programme placement was NOT random and the authors say so. Towns applied on behalf of schools and submitted a within-town priority RANKING reflecting the municipality's judgement of each school's "need" and "ability" to use computers (in practice, schools with some pre-existing computer infrastructure); the Ministry then allocated down the town list, giving priority to towns with more stand-alone middle schools, with a ceiling set by the town's share of national 1-8 enrolment. The paper handles this four ways, and the weaknesses of each should be recorded. (1) OLS of test scores on teacher-reported CAI intensity - confounded, and the authors treat it as such: the positive 4th-grade coefficients shrink to zero once town effects are added, which they read as evidence of omitted school-prosperity variables. (2) REDUCED FORM: regress scores on a Tomorrow-98 receipt dummy, controlling for sex, immigrant and special-education status, a pupil disadvantage index, enrolment, pre-1994 computer stock, the town priority rank, and (in half the models) school-average 1991 test scores from three years before the programme. (3) 2SLS using the receipt dummy as an instrument for a 0-3 teacher-reported CAI-intensity scale. (4) A NONLINEAR IV that instruments with the nonparametrically estimated probability of funding as a function of normalised within-town rank, controlling for a quadratic in that rank - i.e. exploiting the allocation rule itself rather than realised receipt. WEAKNESSES OF THE IDENTIFICATION, stated plainly: the instrument is an administratively allocated programme, not a lottery among schools, so the exclusion restriction is an assumption. Their defences are (a) Tomorrow-98 award status is not systematically associated with pupil characteristics or with schools' 1991 pre-programme test scores; (b) a "T-98 / will-get-T-98" comparison restricting the control group to schools that received computers AFTER June 1996 gives near-identical estimates, which is the strongest of the checks because it compares eventual winners to each other; (c) controlling for schools' 1991 instructional computer USE (not just hardware) barely moves the estimate; (d) the negative score effect appears only in the one grade/subject cell where there was a first stage on teaching practice, which is the pattern a causal chain predicts and a "computers went to weak schools" story does not. Against that: the CAI measure is a single self-reported 4-point survey item ("Which of the following do you use when teaching? ... computer software"), the 2SLS estimates are IMPRECISE (standard errors of 0.19-0.25 on point estimates of -0.31 to -0.44, so only marginally significant at best), the nonlinear-IV estimates are smaller and mostly insignificant, and the Hebrew and 8th-grade results are all null. THIS IS AT THE WEAK END OF GRADE B - a policy-rollout IV with a defensible but unproven exclusion restriction and wide confidence intervals, saved from C by the pre-programme balance checks, the lagged-score controls, the winners-vs-future-winners comparison and the coherent first-stage/second-stage pattern. WHAT THE COMPARISON CLASSROOMS DID WITH THE SAME TIME: this is an ADD-COMPUTERS-TO-AN-EXISTING-SCHOOL design, not a substitution of software for a teacher. Control schools ran their normal timetable; nearly half of 4th graders and about two-thirds of 8th graders already had some computer equipment before Tomorrow-98, so the contrast is more/newer machines plus training versus the existing stock - which makes the negative sign harder to explain away as "no technology at all in the control group". DISPLACEMENT WAS DIRECTLY TESTED AND NOT FOUND: the authors used the teacher survey to check whether programme status changed class size, subject coverage, hours of instruction, frequency of teacher training, use of non-computer audio-visual or TV materials, or teacher satisfaction, and none of these was related to programme status. They conclude "this suggests there was no displacement" - so the 4th-grade Maths decline is not obviously an artefact of computers crowding out instructional time, and their preferred reading is simply that CAI was no better and possibly worse than the teaching it replaced within the lesson. COST: the Ministry budgeted $3,000 per machine including software and set-up; programme schools got about 40 machines, roughly $120,000 per school, which in Israel at the time would pay up to four teachers' wages, or about one teacher per year in flow cost at 25 percent depreciation.
Key findings
Tomorrow-98 clearly changed teaching - 4th-grade teachers in funded schools moved up the CAI-intensity scale by 0.60 points (SE 0.22) out of 3 - and did not raise learning. The reduced-form effect on 4th-grade Maths scores was -0.20 SD (SE 0.089) and -0.24 SD (SE 0.088) with lagged-score controls: negative and at least marginally significant. Scaled by the first stage, 2SLS puts a one-unit increase in CAI intensity at -0.34 to -0.44 SD on 4th-grade Maths across specifications (SEs 0.19-0.25), with the month-dummy instrument set giving -0.24 (SE 0.106). Every other cell is null: 4th-grade Hebrew, 8th-grade Maths and 8th-grade Hebrew reduced forms are all insignificant, and in 8th grade there was no first stage on teaching practice either, which the authors treat as a specification check. The authors' own summary is that the results "do not support the view that CAI improves learning", that the estimates are "consistently negative and marginally significant" for 4th-grade Maths and "mostly negative" elsewhere, and that in Israel "money spent on CAI would have been better spent on other inputs" - explicitly contrasting this with their own class-size and teacher-training results. The main caveat the authors themselves raise: schools had held the machines for about one school year, so a slow-burn benefit cannot be excluded - but enough time had passed for a large, significant change in instructional method, so the short-run cost is real either way.
Genetic confound
Low-to-medium. Not randomised - programme placement followed municipal priority rankings - but funded and unfunded applicant schools were balanced on pupil characteristics and on 1991 (pre-programme) test scores, and the models control for those lagged scores plus a pupil disadvantage index, so heritable pupil composition is unlikely to drive the negative estimate.
Replication notes
Nobody has re-run Tomorrow-98, so this specific estimate has no direct replication. The broader claim it supports - that buying computers at district/national scale does not raise achievement - has been tested repeatedly since with a split record: Goolsbee & Guryan (2006, E-Rate California) and Dynarski et al. (2007, NCEE) find nulls, Leuven et al. (2007, Dutch RD) finds negatives, and Machin, McNally & Silva (2007, UK ICT funding rule) finds positives in English and science. Recorded as `mixed` for that reason, not because the Israeli result was contradicted on its own terms.
DOI / URL
10.1111/1468-0297.00068

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
FIRST STAGE - Tomorrow-98 receipt on teachers' computer-aided-instruction intensity, 4th grade Maths (0-3 self-reported scale)scale points+0.599 (SE 0.224) with basic controls; +0.563 (SE 0.227) with lagged 1991 scores. The probability of any CAI use rose by 0.234 (SE 0.121). The programme did change teaching.self-report surveyJune 1996 teacher survey, ~9 months after computers arrivedbusiness-as-usualend-of-treatmentbehaviour
FIRST STAGE - Tomorrow-98 receipt on CAI intensity, 8th grade Mathsscale points+0.118 (SE 0.152), ns - no first stage in middle school, despite middle schools being the programme's funding priority. The authors treat the absence of an 8th-grade score effect as a specification check rather than a finding.self-report surveyJune 1996 teacher surveybusiness-as-usualend-of-treatmentbehaviour
REDUCED FORM - 4th grade Maths test score, Tomorrow-98 schools vs other applicant schoolsSD-0.204 (SE 0.089) with basic controls; -0.241 (SE 0.088) adding school-average 1991 scores. Statistically significant; the only significant score effect in the paper.standardizednational test, June 1996business-as-usualend-of-treatmentdomain-skill
REDUCED FORM - 4th grade Hebrew test scoreSD-0.052 (SE 0.088); -0.079 (SE 0.088) with lagged scores. Negative, not significant.standardizednational test, June 1996business-as-usualend-of-treatmentdomain-skill
REDUCED FORM - 8th grade Maths and Hebrew test scoresSD8th Maths -0.080 (SE 0.095) and -0.051 (SE 0.096); 8th Hebrew +0.055 (SE 0.072) and +0.070 (SE 0.072). All insignificant, and there was no first stage on teaching practice in 8th grade.standardizednational test, June 1996business-as-usualend-of-treatmentdomain-skill
2SLS - effect of a one-unit increase in CAI intensity on 4th grade Maths scoresSD-0.340 (SE 0.214) applicants; -0.427 (SE 0.252) applicants with lagged scores; -0.417 (SE 0.251) in the "T-98 vs will-get-T-98" sample; -0.309 (SE 0.187) controlling for 1991 instructional computer use; -0.236 (SE 0.106) using month-of-exposure dummies as instruments (the only clearly significant version, over-identification test 8.8 on 12 df). Consistently negative, imprecise, and robust to controlling for the town priority ranking.standardizednational test, June 1996business-as-usualend-of-treatmentdomain-skill
2SLS - effect of CAI intensity on 4th grade Hebrew scoresSD-0.116 to -0.284 across the same five specifications (SEs 0.14-0.31); none significantstandardizednational test, June 1996business-as-usualend-of-treatmentdomain-skill
NONLINEAR IV using the within-town funding-rank allocation rule, 4th grade MathsSDroughly -0.12 to -0.15 in the applicant sample and -0.21 to -0.27 among applicants with lagged scores across bandwidths of 0.2-0.4 (SEs 0.12-0.17); one estimate marginally significant. Restricting to schools with normalised rank above 0.5 gives larger but far noisier estimates. Broadly consistent with, and slightly smaller than, the 2SLS results.standardizednational test, June 1996business-as-usualend-of-treatmentdomain-skill
OLS association between CAI intensity and test scores (the confounded estimate, for contrast)SD+0.047 (SE 0.035) for 4th grade Maths among applicants, falling to +0.007 (SE 0.034) once town effects are added; the only significant OLS estimate is a NEGATIVE effect on 8th grade Maths with town effects (-0.136, SE 0.070). Effects shrink monotonically as controls are added - the authors read this as omitted-variable bias from school prosperity, since private fundraising for school technology is common in Israel.standardizednational test, June 1996nonenot-applicabledomain-skill
DISPLACEMENT - other school inputs tested for change under the programmenone detectedprogramme status was unrelated to class size, subject coverage, hours of instruction, frequency of teacher training, use of non-computer audio-visual or TV materials, and teacher satisfaction with training and class size. The authors conclude "there was no displacement", which rules out the most convenient benign explanation for the negative maths estimate.self-report surveyJune 1996 teacher surveybusiness-as-usualend-of-treatmentbehaviour

Cited by