The Evidence on Teaching

Remedying Education: Evidence from Two Randomized Experiments in India

Banerjee, A. V., Cole, S., Duflo, E., & Linden, L. · 2007

grade Brctindependentreplicatednumbers spot-checked
Sample
Vadodara CAL arm: 111 municipal primary schools (55 treatment, 56 comparison) in year 2, 56 vs 55 in year 3, all grade four; 5,945 grade-four children in the year-2 sample and 5,945 in year 3; regression samples of 5,732 (year 2) and 4,688 (one-year follow-up). Over 15,000 students across both experiments and three years.
Population
Grade-four children in Vadodara Municipal Corporation primary schools, urban Gujarat, India — poor urban-slum government schools with very low baseline achievement. (The paper's other arm, the Balsakhi remedial tutor, covered grades three and four in Vadodara and Mumbai and is not an ed-tech intervention.)
Design
INSTRUCTIONAL intervention, not an access one — this is the paper that separates the two. The computers were largely ALREADY IN THE SCHOOLS; what the programme added was structured, level-targeted content plus a paid instructor to keep children on task. Comparison schools were free to use their own computers and were never observed doing so for instruction, which is exactly the Barrera-Osorio/Cueto mechanism showing up as the counterfactual here. DOSE AND THE INSTRUCTIONAL-TIME QUESTION, which the archive must get right: two hours per week of shared computer time, two children per machine — ONE HOUR DURING CLASS TIME AND ONE HOUR IMMEDIATELY BEFORE OR AFTER SCHOOL. So half the dose substituted for regular instruction and half was added instructional time. The +0.35/+0.47 SD is therefore not a like-for-like productivity gain against the same hour; roughly half of it is bought with extra time. `baseline` is recorded as business-as-usual because the comparison group got the normal school day and nothing else, but a fair reading discounts for the added hour. Design: school-level randomization stratified on Balsakhi treatment status, gender, language of instruction and prior-year math scores; the two experiments were crossed in year 2 so their effects and interaction could be separately estimated (the interaction is small, negative and insignificant). Treatment and comparison groups were SWITCHED between year 2 (2002-03) and year 3 (2003-04), which is why the "second year" estimate is not a two-year-exposure estimate for the same children. Attrition 3.4-3.8% in the treatment years and about 20% at the one-year follow-up, balanced across arms. MEASURE TYPE IS THE MAIN DISCOUNT. The tests were designed and administered by Pratham, the implementing NGO, and built around the competencies the Vadodara Municipal Corporation prescribes for grades one to four — the same competency list the software's games were selected to emphasise. That is a researcher-designed, treatment-aligned measure in the sense METHODOLOGY.md warns inflates effects roughly twofold against independent standardized tests. Pratham took real care with test administration (three staff per room against cheating, home make-up tests to hold down attrition), but the alignment issue is structural, not procedural. Numbers below are from the full text of NBER working paper w11904; the published QJE 122(3):1235-64 version carries the same headline estimates. Funded by ICICI, the World Bank, the Sloan Foundation and the MacArthur Network; Pratham developed and ran the programme, the academic authors designed and ran the evaluation. GRADE: B rather than A. The paper as a whole is a large multi-site experiment, but the CAL ARM — the part this record is about — is a single-city, single-grade RCT in 111 Vadodara schools whose outcome measure was built by the implementing NGO around the same competency list the software targeted. Strong single RCT, discounted for treatment-aligned measurement; the Balsakhi arm's two-city replication does not transfer to the CAL estimate.
Key findings
A shared-computer math game programme run inside Indian municipal primary schools raised math scores by 0.35 SD in its first year and 0.47 SD in its second, with no effect on language (point estimates near zero, as expected since the software was math-only) — and, unlike the remedial-tutor arm, it helped children at every point in the ability distribution roughly equally. Then it went away. One year after the programme ended, the math effect was +0.097 SD (SE 0.053) and the combined score effect was +0.008; roughly 80% of the gain had evaporated in twelve months. The authors read the residual 0.10 SD as encouraging; the honest reading for a database organised around fadeout is that a large end-of-treatment effect on an implementer-designed, curriculum-aligned test decayed to something indistinguishable from noise within a year of withdrawal. Two further facts belong on the record: half the dose was extra instructional time outside the school day, and the comparison schools already had computers and never used them for teaching.
Genetic confound
Minimal (school-level randomization, stratified, with the crossed Balsakhi design and balance checks).
Replication notes
The CAL result is the founding positive of the developing-country computer-assisted-learning literature and has been broadly reproduced in direction: Muralidharan, Singh & Ganimian (2019) found +0.37 SD in math from an adaptive after-school programme in Delhi, and Muralidharan & Singh (2025) found +0.22 SD after 18 months from the same product delivered inside rural government schools. It contrasts sharply with the null CAL/computer results from richer settings and from access-only interventions (Angrist & Lavy 2002; Barrera-Osorio & Linden 2009; Cristia et al. 2017). The FADEOUT recorded here has itself been replicated in the Chinese CAL literature (Mo et al. 2015). What has never been replicated is a durable CAL effect.
DOI / URL
10.1162/qjec.122.3.1235

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
CAL — mathematics, first year of the programme (2002-03)SD+0.35 SD (value-added specification). Raw post-test difference +0.32 SD (SE 0.087). Effect on the combined math+language score +0.21 SD.researcher-designedend of the first programme yearbusiness-as-usualend-of-treatmentdomain-skill
CAL — mathematics, second year of the programme (2003-04, treatment/comparison switched)SD+0.47 SD (value-added specification). Raw post-test difference +0.58 SD, but pre-test scores were already 0.13 SD higher in the treatment group that year, which is why the difference-in- difference/value-added estimate is the one to use. Combined score +0.23 SD.researcher-designedend of the second programme yearbusiness-as-usualend-of-treatmentdomain-skill
CAL — FADEOUT, mathematics one year after leaving the programmeSD+0.097 (SE 0.053) difference-in-difference / +0.092 (SE 0.045) value-added, i.e. about a fifth of the 0.47 SD end-of-treatment effect and significant only at the 10% level. By initial ability: bottom third +0.085 (SE 0.050), middle third +0.103 (SE 0.061), top third +0.111 (SE 0.079). Tested March 2004, one year after the programme was withdrawn; attrition 20%, balanced.researcher-designedMarch 2004, one year after the end of exposurebusiness-as-usual1-2yrdomain-skill
CAL — FADEOUT, language and combined score one year after leaving the programmeSDlanguage -0.078 (SE 0.054), combined total +0.008 (SE 0.050). The overall test-score effect one year out is a zero; only the math component retains anything.researcher-designedMarch 2004, one year after the end of exposurebusiness-as-usual1-2yrdomain-skill
CAL — language/verbal achievement during treatmentSDno discernible effect, point estimates near zero in both years (year 2 +0.014, SE 0.073). The software targeted math only, and the effect stayed inside the trained domain — a clean demonstration of the absence of even one-subject-over transfer.researcher-designedend of each programme yearbusiness-as-usualend-of-treatmentfar-transfer
CAL — distribution of gains by initial abilitySDyear 2 math gains: bottom third +0.417 (SE 0.107), middle third +0.341 (SE 0.088), top third +0.319 (SE 0.086). Unlike the Balsakhi tutor arm (whose effect was twice as large at the bottom as at the top), the software helped everyone about equally. Relevant to the claim that adaptive software is uniquely good for weak students — here it simply was not differential.researcher-designedend of the first programme yearbusiness-as-usualend-of-treatmentdomain-skill
CAL — specific mathematics competencies masteredppabout +13 percentage points each on grade-one and grade-two competencies in year 2, and +7.7 points on grade-three competencies in year 3 (only 1.3% of fourth graders passed those at pre-test, and 8.2% of comparison children at post-test). The gains are concentrated in below-grade-level basics, the same pattern later found for Mindspark.researcher-designedend of each programme yearbusiness-as-usualend-of-treatmentdomain-skill
DISPLACEMENT / dose: instructional time added versus substitutedhours per weektwo hours per week of shared computer time (two children per machine) — one hour taken from class time and one hour before or after school. So the intervention ADDED roughly one hour a week of instruction on top of a four-hour school day and substituted for another. Attendance effects were null in year 2 and +2.5pp (SE 1.5, p<0.10) in year 3, so the extra hour was largely genuine additional exposure rather than recovered absence.administrativethroughout the programme yearsbusiness-as-usualend-of-treatmentbehaviour
Counterfactual computer use in comparison schoolsqualitative observationcomparison schools already had computers and were free to use them; the researchers "did not observe those schools employ them for instructional purpose". The contrast being estimated is therefore structured software plus a paid instructor versus idle hardware — the same failure mode Barrera-Osorio & Linden (2009) and Cueto et al. (2025) documented as the reason access interventions do nothing.mixedthroughout the programme yearsbusiness-as-usualend-of-treatmentbehaviour

Cited by