Disrupting Education? Experimental Evidence on Technology-Aided Instruction in India
Muralidharan, K., Singh, A., & Ganimian, A. J. · 2019
grade Brctindependentmixednumbers spot-checked
Sample
619 students completing the baseline (314 assigned treatment, 305 control), individually randomized; 533 (86%) took the endline; 595 matched to school-exam records. 589 of the 619 were in grades 6-9.
Population
Self-selected volunteer students from public middle schools in low-income neighbourhoods of Delhi, India, mostly grades 6-9, whose families applied for a place at a Mindspark after-school centre. Average grade-6 student was ~2.5 grade levels behind in math; by grade 9 the deficit was ~4.5 grade levels.
Design
THE STEELMAN OF ADAPTIVE SOFTWARE, and it should be recorded as one — the design is strong, the tests are independently built, and the effects are large. All the more reason to be exact about what was actually tested. What it was: an AFTER-SCHOOL programme, six days a week, 90 minutes a day, split 45 minutes of individual work on the Mindspark adaptive platform and 45 minutes of group instructional support from a teaching assistant working with 12-15 students. So the treatment is a bundle of (a) adaptive software, (b) small-group human support, and (c) ADDITIONAL INSTRUCTIONAL TIME on top of the school day — nine hours a week of it. The authors say plainly they cannot experimentally separate the three channels. Their argument that the CAL component is doing the work leans on a contemporaneous Delhi after-school group-tutoring trial with an even longer duration and no test-score impact (Berry & Mukherjee 2016), which is a reasonable but non-experimental inference within this study. `baseline` is therefore recorded as business-as-usual with the caveat that the control group's counterfactual hour was free time, not another lesson. Who it was: 619 volunteers who applied for a subsidised place (INR 200/month) — a self-selected, urban, motivated sample. Attendance among lottery winners averaged 58% (about 50 of a possible 86 days). The headline +0.37/+0.23 are ITT estimates at that attendance rate; the paper's IV estimate for 90 days of attendance (+0.59-0.60 math, +0.36-0.39 Hindi) requires additional assumptions and should not be treated as the observed effect. Measurement is a genuine strength: paper-and-pencil tests DESIGNED INDEPENDENTLY BY THE RESEARCH TEAM, IRT-linked across grades and across baseline/endline so they measure across the whole achievement range rather than at grade level. The paper also checks the item-overlap concern (some test items exist in the 45,000-item Mindspark bank) and finds ITT effects of similar size on items from EI assessments and from other sources. The independent-test-versus-school-exam contrast is one of the most instructive things in the paper and is recorded as its own row. INDEPENDENCE: Educational Initiatives (EI) built and operated Mindspark and supplied the software and the usage logs; EI staff are thanked, not co-authors. The evaluation was designed, run and written by academics at UCSD, UCL and J-PAL, funded by J-PAL's Post-Primary Education initiative; EI's centre operations were funded by the Central Square Foundation, Tech Mahindra Foundation and Porticus. Under this archive's settled distinction that is `independent`, with the vendor relationship recorded here rather than in the enum. Pre-registered (AEA RCT Registry trial 980). Numbers below are from the full text of NBER working paper w22923 (Dec 2016, rev. July 2017), which reports 0.36/0.22; the published AER 109(4) version reports 0.37/0.23. Both are recorded. GRADE: B, not A, and the reason is measurement rather than identification. The randomization, pre-registration, balance and 86% follow-up are all A-grade. But the headline effect rests on tests the research team built for the study, and the only externally-generated measure available — the schools' own examinations — shows +0.19 SD in Hindi and nothing in math. The paper has a good explanation for that gap (the software correctly worked several grade levels below the enrolled grade, which a grade-level exam cannot see), and this archive accepts it; but a single-site trial whose headline number exists only on the researchers' own instrument does not sit in the A row. The larger multi-site scale-up, Muralidharan & Singh (2025), does.
Key findings
A 4.5-month after-school programme combining adaptive software with small-group support raised independently-measured math scores by 0.37 SD and Hindi by 0.23 SD — roughly double and 2.5 times the control group's own value-added over the same period, and the largest credible technology effect in the literature. Absolute gains were the same for strong and weak students, but the relative gain was far larger at the bottom because control-group students in the bottom third made no measurable progress at all across the school year. The caveats are not small: it was an after-school supplement adding nine hours a week of instruction, attendance was 58%, the sample was 619 self-selected urban volunteers, and the treatment bundles software with a human teaching assistant. On the schools' own grade-level exams the same students gained 0.19 SD in Hindi and nothing in math — because the software, correctly targeting students' actual level, spent the math time several grades below grade level. The same product delivered inside government schools six years later produced 0.22 SD in eighteen months.
Genetic confound
Minimal (individually randomized by lottery within a self-selected applicant pool; balance shown on age, gender, SES and baseline scores, and among endline attriters).
Replication notes
Directionally consistent with Banerjee et al. (2007), the other large positive CAL result from India, and with the Chinese CAL trials (Mo et al. 2014, 2015). But the decisive test is the same product moved into the setting it would actually be deployed in, and that test exists: Muralidharan & Singh (2025, NBER w34205) evaluated the adapted IN-SCHOOL Mindspark model in 40 treated Rajasthan government schools (~6,500 students annually, a sample over twenty times larger, not self-selected), and found +0.22 SD in math and +0.20 SD in Hindi AFTER 18 MONTHS — against +0.37 and +0.23 after 4.5 months here. Per unit of time the effect fell by roughly three-quarters, and it did so precisely where the added-instructional-time channel was removed (the school model SUBSTITUTES 25-50% of math and Hindi class time rather than adding an after-school block). That is the efficacy-to-effectiveness curve inside a single product, and it is the reason this result should not be quoted at 0.37 SD as a general property of adaptive software. Recorded separately as edt-muralidharan-singh-2025-mindspark-rajasthan-scale.
DOI / URL
10.1257/aer.20171112
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Mathematics, independent IRT-linked test | SD | +0.37 SD (published AER; +0.36 in the working paper), ITT over 4.5 months at 58% average attendance. Roughly twice the control group's own test-score value-added over the period. | researcher-designed | endline, February 2016, after 4.5 months | business-as-usual | end-of-treatment | domain-skill |
| Hindi (language), independent IRT-linked test | SD | +0.23 SD (published AER; +0.22 in the working paper), ITT over 4.5 months. About 2.5 times the control group's value-added. | researcher-designed | endline, February 2016 | business-as-usual | end-of-treatment | domain-skill |
| Mathematics and Hindi, IV estimate for 90 days of attendance | SD | +0.59-0.60 SD math and +0.36-0.39 SD Hindi. This is an extrapolation from a dose-response relationship under additional assumptions, not an observed treatment arm; 90 days corresponds to 80% attendance for half a school year, well above the 58% actually achieved. | researcher-designed | extrapolated to 90 days of attendance | business-as-usual | end-of-treatment | domain-skill |
| School's own end-of-year grade-level examinations | SD | Hindi +0.19 (SE 0.089, p<0.05); math +0.058 (SE 0.076, ns); science +0.077, social science +0.10, English +0.080, aggregate +0.097 — all insignificant. The contrast with the independent test is the paper's own headline methodological point: because Mindspark correctly targeted math content several grade levels below the enrolled grade, the math gains were invisible on a grade-level test. It cuts both ways for this archive — the independent IRT test is the better measure of what was learned, and the school exam is the better measure of what the school system will notice. | administrative | annual school examinations, March 2016 | business-as-usual | end-of-treatment | domain-skill |
| Heterogeneity by baseline achievement | SD | absolute ITT effects do not vary by baseline test score, gender or household SES. But control-group value-added for the bottom third of the within-grade baseline distribution is statistically indistinguishable from ZERO across the whole school year, so the same absolute gain represents a far larger relative gain for weak students. | researcher-designed | endline, February 2016 | business-as-usual | end-of-treatment | domain-skill |
| DOSE and instructional time added: attendance | days / hours | six days a week, 90 minutes a day, of which 45 minutes on software and 45 minutes with a teaching assistant in groups of 12-15 — about nine hours a week ADDED on top of the regular school day. Mean attendance among lottery winners 58%, roughly 50 of a possible 86 days. This is the single most important qualifier on the headline number: it is not a better use of an existing hour, it is an extra hour that is also better used. | administrative | over the 4.5-month intervention | none | end-of-treatment | behaviour |
| DISPLACEMENT: private tutoring outside the programme | pp | no differential change in the incidence of paid private tutoring among lottery winners in the post-intervention period (October 2015 - March 2016), measured by monthly parent phone surveys. Both arms increased tutoring toward exam season. So the programme's instructional time was net added, not substituted for tutoring the family was already buying. | self-report survey | monthly, July 2015 - March 2016 | business-as-usual | end-of-treatment | behaviour |
| SCALE-UP / EFFECTIVENESS: the same product delivered in government schools (Muralidharan & Singh 2025) | SD | +0.22 SD math and +0.20 SD Hindi after EIGHTEEN months in 40 treated Rajasthan government schools (~6,500 students annually), versus +0.37/+0.23 after 4.5 months here. Year-1 (~6 months) effects in Rajasthan were +0.15 math and +0.11 Hindi. Per unit of exposure time the effect is roughly a quarter of the Delhi estimate, and the school model substitutes 25-50% of math and Hindi instructional time rather than adding an after-school block. Recorded here so the efficacy figure is never read without its effectiveness counterpart. | researcher-designed | 18 months, Rajasthan government schools, 2017-2019 | business-as-usual | end-of-treatment | domain-skill |