From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria
De Simone, M., Tiberti, F., Barron Rodriguez, M., Manolio, F., Mosuro, W., & Dikoru, E. J. · 2025
grade Crctdeveloper-ledunclearnumbers spot-checked
Sample
1,328 randomized (657 treatment, 671 control) from among the 52% of eligible students who volunteered. Only 422 treatment and 337 control completed the endline assessment; regression n ranges 636-654.
Population
First-year senior secondary students (typically age 15) in nine public secondary schools in Benin City, Edo State, Nigeria. English language, following the national curriculum. Six weeks, June-July 2024.
Design
Student-level lottery among volunteers - all first-year senior secondary students in the nine schools were told about the programme and given ten days to express interest; only those who did entered the randomization. So internal validity is decent and external validity is restricted to volunteers (the paper checks and finds no consistent academic-selection pattern between interested and uninterested students across two prior terms). THE BASELINE IS NOT BUSINESS-AS-USUAL IN THE USUAL SENSE. The treatment was TWELVE 90-MINUTE AFTER-SCHOOL SESSIONS - roughly 18 additional hours of supervised, curriculum-aligned instruction in a computer lab with a trained teacher circulating - and the control got nothing additional. Extra instructional time, teacher supervision, computer access and the chatbot all move together and the design cannot separate them; the authors themselves list "additional arms to disentangle the multiple causal mechanisms, including additional instructional time" as future work. Students worked IN PAIRS sharing a computer, so this is not one-to-one AI tutoring either. The model was off-the-shelf Microsoft Copilot (GPT-4 at the time), free, with no fine-tuning and no question bank; the pedagogy lived in teacher-provided session prompts derived partly from Mollick & Mollick (2023) and from desirable-difficulty, retrieval-practice and elaborative-interrogation principles. ATTRITION IS SEVERE AND DIFFERENTIAL: 64% of the treatment arm but only 50% of the control arm completed the endline. The paper acknowledges this is significant and responds with Lee (2005) bounds and inverse-probability weighting, reporting that effects stay positive and significant under both - a genuine and appropriate response, but a 14-point differential is still the largest single threat to the estimate. CONTAMINATION: "some control group students inadvertently gained access to the after-school sessions" because some teachers would not enforce the split, and student-level randomization within schools allows peer spillover; both bias toward zero. OUTCOME: the headline is a RESEARCHER-DESIGNED paper-and-pencil multiple-choice assessment written by experts on the Nigerian curriculum, covering English AND artificial-intelligence knowledge AND digital skills - two of the three sections cover content the treatment group was directly taught and the control group never saw, which is teaching to the test by construction. The more credible outcome is the school's OWN third-term English exam, administered independently by the school and covering the entire year rather than the six weeks of the programme: that gives 0.206 SD. No preregistration is mentioned, and randomization was run "without stratification ... [and] did not incorporate a fixed random seed". INDEPENDENCE: the World Bank team designed the programme, wrote the prompts and the teacher toolkit, implemented the pilot with Edo State, designed the assessment and authored the evaluation - the developers evaluated their own programme, so `developer-led`. Note for anyone reweighting: the "developer" here is the programme team, not a vendor; Microsoft and OpenAI had no role and nothing was being sold. Financial support from the Mastercard Foundation.
Key findings
0.31 SD over six weeks, on an assessment the programme team designed that includes sections on AI and digital skills the treatment group was explicitly taught. The cleaner number is the school's own third-term English exam - independently administered, covering the whole year - at 0.206 SD (SE 0.067), and the English subsection of the study's own test at 0.238 SD. All of it is measured against a control group that got NO additional instruction at all, while the treatment got about 18 extra hours in a computer lab with a supervising teacher; the paper cannot say how much of the effect is the chatbot and how much is the extra time, and says so. Differential attrition was 64% vs 50% endline completion, addressed with Lee bounds and IPW. Effects were LARGER for higher-baseline and higher-SES students, which is the opposite of the equity case usually made for this technology. The paper's most-quoted claim - 1.5 to 2 years of business-as-usual schooling, and 1.2-2.2 SD for a full year of the programme - is a LINEAR EXTRAPOLATION from a per-day dose-response coefficient of about 0.031 SD, assuming no diminishing returns over 36 weeks. It is not a measured effect and should never be cited as one.
Genetic confound
Low-to-medium. Assignment was randomized at the student level among volunteers and baseline exams balanced, so genes cannot differ systematically between arms by design; the 14-percentage-point differential endline attrition reintroduces selection, which Lee bounds and IPW are used to address.
Replication notes
No replication located; the paper itself argues that "numerous questions remain unanswered, underscoring the importance of replicating this study, including with small variations." Recorded as `unclear` rather than `unreplicated` because we did not systematically search. It is a World Bank Policy Research Working Paper (No. 11125, May 2025) and has NOT been peer reviewed.
DOI / URL
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Total score on the programme's own end-of-intervention assessment (weighted) | SD | +0.31 (SE 0.068). ITT, with second-term exam control and school fixed effects. The assessment combines English, AI knowledge and digital skills, two of which the treatment arm was directly taught. | researcher-designed | end of the six-week programme | business-as-usual | end-of-treatment | domain-skill |
| School's own third-term English exam (independently administered, whole-year content) | SD | +0.206 (SE 0.067). The least treatment-aligned outcome in the study and therefore the one the archive treats as the headline. | administrative | end of the school term, after the programme | business-as-usual | end-of-treatment | domain-skill |
| English subscore of the programme's own assessment (the stated primary outcome of interest) | SD | +0.238 (SE 0.068), p < 0.01 | researcher-designed | end of the six-week programme | business-as-usual | end-of-treatment | domain-skill |
| IRT-scaled composite of the programme's own assessment | SD | +0.263 (SE 0.068) | researcher-designed | end of the six-week programme | business-as-usual | end-of-treatment | domain-skill |
| AI-knowledge and digital-skills subscores (content taught only to the treatment arm) | SD | AI knowledge +0.309 (SE 0.077, p<0.01); digital skills +0.139 (SE 0.076, p<0.10). The largest single subscore effect is on knowledge the control group was never given any opportunity to acquire. | researcher-designed | end of the six-week programme | business-as-usual | end-of-treatment | near-transfer |
| Dose-response per day attended, and the extrapolation built on it | SD per additional day of attendance | 0.033 (0.031 under the conservative non-complier assumption). The paper then extrapolates linearly to 1.55 SD for 21 weeks at observed attendance and 2.23 SD for a full 36-week year. These are MODEL PROJECTIONS assuming no diminishing returns, not measurements. | researcher-designed | end of the six-week programme | business-as-usual | end-of-treatment | domain-skill |
| Persistence beyond the programme, and unassisted performance after tool withdrawal | none | NOT MEASURED. All outcomes are at or immediately after the six-week endline. The authors list long-term follow-up as future work. | administrative | not-applicable | business-as-usual | not-applicable | domain-skill |