Effective and Scalable Math Support: Experimental Evidence on the Impact of an AI-Math Tutor in Ghana
Henkel, O., Horne-Robinson, H., Kozhakhmetova, N., & Lee, A. · 2024
grade Crctdeveloper-ledunclearnumbers spot-checked
Sample
637 students took the baseline (336 treatment, 301 control); 477 completed both baseline and endline (treatment 236-237, control 241 - the paper's text and its Table 1 disagree by one). The UNIT OF RANDOMIZATION was the school, so the effective n is 11 clusters (5 treatment, 6 control).
Population
Grades 3-8 (abstract says 3-9), 11 schools in the Rising Academies network in Ghana, selected for similarity in geography, demographics, curriculum and teaching methods. Mathematics. Feb 2023 baseline to late Aug 2023 endline - roughly 8 months, the longest exposure window anywhere in this cluster.
Design
SCHOOL-LEVEL cluster randomization (5 treatment schools vs 6 control) analysed with a STUDENT-LEVEL INDEPENDENT-SAMPLES T-TEST AND NO CLUSTERING ADJUSTMENT. That is the single most important fact about this study: with 11 clusters the reported p < 0.001 is not interpretable, and METHODOLOGY.md names cluster-randomization errors as a grade-C condition explicitly. Treatment students used Rori for two 30-minute sessions per week during study hall, supplementing normal math instruction; control students had normal instruction and no Rori during study hall. Teachers were present only for technical support. ATTRITION IS SUBSTANTIAL AND DIFFERENTIAL: 160 of 637 students (25%) did not sit the endline, and it fell harder on the treatment arm (336 to ~236, ~30%) than the control (301 to 241, ~20%). Dropouts also had markedly lower baseline scores (mean 18.53 vs 22.26), so attrition was ability-selective, and the differential direction would inflate a treatment-arm growth score. The paper dismisses this - "the effect size ... is calculated based on growth scores of students who completed both tests, so it would not be impacted by this difference" - which does not follow, and no bounding or weighting analysis is offered. OUTCOME: the SAME 35-item mixed multiple-choice and open-response assessment at both timepoints for all grades, covering grade 3-5 numeracy and algebra from the Global Proficiency Framework. Rori's own curriculum of 500+ micro-lessons is built on the same Global Proficiency Framework, so the test is aligned to the product's content; it is not an independent standardized measure. Ceiling effects in the higher grades are acknowledged. IS RORI ACTUALLY AN LLM TUTOR? Record the ambiguity rather than resolving it: the instructional content is a FIXED curriculum of micro-lessons, each with a pre-written explanation and a scaffolded question series, and a pre-written hint then a pre-written worked solution when a student fails - that is a scripted mastery-learning sequence, not LLM-generated instruction. The paper says Rori uses "various natural language processing (NLP) methods, including specialized language models (LLMs)" to let students chat in natural language rather than navigate a button menu, and to interpret answer attempts. So the language model is largely an INTERFACE layer over a scripted tutor. This study is weak evidence about LLM tutoring as such and better evidence about mobile-delivered scripted practice. INDEPENDENCE: Rising Academies built Rori, a Rising Academies employee (Horne-Robinson) is a co-author, and the trial was run inside Rising Academies' own schools on Rising Academies' own students with an assessment aligned to Rising Academies' curriculum. Oxford and J-PAL North America affiliations on other authors do not make this an independent evaluation - no independent body ran or authored it, so `developer-led`. Published in the AIED 2024 POSTERS AND LATE BREAKING RESULTS track (CCIS vol 2150, pp. 373-381), not the main peer-reviewed proceedings. Not preregistered.
Key findings
The 0.36 SD everyone quotes is a raw growth-score difference from a two-arm comparison of 11 schools tested with an unclustered t-test, in the developer's own schools, on a product-aligned assessment, with 25% differential ability-selective attrition and no adjustment for any of it. Treatment growth 5.13 raw points (SD 7.03) vs control 2.12 (SD 6.30), a 3.01-point difference, giving Cohen d = 0.36 on the pooled baseline SD. The design's real strength is duration - roughly 8 months of twice-weekly use, by far the longest exposure in the LLM-tutoring cluster - and its real weakness is that the analysis treats 477 students as 477 independent observations when they are nested in 11 randomized schools. No measurement was ever taken with the tool removed, so this says nothing about whether the gain is learning or tool-assisted performance (compare Bastani et al. 2025). Treat 0.36 as an upper bound on a quantity that has not actually been estimated with correct standard errors.
Genetic confound
Medium. Randomization was at the school level with only 11 clusters and no covariate-balanced stratification reported, so cluster-level differences are not ruled out by design; baseline scores, age and gender were balanced, but 25% differential ability-selective attrition reintroduces selection between baseline and endline.
Replication notes
No replication located. The paper itself is explicitly framed as preliminary ("insight from just the first year of intervention") and calls for study of Rori outside the Rising Academies network, which is the replication that matters and which we have not found. Recorded as `unclear` rather than `unreplicated` because we did not systematically search.
DOI / URL
10.1007/978-3-031-64315-6_34
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Math growth score (endline minus baseline raw score), Rori vs regular instruction | d | 0.36 (computed on the pooled baseline SD per Morris 2008). Treatment growth 5.13 (SD 7.03) vs control 2.12 (SD 6.30); raw difference 3.01 points on a 35-point test. Reported p < 0.001 from an unclustered independent-samples t-test - NOT a valid inference given school-level randomization with 11 clusters. | researcher-designed | endline, late August 2023, after roughly 8 months of exposure | business-as-usual | end-of-treatment | domain-skill |
| Raw assessment scores by arm and timepoint | points out of 35 | control 20.20 (SD 8.81) to 22.32 (SD 8.06); treatment 20.29 (SD 8.72) to 25.42 (SD 7.25). Baseline equivalence t(475) = -0.17, p = 0.87. | researcher-designed | baseline Feb 2023 and endline Aug 2023 | business-as-usual | end-of-treatment | domain-skill |
| Attrition (the threat the paper does not bound) | proportion lost between baseline and endline | 160 of 637 lost (25%); treatment ~30% vs control ~20%; dropouts' mean baseline score 18.53 (SD 7.70) vs completers' 22.26 (SD 7.57). Ability-selective and differential; no Lee bounds, inverse-probability weighting or sensitivity analysis reported. | administrative | baseline to endline | none | not-applicable | behaviour |
| Unassisted performance after the tool is withdrawn | none | NOT MEASURED. The endline was taken while students still had ongoing access, so the study cannot separate learning from tool-supported performance. | researcher-designed | not-applicable | business-as-usual | not-applicable | domain-skill |