The Evidence on Teaching

Do intelligent tutoring systems benefit K-12 students? A meta-analysis and evaluation of heterogeneity of treatment effects in the U.S.

Leite, W. L., Zhang, H., Rana, S., Hao, Y., Hatch, A. D., Kong, L., & Kuang, H. · 2025

grade Cmeta-analysisindependentmixednumbers spot-checked
Sample
18 studies, 77 effect sizes, 11 different ITS; US K-12 schools only
Population
US elementary, middle and high schools; mathematics (22 effect sizes) and reading/writing (55 effect sizes). Restricted to the US so that policy and schooling context are held roughly fixed.
Design
Added to the archive because it is the only meta-analysis of this literature restricted to US K-12 school studies, and because it codes the two moderators that decide the answer - outcome measure type and measurement timing - in the same model. Multivariate multilevel meta-regression with random effects of studies and fixed effects of ITS name, plus a MetaForest (random-forest) moderator-importance analysis over 15 coded moderators (cross-validated R-squared 0.25). Internal validity is assessed against What Works Clearinghouse standards: only 8 of 18 studies reported differential attrition at all (5 meet WWC standards without reservations, 3 with reservations, and the status of the other 10 is unknown), which is a fair description of the state of this literature rather than a defect of the review. Publication bias examined by funnel plot only - the authors report a cluster of precise estimates between 0 and 0.5 and spread outside the funnel that they attribute to genuine heterogeneity; no trim-and-fill or PET-PEESE. Grade C: mixed-quality pool, small k, and preprint status.
Key findings
Across 18 US K-12 studies and 77 effect sizes, intelligent tutoring systems move learning g = 0.271 (SE 0.011, p = .001) - smaller than most prior ITS meta-analyses and, the authors note, the first estimate restricted to US K-12 studies screened against WWC standards. Two moderators do the work the archive cares about. Outcome measure: standardized tests g = 0.200 (95% CI 0.166 to 0.234, k = 61) versus researcher-developed tests g = 0.483 (95% CI -0.154 to 1.120, k = 16, p = .098) - a 2.4x gap, in the same direction as every other meta in this cluster except Ma et al. (2014). Measurement timing: immediately after the intervention g = 0.460 versus at the end of the school year g = 0.155, both significant - the effect loses two-thirds of its size in the months between the posttest and the year-end test. Duration runs the same way: up to two weeks 0.334, two weeks to three months 0.449, three to six months 0.225, more than six months 0.110 and not significant. After-school use produced essentially nothing (g = 0.014).
Genetic confound
Medium - the pool is mixed randomized and quasi-experimental, and only 8 of 18 studies reported differential attrition, so differential dropout between arms is unassessed in more than half the pool.
Replication notes
This is the study that adjudicates the cluster's central disagreement, and it lands between the camps while confirming the mechanism both sides argued about. Its standardized-measure estimate (g = 0.200) is close to Kulik & Fletcher's standardized-measure figure (0.13) and to Cheung & Slavin's large randomized figure (+0.06 to +0.12); its researcher-developed-measure estimate (0.483) reproduces the alignment inflation Kulik & Fletcher, Cohen/Kulik/Kulik (1982) and Xu et al. (2019) all report, and directly contradicts Ma et al. (2014), who found no gap. Its end-of-school-year estimate (0.155) versus immediate-posttest estimate (0.460) reproduces the archive's standard fadeout signature. NOT peer-reviewed: presented at AERA 2024 and posted to arXiv in November 2025; as of July 2026 no journal version was located.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
ITS effect on US K-12 learning outcomes, overallg0.271 (SE 0.011, p = .001), accounting for random effects of studies and fixed effects of ITS namemixedend of intervention or end of school yearbusiness-as-usualend-of-treatmentdomain-skill
Effect measured on standardized testsg0.200 (SE 0.009, 95% CI 0.166 to 0.234, p = .001), k = 61standardizedend of intervention or end of school yearbusiness-as-usualend-of-treatmentdomain-skill
Effect measured on researcher-developed testsg0.483 (SE 0.213, 95% CI -0.154 to 1.120, p = .098), k = 16 - 2.4x the standardized estimate and not significant, with an interval an order of magnitude widerresearcher-designedend of interventionbusiness-as-usualend-of-treatmentdomain-skill
Measurement timing (immediate posttest vs end of school year)gimmediately after the intervention 0.460 (SE 0.046, 95% CI 0.244 to 0.677, k = 23) vs end of school year 0.155 (SE 0.010, 95% CI 0.125 to 0.185, k = 47); both significant. Two-thirds of the effect is gone by the year-end test.mixedimmediate posttest vs end of school yearbusiness-as-usualunder-1yrdomain-skill
Intervention durationgup to 2 weeks 0.334 (ns); 2 weeks to 3 months 0.449 (ns); 3 to 6 months 0.225 (p = .006); more than 6 months 0.110 (ns) - longer exposure, smaller effect, the standard hothouse gradientmixedend of interventionbusiness-as-usualend-of-treatmentdomain-skill
How the ITS was usedgas the main instructional method 0.294 (p < .001, k = 45); as individual homework 0.317 (p < .001, k = 10); as a separate activity 0.235 (ns); AFTER SCHOOL 0.014 (ns, k = 6) - the voluntary out-of-school deployment does nothingmixedend of interventionbusiness-as-usualend-of-treatmentdomain-skill
Subject and subgroupgmathematics 0.253 (95% CI 0.224 to 0.281) vs reading/writing 0.323 (95% CI 0.218 to 0.428); effects similar across elementary and middle schools and for low-achieving students (g = 0.270-0.278) but lower in studies including rural schoolsmixedend of intervention or end of school yearbusiness-as-usualend-of-treatmentdomain-skill

Cited by