The Evidence on Teaching

Generative AI without guardrails can harm learning: Evidence from high school mathematics

Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. · 2025

grade Brctindependentunclearnumbers spot-checked
Sample
Nearly 1,000 students across about fifty classrooms (9th, 10th and 11th grade); 2,848 student-session observations in the main regressions; 13,484 problem-level practice observations and 11,392 problem-level exam observations
Population
Grades 9-11 at TED Ankara Koleji, a large (selective, private) high school in Turkey. Mathematics - four topics drawn from the regular course, comprising about 15% of the semester curriculum. Fall semester 2023-24.
Design
Preregistered (AsPredicted #4DL_Q3J) cluster-randomized trial with the PRIMARY registered outcome specified in advance as the unassisted exam. Students at this school are randomly assigned to classrooms (honors classrooms excluded from the main sample), and each classroom was assigned to one of three arms: control (course notes and textbook, no devices), GPT Base (a plain GPT-4 chat interface built to mimic ChatGPT), or GPT Tutor (the same interface with a system prompt containing the worked solution, common student errors, matched hints, and an instruction not to give the answer away). SEs clustered at the classroom level, the unit of randomization; specification includes session, grader, grade-level and teacher fixed effects plus prior-year GPA. Four 90-minute sessions, each in three parts: teacher-led review (identical across arms), an ASSISTED practice period (the only part the treatment touched), then an UNASSISTED closed-book exam in which each problem was built to mirror a practice problem. Both parts counted toward students' final grades, so effort incentives were real. Problems and exams were designed by the teachers; grading was done by independently hired graders on a teacher-designed rubric, with papers balanced across arms per grader to suppress grader bias. No differential attendance attrition across arms or sessions; five class sessions failed to receive their assigned treatment (laptops did not arrive) and the primary analysis is intention-to-treat. Funded by the Wharton AI & Analytics Initiative, the Fishman-Davidson Center and Wharton Global Initiatives; the authors declare no competing interest. INDEPENDENCE: the authors built both GPT tools, but built them as the experimental manipulation rather than as a product they own or sell, there is no vendor in the study, and the headline result is unflattering to the technology - so `independent`, not developer-led. A correction was issued (PNAS 122(34):e2518204122, 20 Aug 2025) but it fixed only a mis-set author affiliation; no result changed. THE HORIZON IS THE SHARPEST LIMIT: the "learning" outcome is an exam taken minutes after the practice period in the same 90-minute session. The paper says plainly that long-term outcomes were out of reach "due to limitations imposed by our partner school".
Key findings
The archive's canonical performance-vs-learning split, and the reason a practice-time effect size may never be quoted as a learning effect size. With the tool in hand, both AI arms crushed the control on practice problems (+48% for GPT Base, +127% for GPT Tutor). On the unassisted exam minutes later, GPT Base students scored 17% BELOW the control (-0.054 of 1, SE 0.022, p<0.05) - i.e. having had unrestricted GPT-4 during practice left students worse off than having had nothing. GPT Tutor, whose prompt withheld answers and was pre-loaded with the correct solution, was statistically indistinguishable from control (-0.004, SE 0.013): the guardrails removed the harm but bought no learning gain despite a 127% practice advantage. Mechanism: GPT Base returned a correct answer only 51% of the time (42% logical errors, 8% arithmetic), yet its error rate did NOT spill over to exam performance - students were not being misled by wrong answers, they were copying answers without reading them. Students could not see any of this: GPT Base students did not perceive that they had learned less, and GPT Tutor students believed they had performed better when they had not.
Genetic confound
Minimal (randomized). Students are randomly assigned to classrooms at this school and classrooms were randomized to arms; balance on pre-collected covariates is reported.
Replication notes
No direct replication located. The authors themselves note the study is a single deployment in a single school, and the closest thing to corroboration in the literature is conceptual - De Simone et al. (2025) cite Lehmann et al. (2024) as finding the same "crutch" pattern for LLM-assisted coding in a lab setting, which we have not read or verified. Recorded as `unclear` rather than `unreplicated` because we have not systematically searched for direct replications; the performance-vs-learning divergence is, as of this writing, resting on one trial.
DOI / URL
10.1073/pnas.2422633122

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Assisted practice-problem performance, GPT Base vs controlnormalized grade (0-1)+0.137 (SE 0.031, p<0.01) on a control mean of 0.284 (SD 0.287) = +48%; approx +0.48 SDresearcher-designedduring the practice period, tool in handbusiness-as-usualend-of-treatmentdomain-skill
Assisted practice-problem performance, GPT Tutor vs controlnormalized grade (0-1)+0.361 (SE 0.032, p<0.01) on a control mean of 0.284 (SD 0.287) = +127%; approx +1.26 SDresearcher-designedduring the practice period, tool in handbusiness-as-usualend-of-treatmentdomain-skill
UNASSISTED exam performance after access removed, GPT Base vs control (preregistered primary outcome)normalized grade (0-1)-0.054 (SE 0.022, p<0.05) on a control mean of 0.321 (SD 0.277) = -17%; approx -0.19 SD. Access to unguarded GPT-4 during practice left students WORSE than students who never had it.researcher-designedclosed-book exam minutes after the practice period, same sessionbusiness-as-usualend-of-treatmentdomain-skill
UNASSISTED exam performance after access removed, GPT Tutor vs control (preregistered primary outcome)normalized grade (0-1)-0.004 (SE 0.013, not significant). The guardrails eliminated the harm; they produced no detectable learning benefit, despite the same students outperforming control by 127% on practice.researcher-designedclosed-book exam minutes after the practice period, same sessionbusiness-as-usualend-of-treatmentdomain-skill
Students' own perception of how much they learned and how they performedpost-session surveyGPT Base students performed worse on the exam but did not perceive that they had performed worse or learned less; GPT Tutor students did not perform better but perceived that they had. The metacognitive signal was absent in both directions.self-report surveyend of each sessionbusiness-as-usualend-of-treatmentnon-cognitive
GPT-4 (GPT Base) answer accuracy on the practice problemsproportion correct over 10 queries per problem, 57 problemscorrect 51% of the time; logical errors 42%, arithmetic errors 8%. Error rate depressed PRACTICE performance for the GPT Base arm but did not carry over to the paired exam problems - evidence that students were copying rather than reading.researcher-designedmeasured post hoc by re-querying the deployed systemnonenot-applicabledomain-skill

Cited by

Generative AI without guardrails can harm learning: Evidence from high school mathematics · The Evidence on Teaching