AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting
Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. · 2025
grade Crctdeveloper-ledunclearnumbers spot-checked
Sample
233 enrolled in the course; 194 eligible (consent + participation in both conditions + all four tests). Post-test observations 142 (AI) and 174 (in-class), 316 pre-test observations - these are lesson-level observations across the two crossover weeks, not 316 students.
Population
Undergraduates in Physical Sciences 2 (PS2), an introductory physics-for-life-sciences course at Harvard University, Fall 2023. TWO lesson topics only - surface tension (week 1) and fluid flow (week 2) - chosen for being self-contained and requiring no more than high-school maths.
Design
Crossover RCT over two consecutive weeks: students randomized to two groups (with the constraint that established peer-instruction groups of 2-3 students stayed together, so the effective randomization unit is the peer group, not the student); each group got the AI lesson one week and the in-class active-learning lesson the other; pre-test before each lesson, post-test after. IRB23-0797. NOT PREREGISTERED, and no follow-up of any kind - each post-test was taken immediately after its lesson, so nothing here speaks to retention, and nothing tests unassisted performance after the tool is removed (contrast Bastani et al. 2025, which does exactly that and gets the opposite sign). WHAT THE COMPARISON ACTUALLY CONTRASTS is not "AI tutor vs teacher": the AI condition was asynchronous, AT HOME, self-paced, with unlimited time, and opened with professionally produced video (Harvard Bok Center production studio, instructor with documentary experience); the control was a live, synchronous, fixed-length 60-minute class. Several ingredients move together and the design cannot separate them. The AI tutor ("PS2 Pal") was not a general-purpose LLM tutor either - the authors state that a system prompt could not reliably scaffold multi-part problems, so the PLATFORM walks students question-by-question, and the prompts were pre-loaded with comprehensive step-by-step human-written solutions ("accuracy of our AI tutor relied on pre-written answers"). OUTCOME MEASURE: a short researcher-designed post-test quiz with an acknowledged CEILING EFFECT, which is why the paper reports a linear-regression estimate (0.63) and a larger quantile-regression range (0.73-1.3) that avoids the ceiling. Tests were written by a team member not involved in building the AI or teaching the lessons - a real but partial safeguard, since that person is still on the author team. INDEPENDENCE: the authors conceived, built and engineered the tutor platform (GK: software conceptualization, design and engineering), wrote its prompts, taught the course as its co-instructors, wrote the tests, ran the trial and authored the paper. That is developer-led under this archive's definition, not merely developer-involved; declared competing interests are none, and no external funder is named. Review turnaround was 25 March to 7 April 2025 (13 days) with publication 3 June 2025. Students received equal participation credit for both conditions and were told test performance would not affect their grade.
Key findings
Real headline, unreal generalization. Students scored substantially higher after the AI-tutored lesson than after the equivalent in-class active-learning lesson: median post-score 4.5 vs 3.5 against a combined pre-test median of 2.75, Mann-Whitney z = -5.6, p < 1e-8; regression estimate 0.63 SD, quantile-regression estimate 0.73-1.3 SD once the post-test ceiling is handled. They also did it in less time (median 49 minutes vs a 60-minute class), and reported higher engagement and motivation. But the estimate covers ONE course, TWO topics, ONE session each, at Harvard, on a short researcher-designed quiz with a ceiling, with no retention measure and no test of unassisted performance later - and the "AI tutor" is a bespoke platform carrying human-written step-by-step solutions, sequenced by software, whose prompts were authored by the instructors who ran and wrote up the evaluation. The number is what it is; it is not evidence that giving students an LLM improves learning, and it is not comparable to a semester-long standardized-test effect.
Genetic confound
Minimal (randomized crossover; every student experiences both conditions, and pre-tests establish baseline for each lesson).
Replication notes
We have not located a replication, and the authors themselves call for one - their own discussion asks for "studies that explicitly replicate known in-class active learning results" before the design is taken as transferable. Recorded as `unclear` rather than `unreplicated` because we did not systematically search. Note that the result points the opposite way from Bastani et al. (2025) on the same technology; that is not a failed replication (different tools, subjects, ages and outcomes) but it is the reason no single number here should travel.
DOI / URL
10.1038/s41598-025-97652-6
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Post-test score, AI tutor vs in-class active learning (linear regression with controls) | SD | 0.63 (p < 1e-8), controlling for pre-test, midterm, FCI, prior ChatGPT experience, topic, test version and time on task; clustered at the student level. The authors call this an UNDERESTIMATE because of a ceiling effect in the post-test. | researcher-designed | immediately after the lesson | active-alternative | end-of-treatment | domain-skill |
| Post-test score, AI tutor vs in-class active learning (quantile regression avoiding the ceiling) | SD | 0.73 to 1.3 - the source of the widely quoted "d = 0.73". This is a range from a ceiling-correction analysis on a short in-house quiz, not a confidence interval. | researcher-designed | immediately after the lesson | active-alternative | end-of-treatment | domain-skill |
| Post-test score distribution (nonparametric) | median score and rank-sum test | AI median 4.5 (N=142) vs in-class median 3.5 (N=174), against a combined pre-test median of 2.75 (N=316); Mann-Whitney z = -5.6, p < 1e-8 | researcher-designed | immediately after the lesson | active-alternative | end-of-treatment | domain-skill |
| Time on task (what the technology displaced) | minutes | AI group median 49 minutes, 70% under 60 minutes, against a fixed 60-minute in-class lesson (of which 15 minutes went to the pre/post tests). Notably, time on task did NOT correlate with post-test score in the AI group despite a spread from under 10 to over 200 minutes. | administrative | during the lesson | active-alternative | end-of-treatment | behaviour |
| Engagement, motivation, enjoyment, growth mindset | 5-point Likert agreement | engaged 4.1 (SD 0.98) vs 3.6 (SD 0.92), t(311) = -4.5, p < 0.0001; motivated 3.4 (SD 1.0) vs 3.1 (SD 0.86), t(311) = -3.4, p < 0.001; enjoyment and growth mindset not significantly different | self-report survey | immediately after the lesson | active-alternative | end-of-treatment | non-cognitive |
| Retention, and unassisted performance after the tool is withdrawn | none | NOT MEASURED. Every outcome is an immediate post-test taken with the lesson still in working memory. The study cannot distinguish learning from performance, which is precisely the distinction Bastani et al. (2025) found reverses the sign. | researcher-designed | not-applicable | active-alternative | not-applicable | domain-skill |