The Evidence on Teaching

A Meta-Analysis of ChatGPT's Influence on Learning Achievement

Doo, M. Y., & Park, Y. · 2026

grade Cmeta-analysisindependentmixednumbers spot-checked
Sample
22 publications yielding 86 effect sizes, total N = 6,992. Individual study sample sizes ranged from 31 to 269, with a MEAN of 86.6.
Population
Mostly undergraduates (70.9% of effect sizes), plus adults and clinical teachers (16.3%), middle and high school students (10.5%) and elementary students (2.3%, k = 2). Twelve countries, heavily concentrated - Taiwan 31.8% of publications, mainland China 18.2%, Turkey 9.1%. Subjects span languages, STEM, medicine, business, music and ethics. Studies published Nov 2022 to Dec 2024; 5 papers from 2023 and 17 from 2024.
Design
GRADE THIS ON WHAT IT POOLS, per METHODOLOGY.md ("a meta-analysis of confounded studies inherits the confounding"). The pool, by the authors' own coding of the 86 effect sizes: 55.8% QUASI-EXPERIMENTAL (k=48); 20.9% one-group pretest-posttest with NO CONTROL GROUP (k=18); 11.6% true experimental with random assignment (k=10); 9.3% posttest-only experimental (k=8); 2.3% CORRELATIONAL (k=2). Random assignment was used for only 44.2% of effect sizes. Roughly a quarter of the pooled evidence comes from designs the archive's hierarchy grades D on their own. NO QUALITY APPRAISAL WAS PERFORMED AT ALL - the authors state that "since the number of studies that met our inclusion criteria was limited, we did not exclude any based on quality," and no risk-of-bias instrument was applied. NOT ONE STUDY IN THE POOL USED AN INDEPENDENT STANDARDIZED OUTCOME MEASURE: the coded outcomes are things like "learning achievement," "test scores," "project performance," "argumentation skills," "music knowledge" - course-level, researcher- or instructor-designed instruments throughout. TREATMENT DURATION WAS NOT CODED AT ALL, so this meta-analysis cannot say whether any effect outlasts a single session, and neither can anyone citing it. Nor was durability: every effect is a post-test at end of treatment. The mean study has 87 participants. Heterogeneity is extreme (Q = 511.1, df = 85, p < .001, I2 = 83.37), individual effects range from -0.84 to +2.94, and THE 95% PREDICTION INTERVAL IS [-0.47, 1.61] - the expected effect for a new study in this literature includes substantial harm. Publication bias: the funnel plot is visibly skewed right, Egger's regression returns t = .53, p = .59 and the authors conclude bias was "not detected"; no trim-and-fill or PET-PEESE correction was run, and Egger's test is applied to 86 effect sizes nested within 22 publications without accounting for that dependency, so its power is overstated. English-language publications only, which the authors flag as a limitation given that most of the primary work is Asian. The single most diagnostic result in the paper is its own moderator analysis by research method: the weaker the design, the bigger the effect. INDEPENDENCE refers to the meta-analysts (two Korean university academics with no product stake); the independence of the pooled primary studies is largely unknown and typically involves instructors evaluating their own course designs. Published open access in IRRODL 27(1), 265-289.
Key findings
Pooled random-effects g = 0.573 (SE 0.063, 95% CI [0.451, 0.696], Z = 9.17, p < .001) from 86 effect sizes in 22 publications, N = 6,992 - and the number should not be used as an estimate of what ChatGPT does to learning. The prediction interval is [-0.47, 1.61]: the pool is so heterogeneous (I2 = 83) that it makes no prediction at all about a new study. Its own moderator table shows effect size tracking design weakness almost monotonically - correlational studies g = 1.698 (k=2), true experimental g = 0.762 (k=18), mixed methods g = 0.520, quasi-experimental g = 0.529 (k=48), and posttest-only designs g = 0.113 with a confidence interval spanning zero (k=8). Instructional-approach moderators run the same way: self-regulated learning g = 1.115 (the least controlled condition) versus game-based learning g = 0.092. The composition is small (mean n = 86.6), short (duration never coded), geographically narrow (half the studies from Taiwan and mainland China), overwhelmingly undergraduate, entirely researcher-designed in outcome, unappraised for quality, and - to the extent duration is knowable - measured in weeks at most. What this meta-analysis actually establishes is that the ChatGPT-and-learning literature is not yet capable of supporting a pooled causal claim.
Genetic confound
Medium-to-high for the pool as a whole. Only 44.2% of effect sizes came from randomized assignment; the remainder are quasi-experimental, single-group or correlational, so selection into treatment (and into the courses that adopted ChatGPT) cannot be ruled out for most of the pooled evidence.
Replication notes
Several meta-analyses of a heavily overlapping primary pool report pooled effects in the same direction but at very different magnitudes - Deng et al. (2025, Computers & Education 227:105224, 69 experimental studies), Wu & Yu (2024, 24 studies on AI chatbots broadly), Sun & Zhou (2024, college students), and this one at g = 0.573. That agreement is NOT independent replication: they re-pool largely the same small, short, unregistered quasi-experiments. And the highest-profile member of that set has been withdrawn - Wang & Fan (2025, Humanities and Social Sciences Communications, 51 studies, g = 0.867 for learning performance) was RETRACTED on 22 April 2026 over "discrepancies in the meta-analysis," with the authors not responding to correspondence (see edt-wang-2025-chatgpt-meta-retracted). Recorded as `mixed` for that reason.
DOI / URL
10.19173/irrodl.v27i1.8775

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Overall effect of ChatGPT on learning achievement (random-effects pooled estimate)g0.573 (SE 0.063, 95% CI [0.451, 0.696], Z = 9.170, p < .001), k = 86 effect sizes from 22 publications, N = 6,992researcher-designedpost-test at end of treatment in every included studyunclearend-of-treatmentdomain-skill
Heterogeneity and prediction interval (the number that matters more than the pooled mean)I2 and 95% prediction intervalQ = 511.1, df = 85, p < .001, I2 = 83.37; individual effects range -0.84 to +2.94; PREDICTION INTERVAL [-0.47, 1.61]. A new study drawn from this literature could plausibly show substantial harm.researcher-designednot-applicableunclearnot-applicabledomain-skill
Effect size by research design (design quality vs effect size)g by subgroupcorrelational 1.698 (k=2, CI [1.007, 2.389]); true experimental 0.762 (k=18, CI [0.516, 1.009]); mixed methods 0.520 (k=10); quasi-experimental 0.529 (k=48, CI [0.374, 0.683]); posttest-only 0.113 (k=8, CI [-0.268, 0.494], ns). Between-group QB(4) = 18.474, p < .05.researcher-designedpost-test at end of treatmentunclearend-of-treatmentdomain-skill
Effect size by school levelg by subgroupmiddle and high school 0.928 > undergraduate 0.538 > elementary (smallest, k = 2 only); between-group difference NOT significant, QB(4) = 4.281, p = .369researcher-designedpost-test at end of treatmentunclearend-of-treatmentdomain-skill
Publication-bias diagnosticsfunnel plot and Egger's regressionfunnel plot "slightly skewed to the right"; Egger's test t = .53, p = .59, reported as bias not detected. No trim-and-fill, no PET-PEESE, and the test ignores the nesting of 86 effect sizes within 22 publications.researcher-designednot-applicablenonenot-applicabledomain-skill
Durability, and treatment durationnoneNEITHER WAS CODED. No included study measured retention past end of treatment, and the meta-analysis did not extract intervention length, so no statement about persistence or about dose is available from this pool.researcher-designednot-applicableunclearnot-applicabledomain-skill

Cited by