RETRACTED ARTICLE: The effect of ChatGPT on students’ learning performance, learning perception, and higher-order thinking: insights from a meta-analysis
Wang, J., & Fan, W. · 2025
grade Dmeta-analysisindependentnot-applicablenumbers spot-checked
Sample
51 studies published between November 2022 and February 2025 (as reported in the retracted article)
Population
Students across education levels and subjects; the pool was the same short, small, mostly quasi-experimental ChatGPT literature synthesized by the other meta-analyses of this period.
Design
Published in Humanities and Social Sciences Communications 12, article 621, 6 May 2025. RETRACTED 22 April 2026 (retraction note DOI 10.1057/s41599-026-07310-z). The Editor's stated reason: "concerns relating to discrepancies in the meta-analysis," raised initially by Magnus Ingebrigtsen and Marko Lukic, and "taken together, the identified issues undermine the Editor's confidence in the validity of the analysis and the conclusions drawn from it." THE AUTHORS DID NOT RESPOND TO CORRESPONDENCE REGARDING THE RETRACTION. The retraction notice was itself amended on 2 July 2026 to credit the people who raised the concerns. This file exists because the archive's rule is that dropped studies stay visible: the figure g = 0.867 for "learning performance" is already loose in the ed-tech discourse and needs a record attached to it saying where it came from and what happened to it. It should not be cited, and any downstream claim resting on it inherits the retraction.
Key findings
Reported a large positive pooled effect of ChatGPT on learning performance (g = 0.867), moderate effects on learning perception (g = 0.456) and higher-order thinking (g = 0.457), from 51 studies. RETRACTED by the editor on 22 April 2026 over unresolved discrepancies in the meta-analysis, with the authors not responding. None of these numbers may be used. The episode is the cleanest available illustration of how fast this literature moved and how little of it was checked: an unreviewed synthesis of small unregistered quasi-experiments became the field's most-circulated effect size within a year of publication, and was withdrawn within another.
Genetic confound
Not assessable - the analysis has been withdrawn. The underlying pool was predominantly non-randomized, so selection into treatment would not have been excluded regardless.
Replication notes
Retracted, so there is nothing to replicate. Recorded because this was the single most widely circulated quantitative claim about generative AI and learning - 58,000+ accesses and an Altmetric score above 500 at the time of retraction - and it will keep being cited from secondary sources long after the withdrawal. Overlapping meta-analyses of largely the same primary literature report smaller pooled effects (Doo & Park 2026, g = 0.573; see edt-doo-2026-chatgpt-learning-meta), which is itself informative about how much of the 0.867 depended on the extraction.
DOI / URL
10.1057/s41599-025-04787-y
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Learning performance (as reported before retraction) | g | 0.867 - RETRACTED, do not cite | researcher-designed | post-test at end of treatment | unclear | end-of-treatment | domain-skill |
| Learning perception (as reported before retraction) | g | 0.456 - RETRACTED, do not cite | self-report survey | post-test at end of treatment | unclear | end-of-treatment | non-cognitive |
| Higher-order thinking (as reported before retraction) | g | 0.457 - RETRACTED, do not cite | researcher-designed | post-test at end of treatment | unclear | end-of-treatment | far-transfer |