RETRACTED: Role of test motivation in intelligence testing
Duckworth AL, Quinn PD, Lynam DR, Loeber R, Stouthamer-Loeber M · 2011
grade Dmeta-analysisindependentfailed
Sample
meta-analysis of randomised incentive experiments plus a longitudinal sample; retracted 1 May 2025
Population
Predominantly below-average-IQ samples in the incentive meta-analysis
Design
Retracted at the authors' request on 1 May 2025 (PNAS retraction notice 10.1073/pnas.2508040122, preceded by an Editorial Expression of Concern, 10.1073/pnas.2502439122). The reason is specific and disqualifying: the Study 1 meta-analysis included data from a Breuning & Zella article that has itself been retracted, contributing three large effect sizes and 24% of the meta-analytic N. Reanalysis without it moved the pooled effect from g = 0.64 to g = 0.47 and, far more seriously, moved the publication-bias diagnostics from 1 of 4 tests indicating bias to STRONG evidence of bias on ALL FOUR.
Key findings
Recorded as a cautionary entry, not as evidence. The original claim - that material incentives raise IQ test scores by roughly two-thirds of a standard deviation, so IQ scores partly measure motivation - was the most-cited support for the proposition that a low test score can be a motivation artefact. That proposition is one this archive has independent reason to take seriously (Wise & DeMars 2005, Wise 2017, Kirkwood 2012 all point the same way), which is exactly why the retraction has to be recorded rather than quietly replaced: the archive must not carry a conclusion whose headline evidence was withdrawn. The residual estimate of g = 0.47 is not usable either, because the same reanalysis found strong publication bias on every test.
Genetic confound
Not applicable to the retraction itself. Note the direction of the original claim - it argued AGAINST a pure-ability reading of IQ scores - so its withdrawal does not favour a hereditarian reading; it simply removes an estimate.
Replication notes
The pooled estimate did not survive removal of a single retracted input, and the bias diagnostics collapsed. That is the definition of a failed finding rather than an unreplicated one.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Original pooled effect of material incentives on IQ scores | Hedges g | 0.64 (as published in 2011; not usable) | standardized | single incentivised administration | business-as-usual | end-of-treatment | g |
| Reanalysis excluding the retracted input study | Hedges g | 0.47, still significant, but with strong publication bias on all four tests | standardized | single incentivised administration | business-as-usual | end-of-treatment | g |
| Contribution of the retracted input | share | 3 large effect sizes and 24% of the meta-analytic N | standardized | not applicable | none | not-applicable | g |