Judging a text by its author — A meta-analysis of interventions to foster source credibility assessment
Fendt M, Muth X, Edelsbrunner PA · 2025
grade Cmeta-analysisindependentnot-applicablenumbers spot-checked
Sample
64 studies from 60 articles, 246 effect estimates, 17,120 participants
Population
Mixed: 32 studies with adolescents, 24 with adults, 8 with children; mean sample ages 9.8 to 45.5 years; 39 studies in schools, 20 in universities, 5 informal. Mostly Western (US 24, Germany 10, Israel 9).
Design
The only quantitative synthesis that pools the four literatures this cluster spans - historical thinking/historical reasoning, multiple-document literacy, sourcing, and lateral reading - into one random-effects three-level model. PRISMA, preregistered materials on OSF, metafor. Coded moderators include intervention approach, education level, control type, delivery, technology, and dependent-variable type. Graded C rather than B because it aggregates 36 quasi-experiments alongside 28 experiments, includes university and adult samples well outside this archive's 4-18 scope, and - the decisive limitation for the archive's purposes - almost every pooled outcome is a proximal, researcher-built credibility-judgment or reasoning task administered immediately after the intervention. There is no pooled estimate of an effect on any standardized achievement measure, and no pooled estimate at any follow-up horizon. Publication-bias checks were unusually thorough: symmetric funnel plot, non-significant Egger test, trim-and-fill imputing nothing, PET-PEESE correcting the estimate down from g = 0.42 to g = 0.36. The authors conclude bias is unlikely; note against that conclusion that their own z-curve reports an observed discovery rate of 51% against an EXPECTED discovery rate of 16% (95% CI .06-.32), a gap that is the classic signature of selective reporting and is not discussed.
Key findings
Overall g = 0.42 (95% CI 0.35-0.49, p < .001), corrected to g = 0.36 by PET-PEESE. Heterogeneity is severe: I-squared = 87.5% and a 95% prediction interval of [-0.33, 1.17], meaning the expected effect in a new study ranges from small negative to large positive - the archive's standard caution about average effects applies with full force. By approach: lateral reading g = 0.55 (k = 43 effects), historical reasoning g = 0.42 (k = 30, from only 7 studies), sourcing g = 0.41 (k = 101), multiple-document literacy g = 0.35 (k = 72); lateral reading beat multiple-document literacy (p = .067) and the combined others (p = .093), both marginal. Secondary-education samples g = 0.46 (k = 109), primary g = 0.41. Participant AGE did not moderate effects (p = .636) and neither did gender composition. Critically, the dependent variable mattered and ran against the field's framing: source knowledge g = 0.52 and search behaviour g = 0.51 - what students know and do - but actual CREDIBILITY JUDGMENT, the thing the whole enterprise is for, was the weakest outcome at g = 0.29 (95% CI 0.19-0.40). Interventions against an instructed control did BETTER (g = 0.48) than against a passive control (g = 0.31), which is anomalous and unexplained. Total intervention length did not moderate effects (p = .977) - median dose was 90 minutes on a single day - though length interacted positively with lateral reading (+0.10 g per extra 100 min) and NEGATIVELY with sourcing and historical reasoning (roughly -0.10 g per extra 100 min), i.e. longer historical-reasoning interventions in this pool did worse, not better.
Genetic confound
Medium. Pooled designs are mostly randomised or quasi-experimental with control groups, which limits genetic confounding within studies, but 36 of 64 are quasi-experimental and selection into condition is not always documented.
Replication notes
not-applicable - this is the synthesis rather than a study to be replicated. It does establish that the field's positive signal is not an artefact of a couple of famous trials: 64 controlled studies point the same way on proximal source-evaluation tasks. What it cannot establish, because no pooled study measured it, is durability or transfer to any standardized outcome.
DOI / URL
10.1016/j.lindif.2025.102782
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Source credibility assessment, all interventions pooled | Hedges g, three-level random-effects model | 0.42 (95% CI 0.35-0.49); PET-PEESE corrected 0.36; 95% prediction interval -0.33 to 1.17 | researcher-designed | immediately post-intervention in the large majority of studies | business-as-usual | end-of-treatment | domain-skill |
| Historical reasoning approach specifically | Hedges g | 0.42 (95% CI 0.31-0.52) from 30 effects across only 7 studies | researcher-designed | immediately post-intervention | business-as-usual | end-of-treatment | domain-skill |
| Lateral reading approach specifically | Hedges g | 0.55 (95% CI 0.35-0.75), largest of the four approaches | researcher-designed | immediately post-intervention | business-as-usual | end-of-treatment | domain-skill |
| Actual credibility judgments (as against knowledge or search behaviour) | Hedges g by dependent-variable type | 0.29 (95% CI 0.19-0.40) - the weakest outcome class, versus source knowledge 0.52 and search behaviour 0.51 | researcher-designed | immediately post-intervention | business-as-usual | end-of-treatment | domain-skill |
| Secondary-education subgroup (the age band relevant to this archive) | Hedges g | 0.46 (95% CI 0.35-0.57), k = 109 effects; age itself did not moderate (p = .636) | researcher-designed | immediately post-intervention | business-as-usual | end-of-treatment | domain-skill |
Cited by
- Does document-based source work move history knowledge, reading comprehension, or neither?mixedconf: mediumgc: low