The Evidence on Teaching

Judging a text by its author — A meta-analysis of interventions to foster source credibility assessment

Fendt M, Muth X, Edelsbrunner PA · 2025

grade Cmeta-analysisindependentnot-applicablenumbers spot-checked
Sample
64 studies from 60 articles, 246 effect estimates, 17,120 participants
Population
Mixed: 32 studies with adolescents, 24 with adults, 8 with children; mean sample ages 9.8 to 45.5 years; 39 studies in schools, 20 in universities, 5 informal. Mostly Western (US 24, Germany 10, Israel 9).
Design
The only quantitative synthesis that pools the four literatures this cluster spans - historical thinking/historical reasoning, multiple-document literacy, sourcing, and lateral reading - into one random-effects three-level model. PRISMA, preregistered materials on OSF, metafor. Coded moderators include intervention approach, education level, control type, delivery, technology, and dependent-variable type. Graded C rather than B because it aggregates 36 quasi-experiments alongside 28 experiments, includes university and adult samples well outside this archive's 4-18 scope, and - the decisive limitation for the archive's purposes - almost every pooled outcome is a proximal, researcher-built credibility-judgment or reasoning task administered immediately after the intervention. There is no pooled estimate of an effect on any standardized achievement measure, and no pooled estimate at any follow-up horizon. Publication-bias checks were unusually thorough: symmetric funnel plot, non-significant Egger test, trim-and-fill imputing nothing, PET-PEESE correcting the estimate down from g = 0.42 to g = 0.36. The authors conclude bias is unlikely; note against that conclusion that their own z-curve reports an observed discovery rate of 51% against an EXPECTED discovery rate of 16% (95% CI .06-.32), a gap that is the classic signature of selective reporting and is not discussed.
Key findings
Overall g = 0.42 (95% CI 0.35-0.49, p < .001), corrected to g = 0.36 by PET-PEESE. Heterogeneity is severe: I-squared = 87.5% and a 95% prediction interval of [-0.33, 1.17], meaning the expected effect in a new study ranges from small negative to large positive - the archive's standard caution about average effects applies with full force. By approach: lateral reading g = 0.55 (k = 43 effects), historical reasoning g = 0.42 (k = 30, from only 7 studies), sourcing g = 0.41 (k = 101), multiple-document literacy g = 0.35 (k = 72); lateral reading beat multiple-document literacy (p = .067) and the combined others (p = .093), both marginal. Secondary-education samples g = 0.46 (k = 109), primary g = 0.41. Participant AGE did not moderate effects (p = .636) and neither did gender composition. Critically, the dependent variable mattered and ran against the field's framing: source knowledge g = 0.52 and search behaviour g = 0.51 - what students know and do - but actual CREDIBILITY JUDGMENT, the thing the whole enterprise is for, was the weakest outcome at g = 0.29 (95% CI 0.19-0.40). Interventions against an instructed control did BETTER (g = 0.48) than against a passive control (g = 0.31), which is anomalous and unexplained. Total intervention length did not moderate effects (p = .977) - median dose was 90 minutes on a single day - though length interacted positively with lateral reading (+0.10 g per extra 100 min) and NEGATIVELY with sourcing and historical reasoning (roughly -0.10 g per extra 100 min), i.e. longer historical-reasoning interventions in this pool did worse, not better.
Genetic confound
Medium. Pooled designs are mostly randomised or quasi-experimental with control groups, which limits genetic confounding within studies, but 36 of 64 are quasi-experimental and selection into condition is not always documented.
Replication notes
not-applicable - this is the synthesis rather than a study to be replicated. It does establish that the field's positive signal is not an artefact of a couple of famous trials: 64 controlled studies point the same way on proximal source-evaluation tasks. What it cannot establish, because no pooled study measured it, is durability or transfer to any standardized outcome.
DOI / URL
10.1016/j.lindif.2025.102782

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Source credibility assessment, all interventions pooledHedges g, three-level random-effects model0.42 (95% CI 0.35-0.49); PET-PEESE corrected 0.36; 95% prediction interval -0.33 to 1.17researcher-designedimmediately post-intervention in the large majority of studiesbusiness-as-usualend-of-treatmentdomain-skill
Historical reasoning approach specificallyHedges g0.42 (95% CI 0.31-0.52) from 30 effects across only 7 studiesresearcher-designedimmediately post-interventionbusiness-as-usualend-of-treatmentdomain-skill
Lateral reading approach specificallyHedges g0.55 (95% CI 0.35-0.75), largest of the four approachesresearcher-designedimmediately post-interventionbusiness-as-usualend-of-treatmentdomain-skill
Actual credibility judgments (as against knowledge or search behaviour)Hedges g by dependent-variable type0.29 (95% CI 0.19-0.40) - the weakest outcome class, versus source knowledge 0.52 and search behaviour 0.51researcher-designedimmediately post-interventionbusiness-as-usualend-of-treatmentdomain-skill
Secondary-education subgroup (the age band relevant to this archive)Hedges g0.46 (95% CI 0.35-0.57), k = 109 effects; age itself did not moderate (p = .636)researcher-designedimmediately post-interventionbusiness-as-usualend-of-treatmentdomain-skill

Cited by