The Evidence on Teaching

Does document-based source work move history knowledge, reading comprehension, or neither?

Document-based instruction reliably improves the sourcing and argument tasks it teaches (g ≈ 0.42) and, in the two best-identified trials, moves neither standardized reading nor history knowledge.

mixedconf: mediumgc: low

history-civics · ages 818

Effect summary

The field pools to g = 0.42 across 64 controlled studies and 17,120 participants (0.36 after PET-PEESE) on source-credibility outcomes — every one of them a researcher-built rubric administered at end of treatment. Inside that pool the moderator runs the wrong way: source KNOWLEDGE g = 0.52 and search BEHAVIOUR g = 0.51, but actual CREDIBILITY JUDGEMENT g = 0.29. Students learn to perform the sourcing moves better than they learn to judge sources. The two best-identified trials are nulls. PACT's 48-school effectiveness RCT (7,183 students) raised history knowledge 0.37-0.46 on its own aligned measure and left standardized reading at g = 0.14 ns. A 60-school, 25-state RCT of a document-analysis history curriculum (Mission US) found NO effect on history knowledge, NO effect on document-analysis ability and NO effect on interest — while significantly more treatment teachers reported being better able to build historical perspective-taking. And the sub-skills the whole approach is named for never move: Reisman found nulls on contextualization and corroboration, Nokes found contextualization moved by no condition, Bertram found 'understanding deconstruction' at b = 0.01 (p = .88), and Wilke found no effect on epistemological beliefs about history and none on democratic skills or dispositions. The famous anchor — Reisman 2012's Gates-MacGinitie ~0.30 — is a non-randomised, developer-led, five-school study with unmodelled clustering whose author attributes the reading effect to time on task.

Practical takeaway

Teach with documents because reading and arguing from sources is worth teaching in a history classroom, and because giving students multiple documents rather than a textbook demonstrably raises content knowledge. Do not buy it as a reading intervention, a civics intervention, or a route to epistemological maturity: the two best-identified trials are nulls, and the specific reasoning moves the approach is named after - contextualization, corroboration, recognising a source as a constructed narrative - are the ones that never move in any study that measured them. Budget for real teacher PD and for whole-class text discussion, because the one evaluation that produced effects had four days of training plus follow-ups and still saw discussion attempted by only three of five teachers. And treat every published effect size here as a statement about a rubric the researchers wrote.

Who this applies to

Group size
whole-class
Delivered by
teacher
Ages studied
1318(narrower than the 818 this topic is filed under — outside it is extrapolation)
Dose
The positive effects came from SUSTAINED dosing: Reisman ran 2-3 document-based lessons per week for six months (~58% of instructional time) after four days of summer PD plus two follow-up workshops; PACT embedded the components in three ~15-day units; Wilke used ~12 fifty-minute lessons; the lateral-reading field study used six 50-minute lessons across a semester after teacher PD. Note the warning inside the meta-analysis: the median study is 90 minutes on a single day, total length did NOT moderate effects overall (p = .977), and it interacted NEGATIVELY with sourcing and historical reasoning at roughly -0.10 g per additional 100 minutes.
Cost
free
Moves
domain-skill
Needs first
Adapted or shortened primary sources - Reading Like a Historian ships modified documents, and about a third of teacher modifications are made for reading access. Teacher subject-matter knowledge. And whole-class text-based discussion, which Reisman observed only three of five trained teachers ever attempt and which she names as the likely reason contextualization and corroboration did not move.
Not for
Do not use it to raise general reading comprehension - PACT's g = 0.14 ns across seven trials, and the one positive (Reisman, ~0.30) is confounded and self-attributed to time on task. Do not use it for civic or democratic outcomes - Wilke's explicit null, with the authors cautioning against excessive optimism. Do not expect contextualization, corroboration, deconstruction or epistemological beliefs to move: four independent teams found nothing on these. Do not use live eyewitness testimony as the route to critical source reading - students who met a live witness scored LOWER than those given the same interview on video or transcript while believing they had learned more. And do not deploy it as an untrained free download, which is the modal real-world case: fewer than 20% of about 1,900 active Reading Like a Historian teachers have ever had any PD.

Verdict

Document-based history instruction teaches document-based history tasks. It does that at a respectable size — around g = 0.4 across sixty-odd controlled studies — and there is nothing wrong with that as an outcome, because reading and arguing from sources is a legitimate thing to want a history student to do.

What it does not do is anything else. The two best-identified trials in the literature are both nulls: a 48-school effectiveness RCT that raised history knowledge on its own measure and left standardized reading at 0.14 ns, and a 60-school, 25-state RCT of a document-analysis curriculum that found nothing on history knowledge, nothing on document analysis, and nothing on interest — while its teachers reported the opposite.

And there is a sharper problem underneath. The specific reasoning moves the approach is named for are the ones that never move. Sourcing improves. Close reading improves. Contextualization, corroboration, recognising a source as a constructed narrative, and beliefs about how historical knowledge is made — four independent teams, four designs, nothing.

Hence mixed: a real, replicated effect on the taught tasks sitting beside replicated nulls on everything the taught tasks were supposed to be for.

What the evidence shows

Source Design Grade Key effect
Roberts 2023 (PACT) school-randomised effectiveness RCT, 48 schools / 7,183 students B Content knowledge 0.37-0.45 (researcher-designed, treatment-aligned); standardized reading g = 0.14 ns
EDC 2025 (Mission US) school-randomised RCT, 60 schools / 25 states, grades 8-10 B Null on history knowledge, null on document-analysis ability, null on interest — while treatment teachers reported improved perspective-taking
Bertram 2017 within-school cluster RCT, 35 classes / 900 students, +2-3 mo follow-up B Understanding-reconstruction 0.31-0.35 SD; understanding-deconstruction b = 0.01, p = .88; live eyewitness worse than video or transcript
Wilke 2022 cluster RCT degraded by differential attrition and post-randomisation exclusions, 668 students C Inquiry essay d ≈ 0.9-1.0 (researcher-designed); null on epistemological beliefs, null on democratic skills and dispositions
Fendt 2025 meta-analysis, 64 studies / 246 effects / 17,120 participants C Overall g = 0.42 (PET-PEESE 0.36), PI [−0.33, 1.17]; sourcing 0.41, lateral reading 0.55; actual credibility judgement the weakest at 0.29
Reisman 2012 quasi-experiment, 5 schools, 6 months, non-random, clustering unmodelled, developer-led C Historical thinking ~0.49, transfer ~0.56, factual ~0.29; Gates-MacGinitie ~0.30. Null on contextualization and corroboration
Nokes 2007 2×2 factorial, 8 classrooms randomly assigned, 246 students, 3 weeks C Multiple texts beat textbook on content and on sourcing/corroboration use; explicit heuristic instruction did not raise content; contextualization moved by nothing
Wineburg 2022 (lateral reading) matched-pair school allocation, 499 students, 6 × 50-min lessons C Condition×time = 1.66 of 14 points; treatment 2.9→5.1 vs control 2.0→2.7 (researcher-designed)
De La Paz 2017 quasi-experiment, volunteer teachers, 18-day curriculum, grade 8 C Basic readers: historical writing 0.45, writing quality 0.32, essay length 0.60 (researcher-scored); no standardized outcome administered
Huijgen 2018 small quasi-experiment with control, n = 131, ages 15-16 C Significant gain in historical contextualization and fewer presentist answers; effect size not reported
McCulley & Osman 2015 synthesis of 12 intervention studies, grades 6-12 C Text-processing instruction improves content learning; no content-coverage penalty; comprehension not established
Fogo 2019 curriculum-use survey, 1,927 Reading Like a Historian teachers D **<20% ever had any PD** yet >80% use it monthly; 86% modify lessons, least often the central question or document set
Wineburg 1991 descriptive think-aloud, historians vs students D Historians source, corroborate and contextualise; able students rate the anonymous textbook most trustworthy

The whole field measures what it just taught, immediately. Across every included study the outcome is a researcher-built rubric or item set administered at end of treatment. Only two carry any standardized outcome; only two carry any follow-up at all. Fendt's meta-analysis pools 64 studies and contains zero standardized outcomes and zero follow-up horizons.

Inside that meta-analysis, the moderator that should matter most runs backwards. Source knowledge g = 0.52 and search behaviour g = 0.51 — but actually judging whether a source is credible g = 0.29. Students get better at performing the moves faster than they get better at the judgement the moves exist to support. Worth noting alongside: Fendt's own z-curve reports an observed discovery rate of 51% against an expected rate of 16% [.06, .32] — the classic selective-reporting signature — which the authors do not discuss while concluding publication bias is unlikely.

Two results contradicted expectation outright and both are worth carrying. First, Bertram's head-to-head: students who met a live eyewitness scored significantly lower on understanding oral history and on deconstruction than students given the identical interview on video or as a transcript — while being more convinced they had learned more. The paper is titled "More Fun — Less Learned." The most vivid primary-source encounter was the least critically instructive. Second, the Mission US trial: well-powered, nationally sampled, school-randomised, and null on every measured outcome, while significantly more treatment teachers agreed they were better able to build historical perspective-taking and empathy. That is the same pattern the archive already recorded in the EEF Grammar for Writing trial, where over 90% of teachers found a programme useful that returned −0.02.

Nokes' factorial is the most useful design in the cluster because it separates the two ingredients. Giving students multiple documents instead of a textbook raised content knowledge and raised their use of sourcing and corroboration. Explicitly teaching the heuristics raised strategy use but produced no content main effect. If that separation holds, the active ingredient is the documents, not the metacognitive scaffolding built around them.

The anchor claim is the weakest link. Reisman 2012 is cited everywhere as evidence that Reading Like a Historian raises general reading comprehension. It is a five-school quasi-experiment with non-random teacher assignment, analysed without accounting for its ten classrooms, run and observed by the curriculum's own designer — and its author attributes the reading effect to time on task rather than to disciplinary transfer.

Hereditarian-lens assessment

Risk: low. The verdict rests on two school-randomised RCTs and a cluster RCT; within those, treated and untreated students have the same expected distribution of ability-relevant alleles, so heredity cannot explain the contrast. The nulls are the load-bearing results and nulls are the hardest thing for selection to manufacture.

Where the lens still matters:

  • Independence is unusually bad in this literature. Among the included sources, nine of thirteen are developer-led or developer-involved. Reisman is the developer. De La Paz's teachers volunteered. Wineburg's lateral-reading study evaluates his own group's curriculum. That ratio is itself a finding about the field, and it is exactly the variable METHODOLOGY says the efficacy-to-effectiveness decay is a claim about — so it is not surprising that the two genuinely independent, school-randomised trials are the two nulls.
  • The one place where a heritable trait plausibly drives the result is the moderator pattern. Reisman's struggling readers gained historical thinking and factual knowledge but not comprehension or transfer; De La Paz's clearest effects are on essay length. Reading ability is substantially heritable, and "it worked better for the students who could already read the documents" is a correlation, not an aptitude–treatment interaction.
  • Fogo's survey identifies the selection that matters most in practice: the evaluated version had four days of PD and five observed teachers; the deployed version is a free download used monthly by teachers of whom fewer than 20% have had any training at all. Whatever the trials measured, it is not what most classrooms are running.

Boundaries & what critics say

  • The sub-skills never move, and this is the field's own finding, not an outsider's complaint. Reisman: null on contextualization and corroboration. Nokes: contextualization improved by no condition. Bertram: deconstruction b = 0.01 (p = .879), 0.03 at follow-up. Wilke: no effect on epistemological beliefs, and procedural knowledge about inquiry was not even correlated with inquiry skill at pre-test. Three independent teams, three designs, the same hole.
  • The meta-analytic dose finding is uncomfortable. Total length did not moderate effects overall (p = .977) and interacted negatively with sourcing and historical reasoning at about −0.10 g per extra 100 minutes. Longer history-reasoning interventions did worse. Either the effect saturates immediately or the short studies are measuring something the long ones are not.
  • Nothing here has a horizon longer than a few months. PACT at 9 weeks and Bertram at 2-3 months are the entire durability evidence base.
  • The undergraduate civic-online-reasoning literature is adjacent and mostly out of scope. Where it has school-age samples the effects are real but small and researcher-measured; the adult lateral-reading work shows an asymmetry — better at rejecting false content, no better at accepting true content — that has never been tested on school students and would be a serious boundary if it held.
  • A defender would say the outcome measures are the point. If you want students to reason about sources, a rubric measuring source reasoning is the correct instrument and demanding a Gates-MacGinitie effect is a category error. That is a fair argument and it is why the verdict is mixed rather than no-effect — but it concedes that document-based instruction should not be sold as a literacy intervention, which is how it is usually sold.
  • An independent evaluation is under way. A US Department of Education-funded AIR study of new Reading Like a Historian materials is the single thing most likely to change this verdict, and its absence is why the anchor carries replication: unreplicated.

Practical guidance

  • Use multiple documents instead of a textbook. This is the ingredient with the cleanest causal support — it raised both content knowledge and sourcing behaviour in the one factorial that isolated it — and it costs nothing but preparation.
  • Do not buy it as a reading intervention. Seven PACT trials give g = 0.14 ns on standardized comprehension. The single positive result in the literature is confounded and its own author explains it as time on task.
  • Do not buy it as a civics intervention either. The one trial that measured democratic skills and dispositions found nothing, and its authors explicitly caution against excessive optimism.
  • Budget for real PD and protect the discussion. The evaluation that produced effects had four days of summer training plus two follow-up workshops, and still only three of five teachers attempted whole-class text discussion — which the evaluator names as the likely reason the higher-order skills did not move. A free download with no training is not the intervention that was tested.
  • Adapt the documents. A third of teacher modifications are made for reading access. Students who cannot read the source cannot source it.
  • Prefer video or transcript to a live guest speaker if the goal is critical source reading. The one randomised head-to-head found the live encounter produced less critical understanding and more confidence that learning had occurred.
  • Ignore teacher enthusiasm as evidence. In the 60-school null, significantly more treatment teachers reported improved perspective-taking. Nothing had improved.

Open questions

  • Can contextualization or corroboration be taught at all? Four studies measured them and none moved them. Either the constructs are not teachable at school age, the instruments cannot detect the change, or nobody has yet taught them properly — and these have very different implications.
  • Is the active ingredient the documents or the heuristics? Nokes' factorial says documents; it is three weeks, eight classrooms, and unreplicated.
  • Does anything survive a year? Nothing in this literature has a horizon longer than a few months.
  • Why do longer interventions do worse? The meta-analytic moderator runs the wrong way and nobody has investigated it.
  • What does the independently evaluated version look like? Nine of thirteen included sources are developer-led, and the two independent randomised trials are the two nulls. The AIR evaluation now under way is the test that matters.

Evidence (13 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore