Does document-based source work move history knowledge, reading comprehension, or neither?
Document-based instruction reliably improves the sourcing and argument tasks it teaches (g ≈ 0.42) and, in the two best-identified trials, moves neither standardized reading nor history knowledge.
mixedconf: mediumgc: lowhistory-civics · ages 8–18
The field pools to g = 0.42 across 64 controlled studies and 17,120 participants (0.36 after PET-PEESE) on source-credibility outcomes — every one of them a researcher-built rubric administered at end of treatment. Inside that pool the moderator runs the wrong way: source KNOWLEDGE g = 0.52 and search BEHAVIOUR g = 0.51, but actual CREDIBILITY JUDGEMENT g = 0.29. Students learn to perform the sourcing moves better than they learn to judge sources. The two best-identified trials are nulls. PACT's 48-school effectiveness RCT (7,183 students) raised history knowledge 0.37-0.46 on its own aligned measure and left standardized reading at g = 0.14 ns. A 60-school, 25-state RCT of a document-analysis history curriculum (Mission US) found NO effect on history knowledge, NO effect on document-analysis ability and NO effect on interest — while significantly more treatment teachers reported being better able to build historical perspective-taking. And the sub-skills the whole approach is named for never move: Reisman found nulls on contextualization and corroboration, Nokes found contextualization moved by no condition, Bertram found 'understanding deconstruction' at b = 0.01 (p = .88), and Wilke found no effect on epistemological beliefs about history and none on democratic skills or dispositions. The famous anchor — Reisman 2012's Gates-MacGinitie ~0.30 — is a non-randomised, developer-led, five-school study with unmodelled clustering whose author attributes the reading effect to time on task.
Teach with documents because reading and arguing from sources is worth teaching in a history classroom, and because giving students multiple documents rather than a textbook demonstrably raises content knowledge. Do not buy it as a reading intervention, a civics intervention, or a route to epistemological maturity: the two best-identified trials are nulls, and the specific reasoning moves the approach is named after - contextualization, corroboration, recognising a source as a constructed narrative - are the ones that never move in any study that measured them. Budget for real teacher PD and for whole-class text discussion, because the one evaluation that produced effects had four days of training plus follow-ups and still saw discussion attempted by only three of five teachers. And treat every published effect size here as a statement about a rubric the researchers wrote.
Who this applies to
Verdict
Document-based history instruction teaches document-based history tasks. It does that at a respectable size — around g = 0.4 across sixty-odd controlled studies — and there is nothing wrong with that as an outcome, because reading and arguing from sources is a legitimate thing to want a history student to do.
What it does not do is anything else. The two best-identified trials in the literature are both nulls: a 48-school effectiveness RCT that raised history knowledge on its own measure and left standardized reading at 0.14 ns, and a 60-school, 25-state RCT of a document-analysis curriculum that found nothing on history knowledge, nothing on document analysis, and nothing on interest — while its teachers reported the opposite.
And there is a sharper problem underneath. The specific reasoning moves the approach is named for are the ones that never move. Sourcing improves. Close reading improves. Contextualization, corroboration, recognising a source as a constructed narrative, and beliefs about how historical knowledge is made — four independent teams, four designs, nothing.
Hence mixed: a real, replicated effect on the taught tasks sitting beside replicated nulls on
everything the taught tasks were supposed to be for.
What the evidence shows
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Roberts 2023 (PACT) | school-randomised effectiveness RCT, 48 schools / 7,183 students | B | Content knowledge 0.37-0.45 (researcher-designed, treatment-aligned); standardized reading g = 0.14 ns |
| EDC 2025 (Mission US) | school-randomised RCT, 60 schools / 25 states, grades 8-10 | B | Null on history knowledge, null on document-analysis ability, null on interest — while treatment teachers reported improved perspective-taking |
| Bertram 2017 | within-school cluster RCT, 35 classes / 900 students, +2-3 mo follow-up | B | Understanding-reconstruction 0.31-0.35 SD; understanding-deconstruction b = 0.01, p = .88; live eyewitness worse than video or transcript |
| Wilke 2022 | cluster RCT degraded by differential attrition and post-randomisation exclusions, 668 students | C | Inquiry essay d ≈ 0.9-1.0 (researcher-designed); null on epistemological beliefs, null on democratic skills and dispositions |
| Fendt 2025 | meta-analysis, 64 studies / 246 effects / 17,120 participants | C | Overall g = 0.42 (PET-PEESE 0.36), PI [−0.33, 1.17]; sourcing 0.41, lateral reading 0.55; actual credibility judgement the weakest at 0.29 |
| Reisman 2012 | quasi-experiment, 5 schools, 6 months, non-random, clustering unmodelled, developer-led | C | Historical thinking ~0.49, transfer ~0.56, factual ~0.29; Gates-MacGinitie ~0.30. Null on contextualization and corroboration |
| Nokes 2007 | 2×2 factorial, 8 classrooms randomly assigned, 246 students, 3 weeks | C | Multiple texts beat textbook on content and on sourcing/corroboration use; explicit heuristic instruction did not raise content; contextualization moved by nothing |
| Wineburg 2022 (lateral reading) | matched-pair school allocation, 499 students, 6 × 50-min lessons | C | Condition×time = 1.66 of 14 points; treatment 2.9→5.1 vs control 2.0→2.7 (researcher-designed) |
| De La Paz 2017 | quasi-experiment, volunteer teachers, 18-day curriculum, grade 8 | C | Basic readers: historical writing 0.45, writing quality 0.32, essay length 0.60 (researcher-scored); no standardized outcome administered |
| Huijgen 2018 | small quasi-experiment with control, n = 131, ages 15-16 | C | Significant gain in historical contextualization and fewer presentist answers; effect size not reported |
| McCulley & Osman 2015 | synthesis of 12 intervention studies, grades 6-12 | C | Text-processing instruction improves content learning; no content-coverage penalty; comprehension not established |
| Fogo 2019 | curriculum-use survey, 1,927 Reading Like a Historian teachers | D | **<20% ever had any PD** yet >80% use it monthly; 86% modify lessons, least often the central question or document set |
| Wineburg 1991 | descriptive think-aloud, historians vs students | D | Historians source, corroborate and contextualise; able students rate the anonymous textbook most trustworthy |
The whole field measures what it just taught, immediately. Across every included study the outcome is a researcher-built rubric or item set administered at end of treatment. Only two carry any standardized outcome; only two carry any follow-up at all. Fendt's meta-analysis pools 64 studies and contains zero standardized outcomes and zero follow-up horizons.
Inside that meta-analysis, the moderator that should matter most runs backwards. Source knowledge g = 0.52 and search behaviour g = 0.51 — but actually judging whether a source is credible g = 0.29. Students get better at performing the moves faster than they get better at the judgement the moves exist to support. Worth noting alongside: Fendt's own z-curve reports an observed discovery rate of 51% against an expected rate of 16% [.06, .32] — the classic selective-reporting signature — which the authors do not discuss while concluding publication bias is unlikely.
Two results contradicted expectation outright and both are worth carrying. First, Bertram's head-to-head: students who met a live eyewitness scored significantly lower on understanding oral history and on deconstruction than students given the identical interview on video or as a transcript — while being more convinced they had learned more. The paper is titled "More Fun — Less Learned." The most vivid primary-source encounter was the least critically instructive. Second, the Mission US trial: well-powered, nationally sampled, school-randomised, and null on every measured outcome, while significantly more treatment teachers agreed they were better able to build historical perspective-taking and empathy. That is the same pattern the archive already recorded in the EEF Grammar for Writing trial, where over 90% of teachers found a programme useful that returned −0.02.
Nokes' factorial is the most useful design in the cluster because it separates the two ingredients. Giving students multiple documents instead of a textbook raised content knowledge and raised their use of sourcing and corroboration. Explicitly teaching the heuristics raised strategy use but produced no content main effect. If that separation holds, the active ingredient is the documents, not the metacognitive scaffolding built around them.
The anchor claim is the weakest link. Reisman 2012 is cited everywhere as evidence that Reading Like a Historian raises general reading comprehension. It is a five-school quasi-experiment with non-random teacher assignment, analysed without accounting for its ten classrooms, run and observed by the curriculum's own designer — and its author attributes the reading effect to time on task rather than to disciplinary transfer.
Hereditarian-lens assessment
Risk: low. The verdict rests on two school-randomised RCTs and a cluster RCT; within those, treated and untreated students have the same expected distribution of ability-relevant alleles, so heredity cannot explain the contrast. The nulls are the load-bearing results and nulls are the hardest thing for selection to manufacture.
Where the lens still matters:
- Independence is unusually bad in this literature. Among the included sources, nine of thirteen are developer-led or developer-involved. Reisman is the developer. De La Paz's teachers volunteered. Wineburg's lateral-reading study evaluates his own group's curriculum. That ratio is itself a finding about the field, and it is exactly the variable METHODOLOGY says the efficacy-to-effectiveness decay is a claim about — so it is not surprising that the two genuinely independent, school-randomised trials are the two nulls.
- The one place where a heritable trait plausibly drives the result is the moderator pattern. Reisman's struggling readers gained historical thinking and factual knowledge but not comprehension or transfer; De La Paz's clearest effects are on essay length. Reading ability is substantially heritable, and "it worked better for the students who could already read the documents" is a correlation, not an aptitude–treatment interaction.
- Fogo's survey identifies the selection that matters most in practice: the evaluated version had four days of PD and five observed teachers; the deployed version is a free download used monthly by teachers of whom fewer than 20% have had any training at all. Whatever the trials measured, it is not what most classrooms are running.
Boundaries & what critics say
- The sub-skills never move, and this is the field's own finding, not an outsider's complaint. Reisman: null on contextualization and corroboration. Nokes: contextualization improved by no condition. Bertram: deconstruction b = 0.01 (p = .879), 0.03 at follow-up. Wilke: no effect on epistemological beliefs, and procedural knowledge about inquiry was not even correlated with inquiry skill at pre-test. Three independent teams, three designs, the same hole.
- The meta-analytic dose finding is uncomfortable. Total length did not moderate effects overall (p = .977) and interacted negatively with sourcing and historical reasoning at about −0.10 g per extra 100 minutes. Longer history-reasoning interventions did worse. Either the effect saturates immediately or the short studies are measuring something the long ones are not.
- Nothing here has a horizon longer than a few months. PACT at 9 weeks and Bertram at 2-3 months are the entire durability evidence base.
- The undergraduate civic-online-reasoning literature is adjacent and mostly out of scope. Where it has school-age samples the effects are real but small and researcher-measured; the adult lateral-reading work shows an asymmetry — better at rejecting false content, no better at accepting true content — that has never been tested on school students and would be a serious boundary if it held.
- A defender would say the outcome measures are the point. If you want students to reason about
sources, a rubric measuring source reasoning is the correct instrument and demanding a
Gates-MacGinitie effect is a category error. That is a fair argument and it is why the verdict is
mixedrather thanno-effect— but it concedes that document-based instruction should not be sold as a literacy intervention, which is how it is usually sold. - An independent evaluation is under way. A US Department of Education-funded AIR study of new
Reading Like a Historian materials is the single thing most likely to change this verdict, and its
absence is why the anchor carries
replication: unreplicated.
Practical guidance
- Use multiple documents instead of a textbook. This is the ingredient with the cleanest causal support — it raised both content knowledge and sourcing behaviour in the one factorial that isolated it — and it costs nothing but preparation.
- Do not buy it as a reading intervention. Seven PACT trials give g = 0.14 ns on standardized comprehension. The single positive result in the literature is confounded and its own author explains it as time on task.
- Do not buy it as a civics intervention either. The one trial that measured democratic skills and dispositions found nothing, and its authors explicitly caution against excessive optimism.
- Budget for real PD and protect the discussion. The evaluation that produced effects had four days of summer training plus two follow-up workshops, and still only three of five teachers attempted whole-class text discussion — which the evaluator names as the likely reason the higher-order skills did not move. A free download with no training is not the intervention that was tested.
- Adapt the documents. A third of teacher modifications are made for reading access. Students who cannot read the source cannot source it.
- Prefer video or transcript to a live guest speaker if the goal is critical source reading. The one randomised head-to-head found the live encounter produced less critical understanding and more confidence that learning had occurred.
- Ignore teacher enthusiasm as evidence. In the 60-school null, significantly more treatment teachers reported improved perspective-taking. Nothing had improved.
Open questions
- Can contextualization or corroboration be taught at all? Four studies measured them and none moved them. Either the constructs are not teachable at school age, the instruments cannot detect the change, or nobody has yet taught them properly — and these have very different implications.
- Is the active ingredient the documents or the heuristics? Nokes' factorial says documents; it is three weeks, eight classrooms, and unreplicated.
- Does anything survive a year? Nothing in this literature has a horizon longer than a few months.
- Why do longer interventions do worse? The meta-analytic moderator runs the wrong way and nobody has investigated it.
- What does the independently evaluated version look like? Nine of thirteen included sources are developer-led, and the two independent randomised trials are the two nulls. The AIR evaluation now under way is the test that matters.
- grade CReading Like a Historian: A Document-Based History Curriculum Intervention in Urban High SchoolsReisman A · 2012 · quasi-experiment
- grade BU.S. History Through Young People's Eyes: Evaluating the Impact of Mission US on Student LearningEducation Development Center (Kennedy JL and colleagues) · 2025 · rct
- grade BLearning Historical Thinking With Oral History Interviews: A Cluster Randomized Controlled Intervention Study of Oral History Interviews in History LessonsBertram C, Wagner W, Trautwein U · 2017 · rct
- grade CFostering historical thinking and democratic citizenship? A cluster randomized controlled intervention studyWilke M, Depaepe F, Van Nieuwenhuyse K · 2022 · quasi-experiment
- grade CJudging a text by its author — A meta-analysis of interventions to foster source credibility assessmentFendt M, Muth X, Edelsbrunner PA · 2025 · meta-analysis
- grade CTeaching high school students to use heuristics while reading historical texts.Nokes JD, Dole JA, Hacker DJ · 2007 · rct
- grade CLateral reading on the open Internet: A district-wide field study in high school government classes.Wineburg S, Breakstone J, McGrew S, Smith MD, Ortega T · 2022 · quasi-experiment
- grade CA Historical Writing Apprenticeship for Adolescents: Integrating Disciplinary Learning With Cognitive StrategiesDe La Paz S, Monte-Sano C, Felton M, Croninger R, Jackson C, Piantedosi KW · 2017 · quasi-experiment
- grade CPromoting historical contextualization: the development and testing of a pedagogyHuijgen T, van de Grift W, van Boxtel C, Holthuis P · 2018 · quasi-experiment
- grade DTeacher adaptation of document-based history curricula: results of the Reading Like a Historian curriculum-use surveyFogo B, Reisman A, Breakstone J · 2019 · review
- grade DHistorical problem solving: A study of the cognitive processes used in the evaluation of documentary and pictorial evidence.Wineburg SS · 1991 · review
- grade APromoting adolescents' comprehension of text: A randomized control trial of its effectivenessRoberts G, Vaughn S, Wanzek J, Furman G, Martinez L, Sargent K · 2023 · rct
- grade CEffects of Reading Instruction on Learning Outcomes in Social Studies: A Synthesis of Quantitative ResearchMcCulley LV, Osman DJ · 2015 · review
Related decisions
- Should history be built on content knowledge or on generic historical-thinking skills?moderate supportconf: mediumgc: low