The Evidence on Teaching

Should history be built on content knowledge or on generic historical-thinking skills?

Content-rich history teaching raises history knowledge (g 0.19-0.46) but not standardized reading in under three years; generic 'historical thinking' has no adequately controlled positive result.

moderate supportconf: mediumgc: low

history-civics · ages 518

Effect summary

Six developer-led randomised trials of content-rich US-history instruction and two independent kindergarten RCTs of Core Knowledge agree, and the answer suits neither camp. CONTENT KNOWLEDGE MOVES RELIABLY: PACT's nationally stratified effectiveness trial (48 schools, 7,183 students) returns g = 0.46 on history knowledge, its efficacy siblings 0.31-0.36, team-based learning 0.19, and Core Knowledge kindergarten read-alouds 0.93 on curriculum-aligned social studies knowledge. STANDARDIZED READING DOES NOT: significant in exactly one of eight trials (Vaughn 2013, Gates-MacGinitie 0.21, not replicated in Vaughn 2015 or Swanson 2015), 0.14 ns in the effectiveness trial, and 0.00 (p = .999) on the Woodcock-Johnson Social Studies subtest in the same kindergarten children whose aligned knowledge moved 0.93. The only route to general reading is MULTI-YEAR DOSE: Kim's three-year spiral gives 0.11-0.12 on standardized reading at grade 3 and 14 months later, and Grissmer's six-year Core Knowledge lottery 0.24 on state ELA — but Kim moved MATHEMATICS by the same 0.12-0.16, which no knowledge-transfer mechanism predicts. The rival organising principle is worse off: generic historical-thinking instruction has no adequately controlled positive result, its flagship studies use volunteer teachers or no counterfactual, and the one study with randomly SELECTED treatment teachers found nothing on its own aligned writing measure.

Practical takeaway

Run history as a knowledge-rich, text-based, cumulative curriculum, and justify it on history knowledge - which it reliably delivers at a modest cost - rather than on reading comprehension, which it does not deliver inside a year. If reading comprehension is the goal, the only doses that have ever moved a standardized measure are multi-year: three years of a coherent content spiral, or a whole-school knowledge curriculum run for six. Do not organise the curriculum around generic 'historical thinking skills': the strongest study of that approach with randomly selected teachers found nothing, and the strongest study of content-rich instruction found the aligned measure moving 0.93 while the standardized measure of the same construct moved 0.00. Budget for years, expect knowledge, and stop promising transfer.

Who this applies to

Group size
whole-classschool-wide
Delivered by
teacher
Ages studied
517(narrower than the 518 this topic is filed under — outside it is extrapolation)
Dose
For history KNOWLEDGE: three 10-15-day text-based units, ~45 min four days a week for about six weeks, after roughly two days (~10 hours) of teacher PD - g 0.19-0.46, retained at 8-12 weeks. For taught content knowledge in kindergarten: 45-60 min daily read-alouds, ~72 lessons over a semester - 0.93 aligned, 0.00 standardized. For any movement in GENERAL reading: three years minimum (Kim's grade 1-3 spiral, 0.11-0.12) or a six-year whole-school curriculum (Grissmer, 0.24). Nothing shorter has ever moved a standardized reading measure.
Cost
low
Moves
domain-skillnear-transfer
Needs first
Adequate decoding - every secondary trial assumes students can read the documents. The same content must be covered in both arms, so the gain comes from HOW it is taught rather than from covering more. And the sequence has to actually run across years: every one-semester test of transfer to reading failed.
Not for
Do not expect standardized reading gains inside one year - Cabell 0.00, Roberts 0.14 ns, Vaughn 2015 and Swanson 2015 null. Do not expect gap closing: low-prior-attainment students benefited least in Wanzek 2014 (no benefit at all), lower-vocabulary students least in Cabell (0.53 at -1 SD vs 1.31 at +1 SD), and struggling readers gained facts but not comprehension in Reisman. Do not buy generic historical-thinking PD as a route to better writing - the one trial with randomly selected teachers found nothing. And there is no evidence at all on chronological versus thematic sequencing; do not pretend the question has been answered.

Verdict

Content-rich history teaching works, on history. That is the finding, it is replicated across eight randomised trials, and it is smaller than what its advocates claim and larger than what its critics concede.

The reason this topic matters beyond history is that history is where the archive's most attractive hypothesis — knowledge is the durable lever for reading comprehension — should cash out, because history is the largest block of teachable general knowledge in a school day. The trials say it cashes out only over years, and not at all within a semester.

Meanwhile the rival organising principle for the subject, generic historical thinking taught as a transferable skill, has no adequately controlled positive result at all. Its best studies use volunteer teachers or no comparison group; the one that randomly selected which teachers got the training found nothing on its own aligned writing measure.

What the evidence shows

Source Design Grade Key effect
Roberts 2023 (PACT effectiveness) school-randomised, 48 schools / 7,183 students, nationally stratified A History knowledge 0.46 (researcher-designed), 0.40 at 9 wks. Standardized reading (Gates-MacGinitie) 0.14 ns; content-area comprehension 0.15 ns
Cabell 2025 (CKLA Knowledge) two RCTs (2nd a replication), 47 schools / 1,194 kindergarteners A Aligned social studies knowledge 0.93; standardized WJ Social Studies 0.00 (p = .999); taught vocabulary 0.63 vs standardized null
Vaughn 2013 (PACT) within-teacher cluster RCT, 27 classes, WWC "without reservations" B The only trial where standardized comprehension moved: Gates-MacGinitie 0.21; knowledge 0.31
Vaughn 2015 (PACT replication) direct replication, 1,487 students / 85 classes B Knowledge 0.32, retained 0.26 at 8 wks. Both comprehension effects failed to replicate
Swanson 2015 (PACT, grade 11) within-teacher RCT, 41 US History classes B Knowledge 0.36, 0.24 at 12 wks; no effect on either comprehension measure
Wanzek 2014 component-dismantling RCT, 26 grade-11 classes B Knowledge 0.19; low-prior-attainment students got no benefit
Kim 2024 (MORE) school-randomised, 30 schools / 2,870 students, grades 1→4, active control B Domain-general reading 0.11 standardized at grade 3, 0.12 at 14 months — and mathematics 0.12-0.16
Kim 2023 (MORE) same trial, grades 1→2 B Science content reading 0.18; transfer strongest on passages with taught vocabulary, weakest far
Grissmer 2023 kindergarten lottery, Core Knowledge charters B Grade 3-6 reading on independent state tests, ITT 0.24 / TOT 0.47
Tyner & Kabourek 2020 ECLS-K:2011 longitudinal regression, n = 6,731 C +30 min/day social studies → 0.15 SD grade-5 reading. ELA 0.03, maths 0.06, science −0.01, all ns
Reisman 2012 (Reading Like a Historian) quasi-experiment, 5 schools, 6 months, non-random, developer-led C Historical thinking ~0.49, transfer ~0.56, factual ~0.29 (researcher-designed); Gates-MacGinitie ~0.30. Struggling readers: no comprehension or transfer gain
Steiss 2026 yearlong quasi-experiment, randomly selected treatment/control teachers, n = 867 C No significant difference on the aligned argument-writing score; d = 0.29 only among self-selected second-year teachers
Nokes 2007 2×2 cluster RCT, 8 classrooms, 3 weeks C Multiple texts raised content knowledge and heuristic use; explicit heuristic instruction showed no content main effect
McCulley & Osman 2015 synthesis, 12 studies 1983-2013, grades 6-12 C Content learning improves with text-processing instruction; literacy instruction does not cost content coverage
De La Paz 2014 13 treatment vs 5 control teachers, non-random, reported as pre-post growth D ~0.5 SD growth in historical argumentation (researcher-scored), not in holistic writing quality; no credible counterfactual
Hwang 2022 meta of integrated content-and-literacy instruction C Content-integrated literacy moves standardized reading comprehension +0.25
Recht & Leslie 1988 the "baseball study" (n = 64) D Domain knowledge can outweigh reading ability on a topic-matched passage — an existence proof, unreplicated in 35 years

Cabell 2025 is the measure-alignment experiment this archive has been waiting for, and it is brutal. Same children, same semester, same construct, two instruments. Curriculum-aligned social studies knowledge moved 0.93 SD. The Woodcock-Johnson Social Studies subtest moved 0.00 SD, p = .999. The one apparently significant standardized effect — WJ Picture Vocabulary at 0.09 — collapsed to 0.01 (p = .702) once the authors deleted three items whose words (canoe, moccasins, loom) the curriculum explicitly taught. METHODOLOGY's nominal correction for researcher-designed measures is 2×. On a one-semester dose here it is not a discount; it is the entire effect.

Eight trials, one significant standardized reading result, and it did not replicate. Vaughn 2013's Gates-MacGinitie 0.21 is the single positive. The same team's own direct replication two years later returned nulls on both comprehension measures, the grade-11 version returned nulls, and the 48-school effectiveness trial returned 0.14 ns. That is a failed replication inside one research programme, which under this archive's replication rule caps how far the transfer claim can be pushed.

The only doses that have ever moved standardized reading are measured in years. Kim's three-year spiral gives 0.11-0.12; Grissmer's six-year Core Knowledge lottery gives ITT 0.24. This is consistent with the background-knowledge verdict that knowledge is a slow lever — but note the awkward detail: Kim's trial moved mathematics by 0.12-0.16, the same amount as reading. No knowledge-transfer mechanism predicts that. Something more general — engagement, teacher effort, time on academic text — is a live alternative explanation.

The skills side has no controlled positive result. Reisman's is the famous one and it is a non-randomised, developer-led, five-school study with unmodelled clustering whose author attributes the reading effect to time-on-task rather than to disciplinary transfer. De La Paz 2014 reports pre-post growth with volunteer teachers and no credible counterfactual. Steiss 2026 is the one design that randomly selected which teachers received the training, and it found nothing — with the positive d = 0.29 appearing only among the self-selected teachers who came back for a second year. And Nokes' 2×2 factorial separates the two ingredients directly: giving students multiple documents raised content knowledge; explicitly teaching them the sourcing heuristics did not.

Hereditarian-lens assessment

Risk: low, and this is one of the better-identified topics in the archive. Two grade-A school-randomised trials and a kindergarten admissions lottery carry the verdict; within each, treated and untreated children have the same expected distribution of ability-relevant alleles.

Where the lens still does work:

  • The gap-closing claim fails under the lens, and it fails in the trials too. Content-rich instruction compounds advantage rather than closing gaps. Cabell's effect on aligned knowledge ran 0.53 at −1 SD entering vocabulary, 0.92 at the mean and 1.31 at +1 SD; English learners 0.58 against native speakers' 0.99. Wanzek 2014's low-pretest students got no benefit at all. Reisman's struggling readers gained facts but not comprehension. This is exactly the pattern METHODOLOGY predicts when instruction removes an environmental bottleneck: the children with more prior knowledge and vocabulary extract more from more knowledge.
  • Tyner & Kabourek is the most-cited source in this area and it is observational. Time allocation is a school and teacher decision correlated with everything. Its own internal check is the strongest argument against it: science is content-rich too, and 30 extra minutes of daily science predicts −0.01 SD. A background-knowledge account predicts a positive science coefficient. The authors' reverse-causation robustness check is good and rules out one confound; it does not create identification.
  • Recht & Leslie is an existence proof, not a lever. High-knowledge poor readers out-recall low-knowledge good readers on a topic-matched passage. That is near transfer inside a known domain, from n = 64, unreplicated in 35 years. It shows knowledge matters; it does not show that teaching knowledge raises general reading, and the trials above are the direct test.

Boundaries & what critics say

  • Almost every positive number here is on a researcher-designed, treatment-aligned measure, produced by the people who built the curriculum. The PACT trials are developer-led; so is Reisman; so is De La Paz. The two independent standardized results in the whole cluster are Grissmer's state tests and Cabell's Woodcock-Johnson, and they point in opposite directions at different doses.
  • PACT's effectiveness trial produced a larger knowledge effect (0.46) than its developer-supported efficacy trials (0.31-0.36) — a clean counterexample to this archive's efficacy-to-effectiveness decay through-line. Two caveats: the outcome is the developers' own aligned measure throughout, and PACT classes were 0.13 SD ahead at baseline.
  • Grissmer 2023 remains an unpublished working paper three years on, with no critique and no replication, and it introduces a novel bias correction of the authors' own devising for "two previously undiscovered sources of bias inherent in kindergarten lotteries." The archive's highest-value curriculum result therefore depends on a non-standard correction nobody has independently vetted.
  • Chronology versus thematic sequencing is unevidenced. There is no causal literature on how to order a history curriculum. Strong opinions exist; trials do not.
  • No causal evaluation of England's 2014 knowledge-led national curriculum exists. A whole national curriculum was rewritten around this question and left no identification strategy behind.
  • A defender of historical thinking would say the trials test the wrong outcome — that disciplinary reasoning is worth teaching for its own sake, not as a reading intervention. That is a fair position, and it is the position historical source work evaluates on its own terms.

Practical guidance

  • Build history as a coherent, cumulative, text-heavy content curriculum, and run it for years. Every measurable transfer effect in this topic required three years or more.
  • Justify it on history knowledge. That is the outcome it reliably delivers, at roughly two days of teacher PD and a set of units, and it is a worthy outcome in its own right.
  • Do not promise a reading-comprehension gain inside a year. Eight trials; one significant result; it did not replicate. If a programme is sold to you on a one-year comprehension effect, ask which test.
  • Use multiple documents rather than a textbook — it raised content knowledge in the one factorial that isolated it — but do not expect the explicit heuristic instruction bolted on top to add content learning.
  • Do not buy generic historical-thinking PD as the organising principle. The best-identified study found nothing on its own aligned measure, and the apparent effect lived entirely in teacher self-selection.
  • Plan for the strong to gain more. Knowledge-rich instruction compounds prior knowledge. If closing gaps is the goal, this is not the lever, and pretending otherwise sets the programme up to be judged on a criterion it will fail.
  • Do not cut science to buy social-studies minutes on the strength of Tyner & Kabourek. The same paper finds science time predicting nothing, which should make anyone cautious about reading its social-studies coefficient causally.

Open questions

  • Is Grissmer replicable? It is the single most valuable result in the archive's curriculum area, it is unpublished, uncritiqued, unreplicated, and rests on a bespoke bias correction. Replicating it is the highest-value study in this topic.
  • Why did Kim's content curriculum move mathematics as much as reading? Either the mechanism is not knowledge transfer, or knowledge transfer is not specific — and both readings change the recommendation.
  • Does the knowledge effect ever show on a standardized measure at a dose under three years? Nothing has, and nobody has mapped the dose–response curve between one semester (0.00) and six years (0.24).
  • Chronological or thematic? Unevidenced by any causal design.
  • What happens to low-prior-knowledge students? Three studies find they gain least or not at all, and nobody has tested a version designed for them.

Evidence (17 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore

Should history be built on content knowledge or on generic historical-thinking skills? · The Evidence on Teaching