The Evidence on Teaching

Effects of Integrated Literacy and Content-area Instruction on Vocabulary and Comprehension in the Elementary Years: A Meta-analysis

Hwang, H., Cabell, S. Q., & Joyner, R. E. · 2022

grade Cmeta-analysisindependentnot-applicablenumbers spot-checked
Sample
35 (quasi-)experimental studies, 13,289 students (9,530 treatment / 7,934 control), kindergarten through grade 5; 32 journal articles (231 effect sizes) plus 3 dissertations (9 effect sizes). 33 US, 1 Netherlands, 1 UK.
Population
US-dominated elementary classrooms (K-5); ~66% students of colour, 62% eligible for lunch services, 16% with reading or learning difficulties. Interventions integrate literacy instruction with science and/or social studies content.
Design
Random-effects meta-analysis with Egger's regression and Mathur & VanderWeele sensitivity analysis for publication bias. Crucially it codes outcome measures on TWO dimensions: standardized vs researcher-developed, and — within researcher-developed — 'complete proximal' (fully aligned to the taught content) vs 'less proximal'. Most interventions (k=28) taught content similar or identical to what control students were taught, and k=26 were implemented as intended. Independence of the meta-analysts is genuine (Minnesota/Florida State reading researchers, not curriculum developers) but the POOLED STUDIES are heavily developer-led — Science IDEAS (Romance & Vitale), Seeds/Roots (Cervetti et al.) and CALI (Connor et al.) are all in the pool and all evaluated by their own developers. Journal issue: Scientific Studies of Reading 26(3), 2022; Crossref online date 2021-08-30. Read in full text.
Key findings
This is the best available estimate of the aligned-measure inflation factor in content-area literacy integration, and it is roughly 2x on comprehension with the standardized estimates for vocabulary and knowledge too imprecise to be usable. Overall: vocabulary g = 0.91, comprehension g = 0.40, content knowledge g = 0.89. Split by measure type: COMPREHENSION researcher-developed g = 0.54 (n=123 effects) vs standardized g = 0.25 (n=26, CI 0.04-0.46) — a 2.2x gap, and 0.25 is the number to quote. VOCABULARY researcher-developed g = 0.86 (n=31, p<.01) vs standardized g = 0.64 but NOT SIGNIFICANT (n=5, CI −1.59 to 2.88) — five effect sizes is not an estimate. CONTENT KNOWLEDGE researcher-developed g = 0.94 (n=49) vs standardized g = 0.68, again NOT SIGNIFICANT (n=6, CI −0.38 to 1.74). Within researcher-developed vocabulary measures, HIGHLY aligned measures were significant and LESS aligned ones were not — the alignment gradient is visible inside the researcher-designed category itself. Egger's test flagged possible publication bias for all three outcomes, but the sensitivity analysis found no attainable level of bias would drive the estimates to zero. No moderator among design features or intervention characteristics was significant, which the authors attribute partly to the small number of studies. The founder-facing translation: integrating literacy with science content reliably teaches the science content and the taught words; its effect on INDEPENDENT reading comprehension is about d = 0.25, and its effect on independent vocabulary and knowledge measures is unestablished because almost nobody has measured them.
Genetic confound
LOW for the pooled estimate — the studies are experimental or quasi-experimental with control groups, so passive gene-environment correlation is not the operative threat. Measure alignment and developer-led evaluation are.
Replication notes
Not applicable — this IS the synthesis. It converges with the archive's existing reading evidence: Elleman 2009 found custom d = 0.50 vs standardized d = 0.10 for vocabulary instruction on comprehension, and the Grissmer Core Knowledge lottery found ITT 0.24 on independent state reading tests, both consistent with an honest standardized value around 0.1-0.25 for knowledge-building levers.
DOI / URL
10.1080/10888438.2021.1954005

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Reading comprehension, integrated literacy + content-area instruction (all measures pooled)g0.40mixedpost-interventionbusiness-as-usualend-of-treatmentnear-transfer
Reading comprehension — RESEARCHER-DEVELOPED measuresg0.54 (n=123 effects, p<.01)researcher-designedpost-interventionbusiness-as-usualend-of-treatmentnear-transfer
Reading comprehension — STANDARDIZED measures (the headline number)g0.25 (n=26 effects, CI 0.04-0.46, p<.05)standardizedpost-interventionbusiness-as-usualend-of-treatmentnear-transfer
Vocabulary — researcher-developed vs standardizedgresearcher-developed 0.86 (n=31, p<.01); standardized 0.64 NOT significant (n=5, CI -1.59 to 2.88)mixedpost-interventionbusiness-as-usualend-of-treatmentdomain-skill
Science/social studies content knowledge — researcher-developed vs standardizedgresearcher-developed 0.94 (n=49); standardized 0.68 NOT significant (n=6, CI -0.38 to 1.74)mixedpost-interventionbusiness-as-usualend-of-treatmentdomain-skill

Cited by