The Evidence on Teaching

How Effective Are Family Literacy Programs? Results of a Meta-Analysis

van Steensel R, McElvany N, Kurvers J, Herppich S · 2011

grade Cmeta-analysisindependentmixed
Sample
30 effect studies (1990-2010) covering 47 samples/comparisons; study sample sizes 15 to 781
Population
Children in the preformal (n=15 programmes) and formal (n=16) education phase, at-risk (17 programmes) and non-at-risk (14), in family literacy programmes where trained parents deliver home literacy activities. 21 distinct programmes including Dialogic Reading (5 studies), Paired Reading (3), HIPPY, Project PRIMER, Project EASE.
Design
Review of Educational Research 81(1), 69-96. The direct competitor to Senechal & Young 2008 on the same question, and it is the more conservative of the two: 47 comparisons vs 16 studies, outliers trimmed at 2 SD rather than 3 SD, and vocabulary measures included rather than excluded. The authors explicitly attribute the gap between their d=0.18 and Senechal & Young's d=0.65 (CI 0.53-0.76) and Mol et al.'s d=0.42 (CI 0.16-0.54) to those two decisions plus the imprecision of the earlier estimates: their own CI is 0.11-0.24. Effect sizes weighted by inverse sampling variance. Publication bias checked two ways: fail-safe N = 38 against a 0.10 criterion, and Egger's regression intercept not significant (p=.516). Fixed- and random-effects models gave identical means. 26 of 47 comparisons came from randomized studies.
Key findings
Parents trained to run home literacy activities move children's literacy by about d = 0.18 — "not more than a three-point gain on a standardized test such as the PPVT". The headline moderator finding is that there were NO significant moderators at all: not programme duration, not home visits vs group meetings, not at-risk status, not book provision, and — the result this archive cares about — not who trained the parent. Professional trainers d = 0.21, semiprofessionals d = 0.18, both d = 0.12, Qbetween = 0.59, n.s.; restricted to at-risk families, professionals and semiprofessionals were identical at d = 0.16 (Qbetween = 0.00). Two contrasts did move, neither significantly: randomized studies d = 0.11 vs non-randomized d = 0.22, and short-term d = 0.20 vs follow-up d = 0.04 (n.s.). Shared-reading-only programmes were the only activity type with a null (d = 0.05, CI -0.11 to 0.20); shared reading PLUS other literacy activities gave d = 0.21. Programme focus did not predict which skill moved: code-focused programmes did not beat comprehension-focused programmes on code-related skills, which the authors read as parents not doing what the programme told them to do.
Genetic confound
Medium. Roughly half the comparisons are randomized and those still yield d = 0.11, so the effect is not purely selection. But the non-randomized half runs at d = 0.22, and families who volunteer for and complete a family literacy programme differ from those who do not on traits that are themselves heritable. The follow-up null (d = 0.04) is the more damaging fact: whatever the programme installed did not persist.
Replication notes
Directly contests Senechal & Young 2008 (already in this archive at d = 1.15 / 0.51 / 0.18 by activity type; overall d = 0.65) and Mol et al. 2008 (d = 0.42) on the same literature, and names the analytic choices responsible. Fikrat-Wevers, van Steensel & Arends 2021 — same lead group, low-SES subset, 48 studies — finds d = 0.50 at immediate posttest but d = 0.16 at follow-up, which reproduces this meta's central fadeout result while disagreeing on the immediate magnitude. The shared-reading-only null replicates in Noble et al. 2019 (g = 0.03 against an active control). Read together, the family-literacy metas agree on the ORDERING (instruction > listening > exposure) and disagree by a factor of three on the level.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
General literacy ability, family literacy programmes vs controlCohen's d (weighted, fixed and random effects identical)0.18, 95% CI [0.11, 0.24], z = 5.51, p < .001, 47 comparisonsmixedend of programme (predominantly)business-as-usualend-of-treatmentdomain-skill
Comprehension-related vs code-related skillsCohen's dcomprehension 0.22 [0.15, 0.29]; code 0.17 [0.08, 0.25] — no differential impactmixedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Who trained the parent (staff quality moderator)Cohen's d by categoryprofessionals 0.21 [0.11, 0.30] (k=29); semiprofessionals 0.18 [0.05, 0.31] (k=10); both 0.12 [-0.09, 0.33] (k=3); Qbetween = 0.59, n.s. Within at-risk families professionals and semiprofessionals were identical at d = 0.16 (Qbetween = 0.00, df = 1, p > .05)mixedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Randomized vs non-randomized comparisonsCohen's d by categoryrandomized 0.11 [0.01, 0.21] (k=26); non-randomized 0.22 [0.14, 0.30] (k=21); Qbetween = 2.83, p < .10 trend only, and the trend disappeared under a mixed-effects model (Qbetween = 1.75)mixedend of programmebusiness-as-usualend-of-treatmentdomain-skill
Persistence — short-term vs follow-up measurementCohen's d by categoryshort term 0.20 [0.13, 0.27] (k=38); follow-up 0.04 [-0.14, 0.22], NOT significant (k=9)mixedfollow-up after programme endbusiness-as-usualuncleardomain-skill
Activity type — shared reading alone vs shared reading plus instructionCohen's d by categoryshared reading only 0.05 [-0.11, 0.20], n.s. (k=14); shared reading + other literacy activities 0.21 [0.14, 0.28] (k=27); literacy exercises only 0.17 [-0.06, 0.40], n.s. (k=6); Qbetween = 3.59, n.s.mixedend of programmebusiness-as-usualend-of-treatmentdomain-skill

Cited by