The Evidence on Teaching

Effects of Family Literacy Programs on the Emergent Literacy Skills of Children From Low-SES Families: A Meta-Analysis

Fikrat-Wevers S, van Steensel R, Arends LR · 2021

grade Cmeta-analysisindependentmixed
Sample
48 (quasi-)experimental studies covering 42 different programmes; 65 experimental comparisons
Population
Children aged 0-6 from low-socioeconomic-status families (also coded for ethnic-minority and immigrant status), in family literacy programmes where parents are trained to run home literacy activities.
Design
Review of Educational Research 91(3). The most recent and most heavily moderated of the family-literacy metas, and the one that separates measurement artefact from effect most cleanly. Programmes averaged 23 training sessions. Trainer type coded three ways (professionals — teachers/researchers; paraprofessionals; both). Instrument coded as existing/standardized vs study-specific. Design coded experimental (individual random assignment) vs quasi-experimental. Timing coded immediate posttest vs follow-up. Publication status was NOT a significant moderator (Q(1) = 0.01), so no indication of publication bias.
Key findings
Two facts dominate. First, the effect is a posttest effect: d = 0.50 (SE 0.07) on immediate posttests collapses to d = 0.16 (SE 0.09), marginal, at follow-up. For comprehension- related skills the collapse is near-total (0.51 -> 0.09); code-related skills hold up better (0.48 -> 0.22). Second, roughly half the headline is measurement: study-specific (researcher- developed) instruments returned d = 0.84 against d = 0.37 for existing standardized tests (Q(2) = 8.61, p < .05), and the gap is starkest on comprehension (1.29 vs 0.34). On the question of who delivers, the answer is again nobody in particular: trainer type was NOT a significant moderator (professionals d = 0.65, paraprofessionals d = 0.40, Q(3) = 5.39, n.s.), and neither were programme duration, number of sessions, use of modeling, or use of guided practice. The authors' own summary is that "program effects delivered by professional trainers are not necessarily more effective than those delivered by paraprofessionals". What DID moderate: targeted programmes focused on a limited set of activities in a single training context beat broad ones. Experimental designs paradoxically returned larger effects than quasi-experimental (0.98 vs 0.32, Q(1) = 8.76, p < .01), which the authors trace to confounding — all 12 experimental programmes were home-only, literacy-only, and mostly shared-reading-only, i.e. the targeted kind.
Genetic confound
Medium. Random assignment covers only 12 of 48 studies, and the low-SES samples are volunteers. But the internal comparisons that matter here — standardized vs researcher-made instrument, professional vs paraprofessional trainer, posttest vs follow-up — are within-literature contrasts where family composition is roughly held constant, so those rankings are more trustworthy than the levels.
Replication notes
Reproduces the fadeout finding of van Steensel et al. 2011 (short-term 0.20 -> follow-up 0.04) with a larger, low-SES-restricted corpus, while disagreeing on the immediate magnitude (0.50 vs 0.18). The researcher-developed-measure inflation replicates the archive-wide pattern (Cheung & Slavin; de Boer; and Senechal & Young's own moderator). The trainer-type null is a direct replication of van Steensel et al. 2011's trainer-type null on a different corpus — two independent tests, same answer: professional training of the parent-trainer does not buy anything measurable.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Emergent literacy, immediate posttestCohen's d0.50 (SE 0.07) overall; comprehension-related 0.51 (SE 0.09); code-related 0.48 (SE 0.09)mixedimmediate posttestbusiness-as-usualend-of-treatmentdomain-skill
Emergent literacy, follow-upCohen's d0.16 (SE 0.09) overall, marginal; comprehension-related 0.09 (SE 0.09); code-related 0.22mixedfollow-upbusiness-as-usualuncleardomain-skill
Measure type — existing standardized vs study-specific instrumentCohen's d by categoryexisting/standardized 0.37 (SE 0.08, k=30) vs study-specific 0.84 (SE 0.14, k=10); Q(2) = 8.61, p < .05. On comprehension: 0.34 vs 1.29. On code: 0.21 vs 0.68 (Q(1) = 8.48, p < .01)mixedposttestbusiness-as-usualend-of-treatmentdomain-skill
Trainer type moderator (who trained the parent)Cohen's d by categoryprofessionals 0.65 (SE 0.13, k=21); paraprofessionals 0.40 (SE 0.22, k=8); Q(3) = 5.39, NOT significant. Programme duration, number of sessions, modeling and guided practice were also all non-significant.mixedposttestbusiness-as-usualend-of-treatmentdomain-skill
Design — experimental vs quasi-experimentalCohen's d by categoryexperimental (individual randomization) 0.98 (SE 0.22, k=12) vs quasi-experimental 0.32 (SE 0.06, k=36); Q(1) = 8.76, p < .01 — authors attribute the reversal to confounding, since all 12 experimental programmes were home-only and literacy-onlymixedposttestbusiness-as-usualend-of-treatmentdomain-skill

Cited by