The Evidence on Teaching

Promoting adolescents' comprehension of text: A randomized control trial of its effectiveness

Roberts G, Vaughn S, Wanzek J, Furman G, Martinez L, Sargent K · 2023

grade Arctdeveloper-ledreplicated
Sample
48 schools, 135 teachers, 7,183 students (59 schools randomised, 11 dropped out post-assignment)
Population
US eighth-graders in public middle schools teaching US history; stratified balanced sample drawn to represent the national population of such schools
Design
SCHOOL-randomised effectiveness trial of PACT (Promoting Adolescents' Comprehension of Text) - a set of five text-and-discourse components (comprehension canopy, essential words, critical reading, team-based-learning comprehension checks, TBL knowledge application) layered onto ordinary US history instruction for three 2-week units (Colonial America, Road to Revolution, Revolutionary War), ~45 min 4 days/week for ~6 weeks. The design is unusually strong on external validity: schools were drawn by stratified balanced sampling from the population frame of US middle schools teaching US history, and the paper reports POPULATION average treatment effects (PATE) reweighted by inverse propensity score weights alongside sample effects, with a B-index of .83 indicating high generalisability after reweighting. Effective sample size was only ~22 schools after trimming weights, so cluster-robust small-sample corrections matter. Attrition 14-20% overall, differential 0.01-7.4%, within WWC tolerable limits. CRITICAL MEASURE ASYMMETRY: the knowledge outcome (ASK-MC) is a RESEARCHER-DEVELOPED 42-item test built by the developers from released state and AP items and aligned to the three taught units; the reading-comprehension outcomes are the researcher-developed ASK-RC (content-area passages) and the independent standardized Gates-MacGinitie. Baseline imbalance on ASK-MC favoured PACT by g = 0.13 (and by g = 0.35 on ASK-RC in the reduced pretest subsample), so the unadjusted posttest ES of .46 is an upper bound; significance tests controlled for pretest. Authors are the programme developers.
Key findings
The cleanest large-scale answer in this literature, and it splits exactly along measure type: teaching history content richly raises HISTORY KNOWLEDGE and does NOT reliably raise standardized reading comprehension. Content knowledge (developer-built ASK-MC): sample ATE 0.46 at posttest and 0.40 at 9-week follow-up; population ATE 0.45 (weighted mixed effects) / 0.37 (weighted OLS) at posttest, 0.53 / 0.33 at follow-up, all statistically significant. Content-area reading comprehension (ASK-RC, researcher-developed): g = 0.15, NOT significant. Broad reading comprehension (Gates-MacGinitie, independent standardized): g = 0.14, NOT significant - though the authors note this is larger than the average of their earlier efficacy trials and argue it is non-trivial for older readers. This is an EFFECTIVENESS trial (no researcher coaching of teachers), and the knowledge effect held up or grew relative to the developer-supported efficacy trials, which cuts against the usual efficacy-to- effectiveness decay - but the knowledge measure is the developers' own, aligned to the taught units, and the archive discounts such measures ~2x.
Genetic confound
Low. Randomisation at school level with pretest-controlled models, so allele distributions are balanced in expectation across arms. Residual threats are the 11 post-assignment school dropouts and 14-20% student attrition, not selection on ability.
Replication notes
This is the sixth-plus study in the PACT programme and its findings replicate the earlier efficacy trials' consistent pattern (Vaughn 2013; Vaughn 2015; Swanson 2015 high school; Wanzek 2014 TBL): content acquisition moves every time (g = 0.17 to 0.46), reading comprehension moves in only one of the trials. All of these are by the same developer team, so this is internal, not independent, replication - no outside group has evaluated PACT.
DOI / URL
10.1037/edu0000794

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
US history content knowledge (ASK-MC), sample average treatment effecteffect size0.46 at posttest; 0.40 at 9-week follow-upresearcher-designedimmediately post-units and 9 weeks laterbusiness-as-usualend-of-treatmentdomain-skill
US history content knowledge (ASK-MC), POPULATION average treatment effectPATE (weighted mixed effects / weighted OLS)0.45 / 0.37 at posttest; 0.53 / 0.33 at 9-week follow-up; all significantresearcher-designedimmediately post-units and 9 weeks laterbusiness-as-usualend-of-treatmentdomain-skill
Content-area reading comprehension (ASK-RC)Hedges g0.15, not statistically significantresearcher-designedimmediately post-unitsbusiness-as-usualend-of-treatmentnear-transfer
Broad reading comprehension (Gates-MacGinitie, standardized)Hedges g0.14, not statistically significantstandardizedimmediately post-unitsbusiness-as-usualend-of-treatmentnear-transfer
Baseline imbalance (context for the 0.46)Hedges g at pretestASK-MC 0.13 (CI -0.13 to 0.38); ASK-RC 0.35 (CI -0.26 to 0.96); GM-RC 0.03mixedpretestnonenot-applicabledomain-skill

Cited by