The Evidence on Teaching

An Investigation of Two Approaches to Fluency Instruction in the General Education Classroom: Repeated Reading Versus Varied Practice Reading

Reed, D. K., Zimmermann, L., Reeger, A., Aloe, A. M., & Sidler-Folsom, J. · 2018

grade Brctdeveloper-ledmixednumbers spot-checked
Sample
827 fourth graders analysed (405 Varied Practice, 422 Repeated Reading) from 1,029 consented, in 21 elementary schools across 10 Iowa districts; Iowa Reading Research Center technical report, Iowa City, IA (not peer-reviewed)
Population
US fourth-grade general-education students, spring 2018. 79.4% White, 7.5% Hispanic, 5.9% Black, 54.7% free/reduced-price lunch, 8.1% English learners, 11.1% special education, 8.8% gifted, 51.9% female. Crucially NOT a struggling-reader sample: mean winter FAST oral reading fluency was 143.5 WCPM against a fourth-grade proficiency benchmark of 136, and students scoring ≤43 WCPM at winter were excluded from analysis. The repeated-reading evidence base this study argues with (Lee & Yoon 2017, Morgan et al. 2012) is largely about students with or at risk for reading disabilities.
Design
Two-arm student-randomised trial with no business-as-usual control, so the study can speak to WHICH form of fluency practice is better and not to whether fluency practice works. Arms: Repeated Reading (one passage read three times in succession) vs Varied Practice Reading (three different passages read once each, the first identical to the RR passage and the other two written to share 85% of its unique words). Practice time was equal by design — about 20 min per session, 3-4 times per week, 30 sessions possible over roughly 12 weeks (Feb-Apr), mean 26 completed, range 1-30. Delivery was peer-dyadic inside the general-education classroom: partners alternated reading until each had read three times, the listener timed the reading and marked errors, then helped review missed words; teachers monitored. Self-reported adherence was 96.5% and balanced across arms. Passages ran 110-647 words (most ~300) and spanned 400-1200 Lexile. Outcome: FAST CBMreading spring benchmark — a commercial standardised screener on unpracticed text, NOT a researcher-designed probe and NOT the practiced passage, which is what makes this a real test of transfer rather than of rehearsal. Analysis was a multilevel ANCOVA on the spring score with students nested in classrooms nested in schools, controlling winter FAST, sessions completed, gender, race, EL, FRL, special education and gifted status. Only winter FAST, special education (-3.699, p = .043) and condition were significant. Sample flow, and a reporting error in it: 1,029 consented; 15 withdrew, moved or submitted no logs (1.5% attrition); the report then says "another 202 students (Varied Practice = 100; Repeated Reading = 102)" were removed, but its own reason breakdown sums to 187 — incomplete FAST data 11, missing demographic data 148, winter FAST ≤43 WCPM 10, reading without the assigned partner for >5 sessions 14, and not staying in the assigned condition 4. The numbers only reconcile if 202 is the TOTAL lost (15 attrition + 187 removals), since 1,029 - 202 = 827 while 1,014 - 202 = 812. So the analytic n is right and the label "another 202" is wrong. Net, 19.6% of consented students are not in the analysis. Most of that loss is benign for randomisation — missing demographic covariates (14.4% of the sample) and a baseline-score floor cannot be caused by treatment — but 18 students (2.2%) were removed for post-randomisation, treatment-related reasons (partner non-compliance, condition crossover), so this is per-protocol rather than intention-to-treat. The report states z-tests found no differential loss between arms, and Tables 1-2 show balance on every demographic, on pretest FAST (143.98 vs 143.05) and on fidelity. Grade B: a single well-powered randomised trial with a standardised outcome, appropriate nesting, documented balance and near-zero treatment-related attrition. The demerits that do not touch the grade are recorded where the methodology puts them — `independence: developer-led`, because Varied Practice Reading was designed by the IRRC, which also ran the trial, analysed it, published the report, and now distributes the 30 VPR passage sets for download. It is also grey literature with no preregistration and no external peer review. FULL TEXT READ 2026-08-05 (9pp, IRRC report); every figure in this record checked against the PDF. The PDF has been on disk since 2026-08-01 and was previously summarised only inside db/sources/ardoin-2016-repeated-vs-wide-reading.yaml; this record makes it addressable.
Key findings
The large-sample companion to Ardoin 2016, and it is a developer-led study whose result argues against the developer's own interest in the only way that counts: the IRRC built Varied Practice Reading, ran the trial, and could only produce a 2.018 WCPM edge over Repeated Reading with a negligible effect size (0.052) that it explicitly declines to call practically significant. Under the efficacy-to-effectiveness decay this archive keeps rediscovering, a developer-led estimate is the ceiling, so the true VPR-over-RR advantage in independent hands is ≤0.05 SD and Ardoin's independent null on the same question is the likelier truth. The load-bearing point for the fluency literature is therefore the negative one: with practice time held equal and the outcome measured on unpracticed text via a commercial screener, it makes essentially no difference whether a fourth grader rereads one passage three times or reads three overlapping passages once each. Two corrections a full-text read forces. First, the widely quoted "+18.9 vs +20.9 WCPM" gains — quoted in this archive's own reading-fluency topic — are the model's intercept and intercept-plus- treatment-coefficient from a posttest-on-pretest ANCOVA, not average gains; the report's own unadjusted means imply ~14 and ~16 WCPM, so the "both groups grew near the 90th percentile" framing is not supported by its own data and would land nearer the 75th. Second, there is no control arm at all, so no reading in this report speaks to whether fluency practice beats doing something else with the twenty minutes. Applicability boundary worth carrying: this is general-education fourth graders reading at or above benchmark (mean 143.5 WCPM vs a 136 cut, sub-43 readers excluded), practising in peer dyads — not the struggling-reader, adult-led one-to-one condition that produced the large repeated-reading effect sizes in the Therrien and Lee & Yoon meta-analyses.
Genetic confound
Low for the between-arm contrast. Students were randomly assigned to condition within participating classes and the arms are balanced on every recorded demographic and on pretest fluency, so heritable differences cannot drive the 2.018 WCPM difference. The uncontrolled growth-against-norms comparison carries the usual confounding of any pre-post design, but nothing in this study rests on it.
Replication notes
Marked `mixed` because the two halves of this study's result have different replication records. The direction that matters — repeating the same passage is not the active ingredient, and equal-time non-repetitive practice does at least as well — converges with Ardoin, Binder, Foster & Zawoyski (2016, Journal of School Psychology; db/sources/ardoin-2016-repeated-vs-wide-reading.yaml), an independent three-arm RCT that yoked oral-reading time (451 vs 449 min) and found no significant repeated-vs-wide difference on any achievement, fluency, prosody or eye-movement measure at any skill level. That is the more valuable agreement, and it comes from a genuinely independent team. What is NOT replicated is this study's own headline: a statistically significant advantage for Varied Practice. Ardoin found a null on the same contrast, and the developer's own follow-up work (Reed then took a $2M IES award to extend VPR to middle school) is not independent evidence. So the honest reading is convergence on the practical bottom line — repetition buys nothing detectable — and non-replication of the one directional win, which the developers themselves call not practically significant. Note the two designs differ in the non-repetitive arm: Ardoin's wide reading used unrelated varied passages, while VPR engineers 85% unique-word overlap between passages (following Rashotte & Torgesen 1985), so VPR is closer to repeated practice of the same WORDS in varied contexts than to wide reading.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Varied Practice Reading vs Repeated Reading at equal practice time (the randomised contrast)Adjusted posttest difference on FAST CBMreading, multilevel ANCOVA controlling winter FAST and demographicsVaried Practice ahead by 2.018 WCPM (SE 1.002, t = 2.014, p = .044), ES = 0.052 — in the report's own words "the effect size was negligible" and "statistically significant but not practically significant... teachers likely would not notice any difference". One internal tension the report does not address: it lists the standard error of that effect size as 0.171, which puts a 95% interval at roughly -0.28 to 0.38 and straddles zero, contradicting the p = .044 obtained from the raw coefficient (2.018/1.002). The raw coefficient is internally consistent with the pretest SD (2.018/38.23 = 0.053), so the reported ES is right and its SE is the anomaly — but as printed the two significance statements disagree.standardizedspring FAST screening, after ~12 weeks (mean 26 of 30 possible sessions)active-alternativeend-of-treatmentdomain-skill
Growth in both arms against published FAST growth norms (uncontrolled, and misreported)Winter-to-spring WCPM change vs FAST developer percentile growth rates (1.5 WCPM/week = 90th pct, 1.28 = 75th, 0.98 = 50th); no control groupThe report headlines gains of 18.90 WCPM (Repeated) and 20.91 WCPM (Varied) and concludes both arms grew "near the 90th percentile", "ambitious growth that not many students could achieve". Those two figures are the ANCOVA intercept (18.895) and the intercept plus the treatment coefficient (18.895 + 2.018 = 20.913). The outcome of that model is the spring SCORE, not a gain (winter FAST enters as a predictor with coefficient 0.933, t = 58.15), so the intercept is a predicted spring score at winter FAST = 0 and cannot be an average gain. The report's own unadjusted means contradict it: winter 143.05/143.98 against spring means stated as "about 7" (Repeated) and "about 10" (Varied) WCPM above the 150 spring benchmark, i.e. ~157 and ~160, giving gains of roughly 14 and 16 WCPM — about 1.16 and 1.34 WCPM/week, between the 50th and 75th percentile benchmarks, not near the 90th. The 2.018 between-arm difference is unaffected by this; the "exceptional growth" claim is. Separately, with no business-as-usual arm this comparison is pre-post against published norms and carries no causal weight either way.standardizedwinter to spring FAST screening, ~12 weeksnoneend-of-treatmentdomain-skill

Cited by