Putting computerized instruction to the test: a randomized evaluation of a "scientifically based" reading program
Rouse, C. E., & Krueger, A. B. (with Markman, L.) · 2004
grade Brctindependentreplicatednumbers spot-checked
Sample
512 students randomly assigned (272 to Fast ForWord, 240 control) in grades 3-6, across 4 schools, randomised within grade-by-school blocks, in two sequential "flights". Follow-up coverage is unusually complete - post-test CELF-3-RP for 100 percent of the sample, Reading Edge for 95 percent, the Success for All assessment for 96/97 percent, and the state assessment for 87/90 percent.
Population
Students in grades 3-6 in four schools of one large, poor urban US district (over 20,000 students; 40 percent African American, over 50 percent Hispanic, ~70 percent free-lunch eligible, 56 percent speaking a language other than English at home). Eligibility was restricted to students scoring in the BOTTOM 20 PERCENT statewide on the state reading test - i.e. exactly the struggling readers Fast ForWord targets, with a mean CELF-3 receptive score of about 26 NCE against a norm of 50.
Design
THE INTERVENTION: Fast ForWord, from Scientific Learning Corporation, the paradigm case of a "scientifically based" ed-tech product - built on Tallal and Merzenich's temporal-processing theory, marketed on brain-imaging evidence, and by 2001 drawing 76 percent of SLC revenue from public schools. Students train 90-100 minutes a day on acoustically modified speech exercises; the programme gates progression on proficiency. FUNDING AND INDEPENDENCE: financed by the Smith Richardson Foundation and the Education Research Section at Princeton University, and run and authored by Princeton/NBER economists. SLC cooperated - it trained the instructors, made periodic site visits and gave phone support - but did not run, fund or write the evaluation, so `independence: independent`. (One school's supervising teacher complained the company was unresponsive to repeated requests for help with software failures.) IDENTIFICATION: student-level random assignment within grade-by-school randomisation blocks. Eligible students were those in the bottom statewide reading quintile; before randomisation, principals removed students they judged unable to sit through 90-100 minutes of computerised instruction, students who had already transferred, and students otherwise unavailable - so the randomised sample is a screened subset, but the screening happened BEFORE assignment and therefore does not bias the contrast. Estimates are difference-in-differences with randomisation-pool fixed effects, pre-tests and demographics; treatment-on-the-treated is estimated by IV using assignment as the instrument for three separate SLC-defined completion measures. WHAT THE CONTROL STUDENTS DID WITH THE SAME TIME - AND WHAT THE TREATED STUDENTS GAVE UP: this is the cleanest displacement accounting in the cluster and it matters. Fast ForWord was delivered as a PULL-OUT. Students left their regular classroom for 90-100 minutes a day; the paper tabulates exactly what they missed, school by school and grade by grade - homeroom, mathematics, writing, language arts, spelling, science, social studies, literature, grammar, "specials" (art, music, gym), some lunch and recess, and in one school before- or after-school time. In NO case were students pulled out of Success for All, the district's core reading programme, so the authors describe FFW as "primarily an add-on to regular reading instruction" while noting the counterfactual was mixed. The honest reading: reading instruction was protected, and roughly an hour and a half a day of mathematics, science, writing and the arts was spent instead on the software - for no measurable reading gain. THE MEASURE CONTRAST IS THE POINT OF THIS PAPER. Four outcomes, deliberately spanning the alignment spectrum: (1) READING EDGE, a computerised assessment OWNED AND SOLD BY SCIENTIFIC LEARNING, the vendor of Fast ForWord, testing precisely the phonological skills the programme trains; (2) the CELF-3-RP, the receptive portion of the Clinical Evaluation of Language Fundamentals (3rd ed.) - three subtests plus Listening to Paragraphs - a well-known independent clinical language test, administered one-to-one by three certified evaluators independent of the district; (3) the Success for All reading assessment used by the district; and (4) the state's standardized reading assessment. Effects appear only on the vendor's own instrument, and only barely. IMPLEMENTATION AGAINST THE VENDOR'S OWN PRESCRIPTION, recorded so the "not implemented properly" defence is checkable: SLC's completion criterion is 20 days minimum (30 preferred) with at least 80 percent completion on a majority of exercises and steady progress on both sound and word exercises. Among students who trained at all, 76 percent (flight 1) and 67 percent (flight 2) reached 20 days - but only 51 percent and 38 percent respectively met BOTH the training-days and the exercise-completion criteria. Only 9 students in the whole study reached the 90-percent-of-all-exercises standard. Available training days averaged 37 in flight 1 and 30 in flight 2. The authors explicitly treat this difficulty as partly inherent to running FFW in an urban school, noting Borman & Rachuba hit the same wall. CONTAMINATION AND ATTRITION were both minimal: 4 control students trained on FFW, 8 treatment students transferred out or never appeared, and follow-up coverage was 87-100 percent depending on the measure - unusually clean for this literature. Baseline balance holds on every pre-treatment measure except the state writing score (p = 0.037), which the authors attribute to chance. GRADE B: a properly randomised, low-attrition, independently run trial with multiple outcome measures including two independent standardized ones - the "single well-powered RCT" row - but n = 512 is not large, and the authors are candid that they cannot rule out small effects (the CELF-3-RP 95 percent CI spans -0.24 to +0.31 SD). It is the replication by Borman & Rachuba that turns "no detectable effect" into a durable null. COST: a year-long site licence about $30,000; with computers, printers and hardware roughly $37,000 per school per year for 20 stations; plus a supervising adult at about $55,000 - roughly $770 per student if one adult handles 40 students a day over three rounds, before counting the indirect cost of rescheduling 90-100 minutes a day of everything else. VERSION READ: numbers checked against the full text of NBER Working Paper 10315 (February 2004), the openly available version of the article published in Economics of Education Review 23(4), 323-338, August 2004; abstract and conclusions are identical, but the typeset journal version was not separately re-checked.
Key findings
The archive's cleanest single demonstration of the aligned-versus-independent measure gap in ed-tech. On READING EDGE - the assessment sold by Fast ForWord's own vendor and targeting exactly the phonological skills the programme trains - the intent-to-treat effect is about 0.1 SD and only just significant at the 10 percent level, with the treatment-on-the-treated gains concentrated in non-word recognition and phoneme blending. On EVERY INDEPENDENT MEASURE the effect vanishes: no detectable effect on the CELF-3-RP clinical language test (95% CI -0.24 to +0.31 SD), an effect size of about 0.05 SD on the Success for All reading assessment, and 0.04-0.06 SD with t-ratios under 1.0 on the state reading test (95% CI -0.08 to +0.16 SD). The authors' own summary: the programme "may improve some aspects of students' language skills" but "it does not appear that these gains translate into a broader measure of language acquisition or into actual reading skills." The design also demonstrates why uncontrolled evaluations of this literature are worthless: treated students posted large raw pre-post gains on all four tests (0.7 SD on Reading Edge, 0.3 SD on CELF-3-RP, ~0.2 SD on the state reading test) - and the control students gained almost exactly as much, leaving a difference of essentially zero. Rouse & Krueger conclude that the disappointing computer-in-school literature is not merely an artefact of vague interventions or weak designs; here the use of computers was well defined, the product was the flagship "scientifically based" one, the design was randomised, and the reading effect was still zero.
Genetic confound
Minimal (randomized within grade-by-school blocks, with balance on every pre-treatment measure except state writing, and 87-100 percent follow-up).
Replication notes
Independently replicated. Borman & Rachuba (2001) - the only other randomised evaluation of Fast ForWord conducted independently of Scientific Learning Corporation - ran 415 children in grades 2 and 7 across 8 Baltimore City Public Schools who scored below the 50th percentile on the CTBS/5, training for up to 8 weeks. Their 95 percent CI for the intent-to-treat effect on reading runs from -0.08 to +0.13 SD, which Rouse & Krueger call "remarkably close" to their own -0.08 to +0.16 on the state reading assessment. Two independent randomised trials therefore rule out large reading effects, and both encountered the same problem of students failing to complete the programme as the vendor prescribes. The vendor's own published claims are far larger; Rouse & Krueger attribute that gap to the vendor's studies lacking control groups and reporting only students who successfully progressed through the programme.
DOI / URL
10.1016/j.econedurev.2003.10.005
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| ALIGNED MEASURE - Reading Edge composite (computerised assessment sold by Scientific Learning, the vendor of Fast ForWord), intent to treat | SD | about +0.1 SD, "just barely statistically significant at the 10 percent level" once the pre-test is controlled; smaller and insignificant with randomisation-pool controls alone. The authors note the in-sample SD of about 30 used to scale this "is undoubtedly an underestimate of the standard deviation in the population", i.e. the true effect size is probably smaller still. | mixed | immediately post-training | business-as-usual | end-of-treatment | domain-skill |
| ALIGNED MEASURE - Reading Edge, treatment on the treated (IV on three SLC completion definitions) | SD | consistently positive across all three definitions of "treated", significant at the 10 percent level; gains concentrated in the NON-WORD RECOGNITION and, less so, PHONEME BLENDING subtests - i.e. precisely the trained tasks, which is near-transfer at best | mixed | immediately post-training | business-as-usual | end-of-treatment | near-transfer |
| INDEPENDENT MEASURE - CELF-3-RP, receptive portion of the Clinical Evaluation of Language Fundamentals, administered one-to-one by three independent certified evaluators | SD | no detectable effect, small and insignificant with or without covariates. 95 percent CI on the intent-to-treat effect size runs from -0.24 to +0.31 SD. Treatment-on-the-treated is also undetectable, though the authors note a 0.17 SD point estimate for those who trained 30+ days AND completed 80 percent of exercises, which they flag as upward biased. | standardized | immediately post-training | business-as-usual | end-of-treatment | domain-skill |
| INDEPENDENT MEASURE - state standardized reading assessment (the actual reading outcome) | SD | effect size about 0.04-0.06 SD, t-ratios under 1.0, not close to significant. 95 percent CI -0.08 to +0.16 SD - the tightest interval in the study because of the larger sample. No effect for the treated-on-the-treated either. | standardized | 2002-03 state testing, after training | business-as-usual | end-of-treatment | domain-skill |
| INDEPENDENT MEASURE - Success for All district reading assessment | SD | intent-to-treat effect size about 0.05 SD, not significant (0.02 without randomisation-pool controls, 0.07 with, falling again once covariates are added) | administrative | immediately post-training | business-as-usual | end-of-treatment | domain-skill |
| WHY UNCONTROLLED EVALUATIONS OF THIS PRODUCT LOOK GOOD - raw pre-post gains, treatment vs control | SD (raw gain) | treated students gained 21 points on Reading Edge (0.7 SD), 6.3 NCE on CELF-3-RP (0.3 SD), 0.27 on the SFA assessment (0.18 SD) and 5.7 percentile points on the state reading test (~0.2 SD) - all "relatively large among educational interventions". Control students gained 17.7, ~6.0, 0.25 and 4.4 over the same period. The difference is essentially zero on three of the four. This is the entire pre-post illusion in one table. | mixed | pre to post within the school year | none | end-of-treatment | domain-skill |
| DISPLACEMENT - what the 90-100 minutes a day was taken from | subjects missed | Fast ForWord ran as a pull-out. Treated students missed, depending on school and grade, homeroom, mathematics, writing, language arts, spelling, science, social studies, literature, grammar, "specials" (art, music, gym), some lunch and recess, and in one school before/after-school time; control students spent that time in those activities. Students were never pulled out of Success for All, the core reading programme. So the intervention protected reading time and consumed roughly an hour and a half a day of everything else, for no measurable reading gain. | administrative | across the training period | business-as-usual | end-of-treatment | behaviour |
| IMPLEMENTATION - completion against the vendor's own prescribed protocol | percent of trainees meeting SLC completion criteria | among students who trained at all, 76 percent (flight 1) and 67 percent (flight 2) reached the 20-day minimum; but only 51 percent and 38 percent met BOTH the day requirement and the 80-percent- of-exercises requirement. Only 9 students in the entire study met the 90-percent-of-all-exercises standard. Available training days averaged 37 then 30. The authors judge this difficulty partly inherent to running the programme in an urban school, and Borman & Rachuba hit the same problem. | administrative | across the training period | none | end-of-treatment | behaviour |