The Evidence on Teaching

The Direct Instruction Follow Through Model: Design and Outcomes

Engelmann S, Becker WC, Carnine D, Gersten R · 1988

grade Cquasi-experimentdeveloper-ledunreplicated
Sample
Nine major Follow Through sponsors compared across sites; tens of thousands of disadvantaged children across the national evaluation
Population
Disadvantaged US children from kindergarten through Grade 3, 1968-1977
Design
The developers' own account of the Direct Instruction arm of Project Follow Through, the largest educational experiment ever run in the United States. It is recorded here for one number that nothing else in the archive supplies: the Metropolitan Achievement Test LANGUAGE subtest, which the authors define in the text as "usage, punctuation, and sentence types" - i.e. exactly the conventions construct this topic is about, measured on a norm-referenced standardised instrument. The design caveats are severe and the authors state most of them themselves: Follow Through was PLANNED VARIATION, not randomisation; communities selected their sponsor; the comparison is between sponsors rather than against a clean control; and the authors concede that "Many points of the House et al. (1978) critique are valid, particularly those citing limitations of research designs where students are not randomly assigned to the experimental or control groups." The 0.75 SD figure is a gap read off a percentile figure, not an effect size with a confidence interval. And DISTAR Language I-III taught usage and sentence forms directly, so the outcome is treatment-aligned in content even though the instrument is standardised.
Key findings
Verbatim: "For Language (usage, punctuation, and sentence types), the Direct Instruction program is three-fourths of a standard deviation ahead of all other programs." Direct Instruction students finished "close to or at national norms on all measures", against a baseline expectation of roughly the 20th percentile for disadvantaged children without special help. On Spelling, "the Behavior Analysis program is the only program other than Direct Instruction approaching national norms". On Total Math, Direct Instruction is "at least one half of a standard deviation ahead of all the others". Direct Instruction was the only one of nine models showing consistently positive outcomes across measures; the more open-ended and child-centred models did worst. The Language result is the archive's single strongest datapoint that conventions ARE movable by explicit teaching at scale - and it comes from a non-randomised planned-variation study reported by the programme's own authors, so it is a strong hint rather than a demonstration.
Genetic confound
Medium to high. Communities were not randomly assigned to sponsors, and the disadvantaged populations served by different models were not equivalent; the Abt evaluation's own covariate adjustments are the mitigation and they are the substance of the House et al. critique. What the premise cannot easily explain away is the pattern: the same non-random assignment machinery produced near-national-norm Language scores for one model and near-20th-percentile scores for others serving similar populations, and the winning model is the one that taught the tested content most explicitly.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
MAT Language subtest ("usage, punctuation, and sentence types"), end of Grade 3gap to the next-best Follow Through modelDirect Instruction three-fourths of a standard deviation ahead of all other programs; close to national normsstandardizedend of Grade 3, for children entering at kindergartenactive-alternativeend-of-treatmentdomain-skill
MAT Spelling subtestnormative positiononly Direct Instruction and Behavior Analysis approached national normsstandardizedend of Grade 3active-alternativeend-of-treatmentdomain-skill
MAT Total Mathgap to the next-best modelDirect Instruction at least half a standard deviation aheadstandardizedend of Grade 3active-alternativeend-of-treatmentdomain-skill

Cited by