The Evidence on Teaching

Does strategy instruction (SRSD, the writing process) improve writing?

Strategy instruction is writing's best-supported method — at roughly a fifth of its advertised size once measures are independent (d≈0.8 → ~0.15).

moderate supportconf: mediumgc: low

writing · ages 518

Effect summary

Strategy instruction is the best-supported thing in this subject and its advertised effect size is roughly five times too large. On researcher-scored writing quality the meta-analyses report d = 0.82 for strategy instruction in Grades 4-12 (1.14 for SRSD specifically), 1.02 in the elementary grades, 0.96 in Grades 4-6, and 0.59-1.04 in K-3. On measures INDEPENDENT of the developers, researchers and teachers, the same literature yields +0.18 overall and +0.17 for writing-process programmes across fourteen qualifying studies. The scale-up record makes the same point: an EEF efficacy trial of SRSD returned ~+0.74 (about nine months' progress); the independently run effectiveness trial of the same programme returned +0.11 over two years and -0.09 over one, neither significant, WITH significantly less progress in reading, spelling and maths. The direction is real and consistent; the magnitude on anything an outsider measures is small.

Practical takeaway

Teach children an explicit, named, memorised routine for planning, drafting and revising a piece of writing, model it, and fade the support - this is the highest-value thing in the writing curriculum and it is one of the few writing practices that survives an independent-measure standard at all. But budget +0.15 to +0.20, not +1.00, and treat the difference as the price of a subjective quality rating. Two conditions matter more than the programme you buy: the person teaching it must be properly trained (the effect halved and then vanished when trainers were not the developers), and it must not displace reading, spelling and maths, which is exactly what happened in the only two-year trial.

Who this applies to

Group size
whole-classsmall-groupone-to-one
Delivered by
teachertutorparent
Ages studied
518
Dose
SRSD is a staged routine (develop background knowledge, discuss it, model it, memorise the mnemonic, support it, then independent performance) taught over roughly 8-12 lessons per genre and then maintained. The EEF trials delivered it across a school year to a whole year group. Dose is the weakest-evidenced parameter here: the only meta-analysis to test total instruction time as a moderator (Kim 2021, K-3, 24 studies) found it explained NONE of the variation in effect sizes. Teacher training quality mattered far more than dose in both EEF trials.
Cost
medium
Moves
domain-skill
Needs first
Children must be able to transcribe well enough to get words on the page without the act of writing consuming all their attention - see [handwriting & transcription](handwriting-and-transcription.md). The self-regulation component (goal setting, self-monitoring, self-instruction) assumes enough working memory and metacognitive capacity to run a remembered routine while composing.
Not for
Do not expect the meta-analytic numbers if you measure with anything you did not design. Do not expect it to survive weak training: in the EEF effectiveness trial, training was delivered by newly recruited trainers rather than the programme developers and the effect collapsed from ~0.74 to 0.11. And do not let it eat the timetable - the trial that ran it for two years found pupils made significantly LESS progress in reading, spelling and maths, which the evaluators attribute to class time being redirected away from those subjects.

Verdict

moderate-support, and the interesting part of this topic is the gap between two numbers that describe the same intervention.

Read the meta-analyses and writing strategy instruction is the most powerful lever in education: d = 0.82 for strategy instruction in Grades 4–12, d = 1.14 for Self-Regulated Strategy Development specifically, d = 1.02 in the elementary grades, d = 0.96 in Grades 4–6. Numbers in that range would mean a term of SRSD is worth several years of ordinary schooling. Nothing in this archive moves an independently measured outcome that far.

Read the same literature filtered through independent measurement and the number is +0.18. Slavin et al. applied the Best Evidence Encyclopedia standards to writing programmes in Grades 2–12 — randomised or well-matched, adequate sample and duration, and measures independent of the developers, researchers and teachers — and fourteen studies of twelve programmes survived. Writing-process programmes: +0.17.

The EEF ran the natural experiment on this gap and published both halves. The efficacy trial of SRSD delivered as IPEELL, in 23 Calderdale schools with the North American developers doing the training and a blind-marked bespoke writing test, produced about +0.74 — roughly nine months' additional progress, one of the largest effects the EEF has ever reported. The effectiveness trial — 84 and 83 schools, newly recruited trainers, delivered to whole year groups across the prior-attainment range — produced +0.11 (CI −0.13 to 0.34) over two years and −0.09 (CI −0.34 to 0.15) over one, neither statistically significant, and pupils who did two years of it made significantly less progress in reading, spelling and maths. IPEELL was removed from the EEF's promising-projects list.

So the verdict is not that strategy instruction fails. Every synthesis points the same way, including the independent-measure one, and unlike grammar there is no negative estimate anywhere. The verdict is that this is a real, small, condition-dependent effect wearing a very large number's clothes.

What the evidence shows

Source Design Grade Key effect
Slavin 2019 (BEE) synthesis, independent measures required, 14 studies / 12 programmes, Grades 2–12 B All writing programmes +0.18; writing-process +0.17; cooperative learning +0.16; reading-writing integration +0.19
Torgerson 2018 (EEF IPEELL, effectiveness) cluster RCT, 84 + 83 schools, ~5,400 children B 2 years +0.11 (−0.13, 0.34); 1 year −0.09 (−0.34, 0.15); both ns. Significantly less progress in reading, spelling and maths
Torgerson 2014 (EEF Improving Writing Quality, efficacy) cluster RCT, 23 schools, Year 6, developer-trained, blind-marked bespoke test B ~+0.74, ≈ 9 months' progress
Graham & Perin 2007 meta-analysis, Grades 4–12 C Strategy 0.82 (0.69, 0.95), k = 20; SRSD 1.14, k = 8; non-SRSD 0.62. Process approach 0.32 overall — but 0.46 with PD and 0.03 without
Graham 2012 (elementary) meta-analysis, 115 documents C Strategy 1.02; adding self-regulation +0.50; text structure 0.59; product goals 0.76; peer assistance 0.89
Koster 2015 meta-analysis, Grades 4–6 C Strategy 0.96; feedback 0.88; text structure 0.76; goal setting 2.03
Kim 2021 meta-analysis, K–3, 24 studies, 166 ES, N = 5,589 C Overall 0.31; SRSD 0.59–1.04; larger effect for initially weak writers; dosage explained no variance

Koster's +2.03 for goal setting is the most useful number in the table and it is not evidence of anything. A two-standard-deviation effect from telling children what to aim for is not a causal parameter; it is a measurement artefact of treatment-aligned scoring in small studies, and METHODOLOGY's rule that d > 0.60 from small studies with researcher-designed measures is a red flag rather than a triumph applies to it exactly. Its presence at the top of a respectable meta-analysis is the clearest available warning about the whole table.

The process-approach decomposition is the most practically important finding in the meta-analytic half. Graham & Perin report the process writing approach at 0.32 overall — but 0.46 where teachers received professional development and 0.03 (CI −0.07 to 0.13) where they did not, falling to −0.05 in Grades 7–12. "Do writing process" as a slogan is worth nothing. The training is the intervention.

The EEF effectiveness trial's own explanation for the collapse is the actionable content. Its evaluators name three differences from the efficacy trial: trainers were newly recruited rather than the developers; delivery was in Years 5–6 rather than across the Year 6–7 transition; and it was used with the full range of prior attainment rather than with low attainers only. They also flag that "the training of teachers was lacking in both the provision and breadth of practical examples and in fully modelling some aspects of the approach." That is a mechanism, not an excuse — and it is the same mechanism Graham & Perin's PD split identifies.

The spillover finding deserves more attention than it gets. Two years of IPEELL produced significantly less progress in reading, spelling and maths. Writing time is not free; it comes out of something. Any school adopting a writing programme should count that cost explicitly, and almost no evaluation measures it.

Hereditarian-lens assessment

Risk: low. The verdict rests on two large independently evaluated cluster-randomised trials and on a synthesis restricted to randomised or well-matched designs with independent outcome measures. Children in IPEELL schools and comparison schools have the same expected distribution of ability-relevant alleles.

Two places where the premise still does work:

  • "Larger effects for initially weak writers" is partly regression to the mean. Kim et al. report a larger effect on writing quality for children with initially weak writing skills, and the EEF two-year trial found three months' additional progress for low prior attainers. Both are subgroup contrasts selected on a baseline score. Given that writing ability is substantially heritable with negligible shared-environment influence (Olson 2013), a child scoring low at pre-test is partly there for stable reasons and partly there by measurement noise; the second part regresses upward regardless of treatment. The effect is probably real and probably smaller than reported.
  • Strategy instruction is exactly the kind of intervention that should increase measured heritability if it works. It removes an environmental bottleneck — nobody invents a planning routine spontaneously — so once every child has the routine, the remaining variance in how well they use it is more genetic than before. Per METHODOLOGY that is evidence the instruction worked, not that it failed, and it is the right frame for a founder: this is floor-raising, not gap-closing.

Nothing in this topic bears on g. The outcome throughout is domain skill in composition.

Boundaries & what critics say

  • The measure-type gap here is larger than METHODOLOGY's nominal 2×. Researcher-scored quality gives 0.82–1.14; independent measurement gives 0.17–0.18. That is roughly 5×. Writing quality is a holistic human judgement of overall merit, usually made by people who know what was taught — the most inflatable outcome class in education. Any writing effect size should be read with the scorer's identity in mind.
  • The steelman for the large numbers is that independent measures are the wrong instrument. Standardised writing tests are short, generic and constrained; a child who has learned to plan and revise a persuasive essay may genuinely show that on an essay and not on a test. This is a real argument. It is also unfalsifiable as usually stated, and the EEF effectiveness trial's bespoke two-year measure — built specifically to be sensitive — still returned +0.11.
  • Only fourteen studies in the entire Grades 2–12 range meet an independent-measure standard. For a core school subject that is a startlingly thin base, and it is a better description of the field's problem than any individual effect size.
  • Single-case SRSD research is excluded and should be. Asaro-Saddler et al. (2021) reports large SRSD effects in a multilevel meta-analysis of single-case designs with students with ASD; it is recorded as design-inadequate because those designs have no control groups and their effect metrics are not commensurable with between-group standardised differences.
  • The replication rule was considered and does not bite. METHODOLOGY caps a "collapsed-at-scale" finding at mixed. IPEELL did not collapse to zero: the two-year estimate stayed positive (+0.11), the independent-measure synthesis is positive (+0.18), and every meta-analysis agrees on direction. It shrank by about 85%. That is shrinkage, which moderate-support describes, not collapse.
  • Dose is genuinely unknown. Kim et al. tested total instruction length as a moderator across 24 primary studies and it explained nothing. Anyone quoting a required number of SRSD lessons is quoting a programme manual, not a finding.

Practical guidance

  • Teach an explicit, named, memorised routine for planning and revising, and model it before expecting it. The SRSD staging — develop background knowledge, discuss it, model it, memorise it, support it, then independent performance — is the best-specified version and it is what the positive evidence is about.
  • Budget +0.15 to +0.20, not +1.00. If a programme's marketing quotes an effect above 0.5, check who scored the writing.
  • Spend the money on training, not on materials. The process approach is worth 0.46 with professional development and 0.03 without; SRSD is worth 0.74 with the developers training and 0.11 with newly recruited trainers. This is the most consistent moderator in the topic.
  • Count what it displaces. Two years of IPEELL cost reading, spelling and maths progress. If writing time is coming out of the reading block, that is a trade the reading evidence says you will probably lose.
  • Do not run it as a substitute for transcription work with young children. In K–3 the overall effect of writing instruction is 0.31, and children who cannot form letters fluently cannot run a planning routine while composing — see handwriting & transcription.
  • Ignore goal setting at d = 2.03. Setting product goals is cheap and probably worth doing (Graham & Perin put it at a more believable 0.70), but that number is not a finding.

Open questions

  • Confidence is capped at medium pending the adversarial pass, and also by the honest tension between two grade-B sources: the independent-measure synthesis and the effectiveness trial agree on a small effect, but no independently measured trial has isolated SRSD specifically.
  • Why did the one-year IPEELL trial point negative? −0.09 on the KS2 categorical writing outcome versus +0.11 on the two-year bespoke continuous outcome. The evaluators note the measures differed and say the difference in direction is unexplained. That is unresolved and it matters.
  • Nobody has run SRSD with an independent standardised writing outcome and a proper comparison group at scale. The obvious trial in this subject does not exist.
  • Does the effect persist? Every estimate here is end-of-treatment or end-of-year. Given the archive's general finding that early gains fade, the absence of any follow-up measurement in the writing literature is a serious gap.
  • What is the real dose-response? Currently unknown, and the one test of it found nothing.

Evidence (8 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore

Does strategy instruction (SRSD, the writing process) improve writing? · The Evidence on Teaching