Does strategy instruction (SRSD, the writing process) improve writing?
Strategy instruction is writing's best-supported method — at roughly a fifth of its advertised size once measures are independent (d≈0.8 → ~0.15).
moderate supportconf: mediumgc: lowwriting · ages 5–18
Strategy instruction is the best-supported thing in this subject and its advertised effect size is roughly five times too large. On researcher-scored writing quality the meta-analyses report d = 0.82 for strategy instruction in Grades 4-12 (1.14 for SRSD specifically), 1.02 in the elementary grades, 0.96 in Grades 4-6, and 0.59-1.04 in K-3. On measures INDEPENDENT of the developers, researchers and teachers, the same literature yields +0.18 overall and +0.17 for writing-process programmes across fourteen qualifying studies. The scale-up record makes the same point: an EEF efficacy trial of SRSD returned ~+0.74 (about nine months' progress); the independently run effectiveness trial of the same programme returned +0.11 over two years and -0.09 over one, neither significant, WITH significantly less progress in reading, spelling and maths. The direction is real and consistent; the magnitude on anything an outsider measures is small.
Teach children an explicit, named, memorised routine for planning, drafting and revising a piece of writing, model it, and fade the support - this is the highest-value thing in the writing curriculum and it is one of the few writing practices that survives an independent-measure standard at all. But budget +0.15 to +0.20, not +1.00, and treat the difference as the price of a subjective quality rating. Two conditions matter more than the programme you buy: the person teaching it must be properly trained (the effect halved and then vanished when trainers were not the developers), and it must not displace reading, spelling and maths, which is exactly what happened in the only two-year trial.
Who this applies to
Verdict
moderate-support, and the interesting part of this topic is the gap between two numbers that describe the
same intervention.
Read the meta-analyses and writing strategy instruction is the most powerful lever in education: d = 0.82 for strategy instruction in Grades 4–12, d = 1.14 for Self-Regulated Strategy Development specifically, d = 1.02 in the elementary grades, d = 0.96 in Grades 4–6. Numbers in that range would mean a term of SRSD is worth several years of ordinary schooling. Nothing in this archive moves an independently measured outcome that far.
Read the same literature filtered through independent measurement and the number is +0.18. Slavin et al. applied the Best Evidence Encyclopedia standards to writing programmes in Grades 2–12 — randomised or well-matched, adequate sample and duration, and measures independent of the developers, researchers and teachers — and fourteen studies of twelve programmes survived. Writing-process programmes: +0.17.
The EEF ran the natural experiment on this gap and published both halves. The efficacy trial of SRSD delivered as IPEELL, in 23 Calderdale schools with the North American developers doing the training and a blind-marked bespoke writing test, produced about +0.74 — roughly nine months' additional progress, one of the largest effects the EEF has ever reported. The effectiveness trial — 84 and 83 schools, newly recruited trainers, delivered to whole year groups across the prior-attainment range — produced +0.11 (CI −0.13 to 0.34) over two years and −0.09 (CI −0.34 to 0.15) over one, neither statistically significant, and pupils who did two years of it made significantly less progress in reading, spelling and maths. IPEELL was removed from the EEF's promising-projects list.
So the verdict is not that strategy instruction fails. Every synthesis points the same way, including the independent-measure one, and unlike grammar there is no negative estimate anywhere. The verdict is that this is a real, small, condition-dependent effect wearing a very large number's clothes.
What the evidence shows
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Slavin 2019 (BEE) | synthesis, independent measures required, 14 studies / 12 programmes, Grades 2–12 | B | All writing programmes +0.18; writing-process +0.17; cooperative learning +0.16; reading-writing integration +0.19 |
| Torgerson 2018 (EEF IPEELL, effectiveness) | cluster RCT, 84 + 83 schools, ~5,400 children | B | 2 years +0.11 (−0.13, 0.34); 1 year −0.09 (−0.34, 0.15); both ns. Significantly less progress in reading, spelling and maths |
| Torgerson 2014 (EEF Improving Writing Quality, efficacy) | cluster RCT, 23 schools, Year 6, developer-trained, blind-marked bespoke test | B | ~+0.74, ≈ 9 months' progress |
| Graham & Perin 2007 | meta-analysis, Grades 4–12 | C | Strategy 0.82 (0.69, 0.95), k = 20; SRSD 1.14, k = 8; non-SRSD 0.62. Process approach 0.32 overall — but 0.46 with PD and 0.03 without |
| Graham 2012 (elementary) | meta-analysis, 115 documents | C | Strategy 1.02; adding self-regulation +0.50; text structure 0.59; product goals 0.76; peer assistance 0.89 |
| Koster 2015 | meta-analysis, Grades 4–6 | C | Strategy 0.96; feedback 0.88; text structure 0.76; goal setting 2.03 |
| Kim 2021 | meta-analysis, K–3, 24 studies, 166 ES, N = 5,589 | C | Overall 0.31; SRSD 0.59–1.04; larger effect for initially weak writers; dosage explained no variance |
Koster's +2.03 for goal setting is the most useful number in the table and it is not evidence of anything. A two-standard-deviation effect from telling children what to aim for is not a causal parameter; it is a measurement artefact of treatment-aligned scoring in small studies, and METHODOLOGY's rule that d > 0.60 from small studies with researcher-designed measures is a red flag rather than a triumph applies to it exactly. Its presence at the top of a respectable meta-analysis is the clearest available warning about the whole table.
The process-approach decomposition is the most practically important finding in the meta-analytic half. Graham & Perin report the process writing approach at 0.32 overall — but 0.46 where teachers received professional development and 0.03 (CI −0.07 to 0.13) where they did not, falling to −0.05 in Grades 7–12. "Do writing process" as a slogan is worth nothing. The training is the intervention.
The EEF effectiveness trial's own explanation for the collapse is the actionable content. Its evaluators name three differences from the efficacy trial: trainers were newly recruited rather than the developers; delivery was in Years 5–6 rather than across the Year 6–7 transition; and it was used with the full range of prior attainment rather than with low attainers only. They also flag that "the training of teachers was lacking in both the provision and breadth of practical examples and in fully modelling some aspects of the approach." That is a mechanism, not an excuse — and it is the same mechanism Graham & Perin's PD split identifies.
The spillover finding deserves more attention than it gets. Two years of IPEELL produced significantly less progress in reading, spelling and maths. Writing time is not free; it comes out of something. Any school adopting a writing programme should count that cost explicitly, and almost no evaluation measures it.
Hereditarian-lens assessment
Risk: low. The verdict rests on two large independently evaluated cluster-randomised trials and on a synthesis restricted to randomised or well-matched designs with independent outcome measures. Children in IPEELL schools and comparison schools have the same expected distribution of ability-relevant alleles.
Two places where the premise still does work:
- "Larger effects for initially weak writers" is partly regression to the mean. Kim et al. report a larger effect on writing quality for children with initially weak writing skills, and the EEF two-year trial found three months' additional progress for low prior attainers. Both are subgroup contrasts selected on a baseline score. Given that writing ability is substantially heritable with negligible shared-environment influence (Olson 2013), a child scoring low at pre-test is partly there for stable reasons and partly there by measurement noise; the second part regresses upward regardless of treatment. The effect is probably real and probably smaller than reported.
- Strategy instruction is exactly the kind of intervention that should increase measured heritability if it works. It removes an environmental bottleneck — nobody invents a planning routine spontaneously — so once every child has the routine, the remaining variance in how well they use it is more genetic than before. Per METHODOLOGY that is evidence the instruction worked, not that it failed, and it is the right frame for a founder: this is floor-raising, not gap-closing.
Nothing in this topic bears on g. The outcome throughout is domain skill in composition.
Boundaries & what critics say
- The measure-type gap here is larger than METHODOLOGY's nominal 2×. Researcher-scored quality gives 0.82–1.14; independent measurement gives 0.17–0.18. That is roughly 5×. Writing quality is a holistic human judgement of overall merit, usually made by people who know what was taught — the most inflatable outcome class in education. Any writing effect size should be read with the scorer's identity in mind.
- The steelman for the large numbers is that independent measures are the wrong instrument. Standardised writing tests are short, generic and constrained; a child who has learned to plan and revise a persuasive essay may genuinely show that on an essay and not on a test. This is a real argument. It is also unfalsifiable as usually stated, and the EEF effectiveness trial's bespoke two-year measure — built specifically to be sensitive — still returned +0.11.
- Only fourteen studies in the entire Grades 2–12 range meet an independent-measure standard. For a core school subject that is a startlingly thin base, and it is a better description of the field's problem than any individual effect size.
- Single-case SRSD research is excluded and should be. Asaro-Saddler et al. (2021)
reports large SRSD effects in a multilevel meta-analysis of single-case designs with students with ASD; it is
recorded as
design-inadequatebecause those designs have no control groups and their effect metrics are not commensurable with between-group standardised differences. - The replication rule was considered and does not bite. METHODOLOGY caps a "collapsed-at-scale" finding at
mixed. IPEELL did not collapse to zero: the two-year estimate stayed positive (+0.11), the independent-measure synthesis is positive (+0.18), and every meta-analysis agrees on direction. It shrank by about 85%. That is shrinkage, whichmoderate-supportdescribes, not collapse. - Dose is genuinely unknown. Kim et al. tested total instruction length as a moderator across 24 primary studies and it explained nothing. Anyone quoting a required number of SRSD lessons is quoting a programme manual, not a finding.
Practical guidance
- Teach an explicit, named, memorised routine for planning and revising, and model it before expecting it. The SRSD staging — develop background knowledge, discuss it, model it, memorise it, support it, then independent performance — is the best-specified version and it is what the positive evidence is about.
- Budget +0.15 to +0.20, not +1.00. If a programme's marketing quotes an effect above 0.5, check who scored the writing.
- Spend the money on training, not on materials. The process approach is worth 0.46 with professional development and 0.03 without; SRSD is worth 0.74 with the developers training and 0.11 with newly recruited trainers. This is the most consistent moderator in the topic.
- Count what it displaces. Two years of IPEELL cost reading, spelling and maths progress. If writing time is coming out of the reading block, that is a trade the reading evidence says you will probably lose.
- Do not run it as a substitute for transcription work with young children. In K–3 the overall effect of writing instruction is 0.31, and children who cannot form letters fluently cannot run a planning routine while composing — see handwriting & transcription.
- Ignore goal setting at d = 2.03. Setting product goals is cheap and probably worth doing (Graham & Perin put it at a more believable 0.70), but that number is not a finding.
Open questions
- Confidence is capped at
mediumpending the adversarial pass, and also by the honest tension between two grade-B sources: the independent-measure synthesis and the effectiveness trial agree on a small effect, but no independently measured trial has isolated SRSD specifically. - Why did the one-year IPEELL trial point negative? −0.09 on the KS2 categorical writing outcome versus +0.11 on the two-year bespoke continuous outcome. The evaluators note the measures differed and say the difference in direction is unexplained. That is unresolved and it matters.
- Nobody has run SRSD with an independent standardised writing outcome and a proper comparison group at scale. The obvious trial in this subject does not exist.
- Does the effect persist? Every estimate here is end-of-treatment or end-of-year. Given the archive's general finding that early gains fade, the absence of any follow-up measurement in the writing literature is a serious gap.
- What is the real dose-response? Currently unknown, and the one test of it found nothing.
- grade CA meta-analysis of writing instruction for adolescent students.Graham S, Perin D · 2007 · meta-analysis
- grade CA Meta-Analysis of Writing Instruction for Students in the Elementary Grades.Graham S, McKeown D, Kiuhara S, Harris KR · 2012 · meta-analysis
- grade CTeaching children to write: A meta-analysis of writing intervention researchKoster M, Tribushinina E, de Jong PF, van den Bergh H · 2015 · meta-analysis
- grade CWriting instruction improves students’ writing skills differentially depending on focal instruction and children: A meta-analysis for primary grade studentsKim YSG, Yang D, Reyes M, Connor C · 2021 · meta-analysis
- grade BA Quantitative Synthesis of Research on Writing Approaches in Grades 2 to 12. Best Evidence Encyclopedia (BEE)Slavin RE, Lake C, Inns A, Baye A, Dachet D, Haslam J · 2019 · meta-analysis
- grade BImproving Writing Quality: Evaluation Report and Executive SummaryTorgerson D, Torgerson C, Ainsworth H, Buckley H, Heaps C, Hewitt C, Mitchell N · 2014 · rct
- grade BCalderdale Excellence Partnership: IPEELL. Evaluation report and executive summaryTorgerson CJ, Ainsworth H, Bell K, Elliott L, Fountain I, Gascoine L, Hewitt CE, Kasim A, Kokotsaki D, Torgerson DJ · 2018 · rct
- grade BGenetic and environmental influences on writing and their relations to language and readingOlson RK, Hulslander J, Christopher M, Keenan JM, Wadsworth SJ, Willcutt EG, Pennington BF, DeFries JC · 2013 · twin-adoption
Related decisions
- Does handwriting and transcription instruction matter, and should young children type instead?moderate supportconf: mediumgc: medium
- Does teaching grammar improve children's writing?negativeconf: mediumgc: low
- Does teaching spelling work, and what does it transfer to?moderate supportconf: mediumgc: low