The Evidence on Teaching

Can capitalisation, punctuation and usage be taught directly as their own subject?

Capitalisation and punctuation move when directly taught and directly measured, and don't move when embedded in writing programmes. A teachable subject, not a writing lever.

mixedconf: mediumgc: medium

writing · ages 518

Effect summary

This is NOT the grammar question, and the archive answers it differently. [Grammar as a route to better writing](grammar-instruction.md) is graded `negative`; conventions as a body of knowledge in their own right behave in a way that is conditional but coherent: THEY MOVE WHEN THEY ARE THE DIRECT TARGET AND THE DIRECT MEASURE, AND THEY DO NOT MOVE WHEN THEY ARE EMBEDDED IN A WRITING PROGRAMME OR MEASURED INSIDE EXTENDED COMPOSITION. On the positive side: Direct Instruction in Project Follow Through finished three-quarters of a standard deviation ahead of all eight other models on the Metropolitan Achievement Test LANGUAGE subtest, which the publisher defines as punctuation, capitalisation and usage; a 60-minute RCT with 59 eight-to-ten-year-olds moved Standard English usage into their own free writing, but only in the arm that added guided practice and correction to the rule; nine weeks of explicit teaching plus dictation took full-stop omission in Year 2 from 94% to 28% of pupils against a flat comparison class; and the one randomised trial anywhere targeting punctuation itself produced significant gains that generalised to self-constructed sentences. On the null side, and it is the better-identified side: a 39-school RCT scoring CONVENTIONS as its own rubric dimension returned 0.02; three EEF cluster trials of two different programmes moved the English national grammar-punctuation-spelling test DOWN, not up (-0.06, -0.22, -0.28); five weeks of daily traditional grammar left errors per sentence at 1.3 before and 1.2 after; and automated grammar feedback gives g = 0.31 at post-test and -0.02 at maintenance. NO ADEQUATELY POWERED, INDEPENDENTLY EVALUATED RANDOMISED TRIAL OF DIRECT CONVENTIONS INSTRUCTION AGAINST A CONTROL, MEASURED ON A STANDARDISED CONVENTIONS TEST, EXISTS AT ANY AGE. The literature reaches down to about seven; below that there is nothing.

Practical takeaway

Teach conventions directly, narrowly, and briefly - and be honest that you are teaching a separate subject, not improving writing. Pick one mark or rule, state it, show it, have the child use it immediately, correct it, and check the next thing they write. That specific shape - rule PLUS guided practice PLUS correction - is the only one with a randomised result behind it, it took an hour in the trial that produced it, and rules without practice did nothing in the same experiment. Spend the time on the short list that actually accounts for the errors: full stops, capital letters, and the two or three punctuation marks the child is currently getting wrong. Do not buy a programme, do not run a daily grammar-worksheet routine, and do not expect the payoff to appear in the quality of what they write. Expect it on conventions, and expect a gap between what they can do on a worksheet and what they do while composing - that gap is the finding, not an implementation failure, and the only way anyone has closed it is by teaching the transfer explicitly as its own step.

Who this applies to

Group size
one-to-onesmall-groupwhole-classindependent
Delivered by
teachertutorparent
Ages studied
718(narrower than the 518 this topic is filed under — outside it is extrapolation)
Dose
Everything that worked was small and targeted. Fogel & Ehri moved usage into children's free writing with TWO SESSIONS TOTALLING 60 MINUTES on six specific features - but only in the arm that added guided transformation practice with correction; rules plus exposure alone did nothing. Robinson-Kooi & Hammond ran nine weeks of explicit teaching with short daily sentence dictation at Year 2. Schumaker et al. used six lessons covering eight punctuation marks, about 11 hours. Campbell et al. taught capitalisation in 15-20 minute peer-tutored sessions over 17-28 days. Set against the doses that produced NOTHING or worse: six weeks of daily hour-long whole-class grammar lessons (KS2 GPS -0.06); twenty 15-minute grammar sessions over ten weeks (national grammar test p = 0.91); a year of a writing curriculum (conventions rubric 0.02); two years of a writing programme (KS2 GPS -0.28). More time is not the variable. Narrow target, explicit rule, immediate practice, correction, and then use in the child's own writing is the variable.
Cost
free
Moves
domain-skill
Needs first
The child must be able to transcribe a sentence without the act of writing consuming all their attention - letter formation and enough spelling that the pencil is not the bottleneck (see [handwriting & transcription](handwriting-and-transcription.md) and [spelling](spelling.md)). They must be able to produce a conventional simple sentence orally and in writing; the ability to combine sentences using conventional grammar arrives around age seven for most children.
Not for
Do not use it to improve writing - that is a different claim and the archive grades it `negative`; see [grammar instruction](grammar-instruction.md). Do not expect conventions knowledge to appear in extended writing on its own: every well-identified estimate of that transfer is approximately zero, and the same children make significantly more surface errors when the content load rises. There is NO evidence below about age seven - no controlled study of punctuation or capitalisation instruction has ever measured a child younger than that, and the one robust trial of grammar teaching at ages 6-7 did not measure conventions at all. Do not buy a grammar or mechanics programme: no packaged product has ever beaten a control on a standardised conventions test. Do not rely on grammar-checking or automated-feedback software for durable gains - surface-level effects run g = 0.31 while the tool is in use and -0.02 once it is withdrawn. And do not assume a child who omits full stops does not know the rule: given a reward and no teaching whatsoever, eight-and-nine-year-olds went from 0% and 13% correct to roughly 88% within days.

Verdict

This topic exists because the archive was about to refuse a question it can actually answer.

Grammar instruction is graded negative, and correctly: three independent randomised trials and four meta-analyses agree that teaching grammar does not improve writing. But that file itself separates three senses of the word, and its third sense is this one — "GRAMMAR AS ITS OWN SUBJECT (naming word classes, passing a GPS test) is teachable, but nothing here shows it moves writing." A parent who says "I want better grammar" often means the third thing: capital letters, full stops, apostrophes, subject–verb agreement. That is the CAT Language Mechanics domain, and the What Works Clearinghouse gives it a name — "Mechanics refers to assessments of handwriting, spelling, capitalization, and punctuation. The term usage also may be applied and typically refers to the combination of capitalization and punctuation."

Answering "grammar doesn't work" to that question is a category error. So: can conventions be taught directly, to what effect, at what dose, at what age — measured on conventions themselves?

The verdict is mixed, and the condition under which the effect appears is the whole finding. Conventions move when they are the direct target of instruction and the direct object of measurement. They do not move when they arrive inside a writing programme, and they do not reliably show up inside extended composition. Both halves are evidenced, and — awkwardly for anyone wanting a clean answer — the null half has the better designs behind it.

The single most important structural fact is an absence. No adequately powered, independently evaluated randomised trial has ever tested direct conventions instruction against a control and measured a standardised conventions test, at any age. The instruments exist and have existed for sixty years — the CAT Language Mechanics subtest, the ITBS Language subtest, the Metropolitan Achievement Test Language subtest, England's KS2 GPS test. The field measures writing quality instead. Not one of the writing meta-analyses this archive holds — Graham & Perin 2007, Graham et al. 2012, Koster et al. 2015, Kim et al. 2021, Andrews et al. 2004 — pools a conventions or mechanics outcome. The question is unanswered not because trials returned nulls but because almost nobody asked it.

What the evidence shows

Source Design Grade Key effect
Puma 2007 (Writing Wings) cluster RCT, 39 schools / 21 states / ~3,000 pupils, Grades 3–5 B CONVENTIONS rubric dimension: difference 0.021, ES 0.02, HLM coefficient 0.006 (ns). The only randomised curriculum evaluation that scores conventions separately
Tracey 2019 (EEF Grammar for Writing) cluster RCT, 155 schools, Year 6 B KS2 GPS ES = −0.06 (95% CI −0.10, −0.01), n = 6,662 — CI excludes zero, in the wrong direction
Torgerson 2018 (EEF IPEELL) two cluster RCTs, 84 and 83 schools B KS2 GPS −0.22 (−0.46, 0.01) after one year; −0.28 (−0.49, −0.06) after two — a significant negative
EEF 2025 (Year 7 grammar, examples) three-arm cluster RCT, 55 schools, Year 7 B Grammar test built from KS2 GPS items: χ² = 0.18, p = 0.912, n = 6,642. And the validity result: KS2 GPS ↔ grammar test r = 0.79; KS2 GPS ↔ writing r = 0.56
Engelmann 1988 (Follow Through) planned variation, 9 sponsors, ~180 communities, K–3 C MAT Language ("usage, punctuation, and sentence types"): Direct Instruction three-quarters of an SD ahead of all other models, at national norms
Fogel & Ehri 2000 RCT, 59 pupils aged 8–10, 60 minutes total C Rule instruction + guided practice + correction moved taught usage into children's own free writing; rules alone did not
Schumaker 2019 RCT, 88 students with LD, Grades 6–12, ~11 h C The only RCT anywhere whose target is punctuation. Significant gains; generalised to self-constructed sentences
Robinson-Kooi & Hammond 2020 quasi-experiment, 2 schools, 60 pupils, Year 2, 9 weeks D Full-stop omission 94%→28% and 87%→47% vs comparison 76%→72%; capitals 89%→22% vs 84%→56%
WWC practice guide (Graham 2012) expert panel, WWC evidence standards C Rec 3 rated moderate; the conventions step rests on one 60-minute study contributing "no eligible measures". Roadblock 3.3 states the transfer problem outright
Fearn & Farnan 2007 quasi-experiment, 3 intact Grade-10 classes, 5 weeks D Five weeks of daily traditional grammar: errors per sentence 1.3 → 1.2. Grammar-test scores indistinguishable between arms
Graham & Santangelo 2014 meta-analysis, 53 studies, K–12 C Spelling 0.54 vs none, 0.70 for more-explicit vs less — and no transfer to extended writing
Kim 2021 meta-analysis, K–3, 24 studies, 166 ES C No conventions category exists. Transcription instruction: quality −0.19, "other" (which contains capitalisation) 0.24 — all ns
Scherer 2024 meta-analysis, 200 comparisons C Surface feedback → surface outcomes g = 0.58; surface feedback → FL deep outcomes g = −0.23. No primary-age studies
Scherer 2026 (algorithmic feedback) meta-analysis, 33 studies C Surface outcomes g = 0.31 at post-test, −0.02 at maintenance; deep outcomes 0.31 → 0.54
Van Beuningen 2012 4-condition classroom experiment, N = 268 secondary pupils C Correction beat self-editing and extra writing practice, in new writing at 1 and 4 weeks; non-grammatical accuracy gained most from indirect correction
Truscott & Hsu 2008 controlled study, adult L2 C Correction improved the revision; on a new narrative a week later the groups were "virtually identical"
Kang & Han 2015 / Truscott 2007 duelling meta-analyses, L2 C g = 0.54 vs "a small negative effect on learners' ability to write accurately." 93% of this corpus is adults
Schmidt 1988 (COPS) multiple-probe, n = 7, one school year D Mechanics errors .27 → .04 per word; generalised to classes where the strategy was never taught — for most pupils
Datchuk & Kubina 2013 systematic review, 19 studies C Grammar/usage: 3 studies, 51 participants, mean age 9–12. Handwriting: 10 studies, 394 participants, ages 6–10
Grünke 2019 single-case, n = 3, ages 8–9 D A reward contract with no teaching took punctuation accuracy from 0% and 13% to ~88% within days
Rogers & Graham 2008 meta of 88 single-case designs D "Teaching grammar and usage" and "strategy instruction for editing" among nine supported treatments — no control groups anywhere

Three EEF cluster trials moved the national conventions test in the wrong direction, and the reason is displacement, not harm. Grammar for Writing replaced six weeks of Year 6 literacy teaching with a contextualised grammar programme; KS2 GPS fell 0.06 SD. IPEELL replaced two years of literacy teaching with a composition programme; KS2 GPS fell 0.28 SD, alongside significant losses in reading and maths. What ordinary Year 6 practice contains, in both arms of both trials, is explicit GPS preparation — the Grammar for Writing process evaluation records that in both intervention and control schools teachers used "real texts supplemented with separate, explicit grammar instruction in preparation for the GPS assessment." So the cleanest reading of the strongest evidence in this topic is a displacement result: take the decontextualised conventions time away and conventions scores fall. That is an inference from what the control arm was doing, not a test of conventions instruction — but it is the most direct signal the randomised literature contains, and it points toward direct teaching rather than away from it.

The evaluators of the Grammar for Writing trial reached the same reading independently, attributing the negative GPS effect to "the GPS assessment being a decontextualized assessment whereas the Grammar for Writing intervention advocates a contextual approach." They also caution — correctly — that this was a pre-planned secondary analysis without family-wise error control, so it "cannot be seen as definite evidence for a negative effect." The IPEELL two-year result does not need that caveat.

The knowledge/use gap is the second finding, and four independent lines converge on it. (1) In 6,642 Year 7 pupils, the national conventions test correlated r = 0.79 with another decontextualised grammar test and r = 0.56 with independently marked writing — roughly 62% of variance versus 31%, in the same children. (2) The WWC panel states it in plain English: "Students may be able to correctly circle parts of speech or identify and correct errors in punctuation, but they often do not develop the ability to use these skills in their own work." (3) The teachers in the 2025 EEF trial reported that pupils could use a taught pattern "when this was highly scaffolded" but "rarely transferred use of the taught grammar patterns into more general writing composition tasks." (4) In a peer-editing RCT recorded in the WWC evidence tables, spelling errors fell relative to control while "the intervention produced no changes on students' punctuation errors."

And the mechanism has a name in this archive already: composition load. Wilcox et al. found the same adolescents made significantly more surface errors writing for social studies than for English — knowledge held constant, load raised, errors up. Grünke's three children, who "could not punctuate," punctuated at ~88% within days of being paid to. Neither is a knowledge deficit. Both are performance deficits, and a worksheet cannot detect the difference.

The one place conventions instruction has been shown to reach real writing, someone taught the transfer on purpose. Schmidt et al. tracked mechanics error rate into general-education classes where the strategy had never been taught. Most students generalised — and "the two students who did not generalize... did so quickly after they had been taught to do so." That is n = 7, single-case, no control, and the developers' own report. It is also the only study in any literature that even looked.

Hereditarian-lens assessment

Risk: medium, and the asymmetry is what earns that grade rather than low.

The null side of this verdict is randomised and clean. Writing Wings randomised teachers within 39 schools; the EEF trials randomised 155, 84 and 83 schools with independent evaluators; the Year 7 trial randomised 228 teacher-class units and verified baseline balance on KS2 GPS (χ² = 2.43, p = 0.296, n = 8,446). No selection story produces those nulls.

The positive side is where the premise bites. The strongest pro-teachability datapoint — Direct Instruction's three-quarter-SD lead on the MAT Language subtest — comes from planned variation: communities chose their sponsor, assignment was not random, and the account is the programme developers' own. The authors concede the point themselves, writing that House et al.'s critique is "valid, particularly those citing limitations of research designs where students are not randomly assigned." Robinson-Kooi & Hammond compared two whole schools without randomisation, so any intake difference sits inside the estimate — and the two intervention classes in that study differ from each other by more than one of them differs from the comparison. A verdict resting on a mix of randomised nulls and non-randomised positives is medium by definition.

Where heredity does specific work, it is on the r = 0.79 versus r = 0.56 contrast. Both correlations are inflated by general verbal ability, which is substantially heritable. That is exactly why the gap is the informative quantity rather than either number: whatever ability contributes, it contributes to both, so the extra shared variance between two decontextualised tests is method, not skill.

One consequence of the premise is worth stating because it cuts against the practice. Olson et al. find substantial heritability for spelling, word recognition, handwriting copy and writing samples in 540 twins, with shared environment not significant except for vocabulary. Children of literate parents will absorb a great deal of conventional punctuation from reading, whatever anyone teaches. So a survival argument built on "educated adults were taught this and can punctuate" is exactly the confounded inference the archive's premise forbids, and the practice_record prior is deliberately scoped to the narrow immediate-feedback act to avoid it.

Finally, and per METHODOLOGY: if direct conventions teaching works the way this topic suggests, it is floor-raising. It should get nearly every child over a threshold almost nobody crosses without teaching, and thereby increase the measured heritability of conventions. That would be evidence the teaching worked. Nothing here bears on g.

Boundaries & what critics say

  • The best single counter to this topic is that its positive arm has no grade-A or grade-B source at all. Every well-identified randomised result points at zero or below. A reader who weights design quality strictly and ignores everything below grade B should read this verdict as no-effect, and that reading is defensible. What stops the archive from adopting it is that the grade-B trials all tested writing programmes, not conventions instruction — Writing Wings is the closest, and even it is a whole curriculum with a conventions rubric attached rather than a conventions intervention.
  • Follow Through's Language result is treatment-aligned by construction. DISTAR Language I–III taught usage and sentence forms explicitly and directly; the MAT Language subtest asks about usage, punctuation and capitalisation. That the programme teaching the tested content won is not a surprise, and it is exactly the measure-alignment discount METHODOLOGY applies everywhere else. The instrument is at least an independent published norm-referenced test, which is more than any other positive source here can say.
  • The written-corrective-feedback literature is the largest body of evidence on error correction and it is the wrong population. A critical review of 42 WCF studies found 93% used adult learners and about 9.5% ran in schools. Kang & Han's g = 0.54 and Truscott's "small negative effect" are computed on overlapping corpora of university second-language writers. They are recorded here because they are the only quantitative answers to "does correcting errors teach anything" that exist, and they disagree with each other. Nothing in them licenses a claim about a native-speaking child.
  • Van Beuningen et al. is the best rebuttal to the transfer pessimism and deserves full strength. N = 268 secondary-school pupils, two active control arms including one that got extra writing practice instead of feedback, outcomes on new texts at one and four weeks, and correction won on both. It also falsifies the avoidance objection — corrected pupils did not write more simply. If this replicated in L1 primary children it would move this topic to moderate-support. It has not been tried.
  • The single-case literature is the one that most supports direct conventions teaching, and it has no control groups. Rogers & Graham's meta of 88 single-case designs lists "teaching grammar and usage" and "strategy instruction for editing" among nine supported treatments. That corpus typically means targeted correction of one error class in one struggling writer, measured on that same error class — which is precisely this topic's intervention, precisely this topic's outcome, and precisely the design that cannot separate treatment from maturation or regression to the mean.
  • "Insufficient evidence" is not "does not work", and the archive must not let the first read as the second. Nobody disputes that a six-year-old can be shown where a full stop goes and will then put one there. What is missing is a magnitude, a dose–response, and an age floor — not plausibility.
  • The conventions literature is far smaller than the fuss about it suggests. Datchuk & Kubina's review is the cleanest measure: the whole controlled base for teaching grammar and usage to children is three studies and 51 participants, against ten studies and 394 participants for handwriting. England rewrote its national curriculum around this content in 2014 and introduced a statutory test for it, and no causal design has ever evaluated that change.

Age: what the evidence says about five-, six- and seven-year-olds

It does not reach them, and that is the answer rather than a gap in the search.

  • No controlled study of punctuation or capitalisation instruction has measured a child younger than about seven. The youngest is Robinson-Kooi & Hammond at Year 2 (ages ~7–8) — two schools, not randomised. The youngest capitalisation training study is Campbell et al. 1991, single-case, n = 3, mean age 9. Fogel & Ehri's usage RCT starts at Grade 3.
  • The best developmental account of capitalisation begins at age 8 (Hawkey et al. 2025, ages 8–12), and finds that children capitalise more reliably when cued by word class than by sentence position — i.e. the rule every Year 1 curriculum teaches first (capital at the start of a sentence) is the weaker cue.
  • Punctuation appears at six, years before rule-governed use. Cordeiro et al. found apostrophes and quotation marks appearing in first-graders' writing before reliable full-stop placement: children mark what is interesting, not what is structural. Later work finds seven-to-nine-year-olds still not understanding speech marks. Between "six-year-olds sprinkle marks decoratively" and the rhetorically controlled punctuation studied at nine and ten, no quantitative developmental trajectory exists at all.
  • The flagship trial at exactly this age did not measure conventions. Wyse et al. ran the first robust RCT of grammar teaching on six- and seven-year-olds — 70 schools, 1,246 pupils — and its two outcomes were narrative writing (d = 0.04) and sentence generation (d = 0.14, ns). No punctuation measure, no capitalisation measure. The trial designed to test England's statutory grammar curriculum at Key Stage 1 cannot tell you whether the children learned any of it.
  • What IS evidenced at ages 5–7 is the transcription layer underneath: explicit letter formation (legibility 0.59, fluency 0.63, delivered in groups of three, 20 minutes twice weekly) and explicit systematic spelling (0.54 against no instruction, 0.70 for more-explicit over less). Datchuk & Kubina's asymmetry is the one-line version: the evidence goes that young for transcription. It does not for conventions.

Practical guidance

  • Treat conventions as a separate subject and say so. Justify the time on conventions, never on writing quality. The archive grades the writing claim negative and this topic does not rescue it.
  • Use the shape that has a randomised result behind it: rule → example → immediate guided practice → correction → use it in their own writing. In the one trial that isolated the components, exposure alone failed, rules-plus-exposure failed, and rules-plus-guided-transformation-practice-with-feedback worked — in sixty minutes.
  • Go narrow. Spelling, capitalisation and a handful of punctuation marks account for over half of the errors real students actually make. Teach the marks the child is currently getting wrong, not a curriculum.
  • Keep it short and daily rather than long and blocked. Nine weeks of brief explicit teaching with daily sentence dictation at Year 2 moved full-stop omission from 94% to 28% of pupils. Six weeks of daily hour-long grammar lessons at Year 6 moved the national conventions test down.
  • Check the next piece of writing, not the worksheet. The gap between the two is this topic's central finding. If you only ever measure conventions on exercises you will not find out whether anything transferred, and the honest expectation is that it will not transfer by itself.
  • Teach the transfer as its own step. The only demonstration of conventions gains reaching untaught settings is one in which the students who did not generalise were explicitly taught to generalise, and then did. Treat "use this in everything you write" as a separate thing you teach, not as an automatic consequence.
  • Do not use grammar-checking software as the intervention. Automated surface feedback gives g = 0.31 while the tool is there and −0.02 once it is gone. It is performance support, not instruction. (Neither NoRedInk nor IXL Language Arts has a qualifying controlled study; IXL's own ESSA listing for Language Arts reads "no studies met inclusion requirements.")
  • For three children aged five to seven specifically: teach letter formation, teach spelling systematically, and teach the full stop and the capital letter as two concrete rules inside their own sentences. Do not buy a grammar programme, do not expect a measurable effect on their writing, and know that you are past the edge of the evidence — nothing has been tested at that age, and the developmental record says marks appear long before rules are understood.
  • If it is a test score you want, say so. The KS2 GPS test correlates 0.79 with another grammar test and 0.56 with actual writing. Preparing for a conventions test is preparing for a conventions test. That is a legitimate goal; it is just not a writing intervention, and the trial designed to raise GPS scores lowered them.

Open questions

  • The obvious trial has never been run. Randomise children to direct conventions instruction versus an active control of equal time, measure a standardised conventions subtest AND conventions accuracy in extended writing, and follow up. Every component exists; nobody has assembled them. This is the largest single gap in the writing subject.
  • Nothing exists below age seven. Given that this is where the demand is, and that the transcription literature manages to study six-year-olds routinely, the absence is a choice the field has made rather than a difficulty it has encountered.
  • Does the Van Beuningen transfer result survive in L1 primary children? It is the only well-controlled demonstration that correction reaches new writing, and it is in adolescent second-language writers with two active control arms. Replicating it with nine-year-old native speakers would resolve most of this topic.
  • The Follow Through Language subtest data have never been extracted at model-and-site level. The Abt Volume IV-C appendix tables contain per-model Metropolitan Achievement Test Language and Spelling percentiles across roughly 180 communities — the largest unextracted conventions dataset in existence — and Becker & Gersten's fifth- and sixth-grade follow-up would say whether any of it persisted. Both are retrievable and neither has been read into this archive.
  • Is the knowledge/use gap a dose problem or a kind problem? Nobody has tested whether enough practice automates a convention to the point where it survives composition load, or whether the only route is explicitly taught generalisation. Schmidt et al.'s seven students are the entire evidence base for the second hypothesis.
  • Confidence is capped at medium partly by process. status: surveyed means the adversarial pass METHODOLOGY requires has not been run against this verdict — but here the cap is also substantive: the positive arm has no grade-A or grade-B source, and it would not clear high even after review.

Evidence (28 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore