The Evidence on Teaching

Comprehensible input or explicit grammar teaching — which builds a second language?

Explicit form-focused teaching beats pure exposure, so the archive's negative verdict on L1 grammar does NOT transfer — but the effects are measured on the taught forms, immediately, mostly on adults.

mixedconf: mediumgc: low

foreign-language · ages 418

Effect summary

The L2 answer is genuinely different from the L1 answer, and both halves of that sentence matter. EXPLICIT INSTRUCTION WINS, repeatedly: Norris & Ortega's founding meta-analysis of 49 studies found large target-oriented gains with explicit beating implicit; Goo et al. reproduced the direction on 34 direct comparisons; Spada & Tomita found explicit ahead for BOTH simple and complex features and — the strongest single result here — on spontaneous free production, not merely on controlled knowledge; Kang, Sok & Han pooled 54 studies and 5,051 learners at g = 1.06 (0.84-1.29). THE MEASUREMENT DISCOUNT IS SEVERE. That g = 1.06 moves significantly with mode of outcome measure, onset proficiency, research setting and intensity; Li's corrective-feedback meta finds laboratory effects larger than classroom effects and SHORTER treatments producing LARGER effects than longer ones, which is the signature of a measurement artefact rather than accumulating learning; the outcome measures are overwhelmingly researcher-designed and treatment-aligned, and the corpus is overwhelmingly university students. WHAT SURVIVES INTO A CLASSROOM: Lyster & Saito's 15 classroom-only studies (N=827) find oral corrective feedback significant and DURABLE, larger for prompts than recasts, most apparent on FREE constructed responses, and larger for YOUNGER learners. INPUT ALSO WORKS: extensive reading returns d = 0.46 against a control (0.71 without one) and positive small-to-medium effects across every language domain in the 2025 meta — but larger when text choice was LIMITED and when ACCOUNTABILITY was added, which is not what free voluntary reading means.

Practical takeaway

Teach the forms explicitly and keep the learners producing language, because the L2 evidence — unlike the L1 evidence on grammar-for-writing — consistently favours explicit form-focused teaching, and one meta-analysis finds it improves spontaneous use and not just test performance. Then discount the numbers hard: g = 1.06 is what treatment-aligned tests administered to university students immediately after short treatments produce, and the two syntheses that restrict to real classrooms or to delayed measures give something much smaller. In a classroom, the two best-supported practical moves are oral corrective feedback using PROMPTS rather than recasts (durable, and larger for younger learners) and a lot of extensive reading — with a curated library and some accountability, because the meta-analysis finds constrained choice beats free choice. Treat Krashen's input hypothesis as an unfalsifiable slogan rather than a rival plan: reading a lot works, but 'provide comprehensible input and acquisition follows' specifies no failure condition and cannot be tested against anything.

Who this applies to

Group size
whole-classsmall-groupindependent
Delivered by
teachertutor
Ages studied
1018(narrower than the 418 this topic is filed under — outside it is extrapolation)
Dose
Form-focused instruction: the meta-analytic corpus is dominated by SHORT treatments — a few sessions on a specific structure — with immediate post-tests, and Li 2010 finds shorter treatments produce LARGER effects, so no dose-response can be read off these numbers. Oral corrective feedback in classrooms is the exception: Lyster & Saito find longer treatments larger than short-to-medium ones. Extensive reading: sustained independent reading over a programme of weeks to months, with limited text choice and some accountability.
Cost
free
Moves
domain-skill
Needs first
Enough L2 to read or interact at all. Extensive reading requires a graded-reader library at the learner's level; corrective feedback requires the learner to be producing language, which a purely receptive programme does not generate.
Not for
Do not read these effect sizes as proficiency gains — they are gains on the specific forms taught, measured on tests built around those forms, usually straight after teaching. Do not apply them to children under about 10: the corpus is university students and older adolescents, and there is essentially no controlled trial of input-rich versus form-focused L2 instruction with young children on an independent proficiency measure. Do not run explicit grammar as the whole programme: every meta-analysis here compares instruction against a control or against implicit treatment, none against a well-implemented input-rich course of equal hours. And do not import the archive's L1 grammar verdict — different language, different mechanism, opposite sign.

Verdict

This archive says teaching grammar does not improve writing. It does not follow that teaching grammar does not help you learn a second language, and the evidence says it does. Keeping those two apart is the point of this topic.

Grammar instruction carries a negative verdict because three independent randomised trials and four meta-analyses agree that teaching English grammar to English-speaking children does not make them write better. That is a claim about transfer — from knowing what a subordinate clause is to producing better prose in a language you already speak fluently. The L2 question is not a transfer question at all. The learner does not have the form; instruction supplies it; the test asks whether they now have it. Different mechanism, and the answer comes out the other way.

The verdict is mixed rather than moderate-support because of measurement. The effects are large, they replicate, and they are almost entirely measured on researcher-designed tests of the exact forms that were taught, administered immediately, to university students, after short treatments. Every time someone restricts the corpus — to classrooms, to delayed post-tests, to free production — the number falls.

What the evidence shows

Source Design Grade Key effect
Norris & Ortega 2000 meta-analysis, 49 studies (1980-98) C Large target-oriented gains; explicit > implicit; Focus on Form ≈ Focus on Forms; authors flag that measure type likely drives magnitude
Goo 2015 meta of 34 direct explicit-vs-implicit comparisons C Explicit more effective — the direction survives a 15-year update on a partly new corpus
Spada & Tomita 2010 meta-analysis, 41 studies C Explicit ahead for both simple and complex features, and positive on spontaneous use, not just controlled knowledge
Kang, Sok & Han 2018 meta-analysis, 54 studies / 5,051 learners C g = 1.06 (0.84-1.29); only a minor explicit-implicit difference; significant moderation by measure mode, proficiency, setting, intensity
Sok, Kang & Han 2018 methodological synthesis, 88 studies C Delayed post-testing and multiple measures have increased since 2000; methodological weaknesses persist
Li 2010 meta-analysis, 33 studies incl. 11 dissertations C Medium overall, maintained over time; implicit feedback better maintained; lab > classroom; shorter treatments LARGER than longer; no publication-bias gap
Lyster & Saito 2010 meta of classroom-only studies, 15 studies / N=827 C Significant and durable; prompts > recasts; largest on free constructed responses; younger learners benefit more
Nakanishi 2015 meta-analysis, 34 studies / 3,942 participants C Extensive reading d = 0.46 with a control group, 0.71 without one
Sangers 2025 meta-analysis across 7 outcome domains C Positive small-to-medium in every domain; larger when text choice was limited and when accountability was added
Bryfonski & McKay 2017 meta of TBLT programme implementations, 52 studies D d = 0.93 — but it pools programme evaluations by implementers, with "positive stakeholder perceptions" reported as a finding
Gregg 1984 theoretical critique D Krashen's hypotheses judged incoherent, untestable, or unsupported

Explicit instruction wins, and it wins on the one measure that should be hardest. The standard defence of implicit approaches is that explicit teaching produces declarative knowledge you can use on a grammaticality-judgement test and nothing more. Spada & Tomita tested that directly and found explicit instruction ahead on spontaneous use of both simple and complex forms. That is the single most persuasive finding in the cluster, and it is why the verdict is not a shrug.

Now the discount, and it is large. Consider what Kang, Sok & Han's g = 1.06 is a measurement of. It moves significantly with the mode of the outcome measure, with the setting, and with intensity. Li 2010 reports the tell most clearly: laboratory effects are larger than classroom effects, and shorter treatments produce larger effects than longer ones. A real instructional effect should grow with dose. An effect that shrinks as you teach longer is measuring the freshness of a specific taught form, not the acquisition of a language.

The classroom-only synthesis is the one to trust for a school, and it is smaller and better. Lyster & Saito threw out the laboratory studies and kept 15 classroom trials. Feedback still worked, it lasted, it was largest on free constructed responses, prompts beat recasts, and — unusually for this literature — younger learners benefited more than older ones. That is the most directly usable result in the topic.

Input works too, but not the version Krashen described. Extensive reading is positive across reading comprehension, vocabulary, fluency, motivation, writing, oral proficiency and general proficiency. The moderators are the interesting part: effects were larger when learners' text choice was limited and larger when some accountability was attached. Free voluntary reading works better when it is neither entirely free nor entirely voluntary. Note also the design artefact sitting in plain sight in Nakanishi: 0.46 with a control group, 0.71 without one — a 54% premium for dropping the comparison.

Krashen's framework is not refuted; it is unfalsifiable. Gregg's critique has stood for forty years without being answered on its own terms: the acquisition–learning distinction admits no observation that could separate the two, and the Monitor hypothesis can absorb any result. This matters practically. It means "comprehensible input versus explicit instruction" is not a fair fight in the literature — there is a large corpus of trials of explicit instruction and essentially none of a well-specified input-only programme, because input-only was never specified well enough to trial.

Does the L1 grammar verdict transfer? No.

Explicitly: no, and the reason is structural rather than empirical.

  • grammar-instruction is a transfer claim. The child already speaks English. Teaching them to label a fronted adverbial is supposed to improve their prose. It does not; the meta-analytic estimate against alternative writing instruction is −0.32 and the best randomised trials return zero.
  • L2 form-focused instruction is an acquisition claim. The learner cannot form the past subjunctive. Instruction teaches it. The test asks whether they can now form it. There is no transfer step to fail.
  • The two literatures even agree where they overlap. The L1 verdict is not "explicit teaching is useless" — it is "grammar labelling does not transfer to composition", and the same archive gives moderate-support to explicit phonics, explicit spelling and sentence combining. The through-line is that explicit teaching works on the thing it teaches. That is exactly what the L2 evidence shows.
  • The shared caution is measurement, not direction. In L1 writing the researcher-scored inflation is roughly 5×. In L2 instruction the treatment-aligned inflation is unmeasured but visibly large. Both literatures should be read with the same suspicion of the number and different conclusions about the practice.

Hereditarian-lens assessment

Risk: low. Nearly every study pooled here assigns learners to instructional conditions within an existing class or laboratory session, so treated and untreated learners have the same expected distribution of language-aptitude alleles. This is one of the few areas of language education where the premise does not do much work.

Two places it still matters:

  • Kang et al. find onset L2 proficiency moderates the effect. Prior proficiency is substantially a heritable-aptitude proxy, so "instruction works better for stronger learners" is a correlation between a heritable trait and a response, not a demonstrated aptitude–treatment interaction. It should not be used to allocate instruction.
  • Bryfonski & McKay's TBLT meta-analysis pools programme evaluations rather than controlled comparisons. Which institutions adopt task-based programmes, and which students attend them, is not random. Its d = 0.93 is a selection estimate wearing a meta-analysis's clothes, and it is graded D on that basis.

Boundaries & what critics say

  • The corpus is adults. Norris & Ortega, Goo, Spada & Tomita, Kang et al. and most of Li's studies are dominated by university students. This archive covers ages 4-18. The only synthesis restricted to classrooms — Lyster & Saito — is also the only one that found an age moderator, and it points toward younger learners, but it is 15 studies.
  • There is essentially no controlled trial of input-rich versus form-focused L2 instruction with young children on an independent proficiency measure. That sentence is the honest state of the field and should be repeated whenever someone quotes a d above 1.0 in a primary-school context.
  • Almost every outcome is treatment-aligned. The measures are grammaticality judgements, selected-response items and constrained gap-fills built around the taught structure. Free constructed response — actually using the language — is the rarest measure and, in Norris & Ortega's corpus, the least favourable. Lyster & Saito's classroom studies invert this, which is a genuine tension and worth watching.
  • Nobody has run the comparison the debate is actually about. Every meta-analysis compares instruction against a control, or explicit against implicit instruction. None compares a well-implemented, input-rich course against a form-focused course at equal hours over a school year on an independent proficiency test. The famous argument has no trial behind it.
  • Durability is thinly evidenced but not absent, which is better than this archive usually finds. Li reports the corrective-feedback effect maintained over time with implicit feedback maintained better; Lyster & Saito report durable classroom effects. Sok, Kang & Han confirm delayed post-testing has become more common since 2000 while noting methodological weaknesses persist.
  • Extensive reading's strongest moderators embarrass its own ideology. Limited choice beats free choice; accountability beats none. Whatever is working, it is not the absence of structure.

Practical guidance

  • Teach the forms explicitly. It is the better-evidenced side of a genuinely contested question, it wins on spontaneous production and not merely on tests, and — unlike the L1 case — nothing here suggests it harms anything.
  • Use prompts, not recasts. When a learner produces an error, push them to self-correct rather than reformulating it for them. This is the clearest actionable finding from the only classroom-restricted synthesis, and its effects were durable.
  • Give feedback to younger learners especially. Lyster & Saito's age regression runs toward younger learners benefiting more, which is the opposite of the usual assumption that children should just be bathed in input.
  • Run a large extensive-reading programme, with a curated library and light accountability. Budget d ≈ 0.46, not 0.71, and do not make the library entirely free choice — constrained choice performed better.
  • Discount any L2 effect size above ~0.8 until you know what the test was. If the outcome was built around the structure that was taught and administered the same week, the number is about the test.
  • Do not treat Krashen as the alternative plan. Read a lot, yes. But "comprehensible input is sufficient" specifies no failure condition, has been unanswered as a criticism since 1984, and cannot be planned against.
  • Do not import the L1 grammar verdict. Different language, different mechanism, opposite sign. Someone will try; the argument above is why they are wrong.

Open questions

  • The head-to-head trial does not exist: input-rich versus form-focused L2 instruction, equal hours, school year, independent proficiency test, school-age children. Everything in this topic is a proxy for it.
  • How much of g = 1.06 survives an independent standardized proficiency measure? Nobody knows, because the field almost never uses one.
  • Why do shorter treatments outperform longer ones? Li's finding is either a serious artefact warning or a real fact about how form-focused instruction saturates. It has not been investigated.
  • Does the prompts-over-recasts result hold with young children in a foreign-language classroom, as opposed to the mixed adult/adolescent second-language classrooms it was established in?
  • What is the durable effect at one year? Delayed post-tests in this literature mean weeks, occasionally months. Nothing here measures a school year later.

Evidence (11 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore