The Evidence on Teaching

How effective is feedback for L1, L2, and FL learners’ writing? A meta-analysis

Scherer S, Graham S, Busse V · 2024

grade Cmeta-analysisindependentunreplicated
Sample
200 comparisons from experimental and quasi-experimental studies
Population
Secondary-school and university students learning in a first language (L1), a second language (L2) or a foreign language (FL). Explicitly NOT primary/elementary.
Design
The most useful feedback meta-analysis for this topic because it is the one that keeps SURFACE-LEVEL outcomes (mechanics, grammar, spelling, accuracy) separate from DEEP-LEVEL outcomes (content, organisation, quality) instead of pooling everything into "writing quality", and because it includes an L1 arm rather than being confined to second-language learners. Effect sizes are Hedges g from experimental and quasi-experimental designs. The population limit is the load-bearing caveat for this archive: nothing in it reaches children of primary-school age, so it bounds what can be said about a 5-to-7-year-old at zero. Numbers taken from the published abstract, not the full text.
Key findings
Feedback aimed at surface features moves surface features: surface-level feedback on surface-level outcomes g = 0.58 overall, with FL learners at g = 0.69 and L2 learners at g = 0.34 (the abstract does not report a separate L1 surface-level figure). Instructor feedback on surface outcomes g = 0.72 (FL) and 0.35 (L2). The finding that matters most for a conventions topic is the CROSS effect: surface-level feedback on FL learners' DEEP-level outcomes was g = -0.23, i.e. correcting form may cost composition quality. Deep-level feedback on deep-level outcomes was larger than anything in the surface family (g = 0.80 overall; L1 g = 1.26; FL g = 0.37, ns), and peer feedback on deep-level outcomes reached g = 1.46 for L1 learners. Combined surface-plus-deep feedback gave deep 0.54 and surface 0.36. Algorithm-based and self-feedback showed non-significant medium effects (0.53 and 0.55).
Genetic confound
Low. The pooled corpus is experimental and quasi-experimental, so learners are assigned to feedback conditions rather than selected into them. The L1-versus-L2-versus-FL contrasts are between populations rather than within a randomised design and should not be read as causal moderators.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Surface-level writing outcomes (mechanics, grammar, accuracy) after surface-level feedbackHedges g0.58 overall; FL 0.69; L2 0.34researcher-designedpost-treatmentunclearend-of-treatmentdomain-skill
Deep-level writing outcomes (content, organisation, quality) after SURFACE-level feedback, FL learnersHedges g-0.23researcher-designedpost-treatmentunclearend-of-treatmentdomain-skill
Deep-level writing outcomes after deep-level feedbackHedges g0.80 overall; L1 1.26; FL 0.37 (ns)researcher-designedpost-treatmentunclearend-of-treatmentdomain-skill

Cited by