The Evidence on Teaching

Can algorithm-based feedback help students to write better? A meta-analysis exploring surface- and deep-level outcomes

Scherer S, Graham S, Busse V · 2026

grade Cmeta-analysisindependentunreplicated
Sample
33 studies; 49 comparisons at posttest
Population
Secondary-school and university students. No primary-school studies.
Design
Meta-analysis of automated/algorithm-based writing feedback (automated writing evaluation, grammar checkers and similar), keeping surface-level outcomes separate from deep-level ones and - crucially, and unusually for this literature - reporting MAINTENANCE effects as well as posttest effects. That single design choice produces the most important number in the whole software-and-conventions literature. Numbers taken from the published abstract and reported summary, not the full text.
Key findings
Overall g = 0.36. Surface-level outcomes (grammar, spelling, mechanics): g = 0.31 at POSTTEST but g = -0.02 at MAINTENANCE. Deep-level outcomes ran the other way: 0.31 at posttest rising to 0.54 at maintenance. Eighteen of 49 comparisons were negative at posttest. The signature - surface gains that evaporate once the tool is withdrawn, deep gains that consolidate - is what performance support looks like rather than skill acquisition, and it is exactly the pattern a conventions topic has to worry about: a child using a grammar checker produces more conventional text while using it and is no more conventional afterwards.
Genetic confound
Low. Pooled experimental and quasi-experimental comparisons in which students are assigned to feedback conditions.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Surface-level writing outcomes (grammar, spelling, mechanics) after algorithm-based feedbackHedges g0.31 at posttest; -0.02 at maintenanceresearcher-designedposttest and delayed maintenance testunclearunder-1yrdomain-skill
Deep-level writing outcomes after algorithm-based feedbackHedges g0.31 at posttest; 0.54 at maintenanceresearcher-designedposttest and delayed maintenance testunclearunder-1yrdomain-skill
Overall effect of algorithm-based feedback on writingHedges g0.36; 18 of 49 comparisons negative at posttestresearcher-designedposttestunclearend-of-treatmentdomain-skill

Cited by