The Evidence on Teaching

Does classroom drama build verbal and literacy skills?

The headline effects are a measurement artefact: drama gives d≈0.89 on researcher-made tests and d≈0.29 on standardized ones, and the two randomised trials with standardized outcomes are null.

mixedconf: mediumgc: medium

arts · ages 516

Effect summary

Drama is the one arts-transfer link the REAP meta-analyses endorsed — 80 reports, effects on oral story recall (r .16-.34), written story understanding (.37-.73), reading achievement (.11-.29) and oral language (.20-.41), with vocabulary null and gains that carried to non-enacted texts. Later syntheses reproduce the direction and then explain most of it away. Lee, Enciso and Brown pool 32 studies and publish the moderator that matters: researcher-made tests give d = 0.89 (adjusted 1.27) and STANDARDIZED tests give d = 0.29 (adjusted 0.35); direct outcomes 0.68 versus indirect 0.24; studies using one assessment 0.87 versus more than ten 0.04; and the dose-response runs backwards, 0.87 for 3-10 hours against −0.01 beyond 20. Lee et al.'s earlier 47-study meta says of its own corpus that designs were 'generally weak for making causal inferences' and contains no RCT. The two randomised tests with standardized outcomes are the ones that decide it: the EEF's Speech Bubbles drama trial (n = 821) gives reading −0.05 and oral narrative −0.04, and a Finnish RCT of readers' theatre found the ACTIVE oral-reading control made more progress than the drama arm.

Practical takeaway

Use drama to make a particular text stick — enacting a story reliably improves recall and comprehension of that story, which is a real and useful near-transfer effect. Do not use it as a literacy programme. The published effect sizes are three times larger on tests the researchers built than on published ones, they get smaller the more you measure and the longer you run it, and the two randomised trials with standardized outcomes found nothing — in one of them a plain oral-reading control did better. If you want reading fluency, teach reading fluency.

Who this applies to

Group size
whole-classsmall-group
Delivered by
teacher
Ages studied
516
Dose
Where positives appear, 3-10 hours and more than five lessons; effects are LARGER at low dose and negative beyond 20 hours, which is the signature of a measurement artefact rather than a dose-response. Delivery by the classroom teacher outperforms delivery by a visiting teaching artist (d = 0.89 versus 0.31).
Cost
low
Moves
near-transfer
Needs first
The texts enacted have to be the texts you care about. Every credible effect is on comprehension and recall of material that was dramatised, or on closely adjacent material.
Not for
Vocabulary — null across the REAP meta-analysis (r −.07 to .19). Mathematics — d = 0.24 at best on mixed measures. Silent reading and reading self-efficacy — explicitly null in the one randomised active-control trial. And it is not a substitute for oral reading practice: the plain oral-reading control beat the drama arm on fluency.

Verdict

Mixed. Drama earns a separate topic because it is the one arts-transfer link that a serious meta-analytic project endorsed, and because the story of what happened to it afterwards is the cleanest worked example in this archive of an effect being explained by how it was measured.

The honest position has two halves. Enacting a text improves understanding and recall of that text, and REAP found the gains carried to new, non-enacted texts — that is a real near-transfer effect and it is the reason to use drama. Drama as a route to general literacy attainment does not survive the two things that dissolve most education findings: a standardized outcome measure and a randomised active control.

What the evidence shows

Source Design Grade Key effect
Podlozny 2000 (REAP drama) seven meta-analyses, 80 of 200 reports C Oral recall r .16–.34; written understanding .37–.73; reading achievement .11–.29; vocabulary null
Lee 2015 meta, 47 studies, zero RCTs C "Positive, significant" on achievement; corpus designs "generally weak for making causal inferences"
Lee 2020 meta, 32 studies / 209 effects C Researcher-made 0.89 vs standardized 0.29; 1 assessment 0.87 vs >10 assessments 0.04; >20 hours −0.01
EEF Speech Bubbles cluster RCT, n = 821 A Reading −0.05 [−0.20, 0.10]; oral narrative −0.04 [−0.19, 0.11]
Hautala 2023 RCT with active oral-reading control, n = 146 B Control made more progress than the drama arm; null on comprehension, silent reading and self-efficacy
Hetland & Winner 2001 (REAP) ten meta-analyses C Drama→verbal is one of only three reliable causal links in the whole arts literature

The measure-type moderator is the finding. Lee, Enciso and Brown put the numbers next to each other: nonstandardized researcher-made tests d = 0.89 (adjusted 1.27), standardized tests d = 0.29 (adjusted 0.35). That is a three-fold gap and it matches this archive's general rule that treatment-aligned researcher-designed measures roughly double an effect — here it does more than double it. Two further moderators point the same way and are, if anything, more damning: studies using a single assessment give d = 0.87 while studies using more than ten give 0.04, and the dose-response runs backwards, 0.87 at 3–10 hours against −0.01 beyond 20 hours. Real instructional effects do not shrink as you measure more carefully and practise more.

The randomised trials close it. Speech Bubbles is a well-run drama programme evaluated inside the EEF's Learning About Culture consortium with a pre-specified standardized reading measure and a purpose-chosen oral-narrative measure — the two outcomes drama should move if it moves anything. Both null, at n = 821. Hautala et al. then ran the comparison nobody else ran: readers' theatre against an active control receiving ordinary oral reading instruction. The control group made more progress on oral reading fluency, and the drama arm was null on expressive reading, sentence verification, cloze comprehension, silent reading and self-efficacy. What the drama arm did produce was lower reading anxiety and higher engagement — transiently.

None of this makes REAP wrong about what it measured. Podlozny's strongest cells are story understanding and recall, oral and written — outcomes about the enacted text. That is the effect, it is near transfer, and REAP's observation that it carried to non-enacted texts is the most interesting unresolved claim in the topic. Note also what Podlozny's own moderators show: larger effects for non-experimental designs and for journal-published studies. Design weakness and publication status both inflate the estimate, in her own data.

And the one drama finding with an active control is not about literacy at all. In Greene et al.'s theatre lotteries, students randomised to a live performance gained on tolerance (+0.142) and on plot and vocabulary knowledge (+0.101) while students randomised to a film of the same play gained nothing. Watching drama does something. Doing drama as a reading programme is the claim that fails.

Hereditarian-lens assessment

Risk: medium, but the operative biases here are not hereditary and it is worth being precise about that, because a hereditarian prior is the wrong tool for this topic.

The corpus is almost entirely quasi-experimental — Lee et al. 2015 contains no randomised trial at all — so selection into drama programmes is uncontrolled. But the pattern of the evidence points at different culprits. Effects shrink with standardized measurement, shrink with more assessments, shrink with longer exposure, and shrink when the deliverer is a visiting artist rather than the class teacher (0.31 versus 0.89). That profile is measure proximity, single-outcome selection and unblinded rating — not families passing on verbal ability.

The two designs that do remove selection genetically — random assignment in Hautala and cluster randomisation in Speech Bubbles — are the two that return nulls, which is consistent with selection explaining part of the observational signal. But the archive should not claim more than that: the distinctive thing about drama is that it was mismeasured, not that it was confounded by heredity.

Boundaries & what critics say

  • The near-transfer effect is probably real and is worth having. Acting out a story helps children understand and remember that story. Podlozny's oral and written recall cells are the strongest in the arts literature, and REAP reports transfer to non-enacted texts.
  • Lee et al. 2015's numbers could not be obtained. The field's most-cited drama meta-analysis is paywalled, and its pooled effect size, confidence interval and publication-bias analysis are recorded as not found rather than guessed. That is a real hole in this topic.
  • Podlozny's figures are second-hand. They come from the Critical Links summary rather than the original, and are reported as confidence-interval ranges, not point estimates. The Critical Links reviewers, applying stricter comparison criteria to 19 studies, failed to reproduce the reading-readiness cell.
  • The drama defenders have an argument and it should be stated. Lee et al. explicitly reject the measurement critique, describing standardized tests as "austere, irrelevant" compared with the in-situ measures. That is a claim about what schooling is for, not about what happened in the studies, and this archive does not accept it — but a reader entitled to different weights should know the argument exists.
  • Drama beats music on social outcomes in the one trial that compared them. Schellenberg's randomised drama arm produced d = 0.57 on adaptive social behaviour where music produced nothing. Drama's best-supported effects may be social rather than verbal, and nobody has followed that up.

Practical guidance

  • Dramatise the text you want understood. That is the use the evidence supports: comprehension and recall of enacted material, plus some carry to adjacent texts.
  • Have the class teacher run it, not a visiting artist. Teacher-led d = 0.89 against teaching-artist d = 0.31 across 32 studies. Whatever is happening, buying it in does not deliver it.
  • Do not substitute drama for reading practice. The one randomised active-control trial found plain oral reading did more for fluency than readers' theatre did.
  • Discount any drama effect measured on a test the programme's evaluators built. Multiply by about a third to get the standardized-test equivalent.
  • Keep it short if you keep it. Effects are larger at 3–10 hours and gone beyond 20 — which is a reason to be sceptical of the effect, and, if you use drama anyway, a reason not to build a term around it.

Open questions

  • Whether REAP's transfer to non-enacted texts is real is the most interesting unresolved claim here, and no randomised trial has tested it.
  • Lee et al. 2015's pooled estimate and bias analysis remain unobtained; the topic cannot be closed without them.
  • Why the dose-response runs backwards is unexplained — measurement artefact is the obvious answer, but nobody has demonstrated it.
  • Whether drama's effects are actually social rather than verbal, as Schellenberg's randomised drama arm suggests, has never been followed up.

Evidence (7 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore