The Evidence on Teaching

How Much Mightier Is the Pen than the Keyboard for Note-Taking? A Replication and Extension of Mueller and Oppenheimer (2014)

Morehead, K., Dunlosky, J., & Rawson, K. A. · 2019

grade Creplicationindependentunclearnumbers spot-checked
Sample
Experiment 1 — 193 recruited, 187 analysed (62 longhand, 64 laptop, 61 eWriter). Experiment 2 — 222 undergraduates across four groups (longhand, laptop, eWriter, and a NO-NOTES control of 32). Roughly three times the per-cell sample of the original.
Population
Undergraduates at a large Midwestern university (Kent State), watching TED talks in individual cubicles. Laboratory, not a real course — the same population and setting as the original, which is what makes it a direct rather than a conceptual replication.
Design
A direct replication using Mueller & Oppenheimer's own videos, instructions, question types, distractor interval and scoring, extended in three ways that turn out to matter more than the replication itself. (1) An eWriter arm (write with a stylus on an electronic tablet) to separate the physical act of handwriting from the paper. (2) A NO-NOTES control in Experiment 2 — the condition the original never ran. (3) A delayed test in Experiment 2 preceded by seven minutes of studying one's own notes, which tests the STORAGE function of note-taking rather than only the encoding function. Random assignment to method, video and test group throughout. Notes were transcribed and scored for word count, verbatim overlap, idea units and test relevance. WHY C AND NOT B: the samples are adequate (n = 187 and 222, about three times the original's per cell) and this is a genuine direct replication, but every outcome is still a researcher-designed quiz on TED talks in a lab cubicle, with no standardized measure, no course grade and no stake for the participant — the measure-type discount METHODOLOGY.md applies keeps it off B. Overall performance in Experiment 1 was near floor, which the authors acknowledge and address in Experiment 2 by using only the two easiest videos. Independent academic work; no funder with a stake in either medium. SCOPE NOTE: this pair (this file and the original) is filed under ed-tech rather than with the archive's handwriting sources on purpose. The question here is not whether handwriting builds letter knowledge or transcription fluency — it is whether a DEVICE in a lecture costs learning, which is the argument schools actually invoke when writing a laptop policy.
Key findings
The famous conceptual-question advantage of longhand did not replicate. In Experiment 1 the immediate factual difference favoured longhand and reached the authors' one-tailed threshold (t(61) = 1.79, p = .04, d = 0.45) but the CONCEPTUAL difference — the original's headline — was d = 0.13, p = .31. In Experiment 2 the trend on conceptual questions actually REVERSED toward laptops (d = -0.35, p = .09), and no delayed-test difference between methods survived after students studied their notes. The most damaging extension is the one the original never ran: the NO-NOTES group performed as well as, and sometimes numerically better than, every note-taking group — so in this paradigm there is no detectable encoding benefit of taking notes at all, which makes a medium contrast between two ways of taking them hard to interpret. eWriter notes behaved like longhand and did not differ from either other group on performance. What did replicate cleanly is the MECHANISM without its consequence: laptop notes were longer (d = 0.44 and 0.58) and, in Experiment 1, more verbatim (d = 0.43; in Experiment 2 the overlap difference was not significant, d = 0.20). Pooling all four direct replications gives a small non-significant longhand advantage: conceptual d = 0.17, 95% CI [-0.07, 0.40]; factual d = 0.22, 95% CI [-0.01, 0.46].
Genetic confound
Minimal (randomized allocation to note-taking method in both experiments).
Replication notes
Direction, stated so the flag is not misread: this paper IS the replication — of Mueller & Oppenheimer (2014; `edt-mueller-oppenheimer-2014-longhand-vs-laptop`) — and it is the reason that source is flagged `failed`. Whether this paper's own null has itself been independently reproduced we have not checked, so `unclear` rather than `unreplicated`, which would assert a fact about the literature we cannot source. What it does carry internally is a continuously cumulating meta-analysis pooling all four direct-replication experiments (Mueller & Oppenheimer's Studies 1 and 2 plus these two), which is the best single estimate available: conceptual d = 0.17, 95% CI [-0.07, 0.40], p = .16; factual d = 0.22, 95% CI [-0.01, 0.46], p = .06. The authors' own summary is that "concluding which method is superior for improving the functions of note-taking seems premature" — not that longhand is worse, and not that laptops are fine. The archive should carry that agnosticism rather than flipping the original's conclusion.
DOI / URL
10.1007/s10648-019-09468-2

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Immediate conceptual-application test, longhand vs laptop (Experiment 1) — the direct replication targetdd = 0.13, t(61) = 0.50, p = 0.31 — the original reported this contrast as its headline significant effect; here it is a nullresearcher-designedimmediately after the lecture (same distractor interval as the original)active-alternativeend-of-treatmentdomain-skill
Immediate factual-recall test, longhand vs laptop (Experiment 1)dd = 0.45, t(61) = 1.79, p = 0.04 — significant, but on FACTUAL questions, which is the opposite pattern to the original (where facts showed nothing and concepts showed the effect)researcher-designedimmediately after the lectureactive-alternativeend-of-treatmentdomain-skill
Immediate test, longhand vs laptop (Experiment 2)dfactual d = 0.28, p = 0.15; conceptual d = -0.35, p = 0.09 — the conceptual trend REVERSES, favouring laptopsresearcher-designedimmediately after the lectureactive-alternativeend-of-treatmentdomain-skill
Delayed test after studying one's own notes (Experiment 2, storage function)F / pno significant effect of note-taking method on either factual (F(1,168) = 2.25, p = .14) or conceptual (F(2,168) = 1.17, p = .31) performance; group differences shrank further after studyresearcher-designeddelayed test, after 7 minutes of studying notesactive-alternativeend-of-treatmentdomain-skill
No-notes control vs all note-taking groups (Experiment 2 extension)group comparisonthe no-notes group performed as well as, and sometimes numerically better than, longhand, laptop and eWriter groups — no detectable encoding benefit of note-taking at all in this paradigmresearcher-designedimmediate and delayed testsnoneend-of-treatmentdomain-skill
Pooled estimate across all four direct-replication experiments (continuously cumulating meta-analysis)dconceptual d = 0.17, 95% CI [-0.07, 0.40], p = .16; factual d = 0.22, 95% CI [-0.01, 0.46], p = .06. Both favour longhand; neither is significant.researcher-designedimmediate tests, pooledactive-alternativeend-of-treatmentdomain-skill
What the keyboard changed about the notes (word count and verbatim overlap)dlaptop notes longer in both experiments (d = 0.44, t(122) = 2.43, p < .01; d = 0.58, t(117) = 3.18, p < .01), replicating the original. Verbatim overlap higher for laptop in Experiment 1 (d = 0.43, p = .02) but NOT significantly in Experiment 2 (d = 0.20, p = .14). Pooled across all four experiments — word count d = -0.77 [-0.98, -0.57]; verbatim overlap d = -0.59 [-0.79, -0.39].researcher-designedduring the lectureactive-alternativeend-of-treatmentbehaviour

Cited by