The Pen Is Mightier Than the Keyboard: Advantages of Longhand Over Laptop Note Taking
Mueller, P. A., & Oppenheimer, D. M. · 2014
grade Crctindependentfailednumbers spot-checked
Sample
Study 1 final N = 65 (67 recruited); Study 2 final N = 151; Study 3 final N = 109 (142 attended session 1, 118 returned for session 2). Three separate experiments, all single-site laboratory.
Population
University students from the Princeton subject pool (Study 1) and the UCLA Anderson Behavioral Lab subject pool (Studies 2 and 3), taking notes on TED talks or short prose passages in a laboratory, not in a real course.
Design
Three randomized laboratory experiments contrasting longhand and laptop note-taking on lecture material, with test performance split into FACTUAL-RECALL and CONCEPTUAL-APPLICATION questions. Study 1 tested immediately after a 30-minute distractor. Study 2 added a third arm instructed not to transcribe verbatim (the instruction failed — verbatim overlap was unchanged at 12.07% vs 12.11%). Study 3 crossed medium with an opportunity to review notes and tested a WEEK later. Notes were content-analysed for word count and for verbatim overlap with the lecture using an n-grams program. Crucially, no participant in any study was allowed to use the laptop for anything but notes, so this is a test of the MEDIUM, not of distraction — which is exactly what makes it a different claim from Sana et al. (2013). WHY THIS IS GRADED C, not B: three small single-site lab experiments on university subject pools, no preregistration, and every outcome is a researcher-designed quiz on material the participants had no stake in. The measure-type discount METHODOLOGY.md applies (aligned researcher instruments inflate roughly 2x over independent standardized measures) bites hard here, and there is no achievement, course-grade or standardized outcome anywhere in the paper. A further internal problem, checkable from the article itself: several reported F/p pairs do not cohere with their stated degrees of freedom — Study 1's conceptual effect is given as F(1, 55) = 9.99, p = .03 when that F on those df implies p = .003, and Study 2's as F(1, 89) = 11.98, p = .017 when it implies p = .0008. The errors run in the conservative direction, so they are a reporting failure rather than inflation, but they are the kind of thing the 2018 corrigendum was issued to fix. Independent academic work; no vendor funding. Displacement was measured in one specific sense and it is the paper's best evidence: the notes themselves were counted and compared, so we know exactly what the keyboard changed about the recording behaviour, even though we do not know what it changed about learning.
Key findings
Across three lab experiments, students who typed their notes did no worse than longhand writers on factual recall but worse on conceptual-application questions (Study 1: z-scores +0.154 vs -0.156, F(1,55) = 9.99, eta-p-squared = .13; Study 2: +0.28 vs -0.15, F(1,89) = 11.98, eta-p-squared = .12). The proposed mechanism is shallow encoding through verbatim transcription: laptop notes were about twice as long (309.6 vs 173.4 words in Study 1, d = 1.4) and contained substantially more verbatim overlap with the lecture (14.6% vs 8.8%, d = 0.94), and in mediation models more words predicted BETTER performance while more verbatim overlap predicted worse. Study 3, testing a week later, found the longhand advantage only when students got to review their notes first (longhand-plus-study beat the other three cells, d = 0.64 to 0.97); with no opportunity to study, the medium difference was not significant (t(105) = 1.82, p = .07, d = 0.4). The caveat that most changes how this should be read: a direct, better-powered replication failed, and a meta-analysis of all four direct-replication experiments — including these — puts the conceptual effect at d = 0.17, 95% CI [-0.07, 0.40]. The transcription mechanism is real; the learning consequence is not established.
Genetic confound
Minimal (randomized allocation to note-taking medium in all three studies).
Replication notes
Morehead, Dunlosky & Rawson (2019, Educational Psychology Review; see `edt-morehead-2019-note-taking-replication`) ran two direct replications using this paper's own materials and procedure, with roughly three times the sample, and did not reproduce the result: their conceptual-question effects were d = 0.13 (p = .31) in Experiment 1 and d = -0.35 (i.e. REVERSED) in Experiment 2. They then pooled all four direct-replication experiments — Mueller & Oppenheimer's Studies 1 and 2 together with their own two — in a continuously cumulating meta-analysis. The pooled effect on conceptual questions is d = 0.17, 95% CI [-0.07, 0.40], p = .16, and on factual questions d = 0.22, 95% CI [-0.01, 0.46], p = .06. Neither reaches significance. That analysis also exposes how thin the original was: Mueller & Oppenheimer's own two conceptual effects were d = 0.34 (p = .17) and d = 0.40 (p = .05) — a marginal result in a small sample, not the decisive demonstration the title claims. The MECHANISM survives where the outcome does not: laptop notes are reliably longer (pooled d = -0.77) and more verbatim (pooled d = -0.59), replicated in every experiment. Per METHODOLOGY.md's replication rule, any verdict resting on this source is capped at `mixed` regardless of its fame. A 2018 corrigendum (Psychological Science 29(9), 1565-1568, doi 10.1177/0956797618781773) re-reported the analyses after correcting the z-score computation and excluding 12 participants; the authors state the conclusions are unchanged. We did not obtain the corrigendum full text, so the figures recorded here are the original article's.
DOI / URL
10.1177/0956797614524581
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Conceptual-application test performance, longhand vs laptop (Study 1) | z-score difference / d | longhand +0.154 (SD 1.08) vs laptop -0.156 (SD 0.915); F(1, 55) = 9.99, eta-p-squared = .13 (reported p = .03, which does not correspond to that F on those df — the implied p is .003). Later recomputed as d = 0.34, p = .17 in Morehead et al.'s pooled analysis. | researcher-designed | about 30 minutes after the lecture | active-alternative | end-of-treatment | domain-skill |
| Factual-recall test performance, longhand vs laptop (Study 1) | z-score difference | laptop +0.021 (SD 1.31) vs longhand +0.009 (SD 1.02); F(1, 55) = 0.014, p = .91 — NO difference on facts, which is the half of the result that is usually forgotten | researcher-designed | about 30 minutes after the lecture | active-alternative | end-of-treatment | domain-skill |
| Conceptual-application performance, longhand vs laptop (Study 2) | z-score difference / d | longhand +0.28 (SD 1.04) vs laptop-no-intervention -0.15 (SD 0.85); F(1, 89) = 11.98, eta-p-squared = .12 (reported p = .017; implied p = .0008). The arm told not to transcribe verbatim (-0.11) differed from neither, and the instruction did not reduce verbatim overlap at all. | researcher-designed | about 30 minutes after the lecture | active-alternative | end-of-treatment | domain-skill |
| Test performance one week later, by medium and opportunity to review notes (Study 3) | d | longhand-plus-study outperformed the other three cells, d = 0.64 (t(105) = 3.11, p = .002) and d = 0.97 (t(105) = 4.85, p < .001) on two composites; WITHOUT an opportunity to study, the medium difference was not significant (t(105) = 1.82, p = .07, d = 0.4). The durable advantage is conditional on reviewing the notes. | researcher-designed | one week after the lecture | active-alternative | under-1yr | domain-skill |
| What the keyboard changed about the notes themselves — word count (mechanism/displacement) | d | Study 1 longhand 173.4 words (SD 70.7) vs laptop 309.6 (SD 116.5), d = 1.4; Study 2 155.9 vs 260.9, d = 1.11; Study 3 390.65 vs 548.73, d = 0.77. Pooled across all four direct replications, d = -0.77, 95% CI [-0.98, -0.57] (laptop longer). More words predicted BETTER performance. | researcher-designed | during the lecture | active-alternative | end-of-treatment | behaviour |
| Verbatim overlap between notes and lecture (the proposed mechanism) | percentage / d | Study 1 longhand 8.8% (SD 4.8) vs laptop 14.6% (SD 7.3), d = 0.94; Study 2 6.9% vs 12.11%, d = 1.12; Study 3 4.2% vs 11.6%, d = 1.68. Pooled across all four direct replications, d = -0.59, 95% CI [-0.79, -0.39]. More verbatim overlap predicted WORSE performance. This is the part of the paper that replicates. | researcher-designed | during the lecture | active-alternative | end-of-treatment | behaviour |