The Evidence on Teaching

The 2 Sigma Problem: The Search for Methods of Group Instruction as Effective as One-to-One Tutoring

Bloom BS · 1984

grade Dreviewdeveloper-ledfailednumbers spot-checked
Sample
Narrative essay reporting two supervised dissertations (~30 students per arm, grades 4/5/8, 3-week units)
Population
US grades 4, 5 and 8; probability and cartography units chosen for zero prior knowledge.
Design
A narrative essay, not a study: Bloom summarises two dissertations supervised in his own department (Anania 1983; Burke 1983). The 2.0 figure is a compound of researcher-built tests on 3-week novel units, tutoring FUSED with retest-to-criterion mastery, exceptionally trained tutors, ~1 extra hour per week, a 90% criterion applied only in the tutored arm, tutoring REPLACING rather than supplementing class, and control-group SD denominators. Provenance is broken: Bloom's Table 1 claims adaptation from Walberg (1984), who listed tutoring at 0.40; Bloom printed 2.00 with no computed source.
Key findings
The most famous number in education is a compound design artefact. The honest decomposition (von Hippel): ~1.1 sigma from the mastery/testing component plus ~0.9 tutoring-specific on narrow aligned tests — and on standardized measures modern tutoring lands at 0.2-0.35. Bloom's own cited source put tutoring at d = 0.40.
Genetic confound
Bloom read the falling aptitude-achievement correlation as proof that ability differences are alterable; ceiling-compressed distributions on treatment-aligned tests are the simpler explanation. Arlin's growing time-to-mastery ratios contradict the convergence claim, and heritability of achievement RISES with age.
Replication notes
NEVER replicated: 0 of 96 modern tutoring RCTs reach 2.0 (Nickow); 1 of 65 in Cohen 1982 (a 32-student dissertation). Modern honest anchors: 0.3-0.8 efficacy, 0.2-0.35 standardized. The two underlying dissertations are recorded separately and are unretrievable.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Tutoring + mastery bundle vs conventional classsigma2.0 on a researcher-designed 3-week unit testresearcher-designedend of 3-week treatment; no follow-upbusiness-as-usualend-of-treatmentdomain-skill
Mastery-learning arm alonesigma~1.0-1.1 — half of the 2.0 is the testing/feedback component, not tutoringresearcher-designedend of treatmentbusiness-as-usualend-of-treatmentdomain-skill
Aptitude-achievement correlation under tutoringr0.60 -> 0.25 — a restricted-range/ceiling artefact on aligned tests, NOT ability equalizationresearcher-designedend of treatmentnoneend-of-treatmentdomain-skill

Cited by