The 2 Sigma Problem: The Search for Methods of Group Instruction as Effective as One-to-One Tutoring
Bloom BS · 1984
grade Dreviewdeveloper-ledfailednumbers spot-checked
Sample
Narrative essay reporting two supervised dissertations (~30 students per arm, grades 4/5/8, 3-week units)
Population
US grades 4, 5 and 8; probability and cartography units chosen for zero prior knowledge.
Design
A narrative essay, not a study: Bloom summarises two dissertations supervised in his own department (Anania 1983; Burke 1983). The 2.0 figure is a compound of researcher-built tests on 3-week novel units, tutoring FUSED with retest-to-criterion mastery, exceptionally trained tutors, ~1 extra hour per week, a 90% criterion applied only in the tutored arm, tutoring REPLACING rather than supplementing class, and control-group SD denominators. Provenance is broken: Bloom's Table 1 claims adaptation from Walberg (1984), who listed tutoring at 0.40; Bloom printed 2.00 with no computed source.
Key findings
The most famous number in education is a compound design artefact. The honest decomposition (von Hippel): ~1.1 sigma from the mastery/testing component plus ~0.9 tutoring-specific on narrow aligned tests — and on standardized measures modern tutoring lands at 0.2-0.35. Bloom's own cited source put tutoring at d = 0.40.
Genetic confound
Bloom read the falling aptitude-achievement correlation as proof that ability differences are alterable; ceiling-compressed distributions on treatment-aligned tests are the simpler explanation. Arlin's growing time-to-mastery ratios contradict the convergence claim, and heritability of achievement RISES with age.
Replication notes
NEVER replicated: 0 of 96 modern tutoring RCTs reach 2.0 (Nickow); 1 of 65 in Cohen 1982 (a 32-student dissertation). Modern honest anchors: 0.3-0.8 efficacy, 0.2-0.35 standardized. The two underlying dissertations are recorded separately and are unretrievable.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Tutoring + mastery bundle vs conventional class | sigma | 2.0 on a researcher-designed 3-week unit test | researcher-designed | end of 3-week treatment; no follow-up | business-as-usual | end-of-treatment | domain-skill |
| Mastery-learning arm alone | sigma | ~1.0-1.1 — half of the 2.0 is the testing/feedback component, not tutoring | researcher-designed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| Aptitude-achievement correlation under tutoring | r | 0.60 -> 0.25 — a restricted-range/ceiling artefact on aligned tests, NOT ability equalization | researcher-designed | end of treatment | none | end-of-treatment | domain-skill |
Cited by
- Mastery learning (teach → test → reteach to criterion → advance)mixedconf: highgc: low
- Tutoring — the honest effect, the Bloom 2-sigma myth, and what survives scalestrong supportconf: highgc: low