The Evidence on Teaching

Interactive-engagement versus traditional methods: A six-thousand-student survey of mechanics test data for introductory physics courses

Hake RR · 1998

grade Dquasi-experimentindependentreplicatednumbers spot-checked
Sample
62 introductory physics courses, 6,542 students - 14 high-school courses (N = 1,113), 16 college courses (N = 597), 32 university courses (N = 4,832)
Population
Introductory mechanics students, mostly US. The high-school subsample (14 courses, 1,113 students, of which 10 courses were interactive-engagement) is the only part inside this archive's 4-18 age band.
Design
The most-cited result in physics education research and it is an OBSERVATIONAL SURVEY, not an experiment. Read in full text. No randomisation, no matching, no control for instructor or intake; courses were classified as interactive-engagement (IE) or traditional (T) by the instructors' own descriptions of their methods, and Hake obtained the data by soliciting it in talks, colloquia and e-mail postings on physics listservs. He says so himself, in the paper: "This mode of data solicitation tends to pre-select results which are biased in favor of outstanding courses which show relatively high gains on the FCI... When relatively low gains are achieved (as they often are) they are sometimes mentioned informally, but they are usually neither published nor communicated except by those who (a) wish to use the results from a traditional course at their institution as a baseline for their own data, or (b) possess unusual scientific objectivity and detachment." The outcome is the average NORMALIZED GAIN (post minus pre, over the maximum possible gain) on the Force Concept Inventory or its predecessor - a community-built conceptual instrument that is NOT authored by the intervention designers and is not aligned to any particular curriculum, which makes it more independent than the typical conceptual-change measure but still not a standardized achievement test with norms. Hake takes the teaching-to-the-test objection seriously and argues against it: 50% of surveyed IE instructors said the FCI post-test did not count toward the course grade at all, and the maximum IE gain of 0.69 is "disappointingly low" for a test being taught to.
Key findings
Fourteen traditional courses (N = 2,084) averaged a normalized gain of 0.23 (SD 0.04); forty-eight interactive-engagement courses (N = 4,458) averaged 0.48 (SD 0.14). The 0.25 difference is 1.8 SD of the IE distribution and 6.2 SD of the traditional distribution, and IE courses averaged 2.1 times the gain of traditional ones. The pattern holds within level: high-school IE courses 0.55 (SD 0.11, 10 courses), college IE 0.48 (SD 0.12, 13 courses), university IE 0.45 (SD 0.15, 25 courses), against traditional courses at all three levels sitting close to 0.23. No IE course reached the "high-g" region (gain >= 0.7); 15% of IE courses (7 courses, N = 717) landed in the low-g region alongside the traditional ones, which Hake attributes to implementation failure. Within-institution comparisons at matched class time on mechanics reproduce the gap (Arizona State 0.23, Cal Poly 0.31, Harvard 0.29, Monroe CC 0.33 and 0.25, Ohio State 0.24). WHAT IT DOES NOT SHOW: that IE causes the gap. Self-selected instructors who volunteer FCI data are not a random sample of instructors, and Hake makes no causal claim beyond "the classroom use of IE methods can increase mechanics-course effectiveness well beyond that obtained in traditional practice". Treat this as the strongest available DESCRIPTIVE evidence that the impetus-theory misconception survives traditional lecture instruction almost intact - a normalized gain of 0.23 means traditional courses close under a quarter of the gap to full conceptual understanding - and as weak evidence about what fixes it.
Genetic confound
Medium-high as an explanation of the between-course gap: intake differs across courses and the correlation of gain with pretest is only +0.02, which Hake offers as evidence that normalized gain is intake-insensitive - a partial but not complete defence, since instructor and institutional selection remain.
Replication notes
The traditional-course baseline near 0.20-0.25 has been reproduced many times in physics education research, and the IE advantage has since been supported by randomised and quasi-experimental active-learning studies (e.g. the Freeman et al. 2014 STEM meta-analysis). The specific magnitude in this paper is inflated by self-selected reporting.
DOI / URL
10.1119/1.18809

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Average normalized FCI/MD gain, traditional lecture coursesnormalized gain <g>0.23 (SD 0.04), 14 courses, N = 2,084mixedpre-course to post-coursenoneend-of-treatmentdomain-skill
Average normalized FCI/MD gain, interactive-engagement coursesnormalized gain <g>0.48 (SD 0.14), 48 courses, N = 4,458; difference 0.25 = 1.8 SD of the IE distributionmixedpre-course to post-coursebusiness-as-usualend-of-treatmentdomain-skill
HIGH-SCHOOL subsample only (the in-scope portion)normalized gain <g>IE high-school courses 0.55 (SD 0.11, 10 courses); traditional courses at all levels ~0.23; 14 HS courses total, N = 1,113mixedpre-course to post-coursebusiness-as-usualend-of-treatmentdomain-skill
Within-institution comparisons at equal mechanics class timedifference in <g>Arizona State 0.23, Cal Poly 0.31, Harvard 0.29, Monroe CC 0.33 (non-calc) and 0.25 (calc), Ohio State 0.24mixedpre-course to post-coursebusiness-as-usualend-of-treatmentdomain-skill

Cited by