Evaluation of Evidence-Based Practices in Online Learning: A Meta-Analysis and Review of Online Learning Studies
Means, B., Toyama, Y., Murphy, R., Bakia, M., & Jones, K. · 2010
grade Cmeta-analysisindependentmixednumbers spot-checked
Sample
1,132 abstracts screened from a systematic search of 1996 - July 2008; 176 studies with experimental or quasi-experimental designs and objective learning outcomes; 99 of those had an online-or-blended vs face-to-face contrast; 45 studies yielding 50 independent effect sizes had enough data for the meta-analysis. Of the 50 effects, 7 (from 5 studies) involved K-12 learners and 43 did not.
Population
OVERWHELMINGLY POST-SECONDARY AND ADULT, which is the single most important fact about this report. The included studies are mostly undergraduates, graduate students and adults in professional training; the most common subject matter is medicine or health care. Just 5 published studies with K-12 learners met criteria across a twelve-year search window, contributing 7 of 50 effects: two from eighth-grade social studies, one from eighth- and ninth-grade Algebra I, two from middle-school Spanish, one from fifth-grade science in Taiwan, and one from elementary-age special education. Nineteen of 49 contrasts ran for less than a month. Any use of this report to justify online instruction for children is an extrapolation from adults and medical trainees, and the report says so in its own abstract.
Design
WHY THIS SOURCE IS IN THE ARCHIVE: it is the origin of a great deal of unjustified confidence about online instruction for schoolchildren, it was published by the US Department of Education and read as a federal endorsement, and its own text contains the refutation of the reading it received. THE K-12 CAVEAT, IN THE REPORT'S WORDS: a systematic search from 1994 through 2006 found NO experimental or controlled quasi-experimental studies comparing online and face-to-face learning for K-12 students with sufficient data to compute an effect size; extending the window to July 2008 found five. The abstract itself warns that "caution is required in generalizing to the K-12 population because the results are derived for the most part from studies in other settings (e.g., medical training, higher education)". Of the 7 K-12 contrasts, 3 favoured a blended condition, 1 significantly favoured face-to-face, and 3 were null; the K-12 mean is not significant. WHAT "ONLINE" MEANS IN THE INCLUDED STUDIES: it is not one thing, and the report never claims it is. Contrasts split roughly evenly between purely online (27-28 effects) and blended online-plus-face-to-face (23), and by learning experience into instructor-directed (8), independent/self-directed (17) and interactive/collaborative (23). THE REPORT'S OWN FATAL CAVEATS, all stated in its executive summary: (1) online and face-to-face conditions "generally differed on multiple dimensions, including the amount of time that learners spent on task", so the advantage "may be the product of aspects of those treatment conditions other than the instructional delivery medium per se"; (2) many studies did not attempt to equate curriculum materials, pedagogy or learning time, and some authors said this was impossible; (3) the only significant methodological moderator was equivalence of curriculum and instruction - when the two conditions were judged identical or nearly so the effect was +0.13, and when they differed on multiple aspects it was +0.40 (Q = 6.85, p<.01). Read together, these say the report measured "online courses tend to come with more time and more materials", not "the online medium teaches better". DESIGN QUALITY: a mixture of randomised and controlled quasi-experimental studies with no formal risk-of-bias weighting; publication bias is not corrected; no funnel plot or PET-PEESE. Grade C - a meta-analysis of mixed-quality designs, aggregating incommensurable treatments across ages, subjects and durations. INDEPENDENCE: prepared by SRI International's Center for Technology in Learning under contract ED-04-CO-0040 to the US Department of Education. No product vendor; the sponsor is a policy department that was at the time actively promoting online learning, which is a mild conflict worth naming. READ NOTE: both the May 2009 first release (ERIC ED505824) and the Revised September 2010 edition were read in full; the numbers recorded here are the 2010 revision's, with 2009 figures given where they differ.
Key findings
Across 50 contrasts the mean effect favoured online/blended conditions by +0.20 SD (p<.001) - and that single number is almost the only thing anyone remembers about this report. Decomposed, it says something quite different - blended instruction beat face-to-face by +0.35 (p<.001), while PURELY online instruction beat face-to-face by +0.05 (p = .46), which the report itself describes as statistically equivalent. Where the online condition left learners working independently - the closest analogue to software substituting for a teacher - the mean effect was +0.05 and not significant, against +0.39 for instructor-directed and +0.25 for collaborative online instruction. And when analysts judged the curriculum and instruction to be identical or nearly identical across arms, the effect fell to +0.13 against +0.40 when the arms differed on multiple dimensions, which is as close as a meta-analysis gets to admitting it measured the extra materials rather than the medium. The K-12 evidence base was essentially nonexistent - a search from 1994 to 2006 found no usable studies at all, extending to 2008 found five, and their pooled effect is not significant.
Genetic confound
Mixed and not assessable at the pooled level. The corpus contains both randomised and quasi-experimental studies; the report tested strength of study design as a moderator and found it non-significant, but did not down-weight confounded designs, so the pooled estimate inherits whatever selection is in the quasi-experimental half. Since the substantive conclusion from this file is a NULL for pure online substitution, confounding is not what is doing the work.
Replication notes
The headline that travelled - "online is better, blended is best" - has not survived contact with randomised trials of delivery mode. Figlio, Rush & Yin (2013) explicitly wrote their experiment as a test of this report and found live lecture modestly ahead; Bettinger et al. (2017) found -0.25 to -0.33 SD from online course-taking at DeVry; Kofoed et al. (2024) found -0.215 SD in a preregistered randomised trial. Note that the report's own final numbers do not conflict with those results - the purely-online-vs-face-to-face contrast in this revision is +0.05, p = .46, i.e. nothing. What failed replication is the popular reading of the report, not its arithmetic. VERSION HISTORY MATTERS AND IS ITSELF THE STORY: the May 2009 first release reported 51 effects, an overall +0.24, and purely online +0.14 at p<.05 - a nominally significant online advantage. The Revised September 2010 edition reports 50 effects from 45 studies, an overall +0.20, and purely online +0.05 at p = .46, stating that "the learning outcomes for students in purely online conditions and those for students in purely face-to-face conditions were statistically equivalent". The one quantitative basis for "online beats face-to-face" was removed by the authors themselves in the version most people cite by date and never read.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| All 50 contrasts - online or blended vs face-to-face instruction (THE HEADLINE THAT TRAVELLED) | g+ (Hedges) | +0.20, p<.001 (2010 revision; the May 2009 release reported +0.24 across 51 contrasts). The report attaches its own warning: conditions differed on multiple dimensions including time on task, so the advantage may not be attributable to the medium. | mixed | end of instruction (19 of 49 contrasts ran under one month) | active-alternative | end-of-treatment | domain-skill |
| PURELY ONLINE vs purely face-to-face (the substitution question) | g+ (Hedges) | +0.05, p = .46 - not significant. The report states that outcomes in purely online and purely face-to-face conditions "were statistically equivalent". In the May 2009 release this same contrast was +0.14 at p<.05; the authors' own revision removed the only quantitative support for "online beats face-to-face". This is the row the archive should carry. | mixed | end of instruction | active-alternative | end-of-treatment | domain-skill |
| BLENDED (online plus face-to-face) vs face-to-face alone | g+ (Hedges) | +0.35, p<.001 - the source of "blended is best". The report notes blended conditions routinely included additional learning time and instructional elements the control group did not receive, so the comparison is not equal-resourced and the effect "should not be attributed to the media, per se". Under the archive's baseline discipline this is closer to more-instruction vs less-instruction than to a delivery-mode contrast. | mixed | end of instruction | business-as-usual | end-of-treatment | domain-skill |
| Independent, self-directed online learning vs instructor-directed or collaborative online learning | g+ (Hedges) | independent/self-directed +0.05 (not significant); collaborative +0.25 (significant); instructor-directed +0.39 (significant); moderator test Q = 6.19, p<.05 (and only p = .13 once the five K-12 studies are removed). The condition that most resembles software teaching a student without a teacher is the condition with no measurable advantage. | mixed | end of instruction | active-alternative | end-of-treatment | domain-skill |
| Equivalence of curriculum and instruction across arms (the moderator that explains the headline) | g+ (Hedges) | +0.13 when analysts judged curriculum and instruction identical or almost identical across the online and face-to-face conditions, against +0.40 when the conditions differed on multiple aspects of instruction; Q = 6.85, p<.01. The only significant methodological moderator of six tested. Hold everything but the medium constant and most of the effect disappears. | mixed | end of instruction | active-alternative | end-of-treatment | domain-skill |
| K-12 learners (the caveat that should have travelled with the headline) | g+ (Hedges) / study count | 7 contrasts from 5 studies; mean effect positive but NOT statistically significant, and the report says the corpus is too small to warrant confidence. Of the 7, three significantly favoured a blended condition, one significantly favoured face-to-face, three were null. A systematic search from 1994 to 2006 found ZERO experimental or controlled quasi-experimental K-12 studies with computable effect sizes; extending to July 2008 found five. There was, in 2010, effectively no evidence about online instruction for schoolchildren. | mixed | end of instruction | active-alternative | end-of-treatment | domain-skill |
| Time on task (2009 release only; superseded as a significant moderator in the 2010 revision) | g+ (Hedges) | +0.46 in studies where online learners spent MORE time on task, against +0.19 where the face-to-face learners spent as much or more (Q = 3.88, p<.05; p = .06 with the K-12 contrasts removed). Recorded because it is the clearest statement in the report that extra dose, not delivery medium, is what the meta-analysis is measuring. | mixed | end of instruction | active-alternative | end-of-treatment | domain-skill |