Experimental evaluations of elementary science programs: A best-evidence synthesis
Slavin, R. E., Lake, C., Hanley, P., & Thurston, A. · 2014
grade Bmeta-analysisindependentnot-applicablenumbers spot-checked
Sample
23 studies meeting inclusion criteria, from a much larger screened pool; elementary (K-6) science programmes.
Population
US and UK elementary science classrooms; whole-programme evaluations of inquiry kits, inquiry-oriented professional development, technology programmes and science-literacy integration.
Design
Best-evidence synthesis with THREE inclusion rules that make it the most useful independent counterweight in this tranche: randomized or matched control groups, a study duration of at least four weeks, and — the decisive one — achievement measures INDEPENDENT OF THE EXPERIMENTAL TREATMENT. That third rule alone eliminates most of the science-curriculum literature. The authors are independent of every programme reviewed. They also quantify the alignment problem across fields: among the maths studies they have reviewed, treatment-inherent measures averaged +0.45 while treatment-independent measures averaged −0.03; among ten reading studies, +0.51 vs +0.06. Read in full text.
Key findings
Once you require independent measures and a control group, the elementary science curriculum literature mostly evaporates. Inquiry programmes built around SCIENCE KITS: weighted ES = +0.02 across 7 studies — a clean null on the hands-on-materials bet. Inquiry-oriented programmes emphasising PROFESSIONAL DEVELOPMENT but not kits: +0.36 across 10 studies. Technology programmes integrating video/computer resources with cooperative learning: +0.42 across 6 small matched studies. The within-study alignment numbers are the sharpest thing here: one included inquiry programme showed +0.21 overall on the end-of-year test but +0.58 on the specific units taught, and within the same end-of-year test the taught topics returned +0.43 (evaporation) and +0.29 (forces) while the REMAINING items returned +0.09 — the whole programme effect lived in the taught items. On science-literacy integration specifically the synthesis reports Science IDEAS (Romance & Vitale, matched, developer-led) at ES = +0.66 on MAT-Science with +0.11 on ITBS-Reading (1992 version: +0.90 MAT-Science, +0.40 ITBS-Reading, n = 51 vs 77), and Cervetti et al. 2012 at +0.65 science understanding / +0.22 vocabulary / +0.40 writing / +0.09 n.s. reading comprehension — flagging that Cervetti's treatment teachers taught more science per week (3.66 vs 3.03 hrs, ES +0.53), so the integration effect is confounded with dose. Founder-facing conclusion: the leverage in elementary science is in what TEACHERS do all year (instructional approach, cooperative learning, science-reading integration), not in buying kits, and not demonstrably in the ordering of content.
Genetic confound
LOW for the pooled estimates — randomized or matched control groups throughout, with independent outcome measures. The residual threats are matched-design selection in the smaller studies and developer leadership of several included programmes.
Replication notes
Its central discipline (independent measures only) is the same rule that produces the archive's repeated efficacy-to-effectiveness decay finding, and it converges with Hwang et al. 2022's measure-type split (comprehension 0.54 researcher-developed vs 0.25 standardized). The science-kit null (+0.02) converges with the archive's existing curriculum-choice verdict that brand-picking among mainstream materials buys ~0-0.05 SD.
DOI / URL
10.1002/tea.21139
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Inquiry-based science programmes using science KITS vs control | d | +0.02 (7 studies) | standardized | end of year or longer | business-as-usual | end-of-treatment | domain-skill |
| Inquiry-oriented programmes emphasising professional development, no kits | d | +0.36 (10 studies) | standardized | end of year | business-as-usual | end-of-treatment | domain-skill |
| Technology programmes with cooperative learning (small matched studies) | d | +0.42 (6 studies) | standardized | end of year | business-as-usual | end-of-treatment | domain-skill |
| Treatment-inherent vs treatment-independent measures (cross-field benchmark) | d | maths +0.45 vs -0.03; reading +0.51 vs +0.06 | mixed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| Within-study alignment split in one included inquiry programme | d | taught units +0.58; taught-topic items +0.43 / +0.29; all remaining items +0.09; overall +0.21 | mixed | end of year | business-as-usual | end-of-treatment | domain-skill |
| Science IDEAS (science-reading integration, developer-led, matched) on STANDARDIZED tests | d | MAT-Science +0.66 to +0.90; ITBS-Reading +0.11 to +0.40 | standardized | end of year | business-as-usual | end-of-treatment | domain-skill |
Cited by
- Inquiry-based science teaching vs explicit and textbook science teachingmixedconf: mediumgc: low
- Laboratory and practical work in school sciencemixedconf: mediumgc: low
- Science content sequencing — coherence, prerequisites, and course orderinsufficientconf: mediumgc: medium