Teaching and learning floating and sinking: A meta‐analysis
Schwichow M, Zoupidis A · 2024
grade Cmeta-analysisindependentnot-applicablenumbers spot-checked
Sample
69 intervention studies, 191 effect sizes, 10,532 participants
Population
Learners taught floating and sinking (density/buoyancy) between 1977 and 2021, preschool to undergraduate; 42% of samples (4,845 participants) had a mean age between 8 and 12. Studies in English, German and Greek.
Design
A single-misconception meta-analysis, and therefore unusually informative - floating and sinking is the canonical naive-physics misconception (things float because they are light / have air inside) and this pools five decades of attempts to teach it away. Three-level hierarchical meta-regression handles multiple effect sizes per study. The critical design fact is that 61 of 84 design-codable studies are PRE-POST with no control group; only 23 studies (48 effect sizes) have a treatment-control comparison, and those give a smaller estimate. Read in full. Its unique contribution to the durability question is that it models the delay between intervention and test as a CONTINUOUS moderator, which nobody else in this literature does.
Key findings
Mean effect of teaching floating and sinking g = 0.85, 95% CI [0.71, 0.99], across 191 effect sizes from 69 studies and 10,532 participants. THE DESIGN DISCOUNT: pre-post designs g = 0.89 (61 studies, 143 effect sizes) versus treatment-control designs g = 0.72 (23 studies, 48 effect sizes), F(1,189) = 6.46, p < .01, difference g = 0.17 - and in the multiple moderator model the design effect disappears entirely, leaving only duration and hands-on. THE DURABILITY RESULT: the delay between intervention and test, modelled continuously in days across 60 studies and 170 effect sizes, has a regression coefficient of 0.00 (95% CI 0.00 to 0.00), p = 0.11 - no measurable decay over the delays actually used. THE MEASURE RESULT IS THE MOST USEFUL THING HERE AND IT CUTS AGAINST THE STANDARD INFLATION STORY: test format did not moderate at all (multiple choice 0.95, running experiments 1.06, interview 0.87, open-ended 0.70, mixed 0.66; F(4,186) = 1.74, p = 0.14), and tests explicitly built around students preconceptions gave a SMALLER effect (g = 0.75) than tests that ignored preconceptions (g = 0.90), F(1,189) = 0.93, p = 0.34 - i.e. within a single well-defined topic, misconception-targeted instruments are not the ones producing the big numbers. What does moderate: hands-on experiments g = 0.96 versus virtual experiments g = 0.52, F(1,189) = 8.02, p < .001; and each additional hour of instruction adds 0.06 (95% CI 0.02-0.09). Age does NOT moderate (beta = 0.01, p = 0.61 across ages studied) - the concept is as teachable at 8 as at 18. No publication bias (Egger F(1,189) = 0.15, p = 0.70). Dose warning the authors make explicit: the well-known Hardy et al. 2006 study ran 11 periods of 45 minutes, "about half of all science lessons in a semester".
Genetic confound
Low for the treatment-control subset; the pre-post majority has no arm structure at all and its estimate is a within-child gain, not a causal contrast.
Replication notes
First meta-analysis on this topic. Its treatment-control estimate (0.72) sits between Schroeder & Kucera 2022 (0.41, text only) and Pacaci et al. 2024 (1.10 raw, 0.64 randomised only), which is what a whole-unit intervention on one hard concept should look like.
DOI / URL
10.1002/tea.21909
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Teaching floating and sinking, all designs | Hedges g | 0.85 (95% CI 0.71-0.99), 191 effect sizes, 69 studies | researcher-designed | post-instruction | unclear | end-of-treatment | domain-skill |
| Restricted to treatment-control-group designs | Hedges g | 0.72 (95% CI 0.55-0.89), 23 studies, 48 effect sizes | researcher-designed | post-instruction | business-as-usual | end-of-treatment | domain-skill |
| Restricted to uncontrolled pre-post designs | Hedges g | 0.89 (95% CI 0.75-1.03), 61 studies, 143 effect sizes; F(1,189) = 6.46, p < .01 | researcher-designed | post-instruction | none | end-of-treatment | domain-skill |
| DURABILITY - effect of days elapsed between intervention and test | meta-regression slope per day | 0.00 (95% CI 0.00-0.00), p = 0.11; 60 studies, 170 effect sizes - no detectable decay | researcher-designed | continuous delay moderator | unclear | under-1yr | domain-skill |
| PAIRED MEASURE CONTRAST - tests built around students' preconceptions vs tests that were not | Hedges g by subgroup | preconception-targeted 0.75 (49 studies) vs not preconception-targeted 0.90 (20 studies); F(1,189) = 0.93, p = 0.34 | mixed | post-instruction | unclear | end-of-treatment | domain-skill |
| Test-format moderator | Hedges g by subgroup | multiple choice 0.95, running experiments 1.06, interview 0.87, open-ended 0.70, MC+open 0.66; F(4,186) = 1.74, p = 0.14 (not significant) | mixed | post-instruction | unclear | end-of-treatment | domain-skill |
| Hands-on vs virtual experiments | Hedges g by subgroup | hands-on 0.96 (53 studies) vs virtual 0.52 (17 studies); F(1,189) = 8.02, p < .001 | researcher-designed | post-instruction | active-alternative | end-of-treatment | domain-skill |
| Dose - effect per additional hour of instruction | meta-regression slope | +0.06 per 60 min (95% CI 0.02-0.09), p < .001; intercept 0.59 | researcher-designed | post-instruction | unclear | end-of-treatment | domain-skill |
| Student age as moderator | meta-regression slope per year | 0.01 (95% CI -0.04 to 0.02), p = 0.61 - no age effect across the range studied | researcher-designed | post-instruction | unclear | end-of-treatment | domain-skill |
Cited by
- Science misconceptions and conceptual change — can naive intuitions be taught away?mixedconf: mediumgc: low