Science misconceptions and conceptual change — can naive intuitions be taught away?
Conceptual-change teaching reliably moves the test and nothing shows it removes the intuition. Corrected for design and publication bias, a taught unit is g≈0.64 and a refutation text g≈0.28.
mixedconf: mediumgc: lowscience · ages 6–18
Two claims that this literature routinely conflates, with different answers. TEACHING THE CONCEPT: works. The largest synthesis (218 studies, 18,051 students) reports g=1.10, falling to 0.93 under bias correction and to 0.64 among the 33 randomised studies; a single-topic meta (69 studies) gives 0.72 among controlled designs; refutation text against ordinary expository text gives 0.41, trim-and-fill corrected to 0.28. No decay is detectable — the days-elapsed slope is 0.00 — but only two comparisons in the whole refutation literature exceed one month. REMOVING THE INTUITION: no. The naive conception survives into expertise as a measurable accuracy and latency cost, and MORE expertise means MORE conflict, not less. The measure-type inflation here is only ~1.3x (researcher-built 1.17 vs pre-existing 0.88); the real discounts are design (quasi 1.23 vs randomised 0.64, ~1.9x), publication bias (journals 0.45 vs dissertations 0.11), and region (Turkey 1.28 vs America 0.68).
Name the misconception explicitly, say it is wrong, then give the correct account — that is the whole intervention, it is free, and it beats simply presenting the correct account by about 0.3 SD. Give it high-confidence misconceptions rather than shaky ones. Then keep teaching the concept: budget roughly 12 hours of instruction for a hard one, expect about 0.6 SD on a test of that concept, and do NOT tell yourself the intuition has been deleted. Plan for it to resurface, in your students and in yourself.
Who this applies to
Verdict
The brief this topic was built to answer was: does conceptual-change instruction move the concept, or does it move the test? The honest answer is that it moves the test, that this is genuinely worth something, and that the field's own framing — "conceptual change", "restructuring", "replacing the naive theory" — is not supported by the evidence it cites.
mixed, because the two halves come apart:
- Score gains on the taught concept: well supported. A taught unit is worth roughly 0.6–0.7 SD once you restrict to randomised designs and correct for publication bias. A refutation text — a passage that names the wrong idea, says it is wrong, and gives the right one — is worth about 0.28–0.41 SD against an ordinary expository text saying the same true things. That is a large return on a free intervention.
- Replacement of the intuition: not supported, and actively contradicted. The naive conception is still detectable in adults who have spent a career teaching the correct one, as a slower and less accurate response to exactly the statements where the two theories disagree. And the effect grows with expertise.
What the evidence shows
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Pacacı 2024 | meta, 218 studies / 18,051 students | C | g=1.10 raw → 0.93 bias-adjusted → 0.64 among the 33 randomised studies |
| Schwichow 2024 | single-topic meta, 69 studies / 191 ES / 10,532 | C | g=0.85 all designs; 0.72 treatment-control; days-elapsed slope 0.00 |
| Schroeder & Kucera 2022 | meta, 44 comparisons / 3,869 | C | Refutation vs expository 0.41 → 0.28 trim-and-fill |
| Hardy 2006 | quasi-experiment, 161 third-graders | B | High instructional support beats low at one-year follow-up — the only 1-year test in the set |
| van Loon 2015 | RCT, 114 eighth-graders | C | Refutation corrects high-confidence misconceptions; no benefit for low-confidence ones |
| Hake 1998 | 62 courses / 6,542 students (14 high-school courses, N=1,113) | D | Normalized FCI gain 0.23 traditional vs 0.48 interactive; high-school interactive 0.55 |
| Tippett 2010 | review, two decades | D | Advantage often survives 1 week–2 months, but more readers revert at delay than at immediate test |
| Zengilowski 2021 | critique | D | Single-topic designs, uncontrolled testing effects, durability essentially unestablished |
| Vosniadou & Brewer 1992 | 60 children, grades 1/3/5 | D | Only 23 of 60 held the spherical-earth model; 24 held a "synthetic" model |
| Panagiotaki 2009 | critique, 127 six- and seven-year-olds | C | The synthetic models are largely artefacts of the elicitation method — adults given the same protocol also draw flat and hollow earths |
Where the inflation actually is — and it is not measure type
This is the most useful methodological result in the tranche, because it runs against the prior. Pacacı's meta reports both the aligned and the independent version of the same estimate:
| Contrast, inside one meta-analysis | Aligned | Independent | Ratio |
|---|---|---|---|
| Instrument provenance | researcher-developed 1.17 (k=155, 71%) | pre-existing published test 0.88 (k=40) | 1.33× (Q=3.76, p=0.035) |
| Outcome target | misconception-targeted 1.12 (k=192) | general achievement test 0.87 (k=16) | 1.29× |
| Schwichow's version of the same split | preconception-targeted 0.75 | not preconception-targeted 0.90 | 0.83× — reversed |
So measure-type inflation in this literature is about 1.3×, not 2× and nowhere near the writing tranche's 5× — and within a single well-defined topic it reverses. The discounts that do bite are elsewhere and they are bigger:
- Design — 1.9×. True experiments g=0.64 [0.49, 0.79] (k=33) versus quasi-experiments 1.23 (k=162), Q(3)=23.72, p<0.001. Schwichow finds the same shape: 0.72 controlled versus 0.89 uncontrolled pre-post.
- Publication bias — up to 4×. The only significant moderator in Schroeder & Kucera is publication type: journal articles 0.45 (k=36) versus dissertations 0.11, not significant (k=6). Pacacı's Egger test is t(216)=8.21, p<0.001.
- Region — 1.9×. Turkey g=1.28 (k=147, two-thirds of the corpus) versus America 0.68 and Europe 0.66, explaining 23.5% of between-study variance — the largest single moderator in the field's largest meta-analysis.
A founder reading a conceptual-change effect size should discount for design and provenance of the study, not primarily for who wrote the test.
Suppression, not replacement
The strongest evidence on this page is recorded as an excluded source, because its participants are adults and this archive covers ages 4–18. It should still govern how the topic is read.
Shtulman & Valcarcel gave 150 undergraduates 200 statements across ten domains — astronomy, evolution, genetics, mechanics, thermodynamics and five others — half of which have the same truth value under the naive and the scientific theory ("the moon revolves around the Earth") and half of which reverse ("the Earth revolves around the sun"). On the reversing statements participants were less accurate (F(1,149)=1582.49, p<.001, in all ten domains) and slower by about 415–428 ms, in 43 of 50 concepts. The decisive detail is the direction of the moderation: greater domain expertise produced MORE conflict, not less — top-quartile participants were 465 ms slower on inconsistent items than bottom-quartile ones. Familiarity and syntactic-difficulty explanations were tested and run the wrong way.
Two further excluded sources supply the mechanism: an fMRI study in which experts, not novices, recruit inhibitory regions when answering a classic electricity item; and a chronometry study of 17 chemistry professors with a mean of fourteen years' teaching experience, who took 3,631 ms versus 3,058 ms on conflicting items (z=7.80, p<0.001, r=0.40) and got 31% of the paired items wrong, in their own subject.
Whatever conceptual-change instruction does, it does not delete the intuition. It builds a competing representation that has to win a race every time the question is asked.
Is this even a science finding?
Largely, no — it is a text-structure finding that happens to be studied mostly on science content. Schroeder & Kucera's domain moderator: science 0.45 (k=33), mathematics 0.32 (k=5), social science 0.31 (k=6), Q-between = 1.09, p=0.58. Pacacı finds that which conceptual-change strategy was used — cognitive conflict, bridging analogies, ontological reclassification — does not moderate at all (Q=0.95, p=0.62). And an excluded undergraduate RCT found that adding structured peer argumentation on top of a refutation text added nothing.
That matters practically: the refutation move is a cheap, general writing and explaining technique, not a science pedagogy that needs buying.
Hereditarian-lens assessment
Risk: low. The verdict's positive half rests on randomised and controlled contrasts (Pacacı's 33-study randomised subset, van Loon's RCT, Schwichow's 23 treatment-control studies), where prior ability is balanced by design. The verdict's negative half — that intuitions persist — is a within-subject latency comparison, which nets out every stable individual difference by construction.
The one place heredity is worth flagging is Vosniadou & Brewer's "mental models" tradition, where the inference from a child's answers to an internal theory is drawn from a 60-child interview study. Panagiotaki's re-run with disambiguated instructions recovered substantially more scientific understanding, and the companion study found adults given the original protocol also drew flat and hollow earths. Children's misconceptions are real, but a fair part of the classic taxonomy is a measurement artefact rather than a developmental fact. Nothing here bears on g.
Boundaries & what critics say
- The durability evidence barely exists. Schroeder & Kucera's ">1 month" cell contains two comparisons. Schwichow's days-elapsed slope is 0.00 (p=0.11) over 170 effect sizes, which is reassuring but is a between-study regression, not a follow-up. Tippett's longest child delays are six weeks and two months. Pacacı — 218 studies — ran no retention analysis at all and says so. Only Hardy reaches a year, on eight-year-olds, and its magnitudes were not obtainable.
- Group means hide individual reversion. Tippett reports that where delayed tests exist, more readers have gone back to the misconception at delay than at immediate test even while the mean holds.
- Uncontrolled testing effects. Almost every design pre-tests the misconception, states it, refutes it, and post-tests it. Zengilowski names this as unresolved: some of the effect is the pre-test.
- Almost none of this is on children. Thirty of Schroeder & Kucera's 44 comparisons are post-secondary; K-12 supplies ten. Point estimates run against the adult corpus (primary 0.57, middle 0.71, secondary 0.49, post-secondary 0.33) but the moderator is not significant (p=0.25). Pacacı is 46% high school and 25% undergraduate; Schwichow is the one where 42% of participants are 8–12, and there age does not moderate at all.
- Hake's famous 0.23-versus-0.48 is not a randomised comparison. Courses self-selected into interactive engagement, the outcome is a community-built concept inventory rather than an independent achievement test, and only 14 of 62 courses are in scope for school-age students.
- A live conflict with practical work. Schwichow's meta finds hands-on experiments (0.96, 53 studies) beating virtual ones (0.52, 17 studies), F(1,189)=8.02, p<0.001 — while three K-12 randomised trials find physical and virtual materials exactly equivalent. The randomised within-study contrast should win over the between-study moderator, which compares different studies done by different people in different decades. But it is a real disagreement and it is not resolved.
Practical guidance
- Say the wrong idea out loud, label it wrong, then give the right one. That is the entire refutation move and it is worth ~0.3 SD over presenting only the correct account. It costs nothing.
- Target confident wrong beliefs. The one RCT on eighth-graders found the benefit concentrated entirely in high-confidence misconceptions.
- Budget hours, not tricks. +0.06 per additional hour of instruction from an intercept of 0.59: a hard misconception is a ten-to-twelve-hour teaching job, not a paragraph.
- Do not buy a branded conceptual-change programme. Strategy type does not moderate the effect, and adding argumentation on top adds nothing.
- Re-test later, and expect resurfacing. Plan maintenance retrieval on the concepts you most want to hold, because the group mean staying flat is consistent with individual students sliding back.
- Discount hard for design and origin, not for who wrote the test. Randomised beats quasi by ~1.9×, published beats unpublished by up to 4×, and a single national literature supplies two-thirds of the corpus at nearly twice everyone else's effect size.
Open questions
- Nothing in this literature uses a standardized science achievement test. Schroeder & Kucera's "measure type" moderator is item format, not provenance. Pacacı's "pre-existing instrument" category is published concept inventories, not norm-referenced tests. The obvious study — teach the misconceptions, then measure on a state science test — has not been run.
- Nothing measures durability past two months in children except one 2006 study whose effect sizes could not be obtained.
- Guzzetti's 1993 comparative meta-analysis, the founding document of this literature, is behind JSTOR and its effect sizes have never entered this archive. It is recorded as unresolved.
- The suppression finding has never been tested developmentally. If experts show more conflict than novices, the interesting question is what the curve looks like between ages 6 and 18 — and whether an instructional method exists that reduces the interference rather than just raising the score.
- grade CEffectiveness of conceptual change strategies in science education: A meta‐analysisPacaci C, Ustun U, Ozdemir OF · 2024 · meta-analysis
- grade CRefutation Text Facilitates Learning: a Meta-Analysis of Between-Subjects ExperimentsSchroeder NL, Kucera AC · 2022 · meta-analysis
- grade CTeaching and learning floating and sinking: A meta‐analysisSchwichow M, Zoupidis A · 2024 · meta-analysis
- grade DREFUTATION TEXT IN SCIENCE EDUCATION: A REVIEW OF TWO DECADES OF RESEARCHTippett CD · 2010 · review
- grade DA critical review of the refutation text literature: Methodological confounds, theoretical problems, and possible solutionsZengilowski A, Schuetze BA, Nash BL, Schallert DL · 2021 · critique
- grade BEffects of instructional support within constructivist learning environments for elementary school students' understanding of "floating and sinking."Hardy I, Jonen A, Möller K, Stern E · 2006 · quasi-experiment
- grade CRefutations in science texts lead to hypercorrection of misconceptions held with high confidencevan Loon MH, Dunlosky J, van Gog T, van Merriënboer JJG, de Bruin ABH · 2015 · rct
- grade DInteractive-engagement versus traditional methods: A six-thousand-student survey of mechanics test data for introductory physics coursesHake RR · 1998 · quasi-experiment
- grade DMental models of the earth: A study of conceptual change in childhoodVosniadou S, Brewer WF · 1992 · quasi-experiment
- grade CMental models and other misconceptions in children’s understanding of the earthPanagiotaki G, Nobes G, Potton A · 2009 · critique
Related decisions
- Laboratory and practical work in school sciencemixedconf: mediumgc: low