The Evidence on Teaching

Science misconceptions and conceptual change — can naive intuitions be taught away?

Conceptual-change teaching reliably moves the test and nothing shows it removes the intuition. Corrected for design and publication bias, a taught unit is g≈0.64 and a refutation text g≈0.28.

mixedconf: mediumgc: low

science · ages 618

Effect summary

Two claims that this literature routinely conflates, with different answers. TEACHING THE CONCEPT: works. The largest synthesis (218 studies, 18,051 students) reports g=1.10, falling to 0.93 under bias correction and to 0.64 among the 33 randomised studies; a single-topic meta (69 studies) gives 0.72 among controlled designs; refutation text against ordinary expository text gives 0.41, trim-and-fill corrected to 0.28. No decay is detectable — the days-elapsed slope is 0.00 — but only two comparisons in the whole refutation literature exceed one month. REMOVING THE INTUITION: no. The naive conception survives into expertise as a measurable accuracy and latency cost, and MORE expertise means MORE conflict, not less. The measure-type inflation here is only ~1.3x (researcher-built 1.17 vs pre-existing 0.88); the real discounts are design (quasi 1.23 vs randomised 0.64, ~1.9x), publication bias (journals 0.45 vs dissertations 0.11), and region (Turkey 1.28 vs America 0.68).

Practical takeaway

Name the misconception explicitly, say it is wrong, then give the correct account — that is the whole intervention, it is free, and it beats simply presenting the correct account by about 0.3 SD. Give it high-confidence misconceptions rather than shaky ones. Then keep teaching the concept: budget roughly 12 hours of instruction for a hard one, expect about 0.6 SD on a test of that concept, and do NOT tell yourself the intuition has been deleted. Plan for it to resurface, in your students and in yourself.

Who this applies to

Group size
whole-classindependent
Delivered by
teacherself
Ages studied
618
Dose
Refutation text: a single reading of one short passage that states the misconception, flags it as wrong, and gives the scientific account. Whole-unit conceptual-change teaching: ~11-12 hours of instruction (Hardy's floating-and-sinking unit ran 720 minutes); +0.06 per additional hour of instruction, from an intercept of 0.59.
Cost
free
Moves
domain-skill
Needs first
The learner must already hold the misconception, must be able to read the text unaided, and the misconception must be nameable in one sentence. Effects concentrate on HIGH-confidence misconceptions; low-confidence ones gain nothing from refutation over ordinary text.
Not for
Any claim that the naive conception has been removed — it has not. Any expectation of transfer beyond the specific refuted proposition. Adding peer argumentation on top of a refutation text buys nothing. And there is no evidence at all past about two months in children.

Verdict

The brief this topic was built to answer was: does conceptual-change instruction move the concept, or does it move the test? The honest answer is that it moves the test, that this is genuinely worth something, and that the field's own framing — "conceptual change", "restructuring", "replacing the naive theory" — is not supported by the evidence it cites.

mixed, because the two halves come apart:

  • Score gains on the taught concept: well supported. A taught unit is worth roughly 0.6–0.7 SD once you restrict to randomised designs and correct for publication bias. A refutation text — a passage that names the wrong idea, says it is wrong, and gives the right one — is worth about 0.28–0.41 SD against an ordinary expository text saying the same true things. That is a large return on a free intervention.
  • Replacement of the intuition: not supported, and actively contradicted. The naive conception is still detectable in adults who have spent a career teaching the correct one, as a slower and less accurate response to exactly the statements where the two theories disagree. And the effect grows with expertise.

What the evidence shows

Source Design Grade Key effect
Pacacı 2024 meta, 218 studies / 18,051 students C g=1.10 raw → 0.93 bias-adjusted → 0.64 among the 33 randomised studies
Schwichow 2024 single-topic meta, 69 studies / 191 ES / 10,532 C g=0.85 all designs; 0.72 treatment-control; days-elapsed slope 0.00
Schroeder & Kucera 2022 meta, 44 comparisons / 3,869 C Refutation vs expository 0.41 → 0.28 trim-and-fill
Hardy 2006 quasi-experiment, 161 third-graders B High instructional support beats low at one-year follow-up — the only 1-year test in the set
van Loon 2015 RCT, 114 eighth-graders C Refutation corrects high-confidence misconceptions; no benefit for low-confidence ones
Hake 1998 62 courses / 6,542 students (14 high-school courses, N=1,113) D Normalized FCI gain 0.23 traditional vs 0.48 interactive; high-school interactive 0.55
Tippett 2010 review, two decades D Advantage often survives 1 week–2 months, but more readers revert at delay than at immediate test
Zengilowski 2021 critique D Single-topic designs, uncontrolled testing effects, durability essentially unestablished
Vosniadou & Brewer 1992 60 children, grades 1/3/5 D Only 23 of 60 held the spherical-earth model; 24 held a "synthetic" model
Panagiotaki 2009 critique, 127 six- and seven-year-olds C The synthetic models are largely artefacts of the elicitation method — adults given the same protocol also draw flat and hollow earths

Where the inflation actually is — and it is not measure type

This is the most useful methodological result in the tranche, because it runs against the prior. Pacacı's meta reports both the aligned and the independent version of the same estimate:

Contrast, inside one meta-analysis Aligned Independent Ratio
Instrument provenance researcher-developed 1.17 (k=155, 71%) pre-existing published test 0.88 (k=40) 1.33× (Q=3.76, p=0.035)
Outcome target misconception-targeted 1.12 (k=192) general achievement test 0.87 (k=16) 1.29×
Schwichow's version of the same split preconception-targeted 0.75 not preconception-targeted 0.90 0.83× — reversed

So measure-type inflation in this literature is about 1.3×, not 2× and nowhere near the writing tranche's 5× — and within a single well-defined topic it reverses. The discounts that do bite are elsewhere and they are bigger:

  • Design — 1.9×. True experiments g=0.64 [0.49, 0.79] (k=33) versus quasi-experiments 1.23 (k=162), Q(3)=23.72, p<0.001. Schwichow finds the same shape: 0.72 controlled versus 0.89 uncontrolled pre-post.
  • Publication bias — up to 4×. The only significant moderator in Schroeder & Kucera is publication type: journal articles 0.45 (k=36) versus dissertations 0.11, not significant (k=6). Pacacı's Egger test is t(216)=8.21, p<0.001.
  • Region — 1.9×. Turkey g=1.28 (k=147, two-thirds of the corpus) versus America 0.68 and Europe 0.66, explaining 23.5% of between-study variance — the largest single moderator in the field's largest meta-analysis.

A founder reading a conceptual-change effect size should discount for design and provenance of the study, not primarily for who wrote the test.

Suppression, not replacement

The strongest evidence on this page is recorded as an excluded source, because its participants are adults and this archive covers ages 4–18. It should still govern how the topic is read.

Shtulman & Valcarcel gave 150 undergraduates 200 statements across ten domains — astronomy, evolution, genetics, mechanics, thermodynamics and five others — half of which have the same truth value under the naive and the scientific theory ("the moon revolves around the Earth") and half of which reverse ("the Earth revolves around the sun"). On the reversing statements participants were less accurate (F(1,149)=1582.49, p<.001, in all ten domains) and slower by about 415–428 ms, in 43 of 50 concepts. The decisive detail is the direction of the moderation: greater domain expertise produced MORE conflict, not less — top-quartile participants were 465 ms slower on inconsistent items than bottom-quartile ones. Familiarity and syntactic-difficulty explanations were tested and run the wrong way.

Two further excluded sources supply the mechanism: an fMRI study in which experts, not novices, recruit inhibitory regions when answering a classic electricity item; and a chronometry study of 17 chemistry professors with a mean of fourteen years' teaching experience, who took 3,631 ms versus 3,058 ms on conflicting items (z=7.80, p<0.001, r=0.40) and got 31% of the paired items wrong, in their own subject.

Whatever conceptual-change instruction does, it does not delete the intuition. It builds a competing representation that has to win a race every time the question is asked.

Is this even a science finding?

Largely, no — it is a text-structure finding that happens to be studied mostly on science content. Schroeder & Kucera's domain moderator: science 0.45 (k=33), mathematics 0.32 (k=5), social science 0.31 (k=6), Q-between = 1.09, p=0.58. Pacacı finds that which conceptual-change strategy was used — cognitive conflict, bridging analogies, ontological reclassification — does not moderate at all (Q=0.95, p=0.62). And an excluded undergraduate RCT found that adding structured peer argumentation on top of a refutation text added nothing.

That matters practically: the refutation move is a cheap, general writing and explaining technique, not a science pedagogy that needs buying.

Hereditarian-lens assessment

Risk: low. The verdict's positive half rests on randomised and controlled contrasts (Pacacı's 33-study randomised subset, van Loon's RCT, Schwichow's 23 treatment-control studies), where prior ability is balanced by design. The verdict's negative half — that intuitions persist — is a within-subject latency comparison, which nets out every stable individual difference by construction.

The one place heredity is worth flagging is Vosniadou & Brewer's "mental models" tradition, where the inference from a child's answers to an internal theory is drawn from a 60-child interview study. Panagiotaki's re-run with disambiguated instructions recovered substantially more scientific understanding, and the companion study found adults given the original protocol also drew flat and hollow earths. Children's misconceptions are real, but a fair part of the classic taxonomy is a measurement artefact rather than a developmental fact. Nothing here bears on g.

Boundaries & what critics say

  • The durability evidence barely exists. Schroeder & Kucera's ">1 month" cell contains two comparisons. Schwichow's days-elapsed slope is 0.00 (p=0.11) over 170 effect sizes, which is reassuring but is a between-study regression, not a follow-up. Tippett's longest child delays are six weeks and two months. Pacacı — 218 studies — ran no retention analysis at all and says so. Only Hardy reaches a year, on eight-year-olds, and its magnitudes were not obtainable.
  • Group means hide individual reversion. Tippett reports that where delayed tests exist, more readers have gone back to the misconception at delay than at immediate test even while the mean holds.
  • Uncontrolled testing effects. Almost every design pre-tests the misconception, states it, refutes it, and post-tests it. Zengilowski names this as unresolved: some of the effect is the pre-test.
  • Almost none of this is on children. Thirty of Schroeder & Kucera's 44 comparisons are post-secondary; K-12 supplies ten. Point estimates run against the adult corpus (primary 0.57, middle 0.71, secondary 0.49, post-secondary 0.33) but the moderator is not significant (p=0.25). Pacacı is 46% high school and 25% undergraduate; Schwichow is the one where 42% of participants are 8–12, and there age does not moderate at all.
  • Hake's famous 0.23-versus-0.48 is not a randomised comparison. Courses self-selected into interactive engagement, the outcome is a community-built concept inventory rather than an independent achievement test, and only 14 of 62 courses are in scope for school-age students.
  • A live conflict with practical work. Schwichow's meta finds hands-on experiments (0.96, 53 studies) beating virtual ones (0.52, 17 studies), F(1,189)=8.02, p<0.001 — while three K-12 randomised trials find physical and virtual materials exactly equivalent. The randomised within-study contrast should win over the between-study moderator, which compares different studies done by different people in different decades. But it is a real disagreement and it is not resolved.

Practical guidance

  • Say the wrong idea out loud, label it wrong, then give the right one. That is the entire refutation move and it is worth ~0.3 SD over presenting only the correct account. It costs nothing.
  • Target confident wrong beliefs. The one RCT on eighth-graders found the benefit concentrated entirely in high-confidence misconceptions.
  • Budget hours, not tricks. +0.06 per additional hour of instruction from an intercept of 0.59: a hard misconception is a ten-to-twelve-hour teaching job, not a paragraph.
  • Do not buy a branded conceptual-change programme. Strategy type does not moderate the effect, and adding argumentation on top adds nothing.
  • Re-test later, and expect resurfacing. Plan maintenance retrieval on the concepts you most want to hold, because the group mean staying flat is consistent with individual students sliding back.
  • Discount hard for design and origin, not for who wrote the test. Randomised beats quasi by ~1.9×, published beats unpublished by up to 4×, and a single national literature supplies two-thirds of the corpus at nearly twice everyone else's effect size.

Open questions

  • Nothing in this literature uses a standardized science achievement test. Schroeder & Kucera's "measure type" moderator is item format, not provenance. Pacacı's "pre-existing instrument" category is published concept inventories, not norm-referenced tests. The obvious study — teach the misconceptions, then measure on a state science test — has not been run.
  • Nothing measures durability past two months in children except one 2006 study whose effect sizes could not be obtained.
  • Guzzetti's 1993 comparative meta-analysis, the founding document of this literature, is behind JSTOR and its effect sizes have never entered this archive. It is recorded as unresolved.
  • The suppression finding has never been tested developmentally. If experts show more conflict than novices, the interesting question is what the curve looks like between ages 6 and 18 — and whether an instructional method exists that reduces the interference rather than just raising the score.

Evidence (10 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore