The Evidence on Teaching

Effectiveness of conceptual change strategies in science education: A meta‐analysis

Pacaci C, Ustun U, Ozdemir OF · 2024

grade Cmeta-analysisindependentnot-applicablenumbers spot-checked
Sample
218 primary studies, 18,051 students
Population
Science classrooms taught with an explicit conceptual-change strategy. Education level: elementary k = 13 (6%), middle k = 50 (23%), high school k = 101 (46%), undergraduate k = 54 (25%). Subject: physics 42%, chemistry 39%, biology 19%. Region is the single largest moderator - 147 of 218 studies (67%) come from Turkey.
Design
The largest and most methodologically careful synthesis of conceptual-change instruction (cognitive conflict k = 150, cognitive bridging k = 30, ontological category shift k = 9, unclassified k = 29). Read in full. It is the most valuable source in this tranche because it reports the discount variables separately instead of burying them: measure provenance, design rigour, and small-study bias each get their own number, and each of them cuts the headline roughly in half. 74% of primaries are quasi-experimental and 90% used non-random sampling; 81% have samples under 100; mean intervention was 11 course hours. CRITICALLY, THE META-ANALYSIS CONTAINS NO DELAYED-EFFECT ANALYSIS AT ALL - the authors state in the discussion that effectiveness "may differ as measured by the retention test" and call for future work on delayed effects. So the largest synthesis of conceptual-change instruction in existence has nothing to say about whether the change lasts.
Key findings
Unadjusted pooled effect on science achievement Hedges g = 1.10, 95% CI [1.01, 1.19], k = 218; heterogeneity I-squared = 84.8%, prediction interval [0.19, 2.38]. By strategy: cognitive conflict 1.10 (k = 150), cognitive bridging 1.06 (k = 30), ontological category shift 0.88 (k = 9); strategy type does NOT moderate (Q = 0.95, p = 0.62) - the instructional theory you pick makes no difference. Three discounts, all reported by the authors themselves. (1) SMALL-STUDY BIAS IS SEVERE: Egger t(216) = 8.21, two-tailed p < 0.001; trim-and-fill (fixed-random) cuts g to 0.71; a selection model gives 0.92 [0.75, 1.11]; robust Bayesian meta-analysis gives 0.93. (2) DESIGN RIGOUR HALVES IT: true experiments g = 0.64 [0.49, 0.79] (k = 33) against quasi-experiments 1.23 [1.12, 1.33] (k = 162) and "poor" designs 0.84 (k = 23); Q(3) = 23.72, p < 0.001. (3) MEASURE PROVENANCE MOVES IT: researcher-developed tests g = 1.17 [1.07, 1.28] (k = 155, 71% of the corpus) against pre-existing tests 0.88 [0.68, 1.09] (k = 40) and adapted tests 1.01 (k = 23); Q = 3.76, p = 0.035. The same cut from the other direction: misconception-targeted outcome measures g = 1.12 (k = 192, 88%) against GENERAL ACHIEVEMENT tests g = 0.87 [0.60, 1.14] (k = 16). Region dominates everything - Turkey g = 1.28 (k = 147) versus America 0.68 (k = 28), Europe 0.66 (k = 12), Asia 0.79 (k = 23), and region alone explains 23.5% of between-study variance. Age is a weak moderator running the wrong way for the school-age case: elementary 0.96, middle 1.03, high school 1.24, undergraduate 0.93. Question format matters as much as anything substantive: objective items 1.21 versus open-ended 0.75.
Genetic confound
Low as a threat to the causal claim (arms are classes taught differently), but 90% non-random sampling means selection into treatment classes is uncontrolled in most primaries.
Replication notes
First meta-analysis to pool the three conceptual-change strategy families, so nothing to replicate directly; its headline is consistent in direction with Guzzetti et al. 1993 and Schwichow & Zoupidis 2024 and roughly double Schroeder & Kucera 2022 - the gap is design and measure, not substance.
DOI / URL
10.1002/tea.21887

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Conceptual-change strategies vs comparison instruction, unadjustedHedges g1.10 (95% CI 1.01-1.19), k = 218, prediction interval [0.19, 2.38]researcher-designedend of intervention (mean 11 course hours)business-as-usualend-of-treatmentdomain-skill
Same effect corrected for small-study/publication biasHedges g0.71 (trim-and-fill fixed-random); 0.92 (selection model, 95% CI 0.75-1.11); 0.93 (robust Bayesian)researcher-designedend of interventionbusiness-as-usualend-of-treatmentdomain-skill
Restricted to TRUE experiments (randomised)Hedges g0.64 (95% CI 0.49-0.79), k = 33researcher-designedend of interventionbusiness-as-usualend-of-treatmentdomain-skill
Restricted to quasi-experimentsHedges g1.23 (95% CI 1.12-1.33), k = 162; Q(3) = 23.72, p < 0.001 for the design moderatorresearcher-designedend of interventionbusiness-as-usualend-of-treatmentdomain-skill
PAIRED MEASURE CONTRAST - researcher-developed instrumentsHedges g1.17 (95% CI 1.07-1.28), k = 155 (71% of corpus)researcher-designedend of interventionbusiness-as-usualend-of-treatmentdomain-skill
PAIRED MEASURE CONTRAST - pre-existing (already-published) instrumentsHedges g0.88 (95% CI 0.68-1.09), k = 40; adapted tests 1.01 (k = 23); moderator Q = 3.76, p = 0.035standardizedend of interventionbusiness-as-usualend-of-treatmentdomain-skill
PAIRED MEASURE CONTRAST - misconception-targeted outcome vs general achievement testHedges g by subgroupconceptual-change-targeted 1.12 (k = 192, 88%) vs general achievement 0.87 (95% CI 0.60-1.14, k = 16)mixedend of interventionbusiness-as-usualend-of-treatmentdomain-skill
Education-level moderatorHedges g by subgroupelementary 0.96 (k = 13), middle 1.03 (k = 50), high school 1.24 (k = 101), undergraduate 0.93 (k = 54); Q(3) = 9.27, p = 0.026researcher-designedend of interventionbusiness-as-usualend-of-treatmentdomain-skill
Delayed/retention effectsnot analysedno retention-test moderator was computed; authors call for future work on delayed effectsresearcher-designednot measuredbusiness-as-usualuncleardomain-skill

Cited by