The Evidence on Teaching

Science content sequencing — coherence, prerequisites, and course order

No trial has ever varied the sequence of science content and measured the result. The one quasi-experimental test of course order found nothing, and depth-over-breadth is a retrospective survey worth 0.08-0.13 SD.

insufficientconf: mediumgc: medium

science · ages 518

Effect summary

Sequencing is the part of science education with the most confident expert consensus and the least causal evidence. The national documents that set US science sequence — the NRC Framework and NGSS — contain no trial evidence for their sequencing claims, and the field's own 2009 audit called the question of whether progression-based sequencing improves learning 'among the most hypothetical of questions'; the psychometric premise that a student sits at one progression level frequently fails. The one quasi-experimental test of course ORDER (physics-first vs biology-first, one Texas district, standardized end-of-course exams) found 'little or no impact'. The depth-over-breadth result — students whose high-school course spent a month or more on one topic earn higher college science grades — is retrospective self-report on self-selected college samples and is worth 0.08-0.13 SD. The trustworthy finding in that survey work is a NULL: out-of-discipline high-school science does not predict college performance in biology, chemistry or physics, which is evidence against cross-discipline prerequisite arguments. What DOES have causal support is not sequencing at all — it is teacher professional development over science kits (+0.36 vs +0.02 on independent measures) and running science as a content-rich block inside the literacy hour (+0.25 on standardized comprehension).

Practical takeaway

Sequence is a freeroll, not a lever. Pick a coherent, knowledge-rich sequence because it is cheap and defensible, not because evidence says order matters — it does not say that, and the one test of course order found nothing. Teach fewer topics for longer (weak, correlational, d≈0.1, but unanimous in direction). Then spend the attention on what does have causal evidence: sustained teacher professional development rather than science kits (+0.36 vs +0.02 on independent measures), and protecting science time by running it as a reading-and-writing-heavy content block inside the literacy hour.

Who this applies to

Group size
whole-class
Delivered by
teacheradministrator
Ages studied
518
Dose
Depth threshold in the survey evidence: at least one month on a single topic, at the cost of dropping others. Science-literacy integration in the RCTs: the science block replaces part of the literacy block, ~3.7 vs 3.0 hours of science per week.
Cost
low
Moves
domain-skill
Needs first
Protected science instructional time and materials. The causal evidence is elementary and middle school only; nothing here supports any prerequisite claim about high-school course order. What the survey evidence rules OUT is a cross-discipline science prerequisite — out-of-discipline high-school science does not predict college performance in biology, chemistry or physics.
Not for
Any claim that a particular ORDER of science courses matters — that has been tested once, quasi-experimentally, and returned a null on standardized end-of-course exams in science and in maths. Also not for buying a science-kit programme: with a control group and an independent measure, kits return +0.02 across seven studies. And not for expecting integrated science-and-literacy instruction to move general reading comprehension by much: ~0.25 on standardized measures at best, and the developers' own comprehension measure returned +0.09, not significant.

Verdict

insufficient, and the gap is specific: nobody has ever run an experiment that varies the sequence of science content. Not the order of the disciplines, not the order of topics inside a discipline, not spiral against mastery. The literature that appears to be about sequencing is really about three other things — curriculum quality, teacher professional development, and how much time a topic gets — and it is worth separating them, because two of the three have real causal evidence and the sequencing claim has none.

What the field actually has:

  1. Expert consensus, stated with great confidence. The NRC Framework and NGSS restructured US science around "fewer, deeper core ideas" and K-12 learning progressions. Those documents contain no trial evidence for their sequencing claims; the argument is from disciplinary logic and from descriptive TIMSS curriculum comparisons.
  2. A descriptive comparison that everybody quotes. A typical US grade-8 maths textbook covers ~35 topics against ~7 in the Japanese equivalent, and high-performing countries' standards form an "upper triangular" matrix — topics enter, are taught for a few grades, and exit — while US standards add topics and never drop them. This is a real and striking description of curricula. It is not a study of student learning.
  3. One quasi-experimental test of course order, which found nothing.
  4. A retrospective survey finding that is genuinely interesting and genuinely confounded.

What the evidence shows

Source Design Grade Key effect
Slavin 2014 best-evidence synthesis, 23 studies, independent measures required B Science kits +0.02 (7 studies); inquiry programmes emphasising PD +0.36 (10 studies); technology + cooperative learning +0.42 (6)
Hwang 2022 meta, 35 studies / 13,289 students, K-5 C Comprehension 0.54 researcher-developed vs 0.25 standardized; content knowledge 0.94 vs 0.68 (n.s.)
Cervetti 2012 RCT, 94 grade-4 classrooms / ~2,050 students, developer-led C Science understanding 0.65, writing 0.40, vocabulary 0.22; reading comprehension 0.09, n.s.; treatment got 0.6 more science hours/week
Connor 2017 RCT, 418 students / 40 classrooms, developer-led C Content knowledge 2.10–2.27 (researcher-designed); no standardized literacy effect in K-3, positive only at grade 4
Murphy 2024 RCT, 29 middle schools / 1,953 grade-7 students, independent C NGSS physical-science curriculum g=0.40 on a researcher-built test; 38% of schools left the study
Mary 2015 quasi-experiment, one Texas district offering both sequences D Physics-first vs biology-first: "little or no impact" on end-of-course exams; minimal effect on Algebra I / Geometry
Sadler & Tai 2007 retrospective survey, ~8,474 students / 63 institutions D Per year of high-school maths: chemistry +1.86, biology +1.84, physics +1.28. Out-of-discipline science: not significant anywhere
Sadler & Tai 2001 retrospective survey, 1,933 students / 19 courses D High-school physics worth d=0.24 raw; calculus (+2.58) worth as much as a year of physics (+2.26); coverage coefficient −0.49
Schwartz 2009 longitudinal survey, 8,310 students / 55 institutions D Depth +1.22 biology / +0.89 chemistry / +1.54 physics; breadth −1.15 biology. Standardized magnitude 0.08–0.13
Schmidt 2005 / 1997 descriptive curriculum comparison D "Unfocused, repetitive and undemanding"; ~35 topics vs ~7; 530-page grade-4 textbooks
Shin 2019 quasi-experiment, ~1,225 students, developer-involved C Coherent chemistry sequence positive on an aligned measure; benefit limited in low-performing schools
Corcoran 2009 expert-panel audit D Consequential validity of learning progressions: "among the most hypothetical of questions"
Steedle & Shavelson 2009 psychometric validation D Students respond inconsistently across item contexts; level diagnoses often not validly interpretable
NRC 2012 expert consensus D No students, no comparison group, no effect estimate — and it set the sequence for a generation

The prerequisite question: use the null, not the positive

The most-cited evidence on science course-taking is a pair of retrospective surveys, and the trustworthy half of it is the negative half.

The cross-discipline prerequisite claim — that you need chemistry before biology, or physics first because it underpins the rest — fails in both surveys that tested it. Years of out-of-discipline high-school science do not significantly predict college grades in biology, chemistry or physics. The authors state the inferential asymmetry themselves: their design is "epidemiological rather than experimental" and cannot establish causation, but "a low or negative correlation is evidence against a positive causal relationship." A null in a design biased toward finding effects is worth more than a positive one.

The positive coefficients get no such licence. A year of high-school maths is associated with +1.86 on a 100-point college chemistry grade, +1.84 in biology and +1.28 in physics; residual SD is about 10, so these are d ≈ 0.13–0.19 per year of coursework on a sample already selected into college science. Course choice is driven by prior ability, tracking, what the school offers, and parental push — the authors concede that students who skip high-school physics "are more likely to be academically stronger, with more educated parents, having previously taken calculus." Their own calibration is the best warning available: students' self-estimate of how much high-school physics helped averaged +7.8 points, about three times the measured association, and non-takers rated it higher than takers did.

Then the single quasi-experimental test of course order — one Texas district that ran both the traditional biology-chemistry-physics sequence and physics-first concurrently, compared on independent Texas end-of-course exams — found "little or no impact" on the science exams and "minimal effect" on Algebra I and Geometry.

So the honest practical reading is: there is no evidence for a cross-discipline science prerequisite, and no evidence that course order matters. Mathematics preparation is the strongest association in the data and is worth attending to, but it is the same confounded coefficient as everything else here and should not be sold as a causal lever.

Depth over breadth: the finding to half-believe

This is the most-quoted sequencing result in science education and it deserves both the attention and the discount.

Across 8,310 students at 55 randomly selected institutions, students who reported that their high school course spent a month or more on a single topic earned higher college science grades (+1.22 biology, +0.89 chemistry, +1.54 physics), while students who reported covering all major topics earned lower ones (−1.15 biology). Sadler & Tai's independent sample finds the same shape from the other direction: among students who took high-school physics, the coverage coefficient is −0.49 — more topics covered, worse college grade.

Three reasons not to over-read it:

  • The magnitude is 0.08–0.13 SD. The authors report it. It is a small effect dressed in a large narrative.
  • The outcome is a college professor's grade — teacher-assigned, uncalibrated across 55 institutions, in courses whose professors chose to participate.
  • Course-taking is not assigned. Students who take a deep, slow, honours-style science course are not the same students as those who take a survey course, and the model controls are self-reported survey items. About 25% of AP students place out of introductory science entirely and are missing from the sample. This is exactly the selection channel this archive's premise predicts.

The direction is nonetheless corroborated by the one thing in the topic that is an experiment: the NGSS curriculum trial, where the treatment increased time on reading and writing about science and decreased teacher-reported coverage of maths/computational thinking, data analysis and modelling — fewer things, more deeply — and returned g=0.40 on its own test.

Learning progressions: a policy built ahead of its evidence

The NRC Framework and NGSS are organised around K-12 learning progressions. The field's own most careful audit, published three years before the Framework, found the empirical base at the "initial stages of producing evidence" for construct validity and described the consequential question — does sequencing instruction by a progression actually improve learning? — as "among the most hypothetical of questions". Nothing since has answered it.

Worse, the psychometric premise is shaky. Progression-based sequencing assumes a student can be placed at a level. When ordered multiple-choice items keyed to a force-and-motion progression were analysed with latent class analysis, students responded inconsistently across item contexts, so the level diagnosis frequently cannot be validly interpreted. A sequencing rule whose diagnostic step does not work is not yet an instructional method.

What does have causal evidence (and it is not sequencing)

  • Professional development, not materials. In the one synthesis that required a control group, a four-week minimum duration, and outcome measures independent of the treatment, inquiry programmes built around science kits returned +0.02 across seven studies, while inquiry-oriented programmes emphasising teacher professional development returned +0.36 across ten. That is the most decision-relevant number on this page and it is about who trains the teachers, not about sequence.
  • Science inside the literacy block. Integrated literacy-and-content instruction moves standardized reading comprehension +0.25 (26 effects) and produces science knowledge as a by-product — which connects this topic to background knowledge & vocabulary, where the same bet is the highest-upside wager in the archive. Note the direction of the transfer: content instruction helps reading. The reverse claim is weaker — the Seeds of Science RCT moved science understanding 0.65 and reading comprehension 0.09 (not significant).
  • An NGSS-aligned curriculum package, once, at g=0.40 on a researcher-built test with 38% school attrition. Read as encouraging, not settled.

The measure-type problem, measured inside this literature

This strand supplies the two cleanest paired estimates in the whole science tranche:

Contrast Aligned Independent Ratio
Hwang 2022, reading comprehension researcher-developed 0.54 (123 effects) standardized 0.25 (26 effects) 2.2×
Hwang 2022, science/social-studies content knowledge researcher-developed 0.94 (49) standardized 0.68, not significant (6) 1.4×
Slavin 2014, within one inquiry programme's own end-of-year test taught units +0.58; taught topics +0.43 / +0.29 remaining items +0.09; overall +0.21 ~6×

That last row is the sharpest measurement result in the tranche, because it is within a single test taken by a single set of children: the entire programme effect lived in the items covering the topics the programme taught. Slavin's cross-field benchmark says the same thing more brutally — treatment-inherent measures average +0.45 in maths and +0.51 in reading, against −0.03 and +0.06 on treatment-independent ones.

Hereditarian-lens assessment

Risk: medium, and the risk is concentrated in exactly the part of the topic that generates the confident advice. The experimental leg (Slavin's synthesis, the Cervetti, Connor and Murphy RCTs) is randomised or matched with independent measures, so genes cannot systematically differ between arms.

The sequencing leg is not, and taken on its own it would be graded high. Every positive claim about course order, depth, and prerequisites comes from retrospective self-report surveys of students who chose their own high-school courses and then chose to enrol in college science. Course choice is one of the most ability-selected decisions an adolescent makes, the outcome is a professor's grade, and about a quarter of the strongest students are missing from the physics sample because they placed out with AP credit. The archive's premise predicts precisely this pattern — a robust correlation between demanding preparation and later success — with or without any causal effect of the sequence.

The specifications make it worse rather than better: they are simultaneously over-controlled (SAT-Quantitative and prior grades are themselves downstream of the same coursework) and under-controlled (nothing for motivation or conscientiousness).

The adjacent literature is the strongest available check, and it is damning for course-taking coefficients generally. Klopfenstein & Thomas found Advanced Placement coefficients collapse under full curricular controls; Altonji found returns to high-school course-taking shrink to near-nothing under better identification; and the first randomised study of AP science found AP takers earned worse grades and reported higher stress. This archive should read Sadler & Tai the way its authors asked it to — as epidemiology, licensing the null and not the positive.

Nothing here touches g. The outcome class throughout is domain-skill, with reading comprehension entering as near-transfer.

Boundaries & what critics say

  • The descriptive TIMSS work is being asked to carry a causal load it cannot. "A mile wide and an inch deep" describes curricula, not learning; the countries with tighter curricula differ from the US in many other ways.
  • Coherence has one quasi-experimental test, it is developer-involved, non-randomised, measured on a curriculum-aligned instrument, and its benefit was limited in low-performing schools — the opposite of what a founder serving disadvantaged students would want.
  • The science-literacy integration RCTs are developer-led and confounded with time. The Seeds of Science treatment received 3.66 hours of science a week against 3.03 in control — a 0.53 SD difference in dose. Some of the science effect is simply more science.
  • CALI's standardized literacy effects appear only at grade 4 and are null in kindergarten through grade 3, on samples of 75–109 per grade.
  • Textbook-coherence rubrics are recorded but excluded as design-inadequate: they judge books, not learning.
  • Two of the load-bearing sources were not read in full. The Texas course-order null is a ProQuest dissertation available only as an abstract, so its n, covariates and effect sizes are unknown — and it is this topic's headline null. The 2007 survey is paywalled in Science; its coefficients here come from secondary reporting, and its stated sample (~8,474 students / 63 institutions) could not be reconciled against the parent dataset's 55 institutions. Both are marked as read-depth debt in their source files.
  • The coherence advocates concede the point in print. Schmidt's own paper says "questions of causality are impossible to answer with survey data" and that empirical support "must await additional study" — and then argues for acting anyway because coherence "seems a more logical principle from a subject-matter perspective." That is design plausibility, and it became national policy.
  • Micro-sequencing experiments exist and live elsewhere in the archive. Randomised one-lesson contrasts (interleaved vs blocked, explore-first vs explain-first) are real but tiny (n≈85–94) and at least one reverses direction between the immediate and the two-week test. They are graded under interleaving and the productive-failure evidence in explicit vs inquiry, not here.
  • Physics-first has a practical cost the evaluations do not price. Every published case study reports that coordinating ninth-grade physics with ninth-grade algebra was harder than planned and that biology enrolment collapsed for one to two years during the transition.

Practical guidance

  • Cut the topic list and go slower. It is the one sequencing instinct with converging (if weak) support, and the experimental curriculum that worked did exactly this.
  • Do not reorder the sciences and expect a return. Physics-first has been tested once against its alternative on standardized exams and returned nothing, and it carries a real transition cost.
  • Do not build a prerequisite chain between the sciences. The one thing this evidence base rules out is that taking chemistry helps you in biology or physics. Mathematics preparation is the largest association in the data and worth attending to on general grounds — but it is the same confounded coefficient, so treat it as a scheduling default rather than a lever.
  • Do not buy a science-kit programme. +0.02 across seven studies with independent measures.
  • Buy the professional development instead (+0.36), and read the inquiry topic first, because PD that is handed down a trainer chain lost its entire effect there.
  • Teach science inside the literacy block in primary school. It is the one move here that pays on a standardized measure in a different subject, and it is the same bet as curriculum choice's highest-upside recommendation.
  • Treat NGSS/Framework sequencing as a reasonable default, not as evidence. It is expert consensus, its own field called its central question hypothetical, and the diagnostic step it depends on does not reliably work.

Open questions

  • The experiment has never been run: same content, same time, same teachers, two sequences, independent outcome measure. One quasi-experimental district study is the entire causal literature on course order.
  • Spiral versus mastery in science is completely untested. The debate is imported wholesale from mathematics.
  • Nobody has validated a science learning progression consequentially — i.e. shown that teaching in progression order beats teaching the same content in another order.
  • Whether the depth advantage survives a design with assignment is unknown, and it is the most consequential open question here, because "teach fewer things for longer" is the advice everyone is already giving.
  • Whether the science-literacy integration effect is content or time is unresolved in every trial that reports both.

Evidence (15 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore