The Evidence on Teaching

Lateral reading on the open Internet: A district-wide field study in high school government classes.

Wineburg S, Breakstone J, McGrew S, Smith MD, Ortega T · 2022

grade Cquasi-experimentdeveloper-ledmixed
Sample
499 students in the published paper (271 treatment, 228 control) across 6 high schools; the same team's working-paper account of the same field experiment reports 464 students (265 treatment, 199 control)
Population
High school students in a district-mandated government course, large Midwestern district of ~40,000 students, nearly half eligible for free/reduced lunch
Design
The one causal test of Civic Online Reasoning inside this archive's 4-18 age range. Six schools were matched on demographics and the matched pairs assigned to opposite conditions - so randomisation, such as it is, occurred at the school level with THREE clusters per arm, which is the design's decisive weakness: with n = 3 clusters per condition, cluster-level inference is essentially impossible and the multilevel model borrows almost all its power from student-level variance. The team's own modelling found schools explained only 1% of score variance and dropped the school random effect entirely. Teachers (two per school) got professional development and delivered six 50-minute lessons across a semester; control teachers taught their normal curriculum. Outcome was a researcher-built Civic Online Reasoning assessment scored out of 14 points, taken with a live internet connection, with parallel forms counterbalanced. The intervention, the assessment, the rubric and the analysis all come from the group that created the curriculum and sells its reputation. Graded C: real control group, real pre-post, cluster-randomised in principle, but three clusters per arm, a wholly researcher-designed outcome, and no independent evaluator.
Key findings
Treatment students improved from a mean of 2.9 to 5.1 points out of 14; control students from 2.0 to 2.7. The adjusted condition-by-time interaction was 1.66 points (SE 0.44, p < .05) in the fully-covariate-adjusted robust mixed model - roughly half a standard deviation on the study's own scale, though neither the report nor the paper converts to a standardised effect size. Note the absolute levels: after six lessons the average treated student scored 5.1 out of 14, i.e. the intervention moved students from very bad to bad at evaluating online sources. The team's earlier single-school pilot (441 students, 6 teachers, no control group) found a significant pre-post time effect of 1.71 points, which is uninterpretable without a control arm and is reported here only as the provenance of the design. Covariates that predicted scores independently of treatment - speaking English at home (+0.85), self-reported frequency of checking trustworthiness (+0.61), and negative coefficients for Black students (-1.06) and "other race" (-0.82) - indicate the outcome measure is heavily loaded on general verbal skill.
Genetic confound
Medium. Assignment was by matched-pair school allocation rather than individual randomisation, with only three schools per arm, so school-level composition differences cannot be ruled out; the model's own finding that race and home language predict the outcome shows the measure tracks background characteristics that differ across schools.
Replication notes
Mixed. The same heuristic replicates repeatedly in undergraduate and adult samples (Breakstone et al. 2021 online college course; McGrew et al. 2019 university intervention study; Fendt et al. 2023 adult online sample; and lateral reading is the strongest-performing approach in the Fendt et al. 2025 meta-analysis at g = 0.55). Within the 4-18 band, however, the record is thin and developer-owned: this trial, plus McGrew 2020 and McGrew & Breakstone 2023, both of which are single-group pre-post designs with no control. No independent team has evaluated the COR curriculum with school-age students.
DOI / URL
10.1037/edu0000740

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Civic Online Reasoning assessment (14-point researcher-built online source-evaluation test)adjusted condition-by-time interaction, robust linear mixed model1.66 points of 14 (SE 0.44, p < .05); raw means 2.9 to 5.1 treatment vs 2.0 to 2.7 controlresearcher-designedend of a one-semester, six-lesson sequencebusiness-as-usualend-of-treatmentdomain-skill
Absolute attainment after instructionpost-test mean on the 14-point scale5.1 of 14 in treated classes - a large relative gain from a very low floor, not competenceresearcher-designedend of semesternoneend-of-treatmentdomain-skill

Cited by

Lateral reading on the open Internet: A district-wide field study in high school government classes. · The Evidence on Teaching