The Evidence on Teaching

Do Test Score Gaps Grow before, during, or between the School Years? Measurement Artifacts and What We Can Know in Spite of Them

von Hippel PT, Hamrock C · 2019

grade Bcritiqueindependentreplicated
Sample
Re-analysis of three longitudinal data sets — the Beginning School Study (Baltimore first graders from 1982), the Early Childhood Longitudinal Study Kindergarten Cohort of 1998-99, and NWEA's Growth Research Database of MAP test records
Population
US school children from kindergarten or first grade through eighth grade, across three cohorts spanning 1982 to the 2010s
Design
Sociological Science 6, 43-80 (open access). Graded B rather than as a commentary because it re-estimates gap growth on the original data sets rather than arguing about them, and because its central claim is a measurement claim that it demonstrates directly. The two artifacts it identifies are specific and checkable. (1) SCALING: gaps appear to grow if the score scale itself spreads with age. The Beginning School Study used Thurstone-scaled California Achievement Test Form C; early ECLS-K releases used number-right scores. Neither is a vertical interval measure of ability. (2) CHANGE OF TEST FORM: with fixed-form paper tests, school-year learning is measured within one form (autumn to spring) while summer learning is measured ACROSS two different forms (spring to autumn), so any difference between forms is booked as summer loss. Adaptive tests scored on IRT ability scales — ECLS-K's two-stage adaptive test and NWEA MAP's continuously adaptive test — are immune to the second artifact and much less exposed to the first. The paper is candid that even IRT ability scales do not settle the question: on the ECLS-K scale the white-black reading gap doubled between first and eighth grade, while on the GRD scale the same gap grew by only a third, and IRT theta is interval only in a narrow technical sense that a monotone transformation can undo.
Key findings
The classic summer-loss result is not so much refuted as shown to be unidentifiable at the precision the field claimed. Net of artifacts, "the most replicable finding is that gaps form mainly in early childhood, before schooling begins. After school begins, most gaps grow little, and some gaps shrink." On IRT ability scales NO gap so much as doubled between first and eighth grade, and the average gap growth across all comparisons (black versus white, poor versus non-poor and so on) was just 7 PERCENT over those seven years — against results on Thurstone or number-right scales suggesting gaps grew twofold to sixfold after school entry. On the school-versus-summer question the verdict is agnostic, not contrarian: "Evidence is inconsistent regarding whether gaps grow faster during school or during summer... If summer learning gaps are present, most of them are small and hard to discern through the fog of potential measurement artifacts," and "perhaps it is safest to say that neither schools nor summer vacations contribute much to test score gaps." THE PART USUALLY DROPPED FROM BOTH SIDES OF THIS DISPUTE: summer SLOWDOWN survives even if summer GAP GROWTH does not. The authors state that their figures "do consistently show that summer learning is slow for nearly all children, including children from advantaged groups," and argue that this near-universal slowdown is precisely what gives summer programmes their opportunity. The critique is aimed at the inequality claim, not at the claim that little is learned over the summer.
Genetic confound
Not directly addressed and largely orthogonal — the paper is about measurement, not about causes. But the conclusion it reaches has a hereditarian reading the authors do not make: if gaps are essentially fully formed by school entry and then change by about 7% over nine years of schooling, then whatever produces them operates before and outside school, and the candidates include heritable ability transmitted alongside the early environment. The paper attributes the pattern to early-childhood plasticity and early investment instead. Both readings fit the same fact, and nothing in this design distinguishes them.
Replication notes
The re-analysis is itself a replication attempt across three independent data sets with three different tests, and the artifact-free ones agree with each other far better than either agrees with the Thurstone-scaled original. It converges with Quinn and colleagues' finding that seasonal inequality patterns in ECLS-K:2011 do not reproduce those in ECLS-K:1998, and with the broader psychometric literature on vertical scaling. It stands against Cooper and colleagues' 1996 meta-analysis and the Beginning School Study, whose data it re-analyses directly. Note that the dispute has NOT resolved into a single number: Atteberry and McEachin, using the same NWEA MAP data, report substantial and highly heterogeneous summer losses, so "how large is summer loss" remains genuinely open even among people using artifact-free tests.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Growth in test-score gaps between first and eighth grade, on IRT ability scalespercent change in the gapaverage gap growth across all comparisons was 7% over seven years; no gap so much as doubled; some gaps shrankstandardizedgrade 1 to grade 8noneover-2yrdomain-skill
The same gap growth measured on Thurstone or number-right scalesmultiple of the initial gapgaps appeared to grow twofold to sixfold after school entry — an artifact of scales that spread with age rather than a change in relative abilitystandardizedgrade 1 to grade 8noneover-2yrdomain-skill
Whether summer gap growth replicates on adaptive IRT-scaled testsconsistency of the findingdoes not consistently replicate; evidence is "inconsistent regarding whether gaps grow faster during school or during summer", and any summer gaps present are "small and hard to discern"standardizedseasonal (autumn-spring-autumn)nonenot-applicabledomain-skill
Summer learning slowdown itself, as distinct from summer gap growthdirection, all groupssummer learning is slow for nearly all children INCLUDING advantaged ones — the finding survives the critique and is the authors' stated reason to keep taking summer programmes seriouslystandardizedsummer vacationnonenot-applicabledomain-skill
Disagreement between two artifact-free scales on the same gapgrowth in the white-black reading gap, grade 1 to grade 8ECLS-K IRT scale: the gap doubled. NWEA GRD IRT scale: it grew by only one third. Different populations, different item pools and one- versus three-parameter IRT models all contributestandardizedgrade 1 to grade 8noneover-2yrdomain-skill

Cited by