The Evidence on Teaching

Gifted Today but Not Tomorrow? Longitudinal Changes in Ability and Achievement during Elementary School

Lohman DF, Korb KA · 2006

grade Clongitudinaldeveloper-involvedreplicated
Sample
reanalysis of two large longitudinal datasets - Martin (1985), 6,321 students tested annually grades 3-8 on the ITBS; Gustafson (2002), 2,363 students tested in grades 4, 6 and 9 on ITBS Form K and CogAT Form 5 - plus a 10,000-case simulation of selection rules
Population
US elementary and middle-school students in large district and state samples, reweighted to approximate the state achievement distribution
Design
The A-D rubric is built for causal intervention evidence and fits measurement studies badly; C is this archive's ceiling for a well-powered psychometric reanalysis and should be read as "good evidence of its kind." Two limits matter. First, independence: Lohman co-authored the CogAT, and the conditional standard errors quoted here come from his own test's research handbook. That cuts the way developer involvement usually does not - the paper's conclusion is that scores of high scorers are far less trustworthy than practitioners assume, which is against the developer's interest, so the usual inflation discount runs backwards. Second, the headline instability numbers are estimated from correlation matrices under a bivariate-normality assumption rather than counted directly, which the authors state.
Key findings
The single most important empirical fact for anyone reading one child's test score, and it is not about ability. Using ITBS data on 6,321 students, only about 40% of children who scored in the top 3% of the Composite in grade 3 were still in the top 3% in grade 4 - despite Composite reliability of KR-20 = .98 and a grade 3-to-4 correlation of r = .91. For the individual subtest scores (Reading, Language, Mathematics) the first-year fallout was about 50%. By grade 8 only 35-40% of the grade-3 top-3% group remained. Second: measurement error is NOT constant across the score scale. Conditional SEMs for scale scores are smallest in the middle and largest at the floor and the ceiling, and administering a higher (above-level) test halves the error - on the CogAT Verbal battery, a scale score of 221 carries an SEM of 14.8 at Level A but only 7.4 at Level D. Third: the common practice of taking the HIGHEST of several test scores is the worst possible rule, because the highest of a set of near-parallel scores is by construction the most error-encumbered one; averaging is the statistically optimal rule, with the cut score lowered to compensate.
Genetic confound
Not applicable to the finding itself. This is an argument about instruments and regression to the mean, not about the origins of ability. It does bear on hereditarian reasoning in one direction - much of what looks like a child "losing" or "gaining" ability year to year is measurement and scaling, not development.
Replication notes
The regression phenomenon reproduces across two independent datasets analysed here (Martin 1985; Gustafson 2002) and matches Thorndike's 1933 meta-analysis of 36 Stanford-Binet retest studies, whose correlation matrix shows the same simplex decay. Warne (2012) describes the paper as the landmark demonstration. The specific percentages are dataset-specific; the pattern is not.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Students in the top 3% of ITBS Composite at grade 3 still in the top 3% at grade 4%about 40%, despite KR-20 = .98 and r(grade 3, grade 4) = .91standardizedone yearnoneunder-1yrdomain-skill
Same, for individual ITBS subtest totals rather than the Composite%about 50% fallout in the first year - subtests are less stable than the composite that pools themstandardizedone yearnoneunder-1yrdomain-skill
Students in the top 3% at grade 3 still in the top 3% at grade 8%35-40%standardizedfive yearsnoneover-2yrdomain-skill
Conditional standard error of measurement, on-level vs above-level testSEM in scale-score pointsCogAT Verbal scale score 221 - SEM 14.8 at Level A, 7.4 at Level D; a higher test level halves the errorstandardizedsingle administrationactive-alternativenot-applicabledomain-skill
Agreement of two tests on who exceeds a top-3% cutproportionat r = .80 between tests, 45% of those above the cut on test 1 are above it on test 2; at r = .90, 60%standardizedtwo administrationsnonenot-applicabledomain-skill
Selection rule that minimises later regressioncomparison of rulesaveraging scores beats "highest of several" and beats requiring both; the highest of a set of parallel scores is the most error-encumbered member of that setstandardizedsimulation, 10,000 casesactive-alternativeunder-1yrdomain-skill

Cited by