The Evidence on Teaching

Illusions of Meaning in the Ipsative Assessment of Children's Ability

McDermott PA, Fantuzzo JW, Glutting JJ, Watkins MW, Baggaley AR · 1992

grade Clongitudinalindependentreplicated
Sample
multiple representative data sets including the full WISC-R national standardization sample
Population
US children; the anchor analysis uses the WISC-R national standardization sample, so this is a normative rather than a referred population.
Design
Head-to-head comparison of ipsative (within-child, deviation-from-own-mean) subtest scores against ordinary norm-based scores on the same data, evaluated on convergence, short- and long-term stability, and predictive efficiency. Using the standardization sample removes the referral selection that limits most profile studies. Correlational by construction - it establishes that ipsative scores carry no information, not that any particular instructional decision was harmed.
Key findings
Ipsative ability measures were "uniformly inferior to their normative counterparts, conveying no uniquely useful information and otherwise impeding the versatility of assessment." This is the general result behind the specific one: the practice of reading a child's profile relative to their own average - which is exactly what "he is strong in maths and weak in language mechanics" does - subtracts reliable variance and adds none. The subtest peaks and troughs that look like the most individually meaningful part of a score report are the part with the least information in it.
Genetic confound
Not applicable. A comparison of two scoring metrics on the same data; no causal claim and no attribution of a trait.
Replication notes
The finding has been reproduced repeatedly - the same group's 1990 critique, Watkins & Canivez's 2004 chance-level retest result, and the incremental-validity literature all converge on it. This is one of the better-replicated negative results in educational measurement.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Information conveyed by ipsative subtest scores beyond normative scorescomparative validitynone - uniformly inferior on convergence, stability and predictive efficiencystandardizedcross-sectional and longitudinal within the same samplesactive-alternativenot-applicableg

Cited by