Temporal stability of WISC-III subtest composite: strengths and weaknesses.
Watkins MW, Canivez GL · 2004
grade Clongitudinalindependentreplicated
Sample
579 students tested twice with the WISC-III; 66 subtest composites examined
Population
US school-age children referred for psychoeducational evaluation, test-retest interval of roughly three years.
Design
Test-retest stability study on a large referred sample - the correct design for the question, and well powered for it. The A-D rubric is calibrated for causal intervention evidence and has no natural home for psychometrics; C is the ceiling this archive applies to well-powered empirical measurement studies, and it should be read as "good evidence of its kind" rather than as a demotion. Limitation: a referred sample, not the normative population, and WISC-III cognitive subtests rather than an achievement battery - the mechanism (short unreliable subtests differenced against each other) is general, but the exact chance-level result is established here on an IQ test.
Key findings
The finding that should govern how any score profile is read. Each administration produced 6 or 7 apparently interpretable cognitive strengths and weaknesses - the profile always looks meaningful. But across the two occasions those strengths and weaknesses replicated AT CHANCE LEVELS. The same child, the same test, and the peaks and troughs move. The authors' conclusion is the operative one: "Because subtest-based cognitive strengths and weaknesses are unreliable, recommendations based on them will also be unreliable." A profile that does not survive a retest cannot support an instructional plan.
Genetic confound
Not applicable. The finding is that a measurement is unstable, which is a claim about the instrument; no heritable trait is being estimated or attributed.
Replication notes
Consistent with the broader ipsative-assessment literature (McDermott et al. 1990, 1992) and with the follow-on stability work by Borsuk, Watkins & Canivez (2006); the chance-level replication result is the field's most reproduced negative finding about profile interpretation.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Interpretable cognitive strengths and weaknesses identified per administration | count | 6-7 per WISC-III administration, from 66 subtest composites | standardized | each of two test occasions | none | not-applicable | g |
| Replication of identified strengths and weaknesses across test-retest | agreement | chance levels | standardized | test-retest | none | over-2yr | g |
Cited by
- What a standardized achievement score does and does not licensemixedconf: mediumgc: low