What a standardized achievement score does and does not license
Trust the composite; distrust the breakdown. Subtest strengths-and-weaknesses replicate at chance on retest, and most of what people read off a score report is noise.
mixedconf: mediumgc: lowassessment · ages 5–18 · structure
Split the score report in two. The COMPOSITE is a good measurement and worth acting on. Everything finer-grained that people actually read off it is not. Subtest-based strengths and weaknesses replicate at CHANCE levels on retest (n=579); ipsative profile scores carry no information beyond normative scores; only ~40% of children in the top 3% of a composite in grade 3 are still there in grade 4 despite KR-20=.98 and r=.91, and ~50% fall out on individual subtests. Real measurement error is MORE THAN TWICE the vendor-reported figure, and it is largest at the floor and the ceiling — a higher test level halves it (SEM 14.8 → 7.4). Grade equivalents are not an interval scale and the publishers themselves say not to use them for diagnosis or placement. A score at chance carries no information about the child at all: on an 80-item four-option subtest, ANY raw score from 14 to 26 is statistically indistinguishable from pure guessing.
Read the composite; distrust the profile. Before interpreting any subtest, do three checks that cost nothing: (1) count the items and compute the chance band — on 80 four-option items, anything from 14 to 26 correct is indistinguishable from guessing and the subtest measured nothing; (2) check how many items were actually attempted, because omissions score as wrong and a timed-out or skipped section looks identical to ignorance; (3) treat any subtest-to-subtest gap as noise unless it repeats on a second administration. Never convert a grade equivalent into a teaching level — the publishers say so themselves. If you need to know what a child can actually do, use a criterion-referenced check on the specific skill, not a percentile.
Who this applies to
Verdict
This topic exists because the archive is being built toward a tool that meets a real child, usually with a
test result already in hand, and a recommendation built on a misread score is worse than no
recommendation. The verdict is mixed, and the split is clean enough to state as a rule.
The composite score is a good measurement. It is reliable, it predicts, and a parent or teacher acting on it is acting on something real.
Almost everything finer-grained that people read off a score report is not. The subtest profile, the grade equivalent, the gap between two subscores, the "areas of relative weakness" paragraph — these are the parts of the report that feel most individually meaningful and they are the parts with the least information in them. That is not a soft caveat. Subtest-based strengths and weaknesses replicate at chance levels when the same child takes the same test again, and on a current instrument a statistically significant profile discrepancy corresponds to a real difference on the underlying ability in only 40–74% of cases. Statistical significance on a profile is not evidence that the difference is real in the sense a parent means.
A note on grading before the evidence table. This is a psychometrics literature, not an intervention literature. The A–D rubric is calibrated for causal designs, and most of what follows is measurement theory, reliability estimation, and variance decomposition — for which there is no randomised arm to grade. This archive caps well-powered empirical psychometric studies at C and reviews and critiques at D, and that ceiling should be read as "good evidence of its kind", not as a demotion. Inflating these grades to make the topic look stronger would be the exact failure the methodology exists to prevent.
What the evidence shows
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Watkins & Canivez 2004 | test-retest, 579 students, 66 subtest composites | C | 6–7 interpretable strengths/weaknesses per administration; they replicate across retest at CHANCE levels |
| Lohman & Korb 2006 | reanalysis of 6,321 + 2,363 students; 10,000-case simulation | C | Only ~40% of top-3% grade-3 scorers are top-3% in grade 4 (KR-20 = .98, r = .91); ~50% fallout on subtests; SEM largest at floor/ceiling; above-level test halves SEM (14.8 → 7.4) |
| de Jong 2023 | WISC-V standardization data + simulation | C | A statistically significant index-vs-own-performance discrepancy reflects a real difference on the underlying factor in only 40–74% of cases |
| McDermott 1992 | WISC-R standardization sample + others | C | Ipsative (own-mean-relative) scores uniformly inferior, convey no uniquely useful information |
| Reynolds 1981 | psychometric critique | D | The "two years below grade level" rule fails in opposite directions at different ages — GE units are not equal-interval |
| Boyd 2013 | NYC administrative data, consecutive-grade state tests | C | Real measurement error > 2× the vendor-reported figure; vendor split-half omits day-to-day variation |
| Kanaya 2003 | longitudinal IQ records, 9 sites | C | Same child scores −5.6 points on a renormed test; classification changes with the norm, not the child |
| Pearson (publisher guidance) | publisher technical note | D | GEs not an interval scale, cannot be added/subtracted/averaged; 5 raw points = 4 months at age 4 but ~2 years at 16+; "should not be used for making diagnostic or placement decisions" |
| Hoover 1984 | invited address (ITBS co-author) | D | The counter-case: GEs are the only score type on a developmental scale, so they alone can express growth across grades |
| Frederick & Speed 2007 | statistical exposition | D | A score in the guessing range is evidence about the administration, not the person; "below chance" needs a critical value, not just being under the mean |
| Greve 2009 | 1,032 forensic examinees, 3 instruments | C | Significantly below-chance performance is rare even where most incentivised; a single test can raise the question, never settle it |
| Kirkwood 2012 | 276 children aged 8–16 | C | 19% failed an objective validity test; no background variable predicted who; validity performance explained 38% of ability-test variance |
| Wise 2017 | review | D | A rapid guess is "a choice by the test taker to momentarily opt out of being measured"; such items should not be scored |
| Wise & DeMars 2005 | synthesis of 12 experimental manipulations | C | Low motivation produces a substantial performance decrease — established by manipulation, not correlation |
| Meijer & Sijtsma 2001 | methodology review | D | Person-fit statistics exist but misbehave on short and moderate-length tests — i.e. on subtests |
| Cannell 1988 | survey of reported results | D | All 50 states reported above-average performance on nationally normed tests |
| Linn 1990 | audit of state/district results | C | Confirms it; attributes the inflation substantially to norm age and administration conditions |
| Ho 2008 | statistical demonstration on real distributions | C | Threshold statistics give "limited and unrepresentative" pictures; distortions "unpredictable, dramatic" |
| Popham & Husek 1969 | conceptual | D | Norm-referenced answers "compared to whom"; only criterion-referenced answers "can this child do X" |
| Solomon 1971 | two-page note, 4th graders | D | Null: answer-sheet format had no effect on group mean performance |
| Duckworth 2011 | RETRACTED 2025 | D | Excluded. g = 0.64 for incentives on IQ scores → 0.47 after removing a fabricated input, with strong publication bias on all four diagnostics |
The chance band, and why 21% on a four-option test is not a low score
This is the arithmetic the archive was missing, and it is worth stating exactly because it is the single most decision-changing calculation in the whole intake problem. It is derivation, not a citation: the binomial distribution of a pure guesser.
On an 80-item, four-option subtest, a child who knows nothing and guesses every item scores:
- Expected: 20 correct (25%), SD = √(80 × 0.25 × 0.75) = 3.87 items
- 17 correct → z = −0.77, and P(X ≤ 17 | pure guessing) = 0.26. A guesser scores 17 or lower about a quarter of the time.
- To be significantly below chance (one-tailed, p < .05) requires ≤ 13 correct.
- To be significantly above chance requires ≥ 27 correct (33.8%).
- Therefore every raw score from 14 to 26 — 17.5% to 32.5% — is statistically indistinguishable from pure guessing.
Two consequences, and the second is the one that reverses advice.
First, "below chance" is a claim with a critical value. Frederick & Speed's whole point is that practitioners routinely confuse "under the chance mean" with "worse than guessing", and the confusion matters because the two license completely different inferences. A raw score of 21% on four-option items is at chance, not below it. Calling it below-chance overstates the statistic; the archive should not do that even when the overstatement points in the right direction.
Second, at-chance is not a low score — it is an absent measurement. This is the reframe. A score of 17/80 does not say the child has weak knowledge; it says the subtest produced no information about the child. Frederick & Speed, Wise, and Meijer & Sijtsma all converge on this from different directions: a response pattern indistinguishable from guessing carries no signal about ability, so the correct inference is about the administration. Lohman & Korb supply the corroborating psychometric fact — conditional standard error of measurement is at its maximum at the floor of a scale, so the test is at its least informative precisely where this score sits.
The mechanical candidates all predict roughly this pattern, and they are distinguishable by evidence the archive can ask for:
- An offset or mis-tracked answer sheet makes responses effectively random with respect to the key, which predicts a score at chance — about 20/80. Note the honest complication: Solomon (1971) found answer-sheet format had no effect on group mean performance in fourth graders. A mean null on format is not evidence that no individual child ever mis-tracks a sheet, but the archive records that the direct empirical literature on this mechanism is thin and unsupportive, and it should not be cited as though it were strong.
- A skipped section or running out of time predicts a score below chance, because omissions score as wrong. This is the diagnostic that discriminates: ask how many items were attempted. Seventeen correct out of 25 attempted is a completely different child from 17 correct out of 80 attempted.
- Disengagement predicts near-chance performance too. Kirkwood's is the number to hold onto: 19% of children aged 8–16 failed an objective validity test, and no background variable predicted which ones — and validity performance explained 38% of the variance in their ability scores. Invalid performance in children is common, unpredictable, and dominates the profile.
Hereditarian-lens assessment
Risk: low, and for an unusual reason — almost nothing in this topic is a claim about children at all. The findings are claims about instruments: that a subtest is unreliable, that measurement error is larger than advertised, that a scale is not an interval scale, that a norm sample ages. Genes cannot confound a variance decomposition.
The topic does interact with the archive's premise in two directions worth naming.
It cuts against over-reading a single score as a trait. Kanaya's 5.6-point drop happens to the same child when only the reference group changes. Lohman & Korb's 60% fallout from the top 3% between grades 3 and 4 happens with a composite reliability of .98 — most of what looks like a child "losing giftedness" is regression and scaling. Wise & DeMars's motivation manipulations move scores without moving anything about the child. A hereditarian reading of education has to be more careful about measurement, not less: if you believe individual differences are substantially real and substantially stable, then a score that jumps around is telling you about your instrument.
And it cuts against over-reading a profile as a set of aptitudes. The specific-ability interpretation
that subtest analysis rests on — this child is a "verbal learner", that one has a "language mechanics
deficit" — is the same family of claim as learning styles and multiple intelligences, both of which this
archive has already graded debunked and insufficient. McDermott et al. showed the ipsative version
conveys no unique information; Watkins & Canivez showed it does not survive a retest. The profile is not
a weak signal about specialised aptitudes. It is noise with a narrative attached.
Boundaries & what critics say
- The grade-equivalent argument is narrower than the polemics suggest. Hoover — an ITBS co-author, so read with the conflict in view — makes a real point: within-grade standard scores are re-normed every year, so they cannot express growth across grades at all, and GEs can. Both sides actually agree on the operative rule: a GE is a growth scale, not a placement scale. Reynolds's 1981 critique names the specific failure that follows from ignoring that rule — the "two years below grade level" criterion that dominated learning-disability identification overstates severity at upper grades and understates it in the early grades, because the same raw-score gap buys wildly different grade-equivalent movement depending on how steep the norm group's growth curve is at that point. The publisher's own guidance is the cleanest statement of the limit, and it is the strongest kind of source for it, because Pearson is advising against a score type its own instruments print. A developer arguing against its own product gets the reverse of the usual discount.
- The profile-analysis evidence is on IQ subtests, not achievement subtests. Watkins & Canivez, McDermott et al., and the incremental-validity literature all used Wechsler-family cognitive batteries. The mechanism — differencing two short, individually unreliable subtests compounds their error — is general, and Lohman & Korb reproduce the instability directly on the ITBS, an achievement battery. But the archive should be explicit that the chance-level retest figure is established on IQ subtests and extended to achievement subtests by argument plus one corroborating dataset, not by direct replication.
- The below-chance literature is clinical, and its default explanation does not apply to a 9-year-old. Frederick & Speed, Greve et al., and the Slick-criteria tradition are about detecting deliberate underperformance in adults with financial incentives, mostly on two-alternative forced choice, where chance is 50% and the guessing distribution is wide relative to the scale. Importing the statistics to a four-option achievement subtest is legitimate; importing the interpretation is not. For a child taking a school battery, mechanical and attentional explanations dominate and malingering is close to irrelevant.
- Greve's rarity result is a caution in both directions. Genuinely below-chance performance is rare even in the population most motivated to produce it, and multiple validity tests detect it more often than any single one. A single anomalous subtest can raise the question; it cannot settle it. The remedy is another measurement, not a stronger inference from this one.
- Solomon 1971 is the weakest source here and is recorded honestly as such. A two-page note with an unrecoverable sample size, reporting a null. It is included because it is the most direct empirical test of the mis-bubbling mechanism this topic invokes, and recording a null that cuts against the topic's own hypothesis is what the disposition system is for.
- Duckworth et al. 2011 is excluded, and the exclusion is instructive. It was the most-cited evidence that test motivation confounds measured ability — a proposition this archive independently believes on other grounds. It was retracted in May 2025 after a fabricated-data input was found to supply 24% of the meta-analytic sample, and the reanalysis showed strong publication bias on every diagnostic. The conclusion survives on other evidence; the headline number does not, and the archive does not get to keep it.
Practical guidance
- Interpret the composite. Do not interpret the profile. If you take one thing from this topic: the gap between a child's best and worst subtest is the least trustworthy number on the page.
- Always get the raw score and the item count, not just the percentile. Every check that matters here — the chance band, the completion count, the floor/ceiling position — is impossible from a percentile alone. An intake tool that accepts only percentiles has thrown away its diagnostics.
- Compute the chance band before reading anything into a low subtest score. For k items with m options, chance is k/m with SD √(k(1/m)(1−1/m)). If the score is inside roughly ±1.6 SD of chance, the subtest measured nothing and the next step is to re-administer, not to remediate.
- Ask how many items were attempted. Omissions score as wrong on commercial batteries, so a timed-out or skipped section is indistinguishable from ignorance in the reported score and perfectly distinguishable in the response record.
- Require a second administration before acting on any subtest-level difference. Watkins & Canivez is the whole argument: one administration always yields 6–7 apparently meaningful peaks and troughs, and they do not come back.
- Double the published confidence band. Boyd et al. found real measurement error is more than twice the vendor figure, because split-half reliability cannot see day-to-day variation in a child. Whatever band the score report prints, the honest one is wider.
- Treat extreme scores as the least stable, not the most definitive. Only ~40% of top-3% third graders were still top-3% a year later. This is regression to the mean, not a child regressing. And it is why averaging several scores beats taking the highest — the highest of a set of near-parallel scores is by construction the most error-laden member of the set.
- Check the age of the norms. A percentile is a statement about a comparison group. Kanaya's 5.6 points is what happens when only the group changes; Cannell's all-fifty-states impossibility and Linn's audit are what happens when the group is stale. For families testing with commercially normed batteries at home, this is not a historical curiosity — it is the flattering percentile they were sent.
- If you need to know what a child can do, use a criterion-referenced check. Popham & Husek's distinction is the practical one. No percentile will ever tell you whether a child can use an apostrophe. Twenty minutes of asking them to punctuate sentences will.
Open questions
- Confidence is capped at
mediumby the evidence itself, not only by the pending adversarial pass. There are no grade-A/B sources in this topic and there probably cannot be — the questions are measurement questions, not intervention questions.highrequires two independent A/B sources pointing the same way, which this literature structurally cannot supply. - The chance-level retest result has never been replicated on an achievement battery's subtests. Watkins & Canivez did it on the WISC-III; Lohman & Korb show comparable instability on the ITBS but measure percentile fallout rather than profile replication. The direct study — take the same achievement battery twice, ask whether the identified strengths and weaknesses come back — appears not to exist, and it is cheap.
- Nobody has published subtest-level standard errors of measurement in percentile units for the commercial batteries homeschooling families actually use. The numbers exist in technical manuals; they are not in the peer-reviewed literature in a form an intake tool can cite. This is the largest concrete gap.
- Answer-sheet and transcription failure in young children is essentially unmeasured. The only direct evidence located is a two-page 1971 null. Given that the mechanism reverses advice for an individual child, the absence of modern work is striking.
- No one checks validity on ordinary group-administered school achievement tests. Performance-validity screening is standard in clinical neuropsychology, where Kirkwood found a 19% failure rate in children, and absent from education. If one child in five can produce an uninterpretable score in a clinic, the base rate in an unsupervised home administration is an open and important question.
- Rapid-guessing detection requires response times, which paper batteries do not record. The best available diagnostic for the exact failure this topic is about is unavailable in the exact setting the archive most cares about.
- grade CTemporal stability of WISC-III subtest composite: strengths and weaknesses.Watkins MW, Canivez GL · 2004 · longitudinal
- grade CIllusions of Meaning in the Ipsative Assessment of Children's AbilityMcDermott PA, Fantuzzo JW, Glutting JJ, Watkins MW, Baggaley AR · 1992 · longitudinal
- grade DJust Say No to Subtest Analysis: A Critique on Wechsler Theory and PracticeMcDermott PA, Fantuzzo JW, Glutting JJ · 1990 · critique
- grade DThe fallacy of "two years below grade level for age" as a diagnostic criterion for reading disordersReynolds CR · 1981 · critique
- grade CGifted Today but Not Tomorrow? Longitudinal Changes in Ability and Achievement during Elementary SchoolLohman DF, Korb KA · 2006 · longitudinal
- grade CMeasuring Test Measurement ErrorBoyd D, Lankford H, Loeb S, Wyckoff J · 2013 · quasi-experiment
- grade DOn the interpretation of below-chance responding in forced-choice tests.Frederick RI, Speed FM · 2007 · critique
- grade CRates of below-chance performance in forced-choice symptom validity tests.Greve KW, Binder LM, Bianchini KJ · 2009 · longitudinal
- grade CThe implications of symptom validity test failure for ability-based test performance in a pediatric sample.Kirkwood MW, Yeates KO, Randolph C, Kirk JW · 2012 · longitudinal
- grade DRapid-Guessing Behavior: Its Identification, Interpretation, and ImplicationsWise SL · 2017 · review
- grade CLow Examinee Effort in Low-Stakes Assessment: Problems and Potential SolutionsWise SL, DeMars CE · 2005 · review
- grade DTHE EFFECT OF ANSWER SHEET FORMAT ON TEST PERFORMANCE BY CULTURALLY DISADVANTAGED FOURTH GRADE ELEMENTARY SCHOOL PUPILSSolomon A · 1971 · quasi-experiment
- grade DThe Most Appropriate Scores for Measuring Educational Development in the Elementary Schools: GE'sHoover HD · 1984 · critique
- grade DInterpretation Problems of Age and Grade EquivalentsPearson Assessments (corporate author, no named authors) · 2024 · review
- grade CThe Problem With "Proficiency": Limitations of Statistics and Policy Under No Child Left BehindHo AD · 2008 · critique
- grade CThe Flynn effect and U.S. policies: the impact of rising IQ scores on American society via mental retardation diagnoses.Kanaya T, Scullin MH, Ceci SJ · 2003 · longitudinal
- grade DNationally Normed Elementary Achievement Testing in America's Public Schools: How All 50 States Are Above the National AverageCannell JJ · 1988 · critique
- grade CComparing State and District Test Results to National Norms: The Validity of Claims That "Everyone Is Above Average"Linn RL, Graue ME, Sanders NM · 1990 · quasi-experiment