Placement and mastery diagnosis — deciding what to teach next from evidence of current skill
Diagnose what a child already knows: teachers cut 40–50% of curriculum for high-ability children and achievement ROSE. The payoff is skipping, not monitoring.
moderate supportconf: mediumgc: mediumassessment · ages 5–18 · structure
Diagnosing what a child already knows pays, but not through the mechanism usually claimed. The largest single result is a NULL that saves a year: in a district-randomised trial of 27 districts and 783 high-ability children, teachers eliminated 40–50% of the curriculum as already mastered and achievement did not fall — it rose in maths concepts and science. Progress-monitoring effects have deflated hard, from 0.70 SD (1986, developer-led, special-ed) to g=0.24–0.38 in modern independent syntheses, and the moderator is consistent across both eras: explicit decision RULES ~0.92 vs teacher judgement ~0.42, graphed ~0.70 vs ungraphed ~0.26. Measurement alone is not the intervention. Single-sitting placement instruments are weak: college placement exams severely misassign ~1 in 4 (maths) to 1 in 3 (English), under-placing 2–6× more often than over-placing, and a single running record needs 3+ passages to be reliable. Above-level testing solves a real ceiling problem — a higher test level HALVES the standard error (14.8 → 7.4) — but has almost no direct validation.
Pretest, then delete. The strongest evidence here says a competent child can lose 40–50% of the curriculum with no measurable cost, so the highest-value diagnostic act is finding what to skip. Test high scorers ABOVE level — an on-level test cannot measure someone who cannot get it wrong, and going up one level halves the measurement error. Fix your decision rule and your graph before you start collecting data, because that is where the effect actually lives. And never place from one sitting: a single test misassigns a quarter to a third of the people it sorts, and it errs by holding capable students back.
Who this applies to
Verdict
moderate-support, and the support is for a narrower claim than the field usually advances.
The proposition that survives is: find out what the child already knows, and remove it. The strongest single piece of evidence in this topic is a district-randomised trial in which teachers deleted 40–50% of the curriculum for children who had already mastered it and nothing bad happened — achievement was unchanged in reading, maths computation, social studies and spelling, and higher in maths concepts and science. Between two fifths and half of a high-ability child's school year was purchasing nothing measurable.
The proposition that only partly survives is: measure continuously, and let the data drive instruction. It works, but the modern independent effect is g ≈ 0.24–0.38, not the 0.70 the founding meta-analysis reported — a textbook case of this archive's efficacy-to-effectiveness decay. And the moderator structure is unambiguous in both the 1986 and 2005 syntheses: what carries the effect is the explicit decision rule and the graph, not the measurement. Collecting data and leaving a teacher to interpret it recovers less than half the benefit.
The proposition that is assumed rather than evidenced is above-level testing. The theoretical case is sound and this topic endorses it — but Warne looked for the empirical validation and found that the only published reliability coefficient for an above-level achievement test given to a gifted sample dates from 1951. This archive should say that plainly rather than borrow the confidence of the talent-search tradition.
Grading note, as in the sibling topics: this is largely a measurement literature, and the A–D rubric is built for causal designs. Well-powered psychometric work is capped at C here and reviews at D. Two sources reach B because they are genuinely causal — a cluster-randomised trial and a large-administrative-data placement study with a real downstream criterion.
What the evidence shows
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Reis 1993 (Curriculum Compacting Study) | 27 districts randomly assigned, 436 teachers, 783 students, grades 2–6, out-of-level ITBS | B | 40–50% of curriculum eliminated; no difference in reading, maths computation, social studies, spelling; maths concepts HIGHER in all treatment groups; science higher in one |
| Scott-Clayton 2014 | 2 college systems, ~119,000 students, real course outcomes | B | ~1 in 4 maths / 1 in 3 English severely misassigned; under-placement 2–6× more common; transcripts remove 4–8 severe errors per 100 and raise success ~10pp (76%→89%); the test adds little on top |
| Lohman & Korb 2006 | reanalysis, 6,321 + 2,363 students | C | Only ~40% of top-3% grade-3 scorers still top-3% in grade 4; above-level testing halves SEM (14.8 → 7.4); averaging beats "highest of several" |
| Warne 2012 | comprehensive review | D | Grade-level tests ceiling above ~95th percentile, causing range restriction; only one published reliability coefficient exists for above-level achievement testing of gifted samples (Stanley 1951) |
| Warne 2014 | HLM growth, 224 students / 435 administrations | C | Above-level testing "has not been subject to careful psychometric scrutiny"; substantial ethnic differences in both intercept and growth rate |
| Fuchs & Fuchs 1986 | meta, 21 controlled studies, 96 ES, developer-led | C | 0.70 SD overall; decision rules ~0.92 vs judgement ~0.42; graphed ~0.70 vs ungraphed ~0.26; publication type a significant moderator |
| Filderman 2018 | meta, 15 studies, independent | C | DBDM vs BAU g = 0.24 (0.01–0.46); same intervention with vs without DBDM g = 0.27 (0.07–0.47); "experimental investigation is necessary to establish DBDM as an evidence-based practice" |
| Jung 2018 | meta, 14 studies, 57 ES | C | DBI vs BAU g = 0.37; "DBI-plus" 0.38 — the enhancements buy nothing |
| Gesel 2021 | meta of PD studies, teacher outcomes | C | g = 0.57 on teachers — roughly double the student effect, under "ideal, researcher-supported conditions" |
| Stecker 2005 | review, developer-led | D | Gains associated with systematic data-based decision rules, skills-analysis feedback, explicit programme-modification recommendations |
| Deno 1985 | founding position paper | D | Commercial tests are incongruent with curriculum; teachers' informal observation has UNKNOWN reliability and validity |
| Amendum 2017 | synthesis, 26 studies | C | Comprehension: 7 negative, 5 null, 1 optimum, ZERO positive — "in no study did a higher difficulty relate to greater comprehension"; above-grade text at ≥90% accuracy → significantly lower comprehension |
| Fawson 2006 | generalizability study | C | Raters contribute <1% of variance, but ≥3 passages needed for a reliable level (a later study implies 8–10) |
| Spector 2005 | review of 9 commercial IRIs | D | Fewer than half report ANY reliability evidence |
| Park 2013 | SMPY longitudinal (existing) | C | Above-level identification → skipping: 16.3% vs 7.9% STEM PhDs; selection-fragile by the authors' own sensitivity analysis |
| Bernstein 2021 | SMPY longitudinal + replication sample (existing) | C | Zero relation between acceleration and well-being at ~age 50 |
The three findings that matter for an intake tool
1. The highest-value diagnostic act is finding what to delete, and the evidence for it is the best in the topic. Reis et al. randomised at the district level across 27 districts, had teachers identify already-mastered content, removed 40–50% of it, and measured on out-of-level ITBS tests specifically chosen to avoid ceilings. Nothing fell. Maths concepts rose. This is a developer-led null — Reis and Renzulli invented compacting — which under this archive's rules makes it more credible than a developer-led positive would be, since the bias runs against the finding. Its own weak point is honestly recorded: 95% of the replacement activity was enrichment and only 18% acceleration, so it establishes "removing redundant content is safe" and not "acceleration raises achievement." For that second claim the archive already holds the SMPY corpus, where the upside is selection-fragile and the absence of harm is rock-solid.
2. You cannot place a student from a test they cannot get wrong, and the fix is measurable. Warne's review establishes that grade-level tests sit at or near ceiling above roughly the 95th percentile, which restricts range and attenuates everything computed from it. Lohman & Korb put a number on the remedy: moving to a higher test level halves the conditional standard error — on the CogAT Verbal battery, a scale score of 221 carries an SEM of 14.8 at Level A and 7.4 at Level D. That is the psychometric core of the talent-search method, and it is the same fact from the other end as the sibling topic's finding that measurement error is largest at the floor and the ceiling. A 99th-percentile score and a chance-level score are both, in the strict sense, uninformative — for the same reason.
3. Measurement is not the intervention; the rule is. This is the most consistent result across the whole progress-monitoring literature and it holds across a 40-year deflation in effect size. Fuchs & Fuchs found explicit decision rules worth ~0.92 against ~0.42 for teacher judgement, and graphed data worth ~0.70 against ~0.26 ungraphed. Stecker et al. reached the same conclusion narratively twenty years later: what distinguished effective from ineffective CBM was systematic decision rules, skills-analysis feedback, and explicit recommendations for changing the programme. Filderman's cleanest contrast — the same reading intervention with versus without data-based decision making — isolates the added value at g = 0.27. And Gesel's teacher-outcome meta (g = 0.57) sits at roughly double the student effect, which quantifies how much of the chain from data to learning happens after the adult understands the data.
Hereditarian-lens assessment
Risk: medium, and the reason is that the topic's most attractive evidence sits in an already-selected population.
The compacting trial, the SMPY corpus, and the above-level testing literature all operate on high-ability samples selected on the outcome-adjacent trait. Reis's students were teacher-identified as high ability with advanced prior knowledge; SMPY's participants are the top 1% on quantitative reasoning. Randomisation inside those samples is real and protects the internal contrast — Reis randomised districts, and the compacting null is a comparison between equivalent groups of high-ability children. But nothing here licenses extending "40–50% of the curriculum is redundant" to an unselected child, and the archive should not let the number travel.
Two more specific cautions:
- Warne 2014 found substantial ethnic-group differences in both above-level score intercepts and growth rates, with socioeconomic status adding nothing unique beyond ethnicity. Above-level testing is used as a screening gate by talent searches, which means those differences propagate into who gets identified. The paper reports them descriptively and does not attempt attribution; neither does this archive. It is recorded because a gate's differential yield is a fact a recommender needs.
- Deno's founding premise — that teachers' informal observation has unknown reliability and validity — is the direct hereditarian-adjacent risk for a homeschooling intake tool. A parent's assessment of their own child's mastery is the least independent measurement available: it is made by someone who shares the child's genes, chose the curriculum, taught the lesson, and has a stake in the answer. That is not an argument against parental judgement; it is an argument for pairing it with a standardised probe, which is exactly what curriculum-based measurement was invented to do.
The parts of the topic that are not confounded are the ones about instruments: the misassignment rates, the generalizability coefficients, the SEM reduction from above-level testing, and the text-difficulty synthesis. Those are properties of measurement procedures, and genes cannot differ between two administrations of a running record.
Boundaries & what critics say
- The 0.70 is not the number. Fuchs & Fuchs 1986 is 1980s special-education trials, small samples, developer-led, with publication type a significant moderator — the field's own admission of selection in its corpus. Modern independent syntheses of the same practice return 0.24 and 0.37. Quoting 0.70 as the expected effect of progress monitoring is the single most common overstatement in this area, and this archive's efficacy-to-effectiveness rule predicted the direction of the correction exactly.
- Filderman's own conclusion is that the practice is not yet established. For something as widely mandated as data-based decision making, "experimental investigation is necessary to establish DBDM as an evidence-based practice" is a striking sentence to find in the abstract of its own meta-analysis, and the main-effect confidence interval (0.01–0.46) very nearly touches zero.
- Above-level testing is endorsed here on theory plus one corroborating dataset, not on direct evidence. Warne's search for reliability coefficients turned up one, from 1951, on the Nelson-Denny. Warne 2014 is close to the only modern psychometric study of the practice and its own framing is that the practice "has not been subject to careful psychometric scrutiny." Lohman & Korb's SEM halving is the strongest quantitative support and it comes from a CogAT research handbook authored by Lohman himself — developer-involved, though pointing against the developer's interest.
- The instructional-level rule is not validated by this evidence, only bounded. Amendum et al.'s synthesis never tests the classic 95–98% accuracy threshold, so it cannot endorse it. What it does establish, across 26 studies with no contrary finding, is that harder text never helps: seven negative, five null, one optimum at moderate difficulty, zero positive. Three of the four nulls involved scaffolding, and even then students performed as well as with easier text, never better. So the aggressive placement option is ruled out; the specific cut-off remains unevidenced.
- The instruments used to assign reading level barely report whether they work. Fewer than half of nine commercial informal reading inventories reported any reliability evidence at all (Spector), and the generalizability work shows the rater is not the problem — the sample is. Teachers agree with each other almost perfectly (raters contribute <1% of variance) while needing at least three passages, and possibly eight to ten, for a stable estimate. Two teachers will give you the same wrong answer.
- Scott-Clayton's population is post-secondary and the numbers transfer as a mechanism, not a parameter. Community-college entrants are not fourth graders. What generalises is the structure: a single-sitting cut-score instrument misassigns a large minority, an accumulated performance record beats it, and the errors are asymmetric toward holding capable students back. That asymmetry is the one to carry into any placement decision.
- This topic deliberately does not re-litigate formative assessment, which the archive grades at
mixedunder feedback, or mastery pacing, gradedmixedunder mastery-learning. The scope here is the placement decision — what level to teach at, and what to skip — not the ongoing feedback loop.
Practical guidance
- Pretest before every unit, and act on the top of the distribution first. The compacting result says the biggest available win is deletion, not addition. For a child who already knows the content, up to half the year's material is dead weight and removing it costs nothing measurable.
- For a high scorer, test one grade level up. An on-level test that a child cannot get wrong measures nothing about them, and moving up a level halves the measurement error. This is the same fact as the chance-band problem at the floor, seen from the ceiling.
- Decide the rule before you collect the data. Write down, in advance, what change you will make when the data cross a line, and graph it. Across two independent literatures separated by twenty years, that is where roughly half the effect lives.
- Budget g ≈ 0.25, not 0.70. The honest expected value for adding data-based decision making to instruction you were doing anyway is about a quarter of a standard deviation.
- Never place from one sitting of anything. One placement test misassigns a quarter to a third of students. One running record does not reliably assign a reading level; use at least three passages. Where a performance record exists — past work, prior grades, accumulated observation — it beats the test, and adding the test on top of it adds little.
- When you must err, err toward the higher placement. Scott-Clayton's asymmetry is the actionable finding: under-placement runs two to six times more common than over-placement, so the systematic error of test-based placement is holding capable students back, and correcting for it means leaning up.
- Do not use harder material as an accelerant. Across 26 studies, no study found higher text difficulty produced better comprehension. Place at a level the child can work in and accelerate through content, not through difficulty of text.
- Pair parental judgement with an independent probe. The single most confounded measurement available to a homeschooling intake tool is the parent's own assessment of mastery, and the cheapest correction is a short criterion-referenced check the parent did not write.
Open questions
- The compacting trial has never been independently replicated in thirty years. It is the only randomised K-12 evidence on pretest-based content skipping this archive could locate, it is developer-led, and it carries an enormous share of the verdict. That is the topic's biggest fragility.
- Warne 2014's quantification of how much a grade-level test understates growth for a high scorer is recorded debt — it is in a paywalled full text that was not obtained. That number would materially sharpen the above-level recommendation.
- How far ahead to place remains genuinely unanswered. Above-level testing tells you a child is not at ceiling; it does not tell you which grade's curriculum to hand them. No study in this topic compares placement depths.
- Nothing here was tested in a home setting. Every effect is from schools, mostly special-education or gifted-programme contexts, delivered by teachers with researcher support. Whether a parent running curriculum-based measurement at a kitchen table recovers any of g = 0.24 is untested.
- The rules-versus-judgement moderator has never been randomised. It is a moderator in a 1986 meta-analysis and a narrative conclusion in a 2005 review, both by the method's developers. Given that it carries most of the practical advice in this topic, a direct trial — same measurement, decision rule versus teacher judgement, randomised — is the single most valuable missing study.
- The gap between teacher effects (g = 0.57) and student effects (g ≈ 0.24–0.38) is unexplained. Whatever is lost between an adult understanding the data and a child learning more is where the remaining headroom in this practice sits, and nobody has characterised it.
- grade BWhy Not Let High Ability Students Start School in January? The Curriculum Compacting StudyReis SM, Westberg KL, Kulikowich J, Caillard F, Hébert T, Plucker J, Purcell JH, Rogers JB, Smist JM · 1993 · rct
- grade CUsing Above-Level Testing to Track Growth in Academic Achievement in Gifted StudentsWarne RT · 2014 · longitudinal
- grade CGifted Today but Not Tomorrow? Longitudinal Changes in Ability and Achievement during Elementary SchoolLohman DF, Korb KA · 2006 · longitudinal
- grade BImproving the Targeting of TreatmentScott-Clayton J, Crosta PM, Belfield CR · 2014 · quasi-experiment
- grade CEffects of Systematic Formative Evaluation: A Meta-AnalysisFuchs LS, Fuchs D · 1986 · meta-analysis
- grade DUsing Curriculum-Based Measurement to Improve Student Achievement: Review of ResearchStecker PM, Fuchs LS, Fuchs D · 2005 · review
- grade CData-Based Decision Making in Reading Interventions: A Synthesis and Meta-Analysis of the Effects for Struggling ReadersFilderman MJ, Toste JR, Didion LA, Peng P, Clemens NH · 2018 · meta-analysis
- grade CEffects of Data-Based Individualization for Students With Intensive Learning Needs: A Meta-AnalysisJung P-G, McMaster KL, Kunkel AK, Shin J, Stecker PM · 2018 · meta-analysis
- grade CA Meta-Analysis of the Impact of Professional Development on Teachers' Knowledge, Skill, and Self-Efficacy in Data-Based Decision-MakingGesel SA, LeJeune LM, Chow JC, Sinclair AC, Lemons CJ · 2021 · meta-analysis
- grade CDoes Text Complexity Matter in the Elementary Grades? A Research Synthesis of Text Difficulty and Elementary Students' Reading Fluency and ComprehensionAmendum SJ, Conradi K, Hiebert E · 2017 · review
- grade CExamining the Reliability of Running Records: Attaining Generalizable ResultsFawson PC, Ludlow BC, Reutzel DR, Sudweeks R, Smith JA · 2006 · quasi-experiment
- grade CWhen less is more: Effects of grade skipping on adult STEM productivity among mathematically precocious adolescentsPark G, Lubinski D, Benbow CP · 2013 · longitudinal
- grade CAcademic acceleration in gifted youth and fruitless concerns regarding psychological well-being: A 35-year longitudinal studyBernstein BO, Lubinski D, Benbow CP · 2021 · longitudinal
Related decisions
- Does educational technology raise learning — CAI, adaptive software, devices, screens, and AI tutors?mixedconf: mediumgc: low
- Feedback and formative assessmentmixedconf: highgc: low
- Mastery learning (teach → test → reteach to criterion → advance)mixedconf: highgc: low