The Evidence on Teaching

Placement and mastery diagnosis — deciding what to teach next from evidence of current skill

Diagnose what a child already knows: teachers cut 40–50% of curriculum for high-ability children and achievement ROSE. The payoff is skipping, not monitoring.

moderate supportconf: mediumgc: medium

assessment · ages 518 · structure

Effect summary

Diagnosing what a child already knows pays, but not through the mechanism usually claimed. The largest single result is a NULL that saves a year: in a district-randomised trial of 27 districts and 783 high-ability children, teachers eliminated 40–50% of the curriculum as already mastered and achievement did not fall — it rose in maths concepts and science. Progress-monitoring effects have deflated hard, from 0.70 SD (1986, developer-led, special-ed) to g=0.24–0.38 in modern independent syntheses, and the moderator is consistent across both eras: explicit decision RULES ~0.92 vs teacher judgement ~0.42, graphed ~0.70 vs ungraphed ~0.26. Measurement alone is not the intervention. Single-sitting placement instruments are weak: college placement exams severely misassign ~1 in 4 (maths) to 1 in 3 (English), under-placing 2–6× more often than over-placing, and a single running record needs 3+ passages to be reliable. Above-level testing solves a real ceiling problem — a higher test level HALVES the standard error (14.8 → 7.4) — but has almost no direct validation.

Practical takeaway

Pretest, then delete. The strongest evidence here says a competent child can lose 40–50% of the curriculum with no measurable cost, so the highest-value diagnostic act is finding what to skip. Test high scorers ABOVE level — an on-level test cannot measure someone who cannot get it wrong, and going up one level halves the measurement error. Fix your decision rule and your graph before you start collecting data, because that is where the effect actually lives. And never place from one sitting: a single test misassigns a quarter to a third of the people it sorts, and it errs by holding capable students back.

Who this applies to

Group size
one-to-onesmall-grouphome
Delivered by
teachertutorparent
Ages studied
718(narrower than the 518 this topic is filed under — outside it is extrapolation)
Dose
Pretest before a unit and remove what is already mastered — 40–50% is removable for a high-ability child without loss. For ongoing monitoring, 2–5 measures per week with an explicit decision rule and a graph. For reading level, ≥3 passages before assigning a level. For a high scorer, test one grade level up.
Cost
low
Moves
domain-skill
Needs first
A decision rule fixed in advance for what you will change when the data say so. Every synthesis in this topic finds the response, not the measurement, carries the effect — collecting data and leaving interpretation to judgement recovers less than half the benefit.
Not for
Placing a child from a single sitting of anything, placing from a grade equivalent, or placing a high scorer from an on-level test they cannot get wrong. Also not for pushing a child into harder material as an accelerant: across 26 studies, in NO study did higher text difficulty relate to greater comprehension.

Verdict

moderate-support, and the support is for a narrower claim than the field usually advances.

The proposition that survives is: find out what the child already knows, and remove it. The strongest single piece of evidence in this topic is a district-randomised trial in which teachers deleted 40–50% of the curriculum for children who had already mastered it and nothing bad happened — achievement was unchanged in reading, maths computation, social studies and spelling, and higher in maths concepts and science. Between two fifths and half of a high-ability child's school year was purchasing nothing measurable.

The proposition that only partly survives is: measure continuously, and let the data drive instruction. It works, but the modern independent effect is g ≈ 0.24–0.38, not the 0.70 the founding meta-analysis reported — a textbook case of this archive's efficacy-to-effectiveness decay. And the moderator structure is unambiguous in both the 1986 and 2005 syntheses: what carries the effect is the explicit decision rule and the graph, not the measurement. Collecting data and leaving a teacher to interpret it recovers less than half the benefit.

The proposition that is assumed rather than evidenced is above-level testing. The theoretical case is sound and this topic endorses it — but Warne looked for the empirical validation and found that the only published reliability coefficient for an above-level achievement test given to a gifted sample dates from 1951. This archive should say that plainly rather than borrow the confidence of the talent-search tradition.

Grading note, as in the sibling topics: this is largely a measurement literature, and the A–D rubric is built for causal designs. Well-powered psychometric work is capped at C here and reviews at D. Two sources reach B because they are genuinely causal — a cluster-randomised trial and a large-administrative-data placement study with a real downstream criterion.

What the evidence shows

Source Design Grade Key effect
Reis 1993 (Curriculum Compacting Study) 27 districts randomly assigned, 436 teachers, 783 students, grades 2–6, out-of-level ITBS B 40–50% of curriculum eliminated; no difference in reading, maths computation, social studies, spelling; maths concepts HIGHER in all treatment groups; science higher in one
Scott-Clayton 2014 2 college systems, ~119,000 students, real course outcomes B ~1 in 4 maths / 1 in 3 English severely misassigned; under-placement 2–6× more common; transcripts remove 4–8 severe errors per 100 and raise success ~10pp (76%→89%); the test adds little on top
Lohman & Korb 2006 reanalysis, 6,321 + 2,363 students C Only ~40% of top-3% grade-3 scorers still top-3% in grade 4; above-level testing halves SEM (14.8 → 7.4); averaging beats "highest of several"
Warne 2012 comprehensive review D Grade-level tests ceiling above ~95th percentile, causing range restriction; only one published reliability coefficient exists for above-level achievement testing of gifted samples (Stanley 1951)
Warne 2014 HLM growth, 224 students / 435 administrations C Above-level testing "has not been subject to careful psychometric scrutiny"; substantial ethnic differences in both intercept and growth rate
Fuchs & Fuchs 1986 meta, 21 controlled studies, 96 ES, developer-led C 0.70 SD overall; decision rules ~0.92 vs judgement ~0.42; graphed ~0.70 vs ungraphed ~0.26; publication type a significant moderator
Filderman 2018 meta, 15 studies, independent C DBDM vs BAU g = 0.24 (0.01–0.46); same intervention with vs without DBDM g = 0.27 (0.07–0.47); "experimental investigation is necessary to establish DBDM as an evidence-based practice"
Jung 2018 meta, 14 studies, 57 ES C DBI vs BAU g = 0.37; "DBI-plus" 0.38 — the enhancements buy nothing
Gesel 2021 meta of PD studies, teacher outcomes C g = 0.57 on teachers — roughly double the student effect, under "ideal, researcher-supported conditions"
Stecker 2005 review, developer-led D Gains associated with systematic data-based decision rules, skills-analysis feedback, explicit programme-modification recommendations
Deno 1985 founding position paper D Commercial tests are incongruent with curriculum; teachers' informal observation has UNKNOWN reliability and validity
Amendum 2017 synthesis, 26 studies C Comprehension: 7 negative, 5 null, 1 optimum, ZERO positive — "in no study did a higher difficulty relate to greater comprehension"; above-grade text at ≥90% accuracy → significantly lower comprehension
Fawson 2006 generalizability study C Raters contribute <1% of variance, but ≥3 passages needed for a reliable level (a later study implies 8–10)
Spector 2005 review of 9 commercial IRIs D Fewer than half report ANY reliability evidence
Park 2013 SMPY longitudinal (existing) C Above-level identification → skipping: 16.3% vs 7.9% STEM PhDs; selection-fragile by the authors' own sensitivity analysis
Bernstein 2021 SMPY longitudinal + replication sample (existing) C Zero relation between acceleration and well-being at ~age 50

The three findings that matter for an intake tool

1. The highest-value diagnostic act is finding what to delete, and the evidence for it is the best in the topic. Reis et al. randomised at the district level across 27 districts, had teachers identify already-mastered content, removed 40–50% of it, and measured on out-of-level ITBS tests specifically chosen to avoid ceilings. Nothing fell. Maths concepts rose. This is a developer-led null — Reis and Renzulli invented compacting — which under this archive's rules makes it more credible than a developer-led positive would be, since the bias runs against the finding. Its own weak point is honestly recorded: 95% of the replacement activity was enrichment and only 18% acceleration, so it establishes "removing redundant content is safe" and not "acceleration raises achievement." For that second claim the archive already holds the SMPY corpus, where the upside is selection-fragile and the absence of harm is rock-solid.

2. You cannot place a student from a test they cannot get wrong, and the fix is measurable. Warne's review establishes that grade-level tests sit at or near ceiling above roughly the 95th percentile, which restricts range and attenuates everything computed from it. Lohman & Korb put a number on the remedy: moving to a higher test level halves the conditional standard error — on the CogAT Verbal battery, a scale score of 221 carries an SEM of 14.8 at Level A and 7.4 at Level D. That is the psychometric core of the talent-search method, and it is the same fact from the other end as the sibling topic's finding that measurement error is largest at the floor and the ceiling. A 99th-percentile score and a chance-level score are both, in the strict sense, uninformative — for the same reason.

3. Measurement is not the intervention; the rule is. This is the most consistent result across the whole progress-monitoring literature and it holds across a 40-year deflation in effect size. Fuchs & Fuchs found explicit decision rules worth ~0.92 against ~0.42 for teacher judgement, and graphed data worth ~0.70 against ~0.26 ungraphed. Stecker et al. reached the same conclusion narratively twenty years later: what distinguished effective from ineffective CBM was systematic decision rules, skills-analysis feedback, and explicit recommendations for changing the programme. Filderman's cleanest contrast — the same reading intervention with versus without data-based decision making — isolates the added value at g = 0.27. And Gesel's teacher-outcome meta (g = 0.57) sits at roughly double the student effect, which quantifies how much of the chain from data to learning happens after the adult understands the data.

Hereditarian-lens assessment

Risk: medium, and the reason is that the topic's most attractive evidence sits in an already-selected population.

The compacting trial, the SMPY corpus, and the above-level testing literature all operate on high-ability samples selected on the outcome-adjacent trait. Reis's students were teacher-identified as high ability with advanced prior knowledge; SMPY's participants are the top 1% on quantitative reasoning. Randomisation inside those samples is real and protects the internal contrast — Reis randomised districts, and the compacting null is a comparison between equivalent groups of high-ability children. But nothing here licenses extending "40–50% of the curriculum is redundant" to an unselected child, and the archive should not let the number travel.

Two more specific cautions:

  • Warne 2014 found substantial ethnic-group differences in both above-level score intercepts and growth rates, with socioeconomic status adding nothing unique beyond ethnicity. Above-level testing is used as a screening gate by talent searches, which means those differences propagate into who gets identified. The paper reports them descriptively and does not attempt attribution; neither does this archive. It is recorded because a gate's differential yield is a fact a recommender needs.
  • Deno's founding premise — that teachers' informal observation has unknown reliability and validity — is the direct hereditarian-adjacent risk for a homeschooling intake tool. A parent's assessment of their own child's mastery is the least independent measurement available: it is made by someone who shares the child's genes, chose the curriculum, taught the lesson, and has a stake in the answer. That is not an argument against parental judgement; it is an argument for pairing it with a standardised probe, which is exactly what curriculum-based measurement was invented to do.

The parts of the topic that are not confounded are the ones about instruments: the misassignment rates, the generalizability coefficients, the SEM reduction from above-level testing, and the text-difficulty synthesis. Those are properties of measurement procedures, and genes cannot differ between two administrations of a running record.

Boundaries & what critics say

  • The 0.70 is not the number. Fuchs & Fuchs 1986 is 1980s special-education trials, small samples, developer-led, with publication type a significant moderator — the field's own admission of selection in its corpus. Modern independent syntheses of the same practice return 0.24 and 0.37. Quoting 0.70 as the expected effect of progress monitoring is the single most common overstatement in this area, and this archive's efficacy-to-effectiveness rule predicted the direction of the correction exactly.
  • Filderman's own conclusion is that the practice is not yet established. For something as widely mandated as data-based decision making, "experimental investigation is necessary to establish DBDM as an evidence-based practice" is a striking sentence to find in the abstract of its own meta-analysis, and the main-effect confidence interval (0.01–0.46) very nearly touches zero.
  • Above-level testing is endorsed here on theory plus one corroborating dataset, not on direct evidence. Warne's search for reliability coefficients turned up one, from 1951, on the Nelson-Denny. Warne 2014 is close to the only modern psychometric study of the practice and its own framing is that the practice "has not been subject to careful psychometric scrutiny." Lohman & Korb's SEM halving is the strongest quantitative support and it comes from a CogAT research handbook authored by Lohman himself — developer-involved, though pointing against the developer's interest.
  • The instructional-level rule is not validated by this evidence, only bounded. Amendum et al.'s synthesis never tests the classic 95–98% accuracy threshold, so it cannot endorse it. What it does establish, across 26 studies with no contrary finding, is that harder text never helps: seven negative, five null, one optimum at moderate difficulty, zero positive. Three of the four nulls involved scaffolding, and even then students performed as well as with easier text, never better. So the aggressive placement option is ruled out; the specific cut-off remains unevidenced.
  • The instruments used to assign reading level barely report whether they work. Fewer than half of nine commercial informal reading inventories reported any reliability evidence at all (Spector), and the generalizability work shows the rater is not the problem — the sample is. Teachers agree with each other almost perfectly (raters contribute <1% of variance) while needing at least three passages, and possibly eight to ten, for a stable estimate. Two teachers will give you the same wrong answer.
  • Scott-Clayton's population is post-secondary and the numbers transfer as a mechanism, not a parameter. Community-college entrants are not fourth graders. What generalises is the structure: a single-sitting cut-score instrument misassigns a large minority, an accumulated performance record beats it, and the errors are asymmetric toward holding capable students back. That asymmetry is the one to carry into any placement decision.
  • This topic deliberately does not re-litigate formative assessment, which the archive grades at mixed under feedback, or mastery pacing, graded mixed under mastery-learning. The scope here is the placement decision — what level to teach at, and what to skip — not the ongoing feedback loop.

Practical guidance

  • Pretest before every unit, and act on the top of the distribution first. The compacting result says the biggest available win is deletion, not addition. For a child who already knows the content, up to half the year's material is dead weight and removing it costs nothing measurable.
  • For a high scorer, test one grade level up. An on-level test that a child cannot get wrong measures nothing about them, and moving up a level halves the measurement error. This is the same fact as the chance-band problem at the floor, seen from the ceiling.
  • Decide the rule before you collect the data. Write down, in advance, what change you will make when the data cross a line, and graph it. Across two independent literatures separated by twenty years, that is where roughly half the effect lives.
  • Budget g ≈ 0.25, not 0.70. The honest expected value for adding data-based decision making to instruction you were doing anyway is about a quarter of a standard deviation.
  • Never place from one sitting of anything. One placement test misassigns a quarter to a third of students. One running record does not reliably assign a reading level; use at least three passages. Where a performance record exists — past work, prior grades, accumulated observation — it beats the test, and adding the test on top of it adds little.
  • When you must err, err toward the higher placement. Scott-Clayton's asymmetry is the actionable finding: under-placement runs two to six times more common than over-placement, so the systematic error of test-based placement is holding capable students back, and correcting for it means leaning up.
  • Do not use harder material as an accelerant. Across 26 studies, no study found higher text difficulty produced better comprehension. Place at a level the child can work in and accelerate through content, not through difficulty of text.
  • Pair parental judgement with an independent probe. The single most confounded measurement available to a homeschooling intake tool is the parent's own assessment of mastery, and the cheapest correction is a short criterion-referenced check the parent did not write.

Open questions

  • The compacting trial has never been independently replicated in thirty years. It is the only randomised K-12 evidence on pretest-based content skipping this archive could locate, it is developer-led, and it carries an enormous share of the verdict. That is the topic's biggest fragility.
  • Warne 2014's quantification of how much a grade-level test understates growth for a high scorer is recorded debt — it is in a paywalled full text that was not obtained. That number would materially sharpen the above-level recommendation.
  • How far ahead to place remains genuinely unanswered. Above-level testing tells you a child is not at ceiling; it does not tell you which grade's curriculum to hand them. No study in this topic compares placement depths.
  • Nothing here was tested in a home setting. Every effect is from schools, mostly special-education or gifted-programme contexts, delivered by teachers with researcher support. Whether a parent running curriculum-based measurement at a kitchen table recovers any of g = 0.24 is untested.
  • The rules-versus-judgement moderator has never been randomised. It is a moderator in a 1986 meta-analysis and a narrative conclusion in a 2005 review, both by the method's developers. Given that it carries most of the practical advice in this topic, a direct trial — same measurement, decision rule versus teacher judgement, randomised — is the single most valuable missing study.
  • The gap between teacher effects (g = 0.57) and student effects (g ≈ 0.24–0.38) is unexplained. Whatever is lost between an adult understanding the data and a child learning more is where the remaining headroom in this practice sits, and nobody has characterised it.

Evidence (16 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore