Ability grouping and tracking — four practices, four verdicts
Group within classes or across grades by subject: modest, nearly free wins. Whole-school streaming does nothing, and early between-school tracking harms the bottom.
moderate supportconf: highgc: lowgrouping · ages 5–18 · structure
Never pool the practices. Within-class grouping (g≈0.19-0.30) and cross-grade subject grouping (g≈0.26) are modest ~free wins. Undifferentiated between-class streaming does ~nothing (g≈0.04-0.06). The one large tracking RCT (Kenya) raised scores across the WHOLE distribution (+0.14-0.18, bottom half +0.155) via teach-to-level, not peers. BUT early rigid between-SCHOOL tracking (European style) harms low achievers — B-grade causal evidence. The rule: group by subject-specific readiness with genuinely differentiated instruction, keep groups permeable, reassess often.
Group flexibly by subject-specific measured readiness WITH differentiated curriculum (grouping without instructional adaptation is pointless), reassess placements frequently, and never build permanent tracks that lock children in. Run separate classrooms for high achievers selected on demonstrated achievement (+0.5 SD for minority high achievers, no harm to others).
Who this applies to
Not yet assessed. Nobody has recorded the group size, dose, delivery, or boundary conditions for this decision, so it should not be recommended for a specific situation yet — only read. That is a gap in this record, not a claim that it applies everywhere.
Verdict
"Tracking" is four different practices with four different evidence bases, and the debate stays confused by pooling them (taxonomy anchor):
- Within-class grouping (small groups inside one classroom): g≈0.19–0.30, essentially free.
- Cross-grade subject grouping (Joplin plan — regroup across grades by subject readiness): g≈0.26, essentially free.
- Between-class streaming with undifferentiated instruction: g≈0.04–0.06 — nothing. But the one large RCT (Duflo/Dupas/Kremer, Kenya) shows tracking with instructional adaptation raised scores across the entire distribution — bottom half +0.155 — via teach-to-level (median students gained equally in either section, rejecting peer-effects stories). Grouping works exactly insofar as instruction adapts. (Context caveats: grade 1, ~80-100-pupil baseline classes, contract teachers — don't export the magnitude.)
- Separate high-achiever classrooms selected on demonstrated achievement: +0.5 SD for Black/Hispanic high achievers, persistent, no spillover harm, ~zero cost — while the same district's IQ-labelled gifted margin produces nothing, replicated in a second district by Bui, Craig & Imberman.
The essential scope limit (added by adversarial review): early, rigid, between-school tracking — sorting ten-year-olds into separate school types, European style — does harm low achievers on B-grade causal evidence. Two extra comprehensive years raise math and reading with the gain almost entirely in low achievers and nothing lost by high achievers; the cross-country diff-in-diff (weaker, C) points the same way; and three national detracking reforms agree — Sweden (attainment and earnings up for able children of unskilled fathers), Finland (father-son income elasticity down ~23%), and France (wages +4.7% at 40-45, concentrated in low-SES adults). Within-school flexible grouping and between-school child-sorting are different treatments; the harm evidence attaches to the latter, and where tracks are permeable, marginal track assignment has no long-run effect. Meanwhile the meta-analytic record shows no support for the claim that within-school grouping harms low-ability students — effects don't differ by ability level.
Detracking's flagship failed audit: San Francisco's celebrated reform rested on a repeat-rate statistic that was a policy-change artifact; AP-math participation fell 15% (−6pp) with ethnoracial gaps intact, and the district reversed course (source).
Hereditarian-lens assessment
Risk: low for the verdict (RCT/RD core). This is selection-confound central: raw tracked-vs-untracked comparisons mostly measure who sorts where (on substantially heritable ability), which is why only the experimental cells carry weight. The lens adds the design principle: since ability differences are real and persistent, the choice isn't whether children differ but whether instruction meets them where they are (teach-to-level, the one mechanism with RCT support) — while keeping placements permeable and re-measured, because locking children into tracks at 10 is where causal harm actually shows up.
Practical guidance
- Group by subject-specific measured readiness (a child can be ahead in math, behind in reading), with genuinely differentiated curriculum per group — otherwise don't bother.
- Reassess frequently (each term/year); permeability is what separates benign grouping from harmful tracking.
- Run achievement-selected advanced classrooms — selected on demonstrated performance, not IQ labels or teacher referral.
- Don't detrack on equity grounds — the causal record shows no within-school harm to remove, and the flagship detracking reform failed its audit.
- grade CWhat One Hundred Years of Research Says About Ability Grouping and Acceleration (second-order meta-analysis)Steenbergen-Hu, S., Makel, M. C., & Olszewski-Kubilius, P. · 2016 · meta-analysis
- grade BPeer Effects, Teacher Incentives, and the Impact of Tracking (Kenya RCT)Duflo, E., Dupas, P., & Kremer, M. · 2011 · rct
- grade BCan Tracking Raise the Test Scores of High-Ability Minority Students?Card D, Giuliano L · 2016 · quasi-experiment
- grade BIs Gifted Education a Bright Idea? Assessing the Impact of Gifted and Talented Programs on StudentsBui S, Craig S, Imberman S · 2014 · quasi-experiment
- grade BBetter Together? Heterogeneous Effects of Tracking on Student AchievementMatthewes SH · 2021 · natural-experiment
- grade CDoes Educational Tracking Affect Performance and Inequality? Differences-in-Differences Evidence Across CountriesHanushek EA, Wößmann L · 2006 · natural-experiment
- grade BEducational Reform, Ability, and Family BackgroundMeghir C, Palme M · 2005 · natural-experiment
- grade BSchool tracking and intergenerational income mobility: Evidence from the Finnish comprehensive school reformPekkarinen T, Uusitalo R, Kerr S · 2009 · natural-experiment
- grade BThe Long-Term Effects of Early Track ChoiceDustmann C, Puhani PA, Schönberg U · 2017 · natural-experiment
- grade DAhead of the Game? Course-Taking Patterns Under a Math Pathways ReformHuffaker, E., Novicoff, S., & Dee, T. S. · 2025 · longitudinal
Related decisions
- Do career and technical education tracks help students — and which students?moderate supportconf: mediumgc: low