How should a school run its classrooms and its discipline system?
Classroom order is causally worth a great deal — one disruptive peer costs classmates 3% of adult earnings — yet no branded behaviour programme reliably delivers it and exclusion makes things worse.
mixedconf: mediumgc: lowbehavior · ages 4–18 · structure
Three separate claims, three different answers. (1) ORDER IS CAUSALLY VALUABLE: three independent identification strategies agree that a disruptive classmate lowers peers' test scores, and one disruptive peer in a class of 25 reduces classmates' earnings at ages 24-28 by 3% (~$80,000 PDV per year of exposure). (2) PROGRAMMES UNDERDELIVER: generic classroom management is a stable, small g = 0.22-0.23 across two meta-analyses nine years apart, but the flagship branded programme fails independent replication — the Good Behavior Game's Baltimore trial reported halved young-adult drug disorders in high-aggressive males, while a preregistered 77-school English trial found reading 0.03, disruptive behaviour 0.06, prosocial -0.13, all null, with a quarter of schools abandoning it. School-wide PBIS reliably cuts office referrals (33% less likely) and suspensions, and reliably fails to move achievement. (3) EXCLUSION IS THE WRONG TOOL: strict-discipline schools raise adult arrest and incarceration 15-20%; state zero-tolerance laws raised suspensions 0.5pp with no behaviour improvement; New York City's 2012 removal of suspensions for non-violent conduct raised maths 0.05 SD and reading 0.03 SD with no trade-off for other students.
Spend on the room, not on the programme. Order is worth real money — a single disruptive classmate costs the other children about 3% of their adult earnings — so protecting instructional time is a legitimate, well-evidenced goal. But buy it with ordinary classroom management (clear rules, routines, immediate contingent response, teacher-focused training), which returns a small, replicated g ≈ 0.22, rather than with a branded system: the Good Behavior Game is null in its only independent large trial, restorative practices lowered maths scores in its only randomised achievement test, and school-wide PBIS moves the discipline paperwork more convincingly than it moves the children. Above all, do not run the school on suspension. It lowers the suspended child's achievement, raises his adult arrest risk by 15-20%, and the one city that removed suspensions for non-violent conduct got HIGHER test scores with no cost to anyone else.
Who this applies to
Verdict
mixed, and the mixture is not a hedge — it is three distinct claims with three different
answers. Splitting them is the whole job, because the school-behaviour debate is conducted as
though they were one.
- Does classroom order causally matter? Yes, and this is the best-identified finding in the file. It is also the claim least often supported with evidence by the people who assert it loudest.
- Can a behaviour programme deliver order? A little, and the branded ones deliver less than the generic practice. The most celebrated programme in the field is null in its only large independent trial.
- Is exclusion the way to get it? No. Exclusion is causally harmful to the excluded child, its statutory form raised punishment without improving behaviour, and the one city that removed it for non-violent conduct got higher test scores.
This is a topic where the archive's usual scepticism lands on the programmes and not on the phenomenon. Order is real, valuable, and measurable. What is oversold is the idea that you can buy it in a box.
What the evidence shows
Order is worth money — three instruments, one answer
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Carrell, Hoekstra & Kuka 2018 | Peer-composition quasi-experiment, admin earnings | B | One disruptive peer in a class of 25 → −3% earnings at 24-28; ≈ −$80,000 PDV per year of exposure; explains 5% of the rich-poor earnings gap |
| Carrell & Hoekstra 2010 | Domestic-violence records matched to classrooms | B | Significant reductions in classmates' reading and math; survives within-family differencing and school-by-year FE |
| Figlio 2007 | IV from boys with girls' names | B | Disruption reduces peer test scores and raises peer disciplinary problems — a wholly independent instrument |
Three research designs with nothing in common except the question, all agreeing. Note especially the within-family specification in Carrell & Hoekstra: the estimate survives comparing siblings, so it cannot be the heritable characteristics of the affected children doing the work. In an archive that treats observational education research as presumptively confounded, this is one of the few places where the causal claim is clean and the magnitude is large.
Programmes: a small stable generic effect, and a famous programme that does not replicate
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Korpershoek 2016 | Meta, 54 controlled studies, primary | C | g = 0.22 on academic + behavioural + social-emotional; null on motivation |
| Korpershoek 2025 | Update, 76 studies, 2003-2022 | C | g = 0.23 (RVE 0.22). Broad four-component programmes did worse than targeted ones |
| Oliver 2011 (Campbell) | Campbell review, 12 studies | C | Multi-component teacher CM programmes reduce disruptive/aggressive behaviour; components not separable |
| Humphrey 2018 (EEF GBG) | Preregistered 77-school cluster RCT, independent | A | Reading 0.03 (−0.08-0.16); concentration 0.03; disruptive 0.06; prosocial −0.13. All null. ¼ of schools quit |
| Kellam 2008 (Baltimore GBG) | Classroom-randomised trial, followed to 19-21 | B | Males' drug abuse/dependence 19% vs 38%; high-aggressive males 29% vs 83%. Cohort 2 lost it |
| Ashworth 2020 (GBG CACE) | CACE re-analysis + 1y follow-up | B | ITT null throughout; Δ = 0.10 for moderate compliers only at 1 year |
| [Bowman-Perrott 2014](../sources/bhv-flower-2014-gbg-meta, bhv-bowman-perrott-2015-gbg-single-case-meta.yaml) | Meta of 22 single-case studies | C | Moderate-to-large immediate effects on in-class behaviour; no academic outcomes |
| Bradshaw 2012 | 37-school RCT, 12,344 children | B | 33% less likely to receive an office discipline referral; teacher-rated outcomes also improved |
| Bradshaw 2009 | Same trial, school level, 5 years | B | Significant reductions in suspensions and referrals; achievement not the clear positive |
| Horner 2009 | Wait-list RCT, HI + IL | C | Training → fidelity → perceived safety and 3rd-grade reading; ODRs not causally estimable |
| Lee & Gage 2020 | Meta, 29 studies (7 RCT, 22 QED) | C | Significant reductions in discipline, significant achievement gain, small to medium, individual studies inconsistent |
The Good Behavior Game is the archive's cleanest live example of the replication rule. Kellam's Baltimore trial is genuinely impressive on its face — classroom-level randomisation, follow-up to ages 19-21, halved rates of young-adult drug abuse and dependence in males, and 75% vs 20% high-school graduation among high-aggressive boys. It is also developer-involved, the headline results are subgroup estimates in small cells, the outcomes are self-report diagnostic interviews, and the study's own second cohort did not reproduce them: alcohol and antisocial-personality effects vanished, and the baseline-aggression moderator that the entire theory rests on reversed.
Then England ran the trial the theory deserved: 77 schools, preregistered, independently evaluated by Manchester, an independent standardised reading measure, two full years of delivery, EEF "high security" rating. Everything was null, prosocial behaviour pointed the wrong way, and a quarter of the schools stopped playing.
Under this archive's replication rule, a finding with a failed independent replication cannot
exceed mixed regardless of the original effect size, and the GBG is the case the rule was
written for.
School-wide PBIS has the opposite profile: it is real but it is narrow. The Maryland trials are large, randomised, five years long, and report a 33% reduction in the likelihood of an office discipline referral — an administrative outcome, not a rating. That is worth having. What PBIS does not do is raise achievement, and the achievement claims in the literature come from a mediated analysis in a small wait-list trial (Horner) and from a meta-analysis three-quarters composed of quasi-experiments (Lee & Gage). Note also the measurement hazard PBIS is uniquely exposed to: the intervention explicitly trains staff to respond to behaviour differently, and then the outcome is how often staff record a behaviour incident.
Generic classroom management is the boring winner. Two meta-analyses nine years apart, by the same group, on largely different studies, return g = 0.22 and g = 0.23. That is small, but it is the most stable number in this tranche, and the 2025 moderator result argues directly against the comprehensive-package model: interventions that tried to do all four things at once did worse than targeted ones.
Exclusion: the tool that makes it worse
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Bacher-Hicks 2024 | Boundary change + principal switches, admin arrest records | B | Strict-discipline schools → 15-20% more likely arrested and incarcerated as adults |
| Craig & Martin 2025 | DiD, NYC 2012 removal of non-violent suspensions | B | Maths +0.05 SD, reading +0.03 SD over 3 years; runs through school culture; no trade-off for other students |
| Lacoe & Steinberg 2018 | Student FE + IV, Philadelphia | B | Suspension lowers the suspended student's maths and reading; peer spillovers very small |
| Curran 2016 | State + year FE | C | Zero-tolerance laws → +0.5pp suspensions, no behaviour improvement, wider Black-White gap |
| APA 2008 | Evidentiary review | C | After ~20 years, "surprisingly few data… and the data that are available tend to contradict those assumptions" |
| Valdebenito 2018 (Campbell) | Campbell review, RCTs only, k = 37 | B | Exclusion falls for 6 months only; no effect on antisocial behaviour; independent evaluators report lower effects |
| Sorensen 2021 | SRO placement, 4 linked admin systems | B | SROs cut serious violence AND raise suspensions, transfers, expulsions, police referrals |
Craig & Martin is the study that resolves the standard objection. Everyone accepts that suspension is bad for the suspended child; the interesting question is whether the classroom needs it anyway. New York City's 2012 reform removed suspension for non-violent disorderly conduct, and three years later more-affected schools had higher maths and reading scores — with most of the gain arriving through improved measured school culture rather than through the handful of children who avoided suspension, and with no losers. That is the opposite of what the trade-off argument predicts.
But the Sorensen finding keeps this honest. School resource officers reduce serious violence and escalate disciplinary response. Both are true. "Reduce exclusion" and "keep the school safe" are genuinely different objectives, and the evidence does not license pretending otherwise.
Restorative practices: one partial positive, one significant negative, two nulls
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Augustine 2018 (RAND Pittsburgh) | 44-school pair-randomised RCT, 2 years | B | Suspension days −36% vs −18%; but state maths −0.068 (significant), combined −0.065, worst in grades 6-8 |
| Huang 2023 | 18-school cluster RCT, 5,878 students | B | Null main effect (B = −0.001, p = 0.95); large effect only for previously suspended students |
| Acosta 2019 | 13-school 2-year cluster RCT | B | Null on student-perceived school climate |
The Pittsburgh trial is the one everybody cites, and it is worth reading past the headline. The suspension reduction was real and concentrated in elementary schools; it did not appear in middle grades, for boys, for students with IEPs, for violent incidents, or in arrests. Meanwhile achievement moved the wrong way, significantly in maths. And the authors' own diagnostic is important: the suspension benefit and the achievement harm occurred in different grade bands, so "more time in class" cannot be the mechanism and neither can "less time in class." Something about the implementation cost instruction.
Hereditarian-lens assessment
Risk: low, and this topic is one of the archive's better demonstrations of why the premise does not imply nihilism.
Behaviour and self-regulation are substantially heritable, and the ordinary correlational discipline literature — suspended children do worse, so suspension must cause it — is confounded exactly as the premise predicts. This file deliberately does not rest on that literature. Every load-bearing estimate here comes from a design that breaks the confound: classroom or school randomisation (GBG, PBIS, restorative practices), a school-boundary redraw and principal turnover (Bacher-Hicks), within-student fixed effects (Lacoe & Steinberg), a policy change (Craig & Martin), and — most pointedly — within-family comparison in Carrell & Hoekstra, which means the peer-disruption effect cannot be the heritable characteristics of the affected children.
There is a second, subtler place the premise bites. The peer-effects literature says the
externality is what matters: a child's disruptiveness is largely not a lever the school can pull,
but the school entirely controls how many such children sit in each room and what happens when they
disrupt. That is a structural decision, not a character-formation project — which is why this topic
sits in the structure layer and not under character.
Boundaries & what critics say
- The GBG steelman is Ashworth 2020, which finds Δ = 0.10 on reading at one-year follow-up among moderate compliers. Take it seriously and then look at what it concedes: nothing at post-test, nothing at ITT, and a benefit claimed only for a post hoc middle band of dosage in a trial where a quarter of schools could not sustain any dosage at all.
- Single-case GBG evidence is not wrong, it is answering a different question. [Bowman-Perrott](../sources/bhv-flower-2014-gbg-meta, bhv-bowman-perrott-2015-gbg-single-case-meta.yaml) finds moderate-to-large immediate effects on in-class behaviour, and that is almost certainly true. The game works while it is being played. Whether it changes anything later or elsewhere is what the cluster RCTs test.
- PBIS's discipline reductions may partly be recording changes. The intervention trains staff in how to respond to and log behaviour, and then measures logged behaviour. The suspension result is more robust to this than the office-referral result.
- The Baltimore GBG results might still be right for high-risk boys in high-poverty American urban schools. The English trial's at-risk-boys subgroup did trend the right way (−0.30 on disruptive behaviour, non-significant). This is the single most defensible pro-GBG reading, and it is a targeted claim, not the universal one the programme is sold on.
- "Discipline reform causes chaos" is a real hypothesis and the evidence is against it, not silent. Craig & Martin looked for the trade-off and found gains instead. But it is one city and one reform, and Sorensen shows the safety dimension is genuinely two-sided.
- Almost every behaviour outcome in this literature is a teacher rating, produced by teachers who know which arm they are in, in trials whose mechanism is changing teacher behaviour. Where outcomes are administrative (suspensions, referrals, arrests, earnings), effects are smaller and more credible. Weight accordingly.
Practical guidance
- Train teachers in ordinary classroom management and stop there. g ≈ 0.22, replicated, cheap, and the 2025 update suggests targeted is better than comprehensive. Do not buy the four-component whole-school package.
- Do not buy the Good Behavior Game as an attainment intervention. In its only large independent trial it moved reading by 0.03 SD and a quarter of schools abandoned it as not worth the setup time. If you want it as a classroom-order tool for a specific difficult room, the single-case evidence supports that narrow use.
- PBIS-style school-wide systems are a reasonable buy for behaviour, not for scores. Expect fewer referrals and suspensions; do not put achievement in the business case.
- Run the school so that suspension is rare. The suspended child loses achievement and gains a 15-20% higher chance of adult arrest and incarceration, statutory zero tolerance bought more punishment and no better behaviour, and the city that stopped suspending for non-violent conduct got higher test scores with no cost to anyone else.
- Reserve exclusion for violence. That is the one place where the evidence for removal is not contradicted, and it is where restorative practices demonstrably do not work.
- Do not adopt restorative practices expecting an achievement dividend. The only randomised trial that measured achievement carefully found maths 0.07 SD worse. If you adopt them, protect instructional time explicitly and measure it.
- Protect the room from the small number of genuinely disruptive children — by teaching, seating, and support, not by removal. The externality is large and real ($80,000 of classmates' lifetime earnings per year of exposure); the standard response to it makes the disruptive child worse off without helping the class much.
- Insist on administrative outcomes in any behaviour proposal you are shown. A teacher-rating effect in a non-blind trial of a teacher-behaviour intervention is close to uninformative.
Open questions
- Whether the Baltimore GBG long-run effects are real and simply require American high-poverty urban conditions and high-aggressive boys is unresolved, and no trial has been designed to answer it directly.
- Nobody has explained why the Pittsburgh restorative-practices trial lowered achievement in grade bands where it did not lower suspensions. Instructional time is the obvious suspect and was not measured.
- The exclusion literature has no good estimate of the classroom's counterfactual — what happens to the other 24 children when the disruptive child stays. Lacoe & Steinberg find very small peer spillovers from suspension; Carrell & Hoekstra find large peer costs from disruption. Reconciling those two is the unfinished work in this topic.
- Nothing here is genetically informative about responsiveness to behaviour intervention, the same untested trainability assumption the archive flags in talent and trainability.
- grade BThe Long-Run Effects of Disruptive PeersCarrell SE, Hoekstra M, Kuka E · 2018 · quasi-experiment
- grade BExternalities in the Classroom: How Children Exposed to Domestic Violence Affect Everyone's KidsCarrell SE, Hoekstra ML · 2010 · quasi-experiment
- grade CA Meta-Analysis of the Effects of Classroom Management Strategies and Classroom Management Programs on Students' Academic, Behavioral, Emotional, and Motivational OutcomesKorpershoek H, Harms T, de Boer H, van Kuijk M, Doolaard S · 2016 · meta-analysis
- grade CAn Update of the Meta-Analysis of the Effects of Classroom Management Interventions on Students' Academic, Behavioral, Social-Emotional, and Motivational OutcomesKorpershoek H, de Boer H, Mouw JM · 2025 · meta-analysis
- grade CTeacher classroom management practices: effects on disruptive or aggressive student behaviorOliver RM, Wehby JH, Reschly DJ · 2011 · review
- grade AGood Behaviour Game: Evaluation Report and Executive SummaryHumphrey N, Hennessey A, Ashworth E, Frearson K, Black L, Petersen K, Wo L, Panayiotou M, Lendrum A, Wigelsworth M, Birchinall L, Squires G, Pampaka M · 2018 · rct
- grade BEffects of a universal classroom behavior management program in first and second grades on young adult behavioral, psychiatric, and social outcomesKellam SG, Brown CH, Poduska JM, Ialongo NS, Wang W, Toyinbo P, Petras H, Ford C, Windham A, Wilcox HC · 2008 · rct
- grade BGame On—Complier Average Causal Effect Estimation Reveals Sleeper Effects on Academic Attainment in a Randomized Trial of the Good Behavior GameAshworth E, Panayiotou M, Humphrey N, Hennessey A · 2020 · rct
- grade CEffects of the Good Behavior Game on Challenging Behaviors in School SettingsFlower A, McKenna JW, Bunuan RL, Muething CS, Vega R · 2014 · meta-analysis
- grade CPromoting Positive Behavior Using the Good Behavior GameBowman-Perrott L, Burke MD, Zaini S, Zhang N, Vannest K · 2015 · meta-analysis
- grade CA Randomized, Wait-List Controlled Effectiveness Trial Assessing School-Wide Positive Behavior Support in Elementary SchoolsHorner RH, Sugai G, Smolkowski K, Eber L, Nakasato J, Todd AW, Esperanza J · 2009 · rct
- grade BExamining the Effects of Schoolwide Positive Behavioral Interventions and Supports on Student OutcomesBradshaw CP, Mitchell MM, Leaf PJ · 2009 · rct
- grade BEffects of School-Wide Positive Behavioral Interventions and Supports on Child Behavior ProblemsBradshaw CP, Waasdorp TE, Leaf PJ · 2012 · rct
- grade CUpdating and expanding systematic reviews and meta-analyses on the effects of school-wide positive behavior interventions and supportsLee A, Gage NA · 2020 · meta-analysis
- grade BThe School-to-Prison Pipeline: Long-Run Impacts of School Suspensions on Adult CrimeBacher-Hicks A, Billings SB, Deming DJ · 2024 · natural-experiment
- grade BDiscipline Reform, School Culture and Student AchievementCraig AC, Martin DA · 2025 · natural-experiment
- grade CEstimating the Effect of State Zero Tolerance Laws on Exclusionary Discipline, Racial Discipline Gaps, and Student BehaviorCurran FC · 2016 · quasi-experiment
- grade CAre zero tolerance policies effective in the schools?: An evidentiary review and recommendations.American Psychological Association Zero Tolerance Task Force · 2008 · review
- grade BSchool-based interventions for reducing disciplinary school exclusion: a systematic reviewValdebenito S, Eisner M, Farrington DP, Ttofi MM, Sutherland A · 2018 · meta-analysis
- grade BCan Restorative Practices Improve School Climate and Curb Suspensions? An Evaluation of the Impact of Restorative Practices in a Mid-Sized Urban School DistrictAugustine CH, Engberg J, Grimm GE, Lee E, Wang EL, Christianson K, Joseph AA · 2018 · rct
- grade BThe Impact of Restorative Practices on the Use of Out-of-School Suspensions: Results from a Cluster Randomized Controlled TrialHuang FL, Gregory A, Ward-Seidel AR · 2023 · rct
- grade BEvaluation of a Whole-School Change Intervention: Findings from a Two-Year Cluster-Randomized Trial of the Restorative Practices InterventionAcosta J, Chinman M, Ebener P, Malone PS, Phillips A, Wilks A · 2019 · rct
- grade BMaking Schools Safer and/or Escalating Disciplinary Response: A Study of Police Officers in North Carolina SchoolsSorensen LC, Shen Y, Bushway SD · 2021 · quasi-experiment
Related decisions
- Are self-control and grit teachable levers on achievement?mixedconf: highgc: medium
- Can executive function be trained, and does it transfer to learning?no effectconf: highgc: low
- Should a school buy a social and emotional learning programme?mixedconf: highgc: low
- Talent and trainability — what is heritable, and what that does not licensestrong supportconf: highgc: low