The Evidence on Teaching

Should a school buy a social and emotional learning programme?

SEL programmes move the ratings they are evaluated on far more than the attainment they are sold on: the famous +0.27 is now 0.10, and no large independent trial has reproduced it on an externally marked test.

mixedconf: highgc: low

character · ages 418

Effect summary

The headline number has been walked down by the SEL field itself: academic achievement 0.27 (Durlak 2011, k=35 of 213 interventions) → 0.109 (Cipriano 2023, 424 studies) → 0.101 (Ha 2025). SEL skills fell from 0.57 to 0.219 over the same interval. Publication bias is significant in the modern meta (Egger t = 3.59, p < .001; published 0.220 vs unpublished 0.164), 53% of Durlak's and 53.5% of Cipriano's outcome data is child self-report, and 58.7% of Cipriano's outcomes are rating scales. Against externally marked national tests the effect is not small but absent or negative: England's 45-school PATHS efficacy trial returned +0.026, −0.029, −0.025 and −0.106 (Year 6 English, EEF's judgement: attributable to the intervention), and the 77-school Good Behaviour Game trial returned 0.03. The federal SACD trial randomised 84 schools across SEVEN flagship programmes with an independent analyst and found 2 of 60 combined-programme impacts significant against 3 expected by chance, 0 of 18 in growth-curve analysis, and two significant NEGATIVE domain impacts. MYRIAD, the largest school mental-health RCT ever run (85 schools, 8,376 students), returned SMDs of 0.005, 0.02 and 0.02 with five secondaries favouring the control arm.

Practical takeaway

Do not buy a branded SEL curriculum, and above all do not buy one on its attainment claim. The +4 months in the English toolkit and the +0.27 in the founding meta-analysis have never been reproduced by a large independent trial on a test somebody else marked; when England ran that trial on PATHS, Year 6 English came out two months WORSE and the evaluator judged it attributable to the programme. The federal government randomised 84 schools across seven of the best-known programmes at once, with an independent analyst, and found nothing. What you can keep for free: adults who model calm, explicit norms, and a well-run room — see behaviour management, where the evidence is thinner in effect size but far more honestly measured. If you do adopt something, pre-commit to an outcome the vendor did not design, and read the 24-month result, not the 12-month one.

Who this applies to

Group size
whole-classschool-wide
Delivered by
teacher
Ages studied
416(narrower than the 418 this topic is filed under — outside it is extrapolation)
Dose
PATHS is roughly 40–47 lessons a year over two years; Second Step is 41 lessons over three years; MYRIAD is 10 lessons. Delivery consistently falls short — English teachers delivered about half the PATHS lessons and 26 of 47 in Birmingham — and in both trials that measured it, fidelity was NOT associated with outcomes.
Cost
medium
Moves
non-cognitive
Needs first
An outcome measure the programme's authors did not write, rated by someone who does not know the arm. Berry 2016 is the demonstration: at 12 months the developer's own scale was significant on 6 of 11 subscales while the standard scale was significant on 0 of 7, in the same children on the same day.
Not for
Raising attainment. Two independent national-scale trials measured externally marked tests and found zero or negative. Also not for mental health at scale: the definitive 85-school trial found nothing on depression risk, wellbeing or functioning, with five secondary outcomes favouring the control group.

Verdict

mixed, and the mixture separates cleanly along one line: who wrote the outcome measure.

SEL programmes reliably move rated social-emotional competence. That is a real finding, it survives into the modern meta-analysis at g ≈ 0.22, and — contrary to what this archive expected — it is larger in independent evaluations than in developer-led ones.

What does not survive is the reason schools are given for buying them. The achievement claim has been walked down by the SEL field itself, from 0.27 to 0.109 to 0.101 in fourteen years, and it has never once been reproduced by a large independent trial measuring an externally marked test. The two trials that did measure one returned zero and negative.

And the scale evidence is worse than the meta-analytic evidence. When the US federal government randomised 84 schools across seven flagship programmes and handed the analysis to an outside contractor, the answer was nothing — 2 of 60 significant impacts against 3 expected by chance, and two significant impacts in the wrong direction.

What the evidence shows

The meta-analytic walk-down

Source Design Grade Achievement effect Skills effect
Durlak 2011 213 interventions / 270,034 students; 47% randomised; CASEL-affiliated C 0.27 (CI 0.15–0.39), k = 35 of 213 0.57
Corcoran 2018 40 studies, BEE standards C reading 0.25, maths 0.26, science 0.19
Cipriano 2023 Registered Report, 424 studies / 575,361 students B 0.109 (0.037–0.182) 0.219
Ha 2025 40 studies, achievement only B 0.101
Goldberg 2018 Whole-school approaches only C 0.22
Wigelsworth 2022 Overview of 33 reviews C notes 0.22 is "at least half" the undifferentiated estimates

Cipriano et al. is the honest one and it is the field auditing itself — a Registered Report, with Durlak himself on the author list, reporting that the achievement effect is 40% of what the 2011 paper said, that Egger's test for publication bias is significant overall (t = 3.59, p < .001), that published studies outperform unpublished ones (0.220 vs 0.164), that 53.5% of outcomes are child-reported and 58.7% are rating scales, that only 44.6% of studies reported any implementation data, and that of those only 21% achieved moderate or high fidelity.

Look also at what Durlak's 0.27 actually rested on: 35 of 213 interventions. The single most influential number in school social-emotional policy comes from 16% of the studies in its own meta-analysis.

The independent trials, sorted by what they measured

Source Design Grade Key effect
SACD Consortium 2010 84 schools, SEVEN programmes, 3 years, independent analyst (Mathematica) A 2 of 60 impacts significant (3 expected by chance); 0 of 18 in growth-curve analysis, all ES < 0.07; 2 significant NEGATIVE domain impacts; 4 of 5 fidelity associations ran the wrong way
EEF PATHS 2015 45 schools, 3,336 pupils, KS2 SATs and InCAS A Y5 maths +0.026, Y5 reading −0.029, Y6 maths −0.025, Y6 English −0.106 (≈ −2 months) — EEF: attributable to the intervention
Hennessey & Humphrey 2020 Peer-reviewed report of the same trial + implementation profiles A β = −0.02, −0.01, −0.03, −0.05, all ns; dosage profiles not associated with outcomes
Berry 2016 56 schools, 5,074 pupils, 2 years, explicitly independent of the developer A 12 mo: standard SDQ 0 of 7 significant; developer's own scale 6 of 11. 24 mo: all gains gone; conduct favoured control
Kuyken 2022 (MYRIAD) 85 schools, 8,376 students — largest school mental-health RCT ever A Depression 0.005, functioning 0.02, wellbeing 0.02; 5 of 28 secondaries favoured CONTROL
EEF GBG 2018 77 schools, 3,084 pupils, independent standardized reading A Reading 0.03 (−0.08–0.16); behaviour null; ¼ of schools quit
Espelage 2015 36 middle schools, 3,616 students, 3 years B Year-1 physical aggression −42% (self-report); at 3 years no direct effects at all
Rimm-Kaufman 2014 24 schools, 2,904 children, 3 years B ITT null on maths and reading; positive only in a mediation model conditional on teacher practice use
Jones 2011 (4Rs) 18 NYC schools, 1,184 children — the SACD New York site C Self- and teacher-report improvements; no main achievement effect, only a teacher-nominated at-risk subgroup

Berry 2016 is the single most instructive trial in this file, because it runs the critical comparison inside one study. Same children, same day, same teachers. The developer's own rating scale was significant on 6 of 11 subscales at 12 months. The standard rating scale was significant on 0 of 7. By 24 months both were null and conduct favoured the control group. That is what "SEL works" and "SEL doesn't work" look like when they are the same data.

And SACD is the study the field has not absorbed. Seven programmes — including PATHS, Second Step, Positive Action and 4Rs, which between them account for a large share of the meta-analytic evidence base — randomised across 84 schools, followed three years, with the impact analysis contracted out precisely so that the developers would not be marking their own homework. The result is indistinguishable from chance, and it points slightly negative. Note also the fidelity finding, which forecloses the standard defence: of five significant fidelity–outcome associations, four were detrimental associations with low fidelity, not beneficial associations with high fidelity.

Same schools, two answers: the Positive Action case

Source Analyst Grade Same 14 Chicago schools
Bavarian 2013 Developer team (Flay's spouse holds a significant financial interest) C State-test maths 0.38 (p = .07, one-tailed), reading 0.22 (ns), absenteeism −0.78
SACD 2010, Illinois site Mathematica, independent contract A 0 of 18 impacts significant; all effect sizes ≤ 0.09
Snyder 2010 Developer team, 20 schools, school-level analysis C Effect sizes 0.5 to 1.1 across achievement, absenteeism and suspensions

One randomisation. One cohort. Two analyses. This is the cleanest demonstration in the archive that who analyses a trial can determine its published conclusion, and it deserves more weight than any meta-analytic moderator.

The surprise: the meta-analytic developer gradient does not appear

Wigelsworth et al. 2016 set out to test exactly the hypothesis this archive holds, on 89 studies, and it did not confirm it. For social-emotional competence the developer-led / developer-involved / independent series runs 0.21 / 0.48 / 0.69 — independent evaluations produced the largest effects. For conduct problems it runs 0.20 / 0.25 / 0.37, again the wrong way. The hypothesis held for only two of seven outcomes, neither significantly, and the implementation-quality explanation was also rejected (χ² = .633, p = .718).

This is recorded rather than explained away, because a rule that only reports its confirmations is not a rule. Two things reconcile it with the trial-level evidence without dissolving either:

  • What the same paper did confirm is the efficacy-to-effectiveness decay — mean difference 0.13, significant on four of seven outcomes, and 0.38 → 0.22 on academic achievement. That is this archive's most repeated finding, reproduced inside SEL.
  • And it confirmed a transferability decay: social-emotional competence 0.56 in the programme's home country versus 0.06 abroad. Every English trial in this file is an "abroad" trial of an American programme.

The honest statement is therefore narrower than "developers inflate results, on average." It is: developer-led trials of specific programmes have repeatedly produced results that independent analysts of the same schools could not reproduce, and effects shrink sharply when programmes leave efficacy conditions and leave their country of origin.

Hereditarian-lens assessment

Risk: low. This verdict rests on cluster-randomised trials — SACD (84 schools), MYRIAD (85 schools), PATHS in England (45 and 56 schools), the Good Behaviour Game (77 schools) — where genes cannot differ between arms. Nothing here is explained by selection.

The premise does two other jobs in this topic, though.

First, it predicts the shape of the failure. SEL programmes target non-cognitive traits, and the self-control and grit evidence shows those traits are 37–60% heritable with shared environment indistinguishable from zero — meaning differences between families and schools explain almost none of the variation. A programme aimed at a trait whose between-school variance is near zero is aiming at a small target, and the observed effect sizes are consistent with that.

Second, it warns against the achievement claim's causal chain even where the SEL effect is real. "Programme raises measured social-emotional competence" and "programme raises attainment" are different claims with different evidence; the first has support, the second does not. The English trials measured the second directly and found −0.106 to +0.026.

Boundaries & what critics say

  • The strongest case for SEL is that rated competence really does move, at g ≈ 0.22 in the modern meta, and by more in independent evaluations. This topic does not deny that. It denies the attainment claim, and it notes that "children were rated as behaving better" and "children behaved better" are separated in this literature by exactly one blinded observer, who is usually absent.
  • Corcoran 2018 is the steelman for achievement, reporting 0.25 to 0.26 under Best Evidence Encyclopedia standards. Note its own caveat: the authors state that SEL programmes from the more rigorous, larger randomised studies "might not have as meaningful effects for pre-K–12 students as once thought." The later syntheses quantify what they were flagging.
  • Rimm-Kaufman 2014 is the standard shape of a defended null — no ITT effect, positive mediation effects conditional on how much teachers used the practices. Teachers who implement more are not a randomly selected group, and neither Berry's nor Hennessey and Humphrey's trials found any fidelity–outcome association at all.
  • MYRIAD deserves credit as well as weight. Its senior authors direct mindfulness centres and receive royalties on mindfulness books, and they published a clean null on 8,376 students with confidence intervals that rule out important effects. That is what the field should look like.
  • Control conditions matter and are usually generous to the programme. Nearly three quarters of Cipriano's pool used education-as-usual or waitlist controls. MYRIAD's control was ordinary social-emotional teaching, which makes its null a comparison of SEL against SEL — a more demanding and more realistic test than most.
  • None of this argues against schools caring about children's wellbeing. It argues that a purchased curriculum is not how you get it, and that the evidence offered for the purchase does not say what the brochure says.

Practical guidance

  • Do not buy an SEL curriculum on an attainment business case. The number in the business case is 0.27 from 2011; the field's own current estimate is 0.10, and independent national trials on externally marked tests give zero or negative.
  • Read the 24-month result. Berry's developer-instrument gains were significant at 12 months and gone at 24. Espelage's headline 42% reduction was gone by year 3.
  • Require an outcome measure the vendor did not write, and rated by someone who does not know which arm the school is in. This single requirement would have prevented most of the claims in this literature.
  • Discount any programme whose only evidence is developer-analysed. The Positive Action Chicago case is unanswerable: the same 14 schools, effect sizes of 0.22–0.38 when the developers analyse them and 0.00–0.09 when Mathematica does.
  • Expect the effect to shrink when the programme crosses a border. 0.56 at home, 0.06 abroad. An American programme in an English school is an "abroad" trial by definition.
  • Do not accept "it would have worked with better fidelity." In the three trials that tested it — SACD, Berry, Hennessey and Humphrey — fidelity was unrelated to outcomes, and in SACD four of five significant fidelity associations pointed the wrong way.
  • If the actual problem is a disorderly school, buy classroom management (behaviour management) rather than an emotions curriculum. The effects are smaller on paper and much better measured.

Open questions

  • Why independent evaluations produce larger effects on rated social-emotional competence than developer-led ones (0.69 vs 0.21) is unexplained and cuts against a strong prior. It deserves a direct investigation rather than a rationalisation.
  • No trial has established whether the rated competence gains correspond to anything an outside observer would notice. The one trial with blind classroom observation (Berry) found 3 of 9 composites moving at 12 months, and did not follow them to 24.
  • Nobody has run an SEL trial with a preregistered attainment outcome on an externally marked test in the programme's country of origin. Every such trial so far has been a transported one.
  • The SACD trial's inability to obtain usable grades and test scores — so that its "Academics" domain became teacher-report — means the best-designed study in this literature could not test its central claim. That gap has never been filled.

Evidence (19 sources)

Export all: BibTeX · RIS

Related decisions

← Back to explore

Should a school buy a social and emotional learning programme? · The Evidence on Teaching