The Evidence on Teaching

Using Social-Emotional and Character Development to Improve Academic Outcomes: A Matched-Pair, Cluster-Randomized Controlled Trial in Low-Income, Urban Schools

Bavarian N, Lewis KM, DuBois DL, Acock A, Vuchinich S, Silverthorn N, Snyder FJ, Day J, Ji P, Flay BR · 2013

grade Crctdeveloper-involvedfailed
Sample
14 Chicago public schools in 7 matched pairs (6 pairs retained for state-test analyses), grades 3-8, 1,170 students in the analytic sample; of 624 grade-3 baseline students only 131 (21%) remained at grade 8
Population
Chicago, United States; low-income urban elementary and middle schools.
Design
The single most instructive independence contrast in this archive. THESE ARE THE SAME 14 SCHOOLS as the Illinois site of the federal Social and Character Development trial - same randomisation, same cohort, same years - analysed by two different teams. Many tests here are reported ONE-TAILED, attrition is severe, and the conflict-of-interest notice states that the research used the programme, training and technical support of Positive Action Inc. in which Dr Flay's spouse holds a significant financial interest.
Key findings
On the Illinois state test, reading effect size 0.22 (p = 0.16 one-tailed, NOT significant) and mathematics 0.38 (p = 0.07 one-tailed, marginal). Subgroup results include reading for African-American boys 1.50 (p = 0.02 one-tailed) and mathematics for girls 0.41 and for low-income students 0.42 (both p under 0.10, one-tailed). School-level absenteeism -0.78 (p = 0.015 one-tailed). Teacher-rated academic motivation 0.39 (significant), teacher-rated academic ability 0.14 (not significant), self-reported disaffection with learning -0.19 (significant) and self-reported grades 0.02 (not significant). Set against this, Mathematica's independent analysis of the SAME schools found none of 18 growth-curve impacts significant with all effect sizes at or below 0.09.
Genetic confound
Low as designed. The threats are one-tailed testing, 21% cohort retention, and the declared financial conflict.
Replication notes
Failed under independent analysis of the identical schools. This pair - developer team and Mathematica on the same randomisation - is the clearest single demonstration in the archive that analyst identity can determine the published conclusion.
DOI / URL
10.1111/josh.12093

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Illinois state test mathematics, all studentseffect size0.38, p = 0.07 one-tailed (marginal)standardizedend of trialbusiness-as-usualover-2yrdomain-skill
Illinois state test reading, all studentseffect size0.22, p = 0.16 one-tailed — not significantstandardizedend of trialbusiness-as-usualover-2yrdomain-skill
School-level absenteeismeffect size-0.78, p = 0.015 one-tailedstandardizedend of trialbusiness-as-usualover-2yrbehaviour
Independent (Mathematica) analysis of the same 14 schoolscount significant0 of 18 growth-curve impacts; all effect sizes at or below 0.09mixedend of trialbusiness-as-usualover-2yrnon-cognitive

Cited by