The Evidence on Teaching

The Effectiveness and Micro-costing Analysis of a Universal, School-Based, Social-Emotional Learning Programme in the UK: A Cluster-Randomised Controlled Trial

Berry V, Axford N, Blower S, Taylor RS, Edwards RT, Tobin K, Jones C, Bywater T · 2016

grade Arctindependentfailed
Sample
5,074 pupils aged 4-6 at baseline in 56 primary schools in one large UK city; two academic years of implementation
Population
Birmingham, United Kingdom; infant-school pupils.
Design
The cleanest available demonstration inside a single trial of what happens when you measure the same children on a developer's instrument and on a standard one. Cluster randomised at school level, powered for an effect of 0.23, with the primary outcome the teacher-rated Strengths and Difficulties Questionnaire, a secondary outcome the PATHS Teacher Rating Scale (a developer-authored instrument), and independent observation via the Teacher-Pupil Observation Tool in a third of schools. The paper states explicitly that the study was fully powered and independent of the programme developer.
Key findings
At 12 months the PRIMARY outcome showed no statistically significant differences on total difficulties, impact score, or any of the five SDQ subscales. On the same children at the same time, the DEVELOPER'S OWN instrument was significant on 6 of 11 subscales favouring the intervention. Independent observation was significant on 3 of 9 composites. At 24 months all the developer-instrument gains were gone - no significant differences on any subscale in either complete-case or imputed analysis - the SDQ remained null, and on the imputed model conduct at 24 months significantly favoured the CONTROL group. Teachers delivered a mean of 26 of 47 lessons (55%) and there was no relationship between fidelity and treatment effects. Cost was 12,666 pounds per school and 139 pounds per child.
Genetic confound
Low. School-level randomisation.
Replication notes
A failed independent replication of PATHS, which holds Blueprints Model status on developer-led US efficacy trials. The internal contrast between the developer's instrument and the standard one at 12 months, and the disappearance of both by 24 months, is the most instructive part of the record.
DOI / URL
10.1007/s12310-015-9160-1

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Teacher-rated SDQ at 12 months (primary outcome)significanceno significant differences on total difficulties, impact, or any of 5 subscalesresearcher-designed12 monthsbusiness-as-usualend-of-treatmentbehaviour
Developer's own PATHS Teacher Rating Scale at 12 monthscount significant6 of 11 subscales favouring interventionresearcher-designed12 monthsbusiness-as-usualend-of-treatmentbehaviour
Developer's own instrument at 24 monthssignificanceall gains lost; no significant differences on any subscaleresearcher-designed24 monthsbusiness-as-usualover-2yrbehaviour
SDQ conduct at 24 months (imputed model)directionsignificantly favoured the CONTROL groupresearcher-designed24 monthsbusiness-as-usualover-2yrbehaviour
Independent classroom observation at 12 monthscount significant3 of 9 composites (teacher positive behaviours 0.304; class negative to teacher 0.307; class off-task 0.227)standardized12 monthsbusiness-as-usualend-of-treatmentbehaviour
Implementation fidelity and treatment effectsassociationteachers delivered 26 of 47 lessons; no relationship between fidelity and effectsresearcher-designed2 yearsbusiness-as-usualover-2yrbehaviour

Cited by