The Evidence on Teaching

Reducing withdrawal and failure rates in introductory programming with subgoal labeled worked examples

Margulieux, L. E., Morrison, B. B., & Decker, A. · 2020

grade Cquasi-experimentdeveloper-ledmixed
Sample
265 students across multiple sections of one introductory programming course
Population
Undergraduate CS1 at one US university.
Design
Students self-select into lab sections, which then received subgoal-labelled or standard instruction — not randomised at the student level, so section-level selection is uncontrolled. Outcomes are the course's own quizzes and exams, partially aligned with the treatment. The interrater-reliability procedure for open responses is well described. The contribution over prior subgoal work is duration: a full semester rather than a lab session.
Key findings
Formative quizzes, taken within a week of each new procedure, moved d = 0.44, t(264) = 12.03. Summative exams did not clearly move: total exam d = 0.26, p = .04, with component scores of d = 0.20-0.22 that were not significant. The durable finding is about the tail rather than the mean: about half as many subgoal students as control students had a failing exam average or missed exams (control 7% took one exam and 13% took two; subgoal 5% and 5%), and the subgoal group's exam VARIANCE was significantly lower — fewer students failing badly, not more students excelling.
Genetic confound
Medium. Section self-selection is uncontrolled, and students who choose particular lab sections differ.
Replication notes
The subgoal-label effect has a documented failure inside its own research programme — Morrison, Margulieux & Guzdial (2015) report that student performance gains from being GIVEN subgoal labels did not replicate as expected in the introductory CS task, despite replicating in mathematics and science.
DOI / URL
10.1186/s40594-020-00222-7

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Formative quiz performanceCohen d0.44researcher-designedwithin one week of each new procedureactive-alternativeend-of-treatmentdomain-skill
Total summative exam scoreCohen d0.26 (p = .04); component scores 0.20-0.22, not significantresearcher-designedend of semesteractive-alternativeend-of-treatmentdomain-skill
Rate of missing exams or holding a failing exam averagerate ratioapproximately halved against controladministrativeend of semesteractive-alternativeend-of-treatmentattainment

Cited by