Reducing withdrawal and failure rates in introductory programming with subgoal labeled worked examples
Margulieux, L. E., Morrison, B. B., & Decker, A. · 2020
grade Cquasi-experimentdeveloper-ledmixed
Sample
265 students across multiple sections of one introductory programming course
Population
Undergraduate CS1 at one US university.
Design
Students self-select into lab sections, which then received subgoal-labelled or standard instruction — not randomised at the student level, so section-level selection is uncontrolled. Outcomes are the course's own quizzes and exams, partially aligned with the treatment. The interrater-reliability procedure for open responses is well described. The contribution over prior subgoal work is duration: a full semester rather than a lab session.
Key findings
Formative quizzes, taken within a week of each new procedure, moved d = 0.44, t(264) = 12.03. Summative exams did not clearly move: total exam d = 0.26, p = .04, with component scores of d = 0.20-0.22 that were not significant. The durable finding is about the tail rather than the mean: about half as many subgoal students as control students had a failing exam average or missed exams (control 7% took one exam and 13% took two; subgoal 5% and 5%), and the subgoal group's exam VARIANCE was significantly lower — fewer students failing badly, not more students excelling.
Genetic confound
Medium. Section self-selection is uncontrolled, and students who choose particular lab sections differ.
Replication notes
The subgoal-label effect has a documented failure inside its own research programme — Morrison, Margulieux & Guzdial (2015) report that student performance gains from being GIVEN subgoal labels did not replicate as expected in the introductory CS task, despite replicating in mathematics and science.
DOI / URL
10.1186/s40594-020-00222-7
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Formative quiz performance | Cohen d | 0.44 | researcher-designed | within one week of each new procedure | active-alternative | end-of-treatment | domain-skill |
| Total summative exam score | Cohen d | 0.26 (p = .04); component scores 0.20-0.22, not significant | researcher-designed | end of semester | active-alternative | end-of-treatment | domain-skill |
| Rate of missing exams or holding a failing exam average | rate ratio | approximately halved against control | administrative | end of semester | active-alternative | end-of-treatment | attainment |
Cited by
- Does teaching programming actually teach programming — and does the method matter?moderate supportconf: mediumgc: low