Homework and academic achievement: A meta-analytic review of research
Baş G, Şentürk C, Ciğerci FM · 2017
grade Cmeta-analysisindependentunreplicated
Sample
11 experimental and quasi-experimental studies published 2000-2015; 862 students in total (323 elementary, 287 high school, 252 university)
Population
Mostly Turkish and US students; 3 journal articles, 6 master's theses and 2 doctoral dissertations
Design
Issues in Educational Research 27(1), 31-50. Small, and the corpus is mostly unpublished theses with researcher- or teacher-designed achievement tests, which is why this grades C. Its value is entirely in one moderator table nobody else reports: it separates studies that randomised INDIVIDUAL STUDENTS from studies that randomised whole CLASSES. That is the distinction the whole causal literature turns on, and it is almost never made.
Key findings
The overall effect is d = 0.229 with a 95% confidence interval of -0.116 to 0.573 - it crosses zero. By school level the age gradient reappears: elementary d = 0.151 (95% CI -0.069 to 0.372, p = .179, not significant, k = 7); high school d = 0.479 (0.243 to 0.715, k = 2); university d = 0.446 (0.194 to 0.699, k = 2). THE NUMBER THAT MATTERS MOST IS THE DESIGN SPLIT: the two studies that randomised individual students give d = -0.125 (95% CI -0.443 to 0.194, p = .443) while the nine that randomised classes give d = +0.450 (0.300 to 0.601, p = .000). The authors report the between-design test as non-significant (Q_B(1) = 2.843, p = .092) and conclude the designs 'do not differ', but with k = 2 that test has essentially no power, and the point estimate from proper student-level randomisation is NEGATIVE. Seven of eleven effects were positive and four negative.
Genetic confound
Low within the randomised comparisons and irrelevant to the between-design contrast, which is the part worth citing. The class-randomised studies carry the usual cluster problem - a single class per arm confounds teacher with treatment - which is a design confound rather than a genetic one, but it points the same way: the apparently large positive effects come from the comparisons least able to isolate homework.
Replication notes
Unreplicated as a synthesis. Its central observation - that individual-level randomisation gives a smaller or negative estimate than class-level assignment - matches Cooper's own 2006 corpus, where the three studies with some form of random assignment give d = 0.53 against d = 0.83 for the two non-random ones, and matches Grodner and Rupp's field experiment, where randomly REQUIRING homework moves scores about 2 points while self-selected homework COMPLETION is associated with 4 to 6.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Overall effect of homework on achievement (random effects) | Cohen's d | 0.229, 95% CI -0.116 to 0.573 - the interval crosses zero | researcher-designed | end of intervention | business-as-usual | end-of-treatment | domain-skill |
| Effect by school level | Cohen's d | elementary 0.151 (95% CI -0.069 to 0.372, p = .179, k = 7); high school 0.479 (0.243 to 0.715, k = 2); university 0.446 (0.194 to 0.699, k = 2) | researcher-designed | end of intervention | business-as-usual | end-of-treatment | domain-skill |
| Effect by randomisation unit | Cohen's d | students randomised individually: -0.125 (95% CI -0.443 to 0.194, p = .443, k = 2). Classes randomised: +0.450 (0.300 to 0.601, p = .000, k = 9). Q_B(1) = 2.843, p = .092 - underpowered, but the properly randomised point estimate is negative | researcher-designed | end of intervention | business-as-usual | end-of-treatment | domain-skill |
Cited by
- Homework — effects by age and by dosagemixedconf: mediumgc: medium