Reviewing the Evidence on How Teacher Professional Development Affects Student Achievement
Yoon KS, Duncan T, Lee SWY, Scarloss B, Shapley KL · 2007
grade Creviewdeveloper-ledfailednumbers spot-checked
Sample
9 studies (20 effect sizes) met What Works Clearinghouse evidence standards; 1,343 studies were submitted to prescreening, of which only 907 were unique
Population
US elementary-school teachers and their students. All nine studies were elementary; none addressed middle or high school. Teacher samples ran 5 to 44 per study, student samples 98 to 779.
Design
Systematic review (REL Southwest, Issues & Answers 2007-No. 033) applying What Works Clearinghouse screening to in-service teacher PD in maths, science and reading/ELA. Screening funnel, exactly: 1,343 studies submitted to prescreening, of which 907 were unique (the other 436 were duplicates counted again because they spanned multiple subjects); 132 unique studies passed prescreening; 27 passed full screening; 9 met evidence standards. So the famous "1,300+" denominator is inflated by a third with duplicates - the honest denominator is 907. Of the nine: 5 RCTs meeting standards without reservations, 1 RCT with group-equivalence problems, 3 quasi-experiments. Studies date 1986-2003; 6 peer-reviewed, 3 unpublished dissertations. 7 of 9 used standardized achievement measures, 1 used researcher-developed fractions measures, 1 used Piagetian conservation tasks. Crucially it is NOT a meta-analysis: the report says it "conducts none of the additional data manipulations of traditional meta-analysis, such as differential weighting", so the headline average is an unweighted mean of 20 effect sizes in which one study (McGill-Franzen et al. 1999) contributes 6. The report also states the studies "were generally underpowered and did not address clustering or multiple comparisons". `independence: developer-led` records a fact about the evidence this review passes through, not about the REL review team, who were independent IES contractors: "In all nine studies professional development went directly to teachers rather than through a train-the-trainer approach and was delivered by the authors or their affiliated researchers." IES report with no DOI; ERIC ED498548.
Key findings
The most-cited evidence for professional development is nine small, old, elementary-only studies in which the PD was delivered by the researchers evaluating it, and in which 12 of 20 effects were not statistically significant once clustering and multiple comparisons were handled. Two numbers are widely miscited and both are fixed here. (1) The +21 percentile points is the average improvement index across ALL nine studies and all 20 effects - the report says "Average control group students would have increased their achievement by 21 percentile points if their teacher had received professional development" - it is not the effect of a 14-hour dose. (2) The "1,300+ studies screened" figure counts 436 duplicates twice; 907 unique studies were screened. The 14-hour threshold is a bare direction-only statement with no effect size behind it, and the report itself says its limits "preclude any conclusions about the effectiveness of professional development by form, content, or intensity" - so the popular "14+ hours of PD buys 21 percentile points" is a reader inference the report disclaims. Read this review as documentation of how thin the workshop-PD evidence base always was.
Genetic confound
Teacher-level intervention; student selection is not the issue. The threat is allegiance (the PD was delivered by the study authors), small-study bias, and unaddressed clustering.
Replication notes
Two separate things, kept separate. What the 2007 report itself establishes: nothing about replication - it is a one-shot screening exercise, and it explicitly declines to draw conclusions about which kind of PD works. The `failed` flag is an external judgement about what happened next: the large federally funded IES evaluations of sustained, content-focused workshop PD that followed this report raised teacher knowledge while leaving student achievement flat. TNTP (2015), reviewed separately in this archive, summarises the same record - "two federally funded experimental studies of sustained, content-focused and job-embedded professional development have found that these interventions did not result in long-lasting, significant changes in teacher practice or student outcomes". Anyone recomputing under different weights should note the flag is sourced from that later literature, not from the paper in hand.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Student achievement, average across all nine studies (the canonical +21 figure) | percentile points (improvement index) | +21, range -20 to +49; corresponds to an unweighted mean effect size of d = 0.54, range -0.53 to 2.39, across 20 effect sizes | mixed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| Statistical significance of those same 20 effects | count | 12 of the 20 effects were NOT statistically significant after the report applied corrections for unaddressed clustering and multiple outcomes; 1 was negative (-0.53) and 1 was exactly zero | mixed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| Student achievement, restricted to the five clean randomised controlled trials | Cohen d | 0.51 (15 effects from 5 RCTs), range 0 to 1.11; the four with-reservations studies averaged 0.61 (5 effects), range -0.53 to 2.39, so the weaker designs carried both extremes | mixed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| Effect by contact hours (the "14+ hours" claim) | direction only | studies with more than 14 hours of PD "showed a positive and significant effect"; the three studies with 5-14 hours total showed no statistically significant effects. NO effect size is attached to either group, and the report states its own limits preclude conclusions "about the effectiveness of professional development by form, content, or intensity" | mixed | end of treatment | business-as-usual | end-of-treatment | domain-skill |
| Dose actually delivered in the nine studies | contact hours | average 49 hours; range 5 to 100 hours, spread over four weeks to a year | mixed | during treatment | none | not-applicable | domain-skill |
Cited by
- Teacher quality — selection over credentials and workshopsstrong supportconf: highgc: low