Informing Progress: Insights on Personalized Learning Implementation and Effects
Pane, J. F., Steiner, E. D., Baird, M. D., Hamilton, L. S., & Pane, J. D. · 2017
grade Cquasi-experimentindependentmixednumbers spot-checked
Sample
40 NGLC schools and about 10,600 students in the implementation sample; 32 schools and about 5,500 students in the achievement analysis (5,539 mathematics, 5,474 reading); 6,145 students and 241 teachers surveyed
Population
US schools funded by the Next Generation Learning Challenges (NGLC) Wave IIIa and Wave IV launch grants, mostly urban, mostly new, 78% charter. Median 80% free/reduced-price lunch and 96% students of color. 48% high schools, 30% middle, 13% K-8, 10% elementary. Main achievement window fall 2014 to spring 2015.
Design
NO RANDOMIZATION OF ANY KIND, and RAND says so plainly: "given the portfolio of NGLC schools, it was not possible to create randomly assigned treatment and control groups; nor did we have access to data from neighboring schools." The comparison is a VIRTUAL COMPARISON GROUP - for each NGLC student, NWEA drew up to 51 matched students from its national testing database with similar fall MAP scores, the same grade and gender, and similar school demographics, then coarsened exact matching and school fixed effects were applied. RAND names its own two fatal assumptions: (1) matched students may differ on unobservables - "parents of NGLC students might have greater interest in nontraditional schooling environments and this could be related to how well their children do, independently of the NGLC schools' PL treatment," which is exactly the selection channel this archive's premise predicts; and (2) "the VCG approach also assumes that the students in the comparison group are attending more-traditional schools that are not using PL practices, but there is no way to verify this assumption." RAND's own summary sentence is "the achievement analyses use a research design that does not enable strong causal conclusions." A further definitional problem RAND concedes: the schools were not implementing a defined intervention - each chose its own model, "none of the schools looking as radically different from traditional schools as theory might predict," making it "difficult to draw a clear line separating PL from non-PL schools." So the treatment is really "attended a school that won an NGLC grant." THE ONE GENUINE STRENGTH, and the reason this is graded C rather than D, is the outcome: NWEA MAP is an independent, externally administered, adaptive standardized test - the only such measure anywhere in this cluster. FUNDING AND STRUCTURE: the study was funded by the Bill & Melinda Gates Foundation, which is also the primary funder of the NGLC initiative that created the schools being evaluated. RAND ran and authored the evaluation and did not build or sell anything, so `independent` under this archive's rule - but the funder-is-also-the-programme-sponsor structure is real and is recorded here so it can be reweighted. Report RR-2042-BMGF; RAND's own dose-response claim (more PL implementation associated with larger effects) is presented as "suggestive" and "somewhat speculative" and rests on a nine-school district subsample.
Key findings
This is the empirical basis for most "personalised learning works" marketing, and the real numbers are 0.09 in mathematics (statistically significant after multiplicity adjustment) and 0.07 in reading (not significant) - about 3 percentile points for the median student - from a design its own authors say "does not enable strong causal conclusions." Roughly half of the individual schools have positive treatment estimates and half do not, spanning about -0.5 to +0.6. In absolute terms the students started the year significantly below national norms in both subjects and gained about two percentile points; in mathematics they remained significantly below national norms at the end of the year. The 2015 predecessor report from the same team, on more experienced schools, gave 0.27 and 0.19 over two years; the effects fell by roughly two-thirds between reports, and RAND's language became correspondingly hedged. Recorded as INCLUDED because its role in the archive is as a boundary marker: it is the strongest evidence the personalised- learning claim actually has, and it is a matched-comparison design against a national norm sample producing effects at the edge of detectability.
Genetic confound
High. Non-randomized; families select into new charter and nontraditional schools, and the virtual comparison group matches only on observables (prior MAP score, grade, gender, school demographics). RAND explicitly names parental preference for nontraditional schooling as an unobserved confound that could bias the estimate in either direction.
Replication notes
This is the same research team's own second look at the same programme, and the estimates SHRANK. Pane et al. (2015), "Continued Progress: Promising Evidence on Personalized Learning" (RAND RR-1365), reported two-year effects of 0.27 in mathematics and 0.19 in reading - about 11 percentile points - across 62 schools and roughly 11,000 students. This 2017 report, on 32 newer NGLC schools and roughly 5,500 students over one year, reports 0.09 in mathematics (significant) and 0.07 in reading (NOT significant) - about 3 percentile points. RAND attributes the difference to sample composition (the 2015 sample was more experienced implementers, 92% charter, 69% K-5; this one is newer schools, 75% charter, 79% grades 6-12) rather than to regression toward zero, and the two samples partly overlap. Either way the direction is the archive's standard efficacy-to-effectiveness pattern, and the marketing claim "personalised learning works" is almost always sourced to the 2015 numbers.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Mathematics achievement, NGLC personalized-learning schools vs virtual comparison group, 2014-15 | ES | approximately 0.09, statistically significant (p < 0.05) after adjustment for multiple hypothesis tests; equivalent to about 3 percentile points for the median student. n = 5,539. | standardized | fall 2014 to spring 2015, one academic year | business-as-usual | end-of-treatment | domain-skill |
| Reading achievement, NGLC personalized-learning schools vs virtual comparison group, 2014-15 | ES | approximately 0.07, NOT statistically significant; about 3 percentile points. n = 5,474. | standardized | fall 2014 to spring 2015, one academic year | business-as-usual | end-of-treatment | domain-skill |
| The same team's earlier estimates on a more experienced sample (Pane et al. 2015, "Continued Progress", RAND RR-1365) | ES | 0.27 in mathematics and 0.19 in reading over TWO years, about 11 percentile points, 62 schools and roughly 11,000 students. Roughly a threefold larger estimate than the 2017 report on the newer sample. | standardized | two-year span, 2013-15 | business-as-usual | 1-2yr | domain-skill |
| Absolute standing against national norms (rather than against the matched comparison) | percentile rank | students started fall 2014 significantly BELOW national norms in both subjects and gained about two percentile points over the year; in mathematics they were still significantly below national norms in spring, in reading approximately at norms. | standardized | fall 2014 and spring 2015 | none | end-of-treatment | domain-skill |
| School-level dispersion of the treatment estimate | ES by school | about half of the schools have positive estimates and half negative, spanning roughly -0.5 to +0.6 in both subjects. There is no single "personalized learning effect" here to point at. | standardized | 2014-15 and the 2013-15 two-year span | business-as-usual | end-of-treatment | domain-skill |