The Evidence on Teaching

Math at home adds up to achievement in school

Berkowitz T, Schaeffer MW, Maloney EA, Peterson L, Gregor C, Levine SC, Beilock SL · 2015

grade Crctdeveloper-ledmixed
Sample
587 first-grade families from 22 Chicago-area schools — 420 randomized to the math app, 167 to the reading (control) app; the matched-school analysis uses 226 math vs 167 reading families
Population
Demographically diverse Chicago-area primary caregivers and their first-grade children, 2013-14 school year.
Design
Science 350(6257), 196-198. Classroom-level randomization to a math or reading version of Bedtime Learning Together (BLT), an iPad app built by the authors from the Bedtime Math materials; families were given iPad Minis and app use was logged automatically. Outcome was Woodcock-Johnson-III Applied Problems, administered one-to-one at school at the start and end of the year, analysed on W scores via HLM. Design strengths are real: proper cluster randomization, an ACTIVE control (a reading app, not nothing), an independent standardized outcome, and objective dosage logging. The problem is the inference chain. No overall intent-to-treat main effect of condition on math achievement is reported anywhere in the paper. What is reported is (a) an ITT effect inside a median split on parents' math anxiety, and (b) a dose-response relation between logged app usage and end-of-year achievement, which is NOT randomized — families chose how much to use the app. The 2016 Technical Response discloses that the parent-math-anxiety x condition interaction licensing the split was itself only P = 0.06 two-tailed (P = 0.03 one-tailed, on an a priori directional hypothesis).
Key findings
The claim that a parent-delivered maths app buys "almost 3 months" of extra maths achievement is carried by a subgroup and by a non-randomized dosage analysis, not by the randomized contrast. In the intent-to-treat analysis, children of high-math-anxious parents in the math group beat the reading group by almost 3 months of math achievement (b = 5.25, t = 1.99, P = 0.048); children of low-math-anxious parents showed nothing (b = -0.61, t = -0.27, P = 0.79). The moderating interaction that justifies looking at those two cells separately was P = 0.06. The dose-response results are stronger but observational: within the math group, more app use predicted higher end-of-year math (b = 2.88, t = 4.01, P < 0.001), while within the reading group app use predicted nothing on math (b = 0.22, t = 0.25, P = 0.81), giving a group-by-use interaction of b = 4.03, t = 2.83, P = 0.005. Bin analysis: among children of high-math-anxious parents, moving from almost-no-use to about once a week was worth b = 8.08 (t = 3.14, P = 0.002), and going beyond once a week added nothing (b = 0.52, P = 0.79). The honest summary is that a very light-touch, scripted, parent-delivered maths activity plausibly helps the children whose parents would otherwise supply no maths input at all, and that the trial as randomized does not demonstrate it.
Genetic confound
Medium for the ITT contrast (randomized at classroom level, active control), HIGH for the headline. The dose-response analyses condition on a family-chosen behaviour; parents who open a maths app four nights a week differ from those who never do on traits that predict child maths achievement independently of the app. The authors run an instrumental-variables analysis in the 2016 response using randomization as the instrument for usage and report it as still significant for the high-math-anxiety group, which is the right move but rests on the same underpowered subgroup.
Replication notes
Contested in print within five months. Frank (Science 2016, 10.1126/science.aad8008) reanalysed the authors' own posted data and found no significant effect of the intervention on math performance and no significant condition-by-time interaction on either grade-equivalent or raw scores. Berkowitz et al. (Science 2016, 10.1126/science.aad8555) replied that Frank found no mistakes, that his tests were lower-powered and did not control for school-level effects, and that they stand by the conclusions — while conceding the interaction is only marginally significant. Both papers agree on the underlying arithmetic and disagree on which test is the right one. No independent replication of the BLT trial has been located.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
ITT — math app vs reading app, children of HIGH-math-anxious parentsHLM coefficient on end-of-year W score, controlling for beginning-of-year achievementb = 5.25, t = 1.99, P = 0.048 — described as "almost 3 months" of math achievementstandardizedend of school year (~1 school year of exposure)active-alternativeend-of-treatmentdomain-skill
ITT — math app vs reading app, children of LOW-math-anxious parentsHLM coefficient on end-of-year W scoreb = -0.61, t = -0.27, P = 0.79 — no effectstandardizedend of school yearactive-alternativeend-of-treatmentdomain-skill
The moderating interaction itself (parent math anxiety x condition)significance of the interaction termP = 0.06 two-tailed (P = 0.03 one-tailed given the a priori directional hypothesis) — disclosed only in the 2016 Technical Response, not in the original Reportstandardizedend of school yearactive-alternativeend-of-treatmentdomain-skill
Dose-response — app usage x group interaction (NOT randomized)HLM coefficientgroup-by-use interaction b = 4.03, t = 2.83, P = 0.005 (matched schools) and b = 2.55, t = 2.25, P = 0.03 (full sample); within math group b = 2.88, t = 4.01, P < 0.001; within reading group b = 0.22, t = 0.25, P = 0.81standardizedend of school yearactive-alternativeend-of-treatmentdomain-skill
Dose ceiling among children of high-math-anxious parentsHLM coefficient between usage binsonce-a-week vs almost-never b = 8.08, t = 3.14, P = 0.002; twice-or-more-a-week vs once-a-week b = 0.52, t = 0.27, P = 0.79 — no additional return above one session per weekstandardizedend of school yearactive-alternativeend-of-treatmentdomain-skill

Cited by