The Evidence on Teaching

Comment on “Math at home adds up to achievement in school”

Frank MC · 2016

grade Bcritiqueindependentnot-applicable
Sample
Reanalysis of the full Berkowitz et al. 2015 dataset (587 first-grade families)
Population
N/A — secondary analysis of the Bedtime Learning Together randomized field experiment.
Design
Science 351(6278), 1161. A reanalysis rather than a rhetorical critique: Berkowitz et al. posted their data with the Report, and Frank ran the randomized contrast the trial was designed to support. His objection is twofold — that the original analyses were data-dependent (chosen after seeing the results), and that the primary randomized comparison, tested directly, is null. The authors' Technical Response concedes he found no mistakes in their analyses and characterizes his tests as "less statistically powerful, and in our view less appropriate", specifically because change scores are weaker than controlling for baseline, and because his models do not control for school-level effects. So the disagreement is entirely about which test is the correct one on an agreed set of numbers.
Key findings
The Bedtime Math trial, analysed as a trial, is null. Frank reports no significant effect of the intervention on math performance and no significant condition-by-time interaction for either grade-equivalent or raw scores, concluding that "this well-designed trial failed to find strong evidence for the efficacy of the intervention". This is the archive's standard pattern arriving on schedule: a headline effect that lives in a subgroup defined by a marginal interaction (P = 0.06) and in a non-randomized dose-response relation, with the randomized main effect never reported in the original paper. Frank's comment is the reason parent-delivered maths cannot be scored as established from this trial.
Genetic confound
N/A for the reanalysis itself. The relevant point is that Frank's critique removes the one analysis in Berkowitz et al. that was genuinely protected from confounding (the randomized contrast) by showing it null, leaving the surviving claims resting on the dosage analysis, which is a family-chosen behaviour and therefore fully exposed to selection.
Replication notes
Rebutted in the same issue by Berkowitz et al. (10.1126/science.aad8555), who defend their model specification (end-of-year outcome controlling for baseline, within-school matching) and add an instrumental-variables analysis using randomization as the instrument for app use. The exchange terminated there; no third-party adjudication and no independent replication of the trial has been located.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Main effect of the math app on math performance (reanalysis)significance of the randomized contrastno significant effect of the intervention on math performancestandardizedend of school yearactive-alternativeend-of-treatmentdomain-skill
Condition-by-time interactionsignificancenot significant for either grade-level-equivalent scores or raw scoresstandardizedbeginning to end of school yearactive-alternativeend-of-treatmentdomain-skill
Status of the original analysesverdictno computational errors found; objection is that the reported analyses were data-dependent and that the trial provides "limited support for the effectiveness of the intervention"unknownn/anonenot-applicabledomain-skill

Cited by