The Evidence on Teaching

Length of the school day and year, time-on-task, extended time, and summer

Taking time away hurts measurably; adding it back buys almost nothing. The return per hour is ~0.02-0.03 SD, concave, and near zero in disorderly classrooms.

mixedconf: mediumgc: low

time · ages 418 · structure

Effect summary

The subtraction and addition experiments give different answers, and that asymmetry is the finding. SUBTRACTING time is reliably costly: a four-day week removes 3.6 weekly hours and costs 0.03-0.06 SD; COVID closures cost 0.14 SD (95% CI 0.10-0.17), about 35% of a school year, and have neither widened nor recovered in two and a half years; ten extra days of Swedish schooling buy 1.1% of an SD on crystallised tests and nothing on fluid ones, and non-school days buy nothing at all. ADDING time returns far less than the advocacy literature promises. Lavy's famous 0.058 SD per weekly subject-hour replicates only in the wave he used — Bietenbeck and Collins get 0.014-0.037 SD across eleven other PISA and TIMSS waves, and Rivkin and Schiman get 0.017-0.025 with school-by-subject fixed effects. Returns are concave (the first two hours carry most of it) and conditional on order: 0.025 SD at the 75th percentile of classroom quality against 0.011 at the 25th, with 'little or no benefit' in the bottom tail. Large-dose programmes keep returning nothing — Massachusetts added 300 hours a year for five years for no achievement effect plus measurable student and teacher fatigue, a Dutch RCT added five weekly hours for nothing, Chile added 27% of annual time for 0.05-0.07 SD, and Florida's extra literacy hour gave 0.05 SD that decayed to -0.004 SD by year three. SUMMER: the mean loss is real and survives the measurement critique (17-28% of the school-year ELA gain, 25-34% in maths, across ~18 million students on adaptive tests), but the claim it built its policy career on — that summer is where the achievement gap opens — does not replicate. Summer programmes buy 0.08-0.10 SD and the one randomised estimate dissipated within a year.

Practical takeaway

Defend the instructional time you have and be very sceptical about buying more. Protecting time is cheap and the evidence that losing it costs is strong: four-day weeks, pandemic closures and individual absences all show up in test scores. Adding time is expensive and the evidence that it pays is weak, concave and conditional — 0.02-0.03 SD per weekly subject-hour at best, less in disorderly classrooms, and repeatedly zero when bought in 300-hour annual blocks. If you are going to add time, add it in a specific subject to a specific group (Chicago's double-dose algebra is the archive's best case), not as a longer day for everyone. On summer: expect real forgetting (a fifth to a third of the year's gain) but do not sell summer school as gap-closing — it is not where the gap opens, and randomised summer programmes buy about 0.08 SD that fades within a year.

Who this applies to

Group size
school-wide
Delivered by
administrator
Ages studied
618(narrower than the 418 this topic is filed under — outside it is extrapolation)
Dose
Per weekly hour of subject instruction, budget 0.02-0.03 SD, concave — the first one to two hours in a subject are worth roughly twice the fourth. A full school year is worth 0.14-0.21 SD of crystallised skill (Sweden) or about 0.4 SD of measured achievement growth. Large blocks do worse per hour than small targeted ones: +300 h/yr for five years returned nothing in Massachusetts; +3 h/week for 16 weeks returned 0.15 SD in Denmark.
Cost
high
Moves
domain-skill
Needs first
An orderly classroom. Rivkin and Schiman find the return to an extra hour more than doubles between the 25th and 75th percentile of classroom environment, and that schools in the lower tail 'realize little or no benefit'. Adding hours to a school that cannot use the hours it has is the most expensive way to buy nothing in this archive.
Not for
Closing achievement gaps — the return is not reliably larger for disadvantaged students and in Florida's extra-hour mandate the lowest-achieving students, the ones it targeted, gained least (0.027 SD, not significant, against 0.077 for the middle group). Also not for any setting already at a long school day: returns are concave and Lavy's marginal hour beyond four hours a week in a subject is worth less than half the first.

Verdict

Instructional time is not a production input with a constant return, and the literature only looks contradictory if you assume it is. Two questions get conflated. Does removing school time hurt? Yes, reliably, and it is one of the better-identified facts in this archive. Does adding school time help? Barely, expensively, with sharply diminishing returns, and only where the marginal hour lands somewhere that can use it.

The asymmetry is not a paradox. Removing time removes hours that were already being spent on the highest-value use available; adding time adds hours at the bottom of a school's priority list, into whatever slot is left, with whatever teacher is available at 4pm. Lavy's own concavity says so — the marginal hour in the two-to-three-hour range is worth 4.20 PISA points, the marginal hour past four hours only 2.48 — and Rivkin and Schiman say the rest of it: the return more than doubles between a disorderly and an orderly classroom, and schools in the bottom tail get "little or no benefit."

The headline number in this literature has already failed a replication and almost nobody has noticed. Lavy's 0.058 SD per weekly subject-hour is the most-cited estimate of the value of instructional time. Bietenbeck and Collins ran his exact specification on eleven further waves of PISA and TIMSS in the same 22 countries: they reproduce 0.058 exactly in PISA 2006 and get 0.014 to 0.037 everywhere else, tracing the gap to PISA 2006 having measured instruction time in categorical bands. Rivkin and Schiman, working independently with school-by-subject fixed effects, land at 0.017 to 0.025. The defensible central estimate is about 0.02-0.03 SD per weekly subject-hour, not 0.058, and certainly not the 0.15 that gets quoted when the within-pupil standard deviation is used as the denominator.

What the evidence shows

Taking time away

Source Design Grade Key effect
Betthäuser 2023 (COVID) pre-registered meta, 42 studies / 15 countries, ROBINS-I screened B d = −0.14 (CI −0.17 to −0.10)35% of a school year; arose early, then "neither closed nor widened"; maths −0.18 vs reading −0.09
Engzell 2021 national natural experiment, Netherlands B 0.08 SD lost over an 8-week closure — "students made little or no progress while learning from home"
Thompson 2021 DiD on four-day-week adoption, 1.85M student-years B maths −0.044, reading −0.033; the schedule loses 215.8 min/week; IV: +0.018 SD maths per weekly hour
Carlsson 2015 conditionally random enlistment test dates, n=128,617 B 10 school days = +1.1% SD (synonyms), ≈ 0.21 SD per 180-day year; non-school days ≈ 0; fluid intelligence: nothing
Goodman 2014 student FE + snowfall IV, Massachusetts B own absence −0.008 SD/day (within-student); school closures ≈ 0, ruling out effects >0.01 SD/day
Pischke 2007 DiD, German short school years B losing ⅔ of a year: grade repetition +20-25% relative; Gymnasium entry null; earnings 0.003 (CI −0.019 to 0.026)

Adding time back

Source Design Grade Key effect
Lavy 2015 within-pupil across-subject, PISA 2006, 50+ countries B +5.76 PISA points/weekly hour = 0.15 within-pupil SD = 0.07 between-pupil SD; concave; developing countries half
Bietenbeck & Collins 2023 direct replication, 6 PISA + 6 TIMSS waves B reproduces 2006 at 0.058, gets 0.014-0.037 everywhere else; blames categorical measurement
Rivkin & Schiman 2015 school-by-subject FE, PISA 2009, 72 countries B 0.017-0.025 per weekly minute; 0.025 SD at the 75th percentile of classroom quality vs 0.011 at the 25th
Andersen 2016 (Denmark) cluster RCT, 90 schools / 1,931 pupils A +3 h/week × 16 weeks: +0.15 SD with NO teaching programme, +0.04 (n.s.) WITH an expert programme
Meyer & Van Klaveren 2013 RCT, 7 Dutch schools, +5 h/week × 3 months B "no significant effect on math or language achievement"
Checkoway 2013 (Massachusetts ELT) matched CITS, 24 vs 25 schools, +300 h/yr × 5 years C no significant achievement effects in years 1-3; one science result in year 4; more student and teacher fatigue, less liking of school
Bellei 2009 (Chile) DiD on the full-school-day reform, +27% annual time C language 0.05-0.07 SD; maths 0.00-0.12; larger for the already higher-achieving
Figlio 2018 (Florida) sharp RD, +1 literacy hour in the lowest 100 schools B +0.05 SD year 1+0.07 SD year 2−0.004 SD year 3; the lowest achievers gained least (0.027, n.s.)
Cortes, Goodman & Nomi RD, Chicago double-dose algebra B +0.18-0.24 SD on independent tests, +9.8pp graduation, +3.3pp BA at 12 years — and a null in the next cohort as implementation drifted
Ritchie & Tucker-Drob 2018 meta of quasi-experiments, 42 datasets / 600,000+ B a year of schooling durably raises IQ 1-5 points — but on taught skills, not on g

Summer

Source Design Grade Key effect
Cooper 1996 meta, 39 studies (13 pooled) D "about one month" on grade equivalents; maths worse than reading; SES gap in READING only
von Hippel & Hamrock 2019 re-analysis of BSS, ECLS-K and NWEA GRD B on IRT scales no gap doubled grade 1→8; average gap growth 7%; summer gap growth does not replicate — but "summer learning is slow for nearly all children"
Atteberry & McEachin 2021 ~200M scores, ~18M students, adaptive IRT B 17-28% of the school-year ELA gain and 25-34% of the maths gain lost each summer; 19% of the grade 1-8 pathway is summer; demographics explain ~4% of the variance
Kuhfeld, Condron & Downey 2019 2.5M students, 9 school years / 6 summers B the Black-White gap widens IN SCHOOL and narrows over summer; every group loses ground in summer
Kidron & Lindsay 2014 WWC-screened meta, 30 of 7,000+ studies B summer-school literacy g = 0.16 (CI −0.04 to 0.36), n.s.; "insufficient evidence" overall
Lynch 2023 meta, 37 design-screened studies B summer maths +0.10 SD; 33% of effects negative; no-control-group studies average 0.30 SD
Augustine 2016 (RAND) lottery RCT, 5 districts, ~5,600 randomised A +0.08 SD maths after summer 1, dissipated by the next fall; two summers: nothing

Hereditarian-lens assessment

Risk: low, and unusually cleanly so. Nearly every estimate above comes from a shock imposed on children by an institution: a district changing its calendar for budget reasons, a state mandating an extra hour in its lowest-scoring 100 schools, a government closing every school in the country, a military bureaucracy assigning an enlistment date from a birthdate and a parish. None of those decisions is made by anyone whose genotype is correlated with the outcome. Carlsson and colleagues go furthest and demonstrate conditional randomness directly, showing that days of schooling are uncorrelated with ninth-grade grades, parental education and paternal income once the assignment variables are conditioned on.

Two places where the confound has not been removed, and they are the two places this literature makes its most political claims:

  • The socio-economic gradient in summer loss. Cooper's reading result — middle-class children gaining and disadvantaged children losing — is a between-family comparison of what happens when school stops constraining the environment. Families differ in books, supervision, travel and parental ability, and they differ in alleles, and a seasonal comparison cannot separate them. The strongest datasets now say the gradient is small anyway: race and SES together explain about 4% of the variance in summer loss.
  • The COVID socio-economic gradient, for the same reason. That disadvantaged children lost more is well-established; that home environment caused the difference is not identified by any study in the meta-analysis.

Carlsson's crystallised/fluid split is the archive's premise showing up inside a single well-identified study. Ten extra school days move synonyms and technical comprehension; they move spatial and logic tests not at all, and if anything negatively relative to a non-school day. Schooling buys taught knowledge. It does not buy general ability. That is exactly the distinction Ritchie, Bates and Deary draw from the other direction, and it is why "a year of school raises IQ by 1-5 points" and "no pedagogical intervention durably raises g" are both true.

Boundaries & what critics say

  • The famous number is scale-dependent, and three different numbers describe one coefficient. Lavy's estimate is 5.76 PISA points, which is 0.15 of the within-pupil SD (38.8), 0.07 of the between-pupil SD (84.4), and 0.058 of the 100-point international scale. Advocacy quotes the largest. Any claim about "the effect of an hour" that does not name its denominator is unusable.
  • Snow days do not cost what absences cost, and this is the deepest result about time in the archive. Goodman shows the earlier snow-day literature was mis-identified because snowfall moves both closures and individual absences; separating them, each snow-induced absence costs 0.05 SD of maths while each closure day costs an amount statistically indistinguishable from zero, with effects larger than 0.01 SD/day ruled out. Synchronised time loss is absorbed by the teacher; staggered time loss is not. Time is not a quantity of hours, it is a quantity of coordinated hours.
  • The COVID estimate is not a clean subtraction of instructional hours. Closure substituted remote schooling for in-person schooling and arrived with household disruption, illness and economic shock. It bounds what losing normal schooling costs; it does not price an hour.
  • The Danish RCT's ranking is uncomfortable for everyone. Unstructured extra time (+0.15 SD) beat expert-programmed extra time (+0.04, n.s.), and the authors conclude "a general increase in instruction time is at least as efficient as an expert-developed, detailed teaching program." The arms could not be statistically separated, so this is a null result about programme design rather than a finding that structure hurts — but it should stop anyone claiming the extra hour only works if it is scripted.
  • The summer-loss dispute is genuinely two disputes, and honest reporting has to keep them apart. Von Hippel and Hamrock demolish the measurement of the classic result: the Baltimore study used Thurstone-scaled fixed-form tests whose scale spreads with age, and estimated school-year learning within a test form while estimating summer learning across two different forms. On IRT scales, average gap growth from grade 1 to grade 8 is 7 percent and no gap doubles. But they do not deny summer loss: "the figures do consistently show that summer learning is slow for nearly all children, including children from advantaged groups." Atteberry and McEachin, on adaptive tests immune to the form artifact and with reconstructed district calendars, find mean loss at every grade and put it at a fifth to a third of the school-year gain. The mean loss replicated. The gap story did not — and Kuhfeld, Condron and Downey find the Black-White gap does the opposite of what the canon says, widening during school and narrowing over summer because White students lose more.
  • A third mechanism in that dispute is still unquantified. Kuhfeld notes that students disengage from testing more in autumn than in spring, and that autumn tests are typically administered four to eight weeks into the year. Nobody has bounded how much apparent summer loss is test-taking effort. Nor has anyone bounded regression to the mean, which Kuhfeld shows is the single strongest predictor of summer change: prior-year growth alone explains 22-39% of the variance.
  • Design quality predicts effect size here as reliably as anywhere in the archive. Lynch and colleagues report summer studies with no control group averaging 0.30 SD against 0.09 SD with one. Kidron and Lindsay explain Cooper's 0.26 SD summer-school result as an artifact of "mostly studies with less rigorous study design." The RAND lottery gives 0.08 SD, fading. The pattern is monotone.
  • The one unambiguous win is targeted, not general. Chicago's double-dose algebra doubled instructional time in one subject for below-median freshmen and produced 0.18-0.24 SD on independent tests, +9.8pp graduation and +3.3pp BA attainment twelve years later — and then produced a null in the following cohort when cut-score adherence slipped and peer sorting worsened. Extra time works when it is aimed at a specific skill deficit in a specific group. It does not work as a longer day.

Practical guidance

  • Protect the time you have before buying more. Every subtraction experiment finds a cost; most addition experiments do not find a benefit. A four-day week is a 3.6-hour weekly cut for 0.03-0.06 SD; individual absences cost about 0.008 SD each within-student.
  • If you add time, add it to a subject and a group, not to the day. Chicago's double-dose algebra and Florida's literacy hour are the two positive results, and both are subject-specific and targeted. Massachusetts's 300 extra hours a year for everyone, sustained five years, bought nothing and cost engagement — more students reported being tired and fewer liked being at school.
  • Budget 0.02-0.03 SD per weekly subject-hour, and less if your classrooms are disorderly. If a proposal forecasts more than that, ask which standard deviation it is dividing by.
  • Fix the classroom before extending it. Rivkin and Schiman's interaction is the most policy-relevant number in this topic: a school in the bottom quartile of classroom order gets roughly nothing from extra hours. Extra time is a multiplier on whatever is already happening.
  • Expect real summer forgetting and do not sell summer school as gap-closing. Plan for a fifth to a third of the year's gain to come back off, mostly in maths. But the achievement gap does not open in summer, so a summer programme justified on equity grounds is being justified on a claim that the best data contradict. Randomised summer programmes buy about 0.08-0.10 SD and it fades.
  • Never justify anything here as raising ability. Schooling durably raises measured IQ by 1-5 points a year and it does so entirely through taught skills; Carlsson's fluid-intelligence nulls are the cleanest single demonstration.

Open questions

  • Confidence is capped at medium by process, not evidence. Several independent grade-A/B sources agree on the core asymmetry, which would clear the bar for high; status: surveyed means the adversarial pass METHODOLOGY requires has not been run.
  • Nobody has priced a marginal hour with random assignment at scale. The Danish trial (+0.15 SD for 3 h/week over 16 weeks) and the Dutch trial (nothing for 5 h/week over 3 months) disagree, and both are small and short. The obvious experiment — randomise schools to +2, +4 and +6 weekly hours and follow them two years — does not exist.
  • How much apparent summer loss is test-taking effort? The mechanism is named by the field's own leading measurement expert and has never been bounded. Until it is, the 17-34% figures are upper bounds.
  • Why did extended-time programmes fail where targeted double-dosing succeeded? The candidates — who gets the extra hour, what fills it, whether the teacher is fresh, whether the group is ability-matched — are all testable and none has been tested against the others.
  • Almost nothing here reaches attainment or earnings. Pischke's German earnings null is precisely estimated and is essentially the only long-run evidence on school-year length. Given how often this archive finds attainment effects that outlive test-score fadeout, that is the largest gap.

Evidence (23 sources)

Export all: BibTeX · RIS

← Back to explore