Does educational technology raise learning — CAI, adaptive software, devices, screens, and AI tutors?
Sort by what the software replaces: adaptive drill inside the school day buys +0.05-0.20 SD; a device, a connection or the teacher buys zero to negative; LLM tutors have no usable evidence at all.
mixedconf: mediumgc: lowedtech · ages 5–18 · structure
The word 'ed-tech' names four interventions with four different signs, and every confused claim in this literature comes from averaging them. (1) ACCESS — give a child a computer, a laptop or an internet connection — is the best-identified null in education: a Dutch RD gives language -0.079 and arithmetic -0.061; a Romanian RD gives school grades -1/3 SD in three subjects; a California RCT gives 95% CIs of [-0.15, 0.05] SD on state tests after a 55-point ownership shock; the largest OLPC trial (318 schools) gives math +0.046 and language -0.039 and rules out anything above 0.11 SD; ten years of registry data on 531 of those schools rules out anything above 0.04 SD and finds grade progression 1.0pp WORSE. (2) SUPPLEMENTAL ADAPTIVE SOFTWARE inside the school day is where the positive results live, and they are real but small: Cheung & Slavin's 74-study math meta gives +0.16 overall, +0.08 restricted to randomized trials and +0.06 (CI 0.00-0.13) restricted to LARGE randomized trials; the 16-product federal RCT across 132 schools gives 0.03/0.02/0.07/-0.06; the EEF's eight independently evaluated trials have a median of one month's progress. The steelman is Mindspark, which survived scale-up honestly — +0.37 SD after 4.5 after-school months in Delhi becomes +0.22 SD after 18 in-school months across 40 government schools, deflation rather than collapse. (3) SOFTWARE SUBSTITUTING FOR THE TEACHER is negative wherever it is tested: -0.215 SD in a randomized same-instructor trial, -0.33 SD with the loss travelling into NEXT term's GPA and a 10.6pp drop in re-enrolment, -0.19 SD and 12 fewer percentage points of credit in online credit recovery, and -0.10 reading / -0.25 math per year in full-time virtual schools. (4) LLM TUTORS are unevidenced: five trials, four developer-led, ZERO measurements taken after the tool was withdrawn, and the one independent trial found students who had used unguarded GPT-4 during practice scored 17% BELOW students who had used nothing. The measurement pattern that explains the whole literature: intelligent tutoring systems score 0.73 on researcher-built tests and 0.13 on independent standardized ones, 0.46 immediately and 0.155 at the end of the school year, 0.78 in studies under 80 students and 0.30 in studies over 250.
Buy software, not hardware, and buy it for one specific job: drilling a well-structured skill, at the child's actual level, in a scheduled period inside the school day, on devices you already own. Budget 0.05-0.15 SD, hold the vendor to usage logs, and insist the evidence you are shown comes from an independent standardized test rather than the product's own assessment — that single distinction is worth a factor of five in this literature. Expect the largest gains exactly where a teacher is most stretched (big classes, high absence, wide spread of prior attainment) and nothing where a teacher can already individualise. Never let the software replace the teaching. And on AI tutors specifically: the archive holds no evidence that supports buying one. Nobody has yet measured what a student retains after the model is taken away, except once, and that measurement was negative.
Who this applies to
Verdict
mixed, and the mixture is not noise — it is four different interventions wearing one word. Almost every
argument about educational technology is an argument between people describing different objects. Sorted by
what the technology replaces, the evidence is remarkably consistent within each category and points in
opposite directions across them:
| What is being replaced | Sign | Best estimate |
|---|---|---|
| Nothing — a device, a connection, a laptop to take home | zero to negative | −0.08 to +0.05 SD; the tightest design rules out anything above 0.04 SD |
| Part of the practice hour — supplemental adaptive drill in school | small positive | +0.06 to +0.20 SD on standardized tests; +0.22 SD is the best result ever recorded at scale |
| The teacher — online courses, virtual schools, credit recovery | negative | −0.19 to −0.33 SD, and it travels into later terms |
| Paper — reading a text on a screen | small negative | −0.21 SD, and growing since 2000 |
| Unknown — LLM tutors | insufficient |
no independent, standardized, post-withdrawal estimate exists anywhere |
This structure is why the field's meta-analyses disagree with its trials. A meta-analysis that pools all four categories reports something near +0.16 and is describing nothing in particular. The moment you split by design quality, measure type, and comparison condition, the +0.16 decomposes into a large number from weak studies on aligned tests and a small number from strong studies on real ones.
The archive's grading rules were built for exactly this literature, and they explain most of it before any substantive question is asked. Take Kulik & Fletcher's 50 evaluations of intelligent tutoring systems, the source of the famous 0.66. Inside that single review:
- 0.73 on local researcher-built tests, 0.13 on independent standardized tests (test type is the strongest moderator in the paper, r = −.63)
- 0.78 in studies with ≤80 participants, 0.30 in studies with >250
- 0.60 against conventional classroom controls, 0.18 when controls were handed the ITS's own content as text
- 0.44 under strong implementation, −0.01 under weak
- For Cognitive Tutor alone: 0.86 on local tests, 0.16 on standardized ones — same product, same review
Kulik & Fletcher's own K-12 mathematics cell restricted to standardized measures is 0.10. That is the number the big randomized trials report. The famous disagreement between the ITS literature and the RCT literature dissolves inside the ITS literature's own moderator tables.
1. Access: the best-identified null in education
Six causal designs, four continents, three decades, one answer.
| Source | Design | Grade | What happened |
|---|---|---|---|
| Cueto 2025 (OLPC long run) | randomized, 531 schools, 10 years of registry data | A | Academic index −0.046; rules out anything above 0.04 SD; grade progression −1.0pp; no effect on secondary completion, university application, or years of education |
| Cristia 2017 (OLPC Peru) | cluster RCT, 318 rural schools | A | Math +0.046 (ns), language −0.039 (ns), average +0.003; rules out >0.11 SD. Computers per student went 0.12 → 1.18 |
| Fairlie & Robinson 2013 | RCT, 1,123 students given free computers | B | +55pp ownership, +2.5 h/week use → grades CI [−0.10, 0.06], STAR CIs [−0.15, 0.05] and [−0.16, 0.04] |
| Malamud & Pop-Eleches 2011 | RD on a voucher income cutoff, Romania | B | School grades −1/3 SD in maths, Romanian AND English; computer skills +0.22–0.40; homework, reading and TV all down; games +13–23pp; educational-software use a precise zero |
| Beuermann 2015 (OLPC at home) | in-class lottery, ~1,000 laptops, Lima | B | XO-laptop proficiency +0.88 SD; general PC/internet skills +0.02 (ns); Raven's +0.05 (ns); reading −6pp, teacher-rated effort −5pp |
| Leuven 2007 | RD on earmarked Dutch computer funding | B | Language −0.079 (p<.01), arithmetic −0.061 (p<.10); negative at every bandwidth and every year, and more negative in year two |
| Goolsbee & Guryan 2006 | E-Rate subsidy variation, all California schools | B | Connected classrooms +68% versus counterfactual; math +.001, reading −.025, science −.003; two-year lags no better |
| Barrera-Osorio & Linden 2009 | RCT, 97 Colombian schools, computers plus 20 months of teacher training | B | Spanish +0.077 (ns), math +0.088 (ns); rejects effects above 0.2 SD |
These are not underpowered nulls. Every one has a large, verified first stage — ownership, connectivity, hours of use, computers per student all moved enormously — and several are precise enough to exclude effects a quarter the size of the ones being claimed. Cueto's ten-year follow-up on 531 randomized schools rules out anything above 0.04 SD, which is about as close to proving a zero as education research gets.
Barrera-Osorio & Linden explains the mechanism better than any of them, because it measured it. The machines arrived at 96% of treated schools, 95% of teachers took the training, and students' computer use rose 29 percentage points in computer-science class and 0.02 points (ns) in Spanish, the subject the entire programme targeted. Cueto found the same thing ten years later in Peru: teachers got the training, gained no digital skills, and did not change how they taught; students used the laptops for entertainment (+0.139**) and not for schoolwork (+0.101, ns). The delivered-but-not-integrated failure is the modal outcome of buying hardware, and it is documented, not inferred.
The one apparent exception is a cognitive-test result that did not survive. Cristia reported a pooled index of
Raven's, verbal fluency and coding at +0.110, significant only at the 10% level, which the published abstract
itself calls "inconclusive" — and which rests partly on three-minute timed tests that were administered for
roughly ten minutes in 40% of schools. Four years of exposure later, Cueto found Raven's at +0.046 (ns) and the
pooled cognitive index at +0.136 (ns). Malamud & Pop-Eleches' Raven's coefficient behaves the same way: significant
at wide bandwidths, +0.013 and +0.027 (ns) at the narrowest ones, which is the opposite of what a real
discontinuity does. Per METHODOLOGY's outcome taxonomy this was always a far-transfer/g claim carrying the
burden of proof, and it has not met it.
2. Supplemental software: where the real, small effects live
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Cheung & Slavin 2013 (math) | meta, 74 studies / 56,886 students | C | +0.16 overall → +0.08 randomized only → +0.06 (CI 0.00–0.13, p=.07) large randomized only. Supplemental drill CAI +0.19 > computer-managed +0.09 > flagship adaptive models +0.06 (ns) |
| Cheung & Slavin 2012 (reading) | meta, 84 studies / 60,553 students | C | +0.16 overall, +0.07 large randomized; supplemental CAI +0.11, "not educationally meaningful" in the authors' words; publication bias present (+0.25 published vs +0.14 unpublished) |
| Dynarski 2007 (federal RCT) | 16 products, 132 schools, 439 randomized teachers | A | Grade 1 reading 0.03, grade 4 reading 0.02, grade 6 math 0.07, algebra −0.06 |
| Campuzano 2009 (year 2) | second cohort, same products, experienced teachers | B | Experience did not rescue it: grade-6 math got significantly worse (−2.80 NCE, p=.02), Algebra I better (+2.90, p=.05), other two ns. 9 of 10 products individually null |
| Pane 2014 (Cognitive Tutor) | pair-matched RCT, 73 schools / 11,066 students | B | ~0 in year 1; +0.21 in year 2 (fragile); companion geometry trial significantly negative; teachers routinely overrode the mastery gating |
| Barrow 2009 (I Can Learn) | RCT, 2,278 students / 17 schools | B | +0.17 SD on the study's aligned test (~0.09–0.11 on national norms); on independent state tests, significant in one district of three |
| Rouse & Krueger 2004 (Fast ForWord) | RCT, 512 students | B | ~0.10 on the vendor's own test; 0.04–0.06 (CI −0.08 to +0.16) on the state reading test; CELF-3 null |
| Muralidharan & Singh 2025 (Mindspark at scale) | RCT, 40 government schools, ~6,500 students/yr, 18 months | A | +0.22 SD math, +0.20 SD Hindi — the best result in the literature, from software carved OUT of existing instructional time |
| Muralidharan 2019 (Mindspark Delhi) | RCT, 619 volunteers, after-school | B | +0.37 math / +0.23 Hindi in 4.5 months — but 9 added hours a week, 58% attendance, a human teaching assistant, and +0.058 (ns) on the schools' own maths exams |
| Banerjee 2007 (India CAL) | RCT, 111 municipal schools | A | +0.35 then +0.47 SD during the programme → +0.097 (SE 0.053) one year after it ended; combined score +0.008 |
| EEF 2019 | tabulation of 8 independently evaluated UK ed-tech RCTs | C | Median +1 month, range −1 to +3. Same programme: +3 months in the efficacy trial, +1 in the effectiveness trial |
| Escueta 2020 | review of ~100 experimental studies | C | Hardware ≈ 0; CAL 20 of 29 developed-country RCTs positive, 15 of those 20 in mathematics; nudges positive and very cheap; fully online small negative |
| Pane 2017 (RAND personalized learning) | matched virtual-comparison, 40 schools | C | +0.09 math (sig), +0.07 reading (ns) — down from 0.27/0.19 in the same team's 2015 report; school estimates span −0.5 to +0.6 |
The honest planning number is 0.05–0.15 SD, and the way you get there is by throwing away the inflated ones. Cheung & Slavin's gradient is the cleanest demonstration in the archive that effect size in this field is a function of study design rather than of technology: +0.16 → +0.08 (randomized) → +0.06 (large and randomized), and the doubling of small over large studies holds inside every design category. Quasi-experiments return 2.5× the randomized estimate. And the ordering by product type is the exact reverse of the marketing: cheap supplemental drill (+0.19) beats computer-managed learning (+0.09) beats the flagship comprehensive adaptive platforms at +0.06, not significant.
Mindspark is the steelman and it survives — with deflation, not collapse. This matters, because the archive's default expectation is that efficacy results vanish at scale. Here they did not. Moving from 619 self-selected Delhi volunteers in an after-school centre to 40 Rajasthan government schools and ~6,500 students a year, with lab time carved out of existing maths and Hindi periods rather than added to the day, the effect went from +0.37 SD in 4.5 months to +0.15 SD in six months and +0.22 SD in eighteen. Per unit of exposure that is roughly a quarter of the efficacy estimate — a real efficacy-to-effectiveness curve that lands somewhere useful instead of at zero. It is the one result in this topic that would justify a purchase.
And the reason it works is not the technology. Muralidharan's own account is that the software's value is teaching each child at their actual level, which for Indian middle-schoolers is several grades below their grade level — the mechanism is diagnosis and targeting, not the screen. The corroborating evidence is inside the trial: on the schools' own grade-level exams the same students gained +0.19 SD in Hindi and +0.058 (ns) in maths, precisely because the maths instruction had been correctly aimed several years below the exam. See placement & mastery diagnosis and mastery learning; this is the same finding those topics record, arriving through a different door.
Barrow's mechanism analysis is the most useful applicability result in the whole topic. The CAI advantage was +0.21 SD in a class of 25 and +0.01 SD (p=0.89) in a class of 15; under +0.06 at average class attendance and +0.35 where classmates' attendance was one SD below average; largest where classes were both large and heterogeneous in prior attainment. Software wins exactly where a human teacher is most stretched and ties where a teacher can already individualise. That is a rule a school builder can act on, and it also predicts why the same products do nothing in well-staffed schools. One moderator cuts against the pacing story and is recorded honestly: the effect did not differ by the student's own prior attainment (p > 0.60), which a pure self-pacing account would predict.
Fadeout is real here and is measured only rarely. Banerjee's +0.47 SD became +0.097 (SE 0.053) one year after the programme stopped, and the combined test-score effect was +0.008. Leite's US K-12 ITS meta finds 0.460 measured immediately versus 0.155 at the end of the school year — two-thirds of the effect gone in the months between the posttest and the year-end test — and 0.110 (ns) for interventions running longer than six months.
3. Substituting for the teacher: the ceiling, and it is below the floor
This is the sharpest test in the topic, the causal evidence is unusually good, and the answer is unambiguous.
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Kofoed 2024 (West Point) | randomized online vs in-person, same instructors and materials | B | −0.215 SD final grade (p=.011); negative on invigilated exams AND homework; online students reported studying slightly more |
| Bettinger 2017 (DeVry) | 2.3M student-course observations, instrumented | B | −0.44 grade points (≈ −0.33 SD); next-term GPA −0.173; still enrolled a year later −10.6pp; penalty twice as large for weak students |
| CREDO 2015 (virtual charters) | matched virtual twins, 158 online charters / 17 states | C | −0.10 SD reading, −0.25 SD math per year; identical against brick-and-mortar charter twins, so it is the online delivery not the charter status |
| Heppen 2017 (Chicago credit recovery) | RCT, 1,224 ninth-graders | B | Credit recovered 66% vs 78% (d = −0.35); algebra posttest d = −0.19; adding a maths-certified mentor did not help; all differences gone by year two |
| Rickles 2023 (LA credit recovery) | RCT, 24 LA high schools | B | Algebra posttest +0.15 (ns) with a credentialed teacher kept in the room — but English 9 credit −23pp (≈ −0.58) |
| Means 2010 (US DoE meta) | meta, 50 contrasts | C | Headline +0.20 — but purely online vs face-to-face = +0.05, p = .46, "statistically equivalent"; blended +0.35; +0.13 when curriculum was identical across arms vs +0.40 when it differed |
| Figlio 2013 | randomized live vs online lecture | C | +1.44/100 favouring live (ns); +4.05 for below-median achievers (p<.01) |
The persistence results are what make this a learning finding rather than a grading artefact. Bettinger's online penalty travels forward into the next term's mostly-in-person GPA and into re-enrolment records a year later; a lenient or harsh marker in one online section cannot do that. Kofoed's students studied more minutes per day and still scored lower, on invigilated exams, with the same instructor teaching both arms.
The two credit-recovery RCTs disagree in a way that identifies the active ingredient. Chicago removed the teaching role from the room — a remote online teacher plus an in-room "mentor" with no instructional expectation, only 53% of them maths-certified — and got −0.19 SD and −12pp. Los Angeles kept a credentialed teacher in the room and swapped only the curriculum onto a screen — and algebra came out at +0.15 SD (ns). The harm tracks how much of the teacher was removed, not whether the content was on a screen. The EEF's ABRA trial makes the same point from the other direction: identical literacy content delivered online and on paper, both beat business-as-usual, neither beat the other.
And the single most-quoted pro-online statistic does not say what it is quoted as saying. The US Department of Education's 2010 meta-analysis is universally cited for "online is better, blended is best." Its own revised edition reports purely online versus face-to-face at +0.05, p = .46, calls the two conditions "statistically equivalent", and reports +0.13 when analysts judged curriculum and instruction identical across arms against +0.40 when they differed on multiple dimensions. That is a meta-analysis conceding it measured extra time and extra materials. Its K-12 base was five studies with a non-significant pooled effect; a search back to 1994 had found none at all.
4. Screens as a medium, and screens as a competitor
| Source | Design | Grade | Key effect |
|---|---|---|---|
| Delgado 2018 | meta, 54 studies / 171,055 participants | C | Screen vs paper comprehension g = −0.21 (same in within-participant designs). The paper advantage has GROWN ~0.01/year since 2000 (p=.03). Time-pressured −0.26 vs self-paced −0.09; informational −0.27, narrative +0.01 |
| Altamura 2024 | meta, 39 effects / 469,564 participants | D | Leisure digital reading ↔ comprehension r = .055, against r ≈ .36–.41 for print; negative in primary and middle school (−.025), positive only from high school |
| Beland & Murphy 2016 | DiD on English phone bans, 91 schools | C | +0.064 SD average (p≈.09); lowest achievement quintile +0.142 (p<.01), top quintile −0.025 (ns) |
| Kessel 2020 | same design, 1,086 Swedish schools | C | Precise null: −0.034, CI [−0.092, +0.024] — the upper bound excludes Beland & Murphy's +0.064 |
| Abrahamsson 2024 | event study on Norwegian registers, 477 schools | C | Full-sample null. Girls: GPA +0.08 (p=.064), blind-marked maths +0.07 rising to +0.22 at four years; psychological-symptom consultations down ~60%; bullying −0.25 to −0.35 SD. The girls-vs-boys difference is never statistically established |
| Figlio & Özek 2025 | DiD on treatment intensity, one large Florida district | C | +1.1 percentile on spring tests in year two only; unexcused absences −5 to −10%, mediating about half the gain; year-one suspensions +12%, and +30% for Black students |
| Sana 2013 | two lab experiments, n=40 and n=38 | C | Multitasker −11pp; student merely SITTING IN VIEW of a multitasker −17pp (d ≈ 1.42) — device use is a negative externality, not a private choice |
| Mueller & Oppenheimer 2014 | three lab experiments | C | Longhand beat laptops on conceptual questions (η²ₚ ≈ .12–.13), not on facts; laptop notes ~2× longer and far more verbatim |
| Morehead 2019 (replication) | direct replication + extension | C | Conceptual advantage d = 0.13 (p = .31), and reverses in Exp 2; pooled across all four direct replications d = 0.17, CI [−0.07, 0.40]; a no-notes control matched every note-taking group |
| OECD 2015 (PISA) | cross-sectional international | D | Partial r = −0.26 (computers per student ↔ maths) and −0.65 (school ICT-use index ↔ maths) after GDP and prior PISA; hill-shaped within-country use gradient |
The reading-medium penalty is small, consistent, and moving the wrong way. Delgado's meta-regression on publication year is the finding that matters, because it directly falsifies the standard defence: the paper advantage grew between 2000 and 2017 rather than shrinking as digital-native cohorts arrived. It rests on a single meta-regression and should be held loosely, but Altamura's correlational picture points the same way from the opposite method — leisure digital reading is essentially uncorrelated with comprehension (r = .055) where leisure print reading sits around r = .36–.41, and the digital association is negative in primary and middle school. Note the boundary: the screen penalty is concentrated in informational text under time pressure (−0.27, −0.26) and is absent for narrative (+0.01).
The phone-ban literature is genuinely unresolved and should be presented as such. Beland & Murphy is the
most-cited evidence for bans worldwide; the closest replication, on 1,086 Swedish schools with a tighter design,
returns a null whose confidence interval excludes the original point estimate. Abrahamsson's Norwegian registers
find nothing on the full sample and a coherent bundle of effects on girls — better blind-marked maths, fewer
psychiatric consultations, less bullying — but she states plainly that she cannot reject equality of the girls' and
boys' coefficients. Per METHODOLOGY's replication rule this is a mixed finding, not a lever. Four studies have
now produced four incompatible subgroup stories — low prior achievers of either sex in England, girls only in
Norway, boys and White students in Florida, nobody in Sweden — which is what a noise-shaped subgroup literature
looks like.
Figlio & Özek is the most informative of the four, and it moves the mechanism. In one large Florida district the academic gain appears only in year two (+1.1 percentiles on the spring tests), and an exploratory mediation analysis attributes nearly half of it to a 5–10% fall in unexcused absences. If that holds, the phone in a US high school costs learning largely by keeping students out of the building, not by distracting them inside it — which is neither the attentional story Sana tells nor the in-class story Beland & Murphy assume. It also records a real cost: year-one suspensions rose about 12%, and about 30% for Black students, before settling in year two. That is a transition cost a school should plan for rather than discover.
The laptop-note-taking result should not be cited any more. Mueller & Oppenheimer's "the pen is mightier than
the keyboard" is one of the most-quoted findings in this area and it has a failed direct replication: Morehead's
conceptual-question effect is d = 0.13 (p = .31) and reverses in the second experiment, and pooling all four direct
replications gives d = 0.17, CI [−0.07, 0.40]. What replicates cleanly is the mechanism without its consequence —
laptop notes really are about twice as long and far more verbatim. And Morehead's extension is the more troubling
result: a no-notes control performed as well as every note-taking group, which makes a contrast between two ways
of taking notes hard to interpret at all. Per METHODOLOGY's replication rule, capped at mixed.
The strongest argument for a device policy in this section is not an achievement estimate at all. It is Sana's externality result — the student sitting near the laptop lost nearly twice as much as the student using it — which makes device use a coordination problem rather than a question of individual responsibility. That result comes from two 40-person lab sessions with effect sizes near d = 1, exactly the profile this archive discounts, so it argues for a policy rather than proving one.
5. AI and LLM tutoring: what the archive can and cannot say
This section exists because the archive is being built toward an AI-mediated recommendation tool, and this is the
question that tool will be asked about itself. The answer is insufficient, and it is well-evidenced.
| Source | Design | Grade | Independence | Key effect |
|---|---|---|---|---|
| Bastani 2025 | preregistered RCT, ~1,000 Turkish high-schoolers | B | independent | Practice +48% (GPT Base) / +127% (GPT Tutor). Unassisted exam minutes later: GPT Base −17% BELOW control; GPT Tutor −0.004 (ns) |
| Wang 2024 (Tutor CoPilot) | RCT, 900 tutors / 1,787 students | B | developer-led | Exit tickets +4pp (+9pp for the weakest tutors); NWEA MAP standardized: null, maths trending negative (−0.35) |
| Kestin 2025 (Harvard physics) | crossover RCT, one course, two lessons | C | developer-led | 0.63 SD (0.73–1.3 quantile) vs an active-learning class — one session, in-house quiz, no retention measure |
| De Simone 2025 (Nigeria) | RCT, 1,328 volunteers, 6 weeks | C | developer-led | +0.31 SD on the team's own assessment; +0.206 SD on the school's own English exam; ~18 added lab hours against a control that got nothing |
| Henkel 2024 (Rori, Ghana) | RCT, 11 schools / 477 students | C | developer-led | d = 0.36 — from an unclustered t-test on an 11-cluster design, in the developer's own schools, with 25% ability-selective attrition |
| Doo 2026 | meta, 22 publications / 86 effects | C | independent | g = 0.573 — prediction interval [−0.47, 1.61]; correlational studies 1.698, true experiments 0.762, posttest-only 0.113 (ns) |
Five findings, in descending order of how much they should change a decision:
1. The longest follow-up anywhere in this literature is zero. Not one study has measured anything after the tool was withdrawn. Kestin's outcome is an immediate post-test; Rori's endline was taken with access still live; Nigeria and Tutor CoPilot end at end-of-treatment; the Doo & Park meta did not even code treatment duration, so nobody quoting g = 0.573 can say what dose produced it. There is no evidence, anywhere, about whether any generative-AI tutoring effect persists.
2. The one study that removed the tool found harm. Bastani's students with unrestricted GPT-4 during practice outperformed the control by 48% while they had it and scored 17% below the control on the unassisted exam minutes later. The mechanism is recorded: GPT-4 was correct only 51% of the time, yet its error rate did not carry into exam performance — students were not being misled, they were copying without reading. And they could not tell: GPT Base students did not perceive that they had learned less, and GPT Tutor students believed they had performed better when they had not. The guardrailed version, whose prompt withheld answers and was preloaded with the correct solution, eliminated the harm and bought no learning despite a 127% practice advantage. Tutor CoPilot shows the same divergence across measure distance rather than time: +4pp on the product's own exit tickets, null on NWEA MAP with maths trending negative.
3. Four of the five tutoring trials are developer-led, and the exception is the one that found harm. Kestin's team built the tutor, wrote its prompts, taught the course, wrote the tests and authored the paper. Tutor CoPilot's lead author worked part-time for the tutoring provider. Rori's trial ran inside its maker's own schools on its own curriculum-aligned assessment. The World Bank team designed the Nigerian programme, its prompts and its assessment, then evaluated it. This is not an accusation of bad faith; it is the variable METHODOLOGY says to record, and in this cluster it is almost perfectly confounded with the sign of the result.
4. The two most-cited AI-tutor successes are substantially hand-authored. Rori is 500+ micro-lessons with pre-written explanations, hints and worked solutions, with language models serving as a natural-language interface over a scripted mastery sequence. Kestin's tutor carries human-written step-by-step solutions injected into the prompt, because a system prompt alone "could not reliably provide enough structure." Whatever these trials measure, it is closer to a step-based intelligent tutoring system with a chat front-end than to "an LLM teaches a child."
5. The field's most-circulated effect size was retracted. Wang & Fan's ChatGPT meta-analysis (g = 0.867, 58,000 accesses) was retracted on 22 April 2026 over unresolved discrepancies, with the authors not responding to correspondence; it is recorded as an excluded source so the number has something attached to it. The surviving open meta (Doo & Park) reports g = 0.573 from 22 publications averaging 87 participants each, half of them from Taiwan and mainland China, overwhelmingly undergraduate, entirely researcher-designed in outcome — with a prediction interval of [−0.47, 1.61], meaning it makes no prediction about a new study at all, and a moderator table in which effect size tracks design weakness almost monotonically.
VanLehn's 0.76 does not transfer, and this is the load-bearing point
The archive holds VanLehn 2011, which reports step-based tutoring software at d = 0.76 against human tutoring at 0.79 — statistical parity. That number is going to be reached for every time someone wants to argue that AI tutors work. Four reasons it cannot be:
- Different technology. VanLehn's systems are hand-built domain models with step-level correctness checking. An LLM has no domain model and no verified notion of a correct step. The failure modes are not the same failure modes.
- The 0.76 is itself measure-inflated. Kulik & Fletcher's independent review of the same technology gives 0.13 on standardized tests and Leite's US K-12 review gives 0.200 standardized against 0.483 researcher-designed. VanLehn's evaluations are "STEM-heavy, substantially college samples, experimenter tests."
- The parity claim is thinner than it reads. Ma 2014 puts ITS at −0.11 against individualised human tutoring (ns) and +0.05 against small-group human instruction (ns); Steenbergen-Hu & Cooper 2014 puts ITS −0.25 below human tutoring. "As good as a human tutor" is a non-significant point estimate favouring the human.
- The K-12 version of the claim is near zero anyway. Steenbergen-Hu & Cooper 2013 finds g = 0.01 to 0.09 for ITS in K-12 mathematics, against 0.32–0.37 for the same technology on self-selected college students — same reviewers, same methods, four times the effect on adults who chose to be there.
A 2011 number for a different technology, measured on experimenter-built tests in college samples, cannot underwrite a 2026 claim about language models in schools. The honest statement is that the archive holds zero evidence licensing an LLM-tutoring recommendation, and one preregistered independent trial suggesting that the unguarded version makes things worse.
Hereditarian-lens assessment
Risk: low, and unusually so. This literature is dominated by randomized assignment (Dynarski's 439 randomized teachers, Cristia's 318 randomized schools, Muralidharan's individual lotteries, Kofoed's randomized sections), lotteries (Beuermann's in-class draws), and administrative discontinuities that no family chose (Malamud's income cutoff, Leuven's funding threshold, Goolsbee's E-Rate formula). Genes cannot differ across arms in any of these, and the strongest results in the topic — including all the nulls — come from exactly those designs.
Where heritable confounding does live, it lives in the observational tail, and it is worth naming because that tail is where the enthusiastic claims come from:
- Zheng's one-to-one laptop meta pools 10 of 96 screened sources into +0.16, and what it is mostly comparing is schools that chose to adopt 1:1 against schools that did not. Four randomized trials of the same intervention class return nulls. The most economical reading of +0.16 is selection.
- Khan Academy's usage-achievement associations (recorded as an excluded source) are exactly the confounded shape the premise predicts: students who exceeded their predicted test score had logged 12 more hours and 26 more problem sets. Direction of causation unidentified, and SRI says so itself.
- The cross-national "moderate computer use is best" U-curve (OECD 2015, grade D) is a cross-sectional correlation between self-reported usage and achievement across schools and families that differ on everything. Its headline — that countries which invested heavily in school computers showed no improvement in reading, maths or science — is directionally consistent with every randomized design above and is not independent evidence for it; the within-country hill shape is exactly what selection on student ability and track placement would produce with a true effect of zero, and OECD says as much itself. Held on file because it is the most-quoted statistic in the popular debate and the archive should have its real grade attached.
- The one claim in this topic that would have concerned the premise directly — that laptops raise general
cognitive ability — is the one that failed to replicate (Cristia +0.110 at p<.10 → Cueto +0.046 ns at four
years; Malamud's Raven's coefficient collapsing to 0.013 at the narrowest bandwidth). Nothing here should be read
as a claim about
g.
A second, subtler point the premise sharpens. Adaptive software's equity case is not supported and may be backwards. Steenbergen-Hu & Cooper find K-12 ITS at g = −0.18 for low-achieving samples against +0.04 for general ones; De Simone's Nigerian trial found larger effects for higher-baseline and higher-SES students; Malamud found home computers hurt grades while raising computer skills. Self-directed software rewards the students who can already direct themselves. The counter-examples are instructive about the boundary: Banerjee's CAL helped across the whole ability distribution roughly equally, Beland & Murphy's phone ban helped only the bottom quintile, and Barrow's CAI helped most where teacher attention was scarcest — all three are cases where the software or the policy substituted for a missing adult rather than asking the child to supply the missing structure.
Displacement: the cost almost nobody measures
Technology in a school day displaces something, and the literature's central omission is that it usually does not say what. Of the ~38 primary studies read for this topic, roughly a third measured what the technology replaced, and not one of the fifteen meta-analyses and reviews measured it at all. What the ones that looked found:
Malamud & Pop-Eleches — the reference case. Daily homework down 7–9pp, daily reading for pleasure down 5–18pp (the authors' most robust time-use result), TV down, games up 13–23pp, and educational-software use a precise zero despite the Ministry distributing it free. And the policy finding buried inside it: rules about homework offset roughly half the grade loss without costing any of the computer-skills gain, while rules about computer use removed the skills gain and recovered no academic ground. Protect the displaced activity; do not restrict the device.
Beuermann — reading down 6pp and teacher-rated academic effort down 5pp in Lima. Independent design, independent continent, same sign as Romania, and in neither paper is it the headline.
Dynarski — the classroom-level version. Teacher-as-leader fell from 58% to 29% of observed intervals and individual student work rose from 40% to 84%. The software replaced teaching with seat-work, on-task behaviour did not change, and neither did achievement.
Rouse & Krueger — 90–100 minutes a day pulled out of maths, writing, science, social studies and literature, for a null on reading.
Muralidharan & Singh — the only instructional trial in the archive that priced its own displacement and won: six of 48 weekly periods, about 25% of primary and 40–50% of middle-school maths and Hindi time, traded for +0.22 SD. That is what a defensible purchase looks like — the cost is stated.
Cristia and Fairlie & Robinson — genuine displacement nulls. Home study time and reading did not move in Peru (notable, since the XO shipped with 200 preloaded books and only 26% of control children had books at home), and homework time did not move in California, where the new hours split evenly between schoolwork, games and social networking. Displacement is not automatic.
Kofoed — the useful negative control. Online students studied more minutes per day and still scored lower, which rules out effort substitution as the mechanism for the online penalty.
Figlio & Özek — the only phone-ban study of the four that measured a displacement outcome, and it changes the mechanism: unexcused absences fell 5–10% and mediated about half the test-score gain. Beland & Murphy, Kessel and Abrahamsson measured nothing, so three of the four attribute achievement and mental-health effects to phones without anyone observing what replaced them. OECD names displacement as the decisive mechanism and holds no time-use data at all.
The syntheses are worse than the trials on this. Not one meta-analysis or review in this topic — not Cheung & Slavin, not Kulik & Fletcher, not Zheng, not Delgado, not Means, not Doo & Park — reports what the technology displaced, and OECD names displacement as the decisive mechanism while holding no time-use data whatsoever. Where a source measured nothing, its file carries an explicit "displacement — NOT MEASURED" effects row, so the gap is queryable rather than invisible.
Boundaries & what critics say
- "The nulls are implementation failures." Sometimes, and the archive records the usage data so the claim can be checked rather than asserted. In Dynarski's trial the delivered dose was 15–29 hours a year against a vendor-prescribed 70–135, roughly a fifth. But Campuzano's second cohort tested the rescue directly — the same products, the same teachers, a year of experience — and usage did not rise, grade-6 maths got significantly worse, and only Algebra I improved. Kulik & Fletcher's implementation moderator (0.44 strong vs −0.01 weak) says implementation quality matters enormously; Cheung & Slavin's says the same; and both warn their implementation ratings are contaminated, since 41–53% of studies reported no implementation information and the reviewers rated those themselves. Implementation is a real moderator and a self-sealing excuse, in that order.
- The strongest pro-technology natural experiment is Machin 2007, and it should be read as the steelman. A change in the UK funding rule, used as an instrument for LEA ICT spending, gives Key Stage 2 English +0.020 to +0.022 on the proportion reaching Level 4 (p<.01) and science +0.016 (p<.10), robust to serial-correlation correction and growing with time since the funding shock, with no crowding-out of other spending. It also gives maths +0.007 and +0.002, "very close to zero" — in the subject where ICT is most used — which is hard to reconcile with an instructional mechanism and which the authors flag themselves. Held as the best evidence on the other side.
- The products in the federal null were not a random draw of the market — they were the market's own nominees. 160 products were submitted, screened for prior evidence of effectiveness, reviewed by two external expert panels, and 12 of the 16 selected had won or been nominated for industry awards. The Department of Education explicitly acknowledged this would "tilt the study toward finding higher levels of effectiveness" and accepted the tilt. A null on a pre-screened, award-winning sample is stronger evidence than a null on a random one.
- Dispersion is large and unexplained. Within a single algebra district, school-level effects in Dynarski ran from about −0.75 to +0.75, with 62–63% of the variance between districts and almost none of it explained by measured school or classroom characteristics. Pane 2017's school-level personalized-learning estimates span −0.5 to +0.6. "It works somewhere" is true and, so far, unpredictable.
- Ma 2014 is the one meta that finds no measure-type
gap (standardized 0.41 vs researcher-designed 0.42), in flat contradiction of Kulik & Fletcher (0.13 vs 0.73)
and Leite (0.200 vs 0.483). Recorded, unresolved, and read at
abstractdepth only — a live piece of debt. - Higgins 2019 publishes no retrievable effect size and is recorded as excluded/redundant rather than dropped.
- Behavioural-nudge technology is a different and better bet than instructional technology, and this topic does
not cover it. Escueta's review finds text- and email-based nudges to students and parents generally positive
across the whole education life cycle at very low cost. That belongs with
attendanceandhome-environment-parenting, not here, but a founder shopping for cheap technology levers should know the cost-effectiveness case lives there rather than in instructional software.
Practical guidance
- Buy software, not hardware. Every design that bought access and stopped returned zero or worse. If the pitch is a device programme, the evidence base for it is a decade of well-powered nulls.
- Ask one question of any vendor: was the effect measured on an independent standardized test? In this literature that distinction is worth roughly a factor of five (0.73 → 0.13; 0.483 → 0.200; Fast ForWord's 0.10 on the vendor's assessment → 0.04 on the state test). If the answer is the product's own assessment, the number is not comparable to anything.
- Budget 0.05–0.15 SD. Not 0.4, not 0.7. The large-randomized-trial estimate is +0.06; the best result ever achieved at real scale is +0.22 over 18 months with a fifth of the maths and Hindi timetable handed over.
- Deploy it as a scheduled period inside the school day, 30–75 minutes a week, on foundational procedural content. Voluntary after-school use returns 0.014. Under 30 minutes a week returns 0.06. Over 75 does not return more.
- Buy diagnosis, not adaptivity. The mechanism behind the one large positive result is teaching each child at the level they are actually at, which is often several grades below the one they sit in. A product that cannot place a child accurately is not doing the thing that works. See placement & mastery diagnosis.
- Expect it to help most where your teachers are most stretched. +0.21 SD in a class of 25, +0.01 in a class of 15; +0.35 where classmate attendance is poor. If your classes are small and your teachers can already individualise, this is not your lever — tutoring or class size is a better-evidenced use of the money.
- Do not let software replace the teacher, and be specific about what "replace" means: removing the instructional adult from the room is what causes the harm, not putting content on a screen. Chicago removed the teacher and lost 12 percentage points of credit; Los Angeles kept the teacher and lost nothing on algebra.
- Hold usage logs as the implementation metric. Gains are proportional to logged platform time, and the modal real-world dose is a fifth of the prescribed one. If you cannot see the logs, you cannot tell a product failure from a deployment failure.
- Keep extended informational reading on paper where you have the choice, especially under time pressure. The penalty is small (−0.21) but it is free to avoid, and it has been growing.
- On phones, decide on the externality, not the test scores. The achievement evidence is genuinely contested (+0.064 in England, a precise null in Sweden, a full-sample null in Norway). The better argument is that the student sitting near a device user loses nearly twice what the user loses, which makes it a coordination problem a school can only solve collectively.
- On AI: no recommendation is available, and pretending otherwise would be the exact failure this archive exists to prevent. If you deploy one anyway, the two things the evidence does support are (a) guardrails that withhold answers — they eliminated the measured harm — and (b) using the model to assist an adult rather than to replace one, which is the only AI design here that produced a defensible gain, and even that was null on the standardized measure.
Open questions
- Confidence is capped at
mediumby the missing adversarial pass, not by the evidence. The substantive bar forhighis met several times over — four independent grade-A causal sources (Dynarski's 132-school RCT, Cristia's 318-school RCT, Cueto's ten-year follow-up, Muralidharan & Singh's scale-up) plus twenty grade-B ones, 47 of the 51 new sources read in full text, all agreeing on the conditional structure. Butstatus: surveyedmeans the skeptic review METHODOLOGY requires has not been run, and METHODOLOGY states the requirement as a conjunction. - Three cited sources rest on abstracts and one on secondary reporting — Ma 2014 and both Steenbergen-Hu & Cooper meta-analyses are APA-paywalled, and Zheng's one-to-one laptop meta could not be obtained at all. Ma is the one meta in the topic that finds no measure-type gap, so its abstract-only status is the most load-bearing piece of debt here: it is the single source that would weaken the topic's central claim, and it is the one nobody could read.
- Does any generative-AI tutoring effect survive the withdrawal of the tool? Nobody has looked. This is the single most important unmeasured quantity in the topic and the study that would answer it — independent, powered, standardized outcome, measured a term later with the model switched off — is straightforward to run and does not exist.
- Is there an independent evaluation of any LLM tutor? One, and it found harm. The four positives are all authored by the people who built the product.
- Why did Mindspark deflate rather than collapse? It is the only product in this topic that survived a 20× scale-up into government schools with time carved out rather than added. Whether that is the diagnosis, the Indian counterfactual (control-group students in the bottom third made no measurable progress across a school year), or the implementation regime is untested — and it is the question that decides whether the result travels to a well-functioning school system.
- What explains the enormous between-school dispersion? 62–63% of variance between districts, effects from −0.75 to +0.75 within a single district, and essentially no measured moderator that predicts it.
- Does the archive's
mixedverdict on phone bans survive a better design? Three studies, three different answers, and none of them measured what the removed phone time was spent on. - Nobody has tested adaptive software against the thing it should be tested against. Almost every comparison is against business-as-usual whole-class teaching. The comparisons that matter for a school builder — software versus an equal-cost increment of tutoring, versus small-group instruction, versus a better curriculum — barely exist. Ma's meta gets closest and finds ITS at +0.05 (ns) against small-group human instruction and −0.11 (ns) against individual tutoring.
- Is the screen-reading penalty causal at classroom scale? Delgado's estimate comes from short controlled comparisons; no school-scale trial has randomized the medium of a year's reading and measured comprehension.
- grade BEffectiveness of Cognitive Tutor Algebra I at Scale (RAND RCT)Pane, J. F., Griffin, B. A., McCaffrey, D. F., & Karam, R. · 2014 · rct
- grade CThe Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring SystemsVanLehn, K. · 2011 · review
- grade CThe effectiveness of educational technology applications for enhancing mathematics achievement in K-12 classrooms: A meta-analysisCheung, A. C. K., & Slavin, R. E. · 2013 · meta-analysis
- grade CHow features of educational technology applications affect student reading outcomes: A meta-analysisCheung, A. C. K., & Slavin, R. E. · 2012 · meta-analysis
- grade CEffectiveness of Intelligent Tutoring Systems: A Meta-Analytic ReviewKulik, J. A., & Fletcher, J. D. · 2016 · meta-analysis
- grade CA meta-analysis of the effectiveness of intelligent tutoring systems on K–12 students’ mathematical learningSteenbergen-Hu, S., & Cooper, H. · 2013 · meta-analysis
- grade CA meta-analysis of the effectiveness of intelligent tutoring systems on college students’ academic learningSteenbergen-Hu, S., & Cooper, H. · 2014 · meta-analysis
- grade CIntelligent tutoring systems and learning outcomes: A meta-analysisMa, W., Adesope, O. O., Nesbit, J. C., & Liu, Q. · 2014 · meta-analysis
- grade CDo intelligent tutoring systems benefit K-12 students? A meta-analysis and evaluation of heterogeneity of treatment effects in the U.S.Leite, W. L., Zhang, H., Rana, S., Hao, Y., Hatch, A. D., Kong, L., & Kuang, H. · 2025 · meta-analysis
- grade CUpgrading Education with Technology: Insights from Experimental ResearchEscueta, M., Nickow, A. J., Oreopoulos, P., & Quan, V. · 2020 · review
- grade CLearning in One-to-One Laptop Environments: A Meta-Analysis and Research SynthesisZheng, B., Warschauer, M., Lin, C.-H., & Chang, C. · 2016 · meta-analysis
- grade AEffectiveness of Reading and Mathematics Software Products: Findings from the First Student CohortDynarski, M., Agodini, R., Heaviside, S., Novak, T., Carey, N., Campuzano, L., Means, B., Murphy, R., Penuel, W., Javitz, H., Emery, D., & Sussex, W. · 2007 · rct
- grade BEffectiveness of Reading and Mathematics Software Products: Findings From Two Student CohortsCampuzano, L., Dynarski, M., Agodini, R., & Rall, K. · 2009 · rct
- grade BTechnology's Edge: The Educational Benefits of Computer-Aided InstructionBarrow, L., Markman, L., & Rouse, C. E. · 2009 · rct
- grade BPutting computerized instruction to the test: a randomized evaluation of a "scientifically based" reading programRouse, C. E., & Krueger, A. B. (with Markman, L.) · 2004 · rct
- grade BNew Evidence on Classroom Computers and Pupil LearningAngrist, J. D., & Lavy, V. · 2002 · quasi-experiment
- grade BThe Effect of Extra Funding for Disadvantaged Pupils on AchievementLeuven, E., Lindahl, M., Oosterbeek, H., & Webbink, D. · 2007 · quasi-experiment
- grade BThe Impact of Internet Subsidies in Public SchoolsGoolsbee, A., & Guryan, J. · 2006 · natural-experiment
- grade BNew Technology in Schools: Is There a Payoff?Machin, S., McNally, S., & Silva, O. · 2007 · natural-experiment
- grade BHome Computer Use and the Development of Human CapitalMalamud, O., & Pop-Eleches, C. · 2011 · quasi-experiment
- grade BExperimental Evidence on the Effects of Home Computers on Academic Achievement among SchoolchildrenFairlie, R. W., & Robinson, J. · 2013 · rct
- grade ATechnology and Child Development: Evidence from the One Laptop per Child ProgramCristia, J., Ibarrarán, P., Cueto, S., Santiago, A., & Severín, E. · 2017 · rct
- grade BOne Laptop per Child at Home: Short-Term Impacts from a Randomized Experiment in PeruBeuermann, D. W., Cristia, J., Cueto, S., Malamud, O., & Cruz-Aguayo, Y. · 2015 · rct
- grade ALaptops in the Long Run: Evidence from the One Laptop per Child Program in Rural PeruCueto, S., Beuermann, D. W., Cristia, J., Malamud, O., & Pardo, F. · 2025 · rct
- grade BRemedying Education: Evidence from Two Randomized Experiments in IndiaBanerjee, A. V., Cole, S., Duflo, E., & Linden, L. · 2007 · rct
- grade BDisrupting Education? Experimental Evidence on Technology-Aided Instruction in IndiaMuralidharan, K., Singh, A., & Ganimian, A. J. · 2019 · rct
- grade AAdapting for scale: Experimental Evidence on Technology-aided Instruction in IndiaMuralidharan, K., & Singh, A. · 2025 · rct
- grade BThe Use and Misuse of Computers in Education: Evidence from a Randomized Experiment in ColombiaBarrera-Osorio, F., & Linden, L. L. · 2009 · rct
- grade BVirtual Classrooms: How Online College Courses Affect Student SuccessBettinger, E. P., Fox, L., Loeb, S., & Taylor, E. S. · 2017 · quasi-experiment
- grade BZooming to Class? Experimental Evidence on College Students' Online Learning during COVID-19Kofoed, M. S., Gebhart, L., Gilmore, D., & Moschitto, R. · 2024 · rct
- grade CIs It Live or Is It Internet? Experimental Estimates of the Effects of Online Instruction on Student LearningFiglio, D. N., Rush, M., & Yin, L. · 2013 · rct
- grade BThe Struggle to Pass Algebra: Online vs. Face-to-Face Credit Recovery for At-Risk Urban StudentsHeppen, J. B., Sorensen, N., Allensworth, E., Walters, K., Rickles, J., Taylor, S. S., & Michelman, V. · 2017 · rct
- grade BA Multisite Randomized Study of an Online Learning Approach to High School Credit Recovery: Effects on Student Experiences and Proximal OutcomesRickles, J., Clements, M., Brodziak de los Reyes, I., Lachowicz, M., Lin, S., & Heppen, J. · 2023 · rct
- grade COnline Charter School Study 2015Woodworth, J. L., Raymond, M. E., Chirbas, K., Gonzalez, M., Negassi, Y., Snow, W., & Van Donge, C. (Center for Research on Education Outcomes, Stanford University) · 2015 · quasi-experiment
- grade CEvaluation of Evidence-Based Practices in Online Learning: A Meta-Analysis and Review of Online Learning StudiesMeans, B., Toyama, Y., Murphy, R., Bakia, M., & Jones, K. · 2010 · meta-analysis
- grade CUsing Digital Technology to Improve LearningStringer, E., Lewin, C., & Coleman, R. (Education Endowment Foundation) · 2019 · review
- grade CDon't throw away your printed books: A meta-analysis on the effects of reading media on reading comprehensionDelgado, P., Vargas, C., Ackerman, R., & Salmerón, L. · 2018 · meta-analysis
- grade DDo New Forms of Reading Pay Off? A Meta-Analysis on the Relationship Between Leisure Digital Reading Habits and Text ComprehensionAltamura, L., Vargas, C., & Salmerón, L. · 2025 · meta-analysis
- grade CIll Communication: Technology, distraction & student performanceBeland, L.-P., & Murphy, R. · 2016 · quasi-experiment
- grade CThe impact of banning mobile phones in Swedish secondary schoolsKessel, D., Hardardottir, H. L., & Tyrefors, B. · 2020 · replication
- grade CThe Impact of Cellphone Bans in Schools on Student Outcomes: Evidence from FloridaFiglio, D. N., & Özek, U. · 2025 · quasi-experiment
- grade CLaptop multitasking hinders classroom learning for both users and nearby peersSana, F., Weston, T., & Cepeda, N. J. · 2013 · rct
- grade CThe Pen Is Mightier Than the Keyboard: Advantages of Longhand Over Laptop Note TakingMueller, P. A., & Oppenheimer, D. M. · 2014 · rct
- grade CHow Much Mightier Is the Pen than the Keyboard for Note-Taking? A Replication and Extension of Mueller and Oppenheimer (2014)Morehead, K., Dunlosky, J., & Rawson, K. A. · 2019 · replication
- grade BGenerative AI without guardrails can harm learning: Evidence from high school mathematicsBastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. · 2025 · rct
- grade CAI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational settingKestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. · 2025 · rct
- grade BTutor CoPilot: A Human-AI Approach for Scaling Real-Time ExpertiseWang, R. E., Ribeiro, A. T., Robinson, C. D., Loeb, S., & Demszky, D. · 2024 · rct
- grade CEffective and Scalable Math Support: Experimental Evidence on the Impact of an AI-Math Tutor in GhanaHenkel, O., Horne-Robinson, H., Kozhakhmetova, N., & Lee, A. · 2024 · rct
- grade CFrom Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in NigeriaDe Simone, M., Tiberti, F., Barron Rodriguez, M., Manolio, F., Mosuro, W., & Dikoru, E. J. · 2025 · rct
- grade CA Meta-Analysis of ChatGPT's Influence on Learning AchievementDoo, M. Y., & Park, Y. · 2026 · meta-analysis
- grade CInforming Progress: Insights on Personalized Learning Implementation and EffectsPane, J. F., Steiner, E. D., Baird, M. D., Hamilton, L. S., & Pane, J. D. · 2017 · quasi-experiment
Related decisions
- Class-size reductionmixedconf: highgc: low
- Mastery learning (teach → test → reteach to criterion → advance)mixedconf: highgc: low
- Placement and mastery diagnosis — deciding what to teach next from evidence of current skillmoderate supportconf: mediumgc: medium
- Tutoring — the honest effect, the Bloom 2-sigma myth, and what survives scalestrong supportconf: highgc: low