Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise
Wang, R. E., Ribeiro, A. T., Robinson, C. D., Loeb, S., & Demszky, D. · 2024
grade Brctdeveloper-ledunclearnumbers spot-checked
Sample
900 tutors initially randomized (450/450), 782 at study launch after pre-launch attrition (treatment 386, control 396); 874 full-time tutors assigned to the participating schools; 1,787 students identified by the district. Analysis sample 4,136 tutoring SESSIONS containing 550,000+ chat messages; the end-of-year test analysis has n ~ 1,001 students.
Population
Grades 3-8 mathematics, nine Title I schools (8 elementary, 1 middle) in a large southern US school district serving 30,000+ students; 80% Hispanic, 67% economically disadvantaged. Students were eligible only if they had scored below grade level on the previous spring's state test. Delivery was one-to-one virtual chat tutoring through FEV Tutor's platform.
Design
THE LOAD-BEARING DESIGN FACT: this is an AI ASSISTING A HUMAN TUTOR IN REAL TIME, NOT AN AI TUTORING A STUDENT. Tutor CoPilot is a button inside the tutor's console that generates expert-like suggested moves (ask a guiding question, break the problem down) which the tutor may use, adapt or ignore. The student never talks to a language model. No finding here transfers to "an LLM tutored a child," and the archive should refuse any citation that treats it as such. Preregistered on OSF (osf.io/8d6ha) before data access. Randomization at the TUTOR level; primary model is session-level with strata fixed effects for school x grade, controls for baseline math scores and student demographics, residuals clustered at the student-tutor pair. Balance and attrition checks pass (tutor attrition 11% treatment / 10% control, ns); 6 control tutors were mistakenly given access. Treatment tutors got 2-3 weeks of training on the tool and a buddy system; control tutors got the provider's standard training module - so training dose was NOT equated. Study ran end of March 2024 for TWO MONTHS. PRIMARY OUTCOME is the tutoring platform's own EXIT TICKET pass rate - a within-curriculum mastery check that gates progression to the next lesson. It is administrative, proximal, and internal to the product; it is not a test anyone outside the platform administers. The study also collected NWEA MAP math and reading, which are independent standardized measures, and reported the null on them honestly (Appendix K), noting that most students had near-zero exposure to treated sessions so the analysis was underpowered by construction. INDEPENDENCE: the Stanford team built Tutor CoPilot, and the paper states that "the main author of this study worked part-time with FEV Tutor, the tutoring provider" and collaborated with its engineering, design, operations and curriculum teams to build the system. The people who built the product ran and authored its evaluation - `developer-led`. The mitigations are real and should be weighed: preregistration, an independent standardized outcome collected and reported, and a null reported without spin. Funding acknowledged to the Smith Richardson Foundation and Arnold Ventures/Accelerate. Cost reported at $20 per tutor per year.
Key findings
On the product's own mastery check, students of tutors with Tutor CoPilot passed exit tickets 4 percentage points more often (62% to 66%, SE 0.01, p<0.01), with the gain concentrated in the weakest tutors (+9pp for lower-rated tutors, 56% to 65%; +7pp for less-experienced tutors, 61% to 68%) - the "levels up the bottom" pattern, which is the most interesting thing here. On the INDEPENDENT STANDARDIZED measure collected in the same study - NWEA MAP end-of-year math and reading - there was NO significant effect, with math trending NEGATIVE (-0.35, SE 0.88) and reading positive (+1.44, SE 1.45). The authors attribute the null to thin treatment exposure and a two-month window rather than to absence of effect, which is a fair reading, but the archive records the standardized null as the headline per its own measure-type rule and the exit-ticket gain as the proximal secondary. Message analysis (550,000+ messages) shows treated tutors asked more guiding questions and gave away answers less often - a genuine change in tutor behaviour, which is the mechanism the design was built to test. Student survey outcomes (perceived tutor care, math mindset, session and tutor ratings) all null.
Genetic confound
Minimal (randomized at the tutor level; balance confirmed on baseline covariates, and student baseline math scores are controlled).
Replication notes
No replication located; the authors describe this as "the first randomized controlled trial of a Human-AI system in live tutoring". Recorded as `unclear` rather than `unreplicated` because we have not systematically searched. Still a working paper as of this writing - arXiv 2410.03017 and Annenberg EdWorkingPaper 24-1054 - so it has not been through peer review either.
DOI / URL
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Exit ticket passed (unconditional) - the platform's own topic-mastery check, ITT | pp | +0.04 (SE 0.01, p<0.01); control mean 0.62, i.e. 62% to 66% pass rate. n = 4,136 sessions. | administrative | end of each tutoring session, across a two-month study window | business-as-usual | end-of-treatment | domain-skill |
| Exit ticket passed, students of the LOWEST-rated tutors | pp | +9pp (56% to 65%); +7pp for the least-experienced tutors (61% to 68%). Effects shrink monotonically as baseline tutor quality rises - students of weak treated tutors reached the level of students of strong control tutors. | administrative | end of each tutoring session | business-as-usual | end-of-treatment | domain-skill |
| NWEA MAP end-of-year MATH score (independent standardized measure) | scale-score points per unit of treatment exposure | -0.35 (SE 0.88), not significant; -0.59 (SE 0.86) on the non-imputed-baseline specification. Null, trending negative. n = 1,001 (959 non-imputed). | standardized | end of school year | business-as-usual | end-of-treatment | domain-skill |
| NWEA MAP end-of-year READING score (independent standardized measure) | scale-score points per unit of treatment exposure | +1.44 (SE 1.45), not significant; +1.56 (SE 1.42) non-imputed. Null, trending positive. | standardized | end of school year | business-as-usual | end-of-treatment | domain-skill |
| Exit ticket attempted, and passed conditional on attempting | pp | attempted +0.02 (SE 0.01, p<0.1), control mean 0.84; passed-conditional +0.03 (SE 0.01, p<0.05), control mean 0.73 | administrative | end of each tutoring session | business-as-usual | end-of-treatment | domain-skill |
| Tutor pedagogical behaviour (the mechanism) | classifier-labelled strategy rates over 550,000+ chat messages | treated tutors significantly more likely to use understanding-fostering strategies (asking guiding questions) and less likely to give away the answer. Tutors also flagged that suggestions were sometimes not grade-level appropriate. | researcher-designed | throughout the two-month study | business-as-usual | end-of-treatment | behaviour |
| Student-reported experience (tutor care, math mindset, session and tutor ratings) | 5-point survey items | all null - largest estimate +0.03 (SE 0.04) on tutor rating; n = 1,931-1,952 | self-report survey | post-session | business-as-usual | end-of-treatment | non-cognitive |