Cognitively Guided Instruction cluster RCTs (grades 3-5 efficacy; grades 1-2 replication; 5-year follow-up)
Schoen, R. C., LaVenia, M., Tazaz, A. M., Farina, K., Dixon, J. K., Secada, W. G., Rhoads, C., Perez, A. L. (FSU Learning Systems Institute; IES- and FLDOE-funded) · 2020
grade Brctindependentmixednumbers spot-checked
Sample
Grades 3-5 trial: 149 teachers randomized in 32 schools / 9 FL districts (analytic sample 100 teachers, 2,206 students). Grades 1-2 replication: 22 schools, 2 districts (ITBS N=2,172; MPAC interview N=622). Five-year follow-up: same 22 schools, all K-5 students (i-Ready N~15,000 test records).
Population
Florida elementary students and their teachers; the leading reform-side mathematics PD program.
Design
Aggregates three FSU Learning Systems Institute reports on two independent cluster-randomized trials plus one long-term follow-up, all conducted as third-party evaluations of a programme the evaluators did not build. (1) Grades 3-5 efficacy trial (Research Report 2018-25): individual TEACHERS randomized within 34 school blocks (not schools), wait-list comparison teachers on business-as-usual PD; outcome was the EMSA, a test built by the evaluation team and targeted on the fractions / whole-number content the PD focuses on. Teacher-cluster attrition was 32.9% overall and 11.6% differential, which the report itself says "exceeds the boundary for acceptable threat of bias due to attrition" under WWC 2017; student attrition (9.2% / 4.4%) is acceptable and baseline equivalence (g=0.17-0.21) was established with pretest adjustment. No significant treatment-by-grade or treatment-by-baseline-achievement interactions. (2) Grades 1-2 replication (Research Report 2020-02): matched-pair SCHOOL-level randomization of 22 schools; the comparison arm received a different district-chosen PD programme, so the contrast is against an ACTIVE ALTERNATIVE, not business-as-usual. Outcomes were the researcher-built MPAC one-on-one interview plus the commercially standardized ITBS Math Problems and ITBS Math Computation. Bayesian estimation; every confirmatory credibility interval included zero. (3) Five-year follow-up (working paper 2022; JREE 2024): ITT on the grades 1-2 randomization, all K-5 students in the 22 schools in 2017-18, outcomes the i-Ready Diagnostic and the Florida Standards Assessment — both externally developed standardized tests, which is the strongest measurement in the set. Caveat: comparison-school staff became eligible for the CGI programme from year 2 onward, so the five-year contrast is diluted by treatment diffusion. The PD itself was designed and taught by certified CGI instructors at Teachers Development Group under Linda Levi, a CGI co-author — delivery involvement, not evaluation involvement; the trials were run and analysed by FSU with an external evaluator independently refitting the models.
Key findings
The best modern experimental test of reform-side pedagogy, and it splits by measure content rather than dominating: g=0.18 in grades 3-5 on a test the evaluators built around the programme's own content focus, but a grades 1-2 replication in which nothing was credibly positive and grade-2 computation was credibly NEGATIVE (g=-0.29 to -0.31). The strongest evidence for CGI is the five-year follow-up on independent standardized tests: g=0.16 in grades 3-5 (significant on i-Ready, not on the state exam) and a flat g=0.03 in K-2 — a real but modest durable effect concentrated in the upper elementary grades, an order of magnitude below reform advocacy claims and consistent with guided student-thinking pedagogy trading computational fluency for problem solving.
Genetic confound
Randomized in both trials — teachers within school blocks in grades 3-5, schools in grades 1-2 and the follow-up — so student selection cannot explain the contrasts. The residual story is measure content, not genes: reform-side pedagogy scores positive on problem-solving-flavoured tests and null-to-negative on computation tests within the same trial.
Replication notes
The internal replication weakened and partly reversed the original signal: the grades 1-2 trial produced no credibly non-zero positive effect and one credibly negative one (grade-2 computation). The five-year follow-up then found a durable g=0.16 in grades 3-5 on an externally developed standardized test — but only i-Ready reached significance, the state test did not, and comparison schools had access to the programme from year 2. All five extant CGI RCTs, and every study cited here, come from the same FSU/Wisconsin research community; no independent replication outside it has been located.
DOI / URL
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Grades 3-5 trial, year 1, EMSA mathematics achievement | Hedges g | g=0.18 (p=.007, FIML full sample, N=2,206 students / 100 teachers); complete-case sensitivity analysis g=0.13 (p=.030) | researcher-designed | end of program year 1 (spring 2016) | business-as-usual | end-of-treatment | domain-skill |
| Grades 1-2 replication, year 1, problem-solving / applications / algebraic-thinking measures | Hedges g (Bayesian, 95% credibility interval) | MPAC interview g=0.08 [-0.25, 0.42]; ITBS Math Problems g=0.03 [-0.17, 0.24]. Both intervals include zero. Grade-1 subgroup was the positive half: MPAC g=0.25, ITBS-MP g=0.14 (also not credibly different from zero); grade-2 subgroup MPAC g=-0.01, ITBS-MP g=-0.07. | mixed | end of program year 1 (spring 2014) | active-alternative | end-of-treatment | domain-skill |
| Grades 1-2 replication, year 1, computation (ITBS Math Computation) | Hedges g (Bayesian, 95% credibility interval) | Full sample g=-0.11 [-0.34, 0.11]. Driven by grade 2: subgroup g=-0.29, and g=-0.31 in the grade-2 baseline-test model where the 95% credibility interval EXCLUDES zero — the only credible non-zero main effect anywhere in the trial, and it is negative. Grade-1 computation g=0.03. The report recommends "that the program developers take swift action to adjust the content and delivery of the first year of the program to address important concerns about the potential negative effect on second-grade [students' comput]ational abilities." | standardized | end of program year 1 (spring 2014) | active-alternative | end-of-treatment | domain-skill |
| Five-year follow-up (ITT), grades K-2 mathematics achievement | Hedges g | i-Ready Diagnostic g=0.03, not statistically significant (coefficient 0.89, SE 2.11) | standardized | 5 years after randomization (2017-18) | business-as-usual | over-2yr | domain-skill |
| Five-year follow-up (ITT), grades 3-5 mathematics achievement | Hedges g | i-Ready Diagnostic g=0.16 (coefficient 4.96, SE 1.73, p<.05); Florida Standards Assessment g=0.16 (coefficient 3.63, SE 2.03) — same point estimate, NOT statistically significant on the state test. Grade level was the only significant moderator, effects larger in higher grades. | standardized | 5 years after randomization (2017-18) | business-as-usual | over-2yr | domain-skill |
Cited by
- Explicit vs reform/constructivist math instructionmoderate supportconf: highgc: low