Standardized test outcomes for students engaged in inquiry-based science curricula in the context of urban reform
Geier R, Blumenfeld PC, Marx RW, Krajcik JS, Fishman B, Soloway E, Clay-Chambers J · 2008
grade Dquasi-experimentdeveloper-ledunclear
Sample
Two cohorts of grade 7-8 students participating in LeTUS project-based units, compared with the remainder of the Detroit Public Schools district population
Population
Grades 7-8, Detroit Public Schools, during a district-wide systemic science reform; historically underserved urban students.
Design
This is the study most often cited as evidence that inquiry science raises SCORES ON A REAL TEST, so it matters what it actually is. There is no randomisation and no matched comparison group: participating students are compared with "the remainder of the district population" on the Michigan high-stakes state science test (MEAP). Which teachers adopted the LeTUS project-based units, and therefore which students appear in the treatment group, is a self-selection process inside a district reform. The What Works Clearinghouse reviewed it (March 2020, Project-Based Inquiry Science intervention report) and ruled that it "does not meet WWC group design standards because equivalence of the analytic intervention and comparison groups is necessary and not demonstrated". Every senior author except Clay-Chambers is a developer of the LeTUS/PBIS curriculum being evaluated, so this is developer-led evaluation of a developer's own product with a non-equivalent comparison group. The treatment itself is heavily specified: scripted project-based inquiry units with aligned professional development, learning technologies and administrative support — the developers themselves attribute the result to the curriculum being "highly specified, developed, and aligned", i.e. to structure, not to student autonomy.
Key findings
Effect sizes of about 0.44 (first cohort) and 0.37 (second, scaled-up cohort) on the state standardized science test, with significantly higher pass rates, relative gains persisting up to 18 months after participation, dose-response by number of units taken, and a reduction in the gender gap for African-American boys. These are the largest defensible "inquiry science on a standardized measure" numbers in the literature — and they come from the weakest design in this tranche. Note the direction of the comparison with Harris 2015: the same programme family scored d=0.21-0.25 on researcher-built tests when actually randomised, and d=0.37-0.44 on the state test when not randomised. That is the opposite of what measure-type inflation alone would predict, which is itself the tell that selection, not measurement, is doing the work here.
Genetic confound
Medium-to-high for a causal reading. Participation was not randomised and the comparison group is the rest of a district; family and school selection into the reform-adopting classrooms is unaddressed, and prior achievement equivalence was not demonstrated (WWC).
Replication notes
Its companion randomised trial of the same curriculum family (Harris 2015) was also rejected by WWC. WWC's 2020 review of seven Project-Based Inquiry Science studies found none that met standards, so the programme has no WWC-eligible evidence at all.
DOI / URL
10.1002/tea.20248
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Michigan state standardized science test (MEAP), first cohort | effect size | 0.44 | standardized | end of participation, and up to 18 months after | business-as-usual | 1-2yr | domain-skill |
| Michigan state standardized science test (MEAP), second (scaled-up) cohort | effect size | 0.37 | standardized | end of participation | business-as-usual | end-of-treatment | domain-skill |
Cited by
- Inquiry-based science teaching vs explicit and textbook science teachingmixedconf: mediumgc: low