The Evidence on Teaching

Adaptive e-learning environment based on learning styles and its impact on development students' engagement

El-Sabagh HA · 2021

grade Dquasi-experimentdeveloper-ledunclearnumbers spot-checked
Sample
118 first-year university students in two intact classes - experimental n = 60 (36M/24F), control n = 58 (31M/27F), aged 17-18, one institution, one "learning skills" course, one semester (first term 2019-2020)
Population
Higher-education students, Umm Al-Qura University, Mecca, SAUDI ARABIA (the author's second affiliation is Mansoura University, Egypt, but the study was run in Saudi Arabia)
Design
Recorded because it is the single most-cited recent paper citing Pashler et al. (2008) - an adaptive e-learning system built ON learning styles rather than a test of them. Assignment was of two intact classes to conditions (the paper says the two classes were "selected randomly" and one "randomly assigned as the control group"), i.e. a two-cluster allocation with no student-level randomisation and no clustering adjustment. The outcome is a researcher-adapted self-report engagement scale (Dixson's 48 items translated to Arabic and cut to 27), not learning; no achievement measure was administered at all, despite the discussion claiming the experimental group "had higher learning achievement". No crossover interaction is estimated, and there is no style-matched vs style-mismatched contrast - every experimental-group student got their own matched path, so the design cannot separate matching from simply having a better-built course. THE REPORTED STATISTICS DO NOT COHERE, and this is recorded because it bears on whether any number here can be used at all: (1) the pre-test table (Table 3) reports t-values of 0.32-0.63 and concludes baseline equivalence, but its own means and SDs imply the opposite - Skills 21.07 (SD 1.89) vs 24.25 (SD 1.72) with n ~ 59/group gives t ~ 9.5, not 0.464, and the control group starts 3.2 points AHEAD on Skills; (2) the four pre-test subscale means sum to 58.57 while the "whole engagement scale" mean is given as 26.76, so the subscale and total columns cannot both be right; (3) Table 4 is captioned "pre-test results" but reports post-test data; (4) the reported overall Cohen d = 0.826 is irreconcilable with Table 4's own means and SDs, which imply d between about 6 and 8 on every subscale (e.g. Skills 34.81 (1.34) vs 23.34 (1.79) gives d ~ 7.3); (5) the reported t-values in Table 4 (2.07 to 4.74) are likewise far too small for those means and SDs, which imply t ~ 40 on Skills. Exemplifies the pattern Newton (2015) documented: the literature an educator finds when searching 'learning styles' overwhelmingly endorses the practice. Grade D is if anything generous - researcher-developed self-report outcome only, which is the D row of the grade table on its own.
Key findings
Reports the experimental group statistically significantly higher than control on all four factors of a self-report student engagement scale (skills, participation/interaction, performance, emotional) and on the total, with a claimed overall Cohen d = 0.826 (r = 0.401). The reported effect size, t-values, means and SDs are mutually inconsistent (see design_notes), so no figure here should be propagated as an estimate. Nothing about learning was measured.
Genetic confound
High - non-randomised at the student level (two intact classes), self-report outcome, single site.
Replication notes
The paper establishes nothing about replication of its own result and we have not checked the literature for one. Recorded as `unclear` rather than `unreplicated`, which would assert a fact about the literature we cannot source.

Effects

OutcomeMetricValueMeasureTimingVsHorizonClass
Student engagement, total scale (post-test)group difference, self-report engagement scaleexperimental 38.87 (SD 1.80) vs control 26.13 (SD 2.17); reported t = 4.738, p = .003; reported overall Cohen d = 0.826 - but these means and SDs imply d ~ 6, so the paper's own numbers contradict each other and none is usableself-report surveypost-courseactive-alternativeend-of-treatmentnon-cognitive
Baseline equivalence of the two intact classes (pre-test)reported t-tests vs implied t-testspaper reports all pre-test t-values non-significant (0.321-0.632) and claims equivalence, but its own means/SDs imply t ~ 5-10 (e.g. Skills 21.07 (1.89) vs 24.25 (1.72)); groups were NOT equivalent at baseline on the reported descriptivesself-report surveypre-coursenonenot-applicablenon-cognitive
Learning or achievementnonenot measured - no achievement outcome was administered, despite the discussion asserting higher "learning achievement" for the experimental groupself-report surveynot-applicablenonenot-applicabledomain-skill