Docosahexaenoic acid for reading, working memory and behavior in UK children aged 7-9: A randomized controlled trial for replication (the DOLAB II study)
Montgomery P, Spreckelsen TF, Burton A, Burton JR, Richardson AJ · 2018
grade Breplicationdeveloper-involvedfailednumbers spot-checked
Sample
376 children randomized from 1,230 invited (618 consented and screened), across 84 mainstream primary schools and academies in five UK local/unitary authorities (Oxfordshire, Northamptonshire, Buckinghamshire, Milton Keynes, Swindon); recruitment January 2013 to March 2015, follow-up to July 2015; 372 of 376 reassessed at 16 weeks. Planned n was 400, so the trial fell 24 short.
Population
Healthy UK children aged 7-9 underperforming in reading (below the 20th centile on the recalibrated New BAS II or on BAS 3) - deliberately the exact subgroup in which DOLAB I found its effect. Mean baseline word reading standard score 79.6 (sd 6.5), about 1.3 sd below the norm, roughly 27 months behind chronological age; 62.5% male.
Design
Direct, pre-registered, adequately powered replication of DOLAB I by the ORIGINAL AUTHORS, with a published protocol (ISRCTN48803273; protocols.io), CONSORT-compliant reporting, and anonymised data plus Stata syntax posted on the Open Science Framework (osf.io/9ynjf). Same dose (600 mg/day algal DHA), same duration (16 weeks), same taste- and colour-matched corn/soybean placebo. Randomization by Sealed Envelope Ltd with minimization on school and sex plus a 30% random element; blinding verified post hoc and intact. It recruited MORE children (376 v 362) in MORE schools (84 v 74) and targeted the responsive subgroup directly. Powered at 80% for the d = 0.28 observed in DOLAB I subgroup. This is the strongest design in the omega-3-and-learning literature and it is what a replication should look like. FOUR DEVIATIONS the record must carry, because the authors themselves offer them as reasons the studies disagree. (1) The reading instrument was NOT identical: the 2011 phonics reforms had shifted decoding ability, so DOLAB II used a RECALIBRATED New BAS II plus BAS 3 rather than DOLAB I uncalibrated BAS II, and the authors speculate the recalibrated version may be less sensitive to change. (2) Teacher-rated behaviour was demoted from primary (DOLAB I) to SECONDARY outcome. (3) The pre-planned subgroup here was the sub-10th-centile band (n = 213); the sub-20th-centile band was the whole sample. (4) Capsule shells (colorant and gelatine) were changed in January 2014 mid-trial under a protocol amendment. Two further weaknesses: behaviour questionnaires were the only measures with >15% missing data and lost about half the sample (parent change scores n = 187, teacher n = 196, imputed with treatment-group medians); and post-intervention DHA reached 2.9% of blood fatty acids versus 3.8% in DOLAB I, i.e. lower uptake - though the active group did move from 1.6% to 2.9% while placebo did not (p < 0.001), so the intervention demonstrably reached the bloodstream and the null is not an adherence artefact. Industry ties: funded by DSM Nutritional Products, which also supplied product and placebo and is stated to have had no role in design, data collection, analysis, decision to publish or manuscript preparation ("conducted without direct influence of its funder by way of a robust contract"); PM and AJR declare occasional paid consultancy for companies producing or promoting omega-3. Industry-funded and industry-supplied, but independently executed and analysed - hence developer-involved, not developer-led.
Key findings
The replication FAILED. Reading, working memory and behaviour change scores showed no consistent differences between the DHA and placebo groups; primary-outcome reading change was Active +0.64 (sd 3.7) v Placebo +0.83 (sd 3.6), p = 0.616, and in the pre-planned sub-10th-centile subgroup +1.4 (3.6) v +1.4 (3.7), p = 0.938. The authors report the observed effect size on reading as d = 0.05, implying achieved power of 8% and a required sample of more than 11,500 to detect it. The direction of the residual noise is worth recording because our earlier extraction did not: EVERY statistically significant change-score difference on the behaviour scales favoured PLACEBO, not DHA (parent-rated Anxiety -1.0 v -3.8, p = 0.002; parent-rated Global Emotional Lability -0.9 v -3.2, p = 0.019; teacher-rated Anxiety -0.5 v -3.7, p = 0.005). Conversely, the one signal favouring active was on post-intervention working-memory LEVELS, not change (Recall of Digits Backward difference 1.774 points, 95% CI 0.045 to 3.503) - but Digits Forward was already significantly higher in the active arm AT BASELINE (42.9 v 41.1, p = 0.048), so the post-intervention gap is a baseline-imbalance artefact and the change scores are flat. Under the replication rule this caps any omega-3-for-learning verdict at mixed at best, and given that the target population was the very subgroup DOLAB I identified, the honest reading is that the original subgroup finding was noise.
Genetic confound
None - randomized, double-blind, placebo-controlled.
Replication notes
This IS the replication - of Richardson et al. 2012 (DOLAB I, nutr-richardson-2012-dolab-dha-rct) - and it failed on every primary outcome. Two honest caveats travel with that. (1) It is a SELF-replication: same senior authors, same funder line (Martek, later DSM Nutritional Products), same research group. Under the Replication rule that is not an independent replication, so it does not license a stronger verdict than mixed on its own. (2) Cutting the other way, a self-replication has every incentive - reputational and financial - to succeed, and this one was pre-registered with a published protocol, was larger than the original, targeted the original subgroup directly, verified biological uptake by fingerstick blood DHA, and still found nothing. The authors state plainly in print that the trial did not replicate DOLAB I. No independent group has run the same trial.
Effects
| Outcome | Metric | Value | Measure | Timing | Vs | Horizon | Class |
|---|---|---|---|---|---|---|---|
| Reading change, whole sample below 20th centile - the primary outcome (domain skills) | BAS II age-standardized word reading, change score, n = 376 | Active +0.64 (sd 3.7) v Placebo +0.83 (sd 3.6), p = 0.616 - NULL. Authors report observed d = 0.05, i.e. 8% achieved power; post-intervention group mean difference 0.594 points (95% CI -1.937 to 0.749, placebo minus active) | standardized | 16 weeks | business-as-usual | end-of-treatment | domain-skill |
| Reading change, pre-planned subgroup at or below 10th centile (domain skills) | BAS II age-standardized word reading, change score, n = 213 | Active +1.4 (sd 3.6) v Placebo +1.4 (sd 3.7), p = 0.938 - NULL. Post-intervention mean difference 0.576 points (95% CI -2.019 to 0.867) | standardized | 16 weeks | business-as-usual | end-of-treatment | domain-skill |
| Reading AGE change (domain skills) | BAS II reading age in months, change score | whole sample Active +3.1 (4.4) v Placebo +3.7 (4.9), p = 0.143; sub-10th-centile +2.7 (3.5) v +3.5 (4.3), p = 0.179 - both NULL and both numerically favouring placebo | standardized | 16 weeks | business-as-usual | end-of-treatment | domain-skill |
| Working memory change scores (near transfer) | BAS II Recall of Digits Forward and Backward, T-score change | Forward +0.95 (7.4) v +0.91 (7.7), p = 0.957; Backward +0.4 (9.3) v -0.4 (9.8), p = 0.356 - both NULL | standardized | 16 weeks | business-as-usual | end-of-treatment | near-transfer |
| Working memory post-intervention LEVELS - the one signal favouring DHA, and it is confounded by baseline imbalance (near transfer) | BAS II Recall of Digits, post-intervention group mean difference (placebo minus active) | Digits Backward -1.774 (95% CI -3.503 to -0.045), sub-10th-centile -3.061 (-5.597 to -0.526); Digits Forward -1.797 (-3.665 to 0.071). Favours active and excludes zero for Backward - BUT Digits Forward already favoured active at baseline (42.9 v 41.1, p = 0.048), change scores are null, and the authors note none approaches the 10-point (1 sd) threshold for clinical relevance | standardized | 16 weeks | business-as-usual | end-of-treatment | near-transfer |
| Parent-rated behaviour, the primary behaviour outcome (non-cognitive) | Conners Parent Rating Scale long form, T-score change, 14 scales, ITT | no consistent difference; the only significant contrasts favour PLACEBO - Anxiety -1.0 (7.9) v -3.8 (9.6), p = 0.002 (still significant per protocol, p = 0.007); Global Emotional Lability -0.9 (9.6) v -3.2 (9.8), p = 0.019. Roughly half of parent questionnaires were missing and imputed | standardized | 16 weeks | business-as-usual | end-of-treatment | non-cognitive |
| Teacher-rated behaviour, a secondary outcome here (non-cognitive) | Conners Teacher Rating Scale long form, T-score change, 14 scales, ITT | no consistent difference; the only significant contrast again favours PLACEBO - Anxiety -0.5 (10.8) v -3.7 (11.0), p = 0.005; nothing significant in the per-protocol analysis | standardized | 16 weeks | business-as-usual | end-of-treatment | non-cognitive |
| School absence for illness (health) | school-recorded half-day absences over the 16-week intervention | Active 4.9 (sd 5.3) v Placebo 5.4 (sd 6.2), p = 0.63 - null; no group differences on the Barkley side-effects scale either | standardized | 16 weeks | business-as-usual | end-of-treatment | health |
| Blood DHA uptake - the mechanism check that rules out non-adherence (health) | fingerstick blood DHA as % of total fatty acids | active 1.6% to 2.9%, placebo unchanged (p < 0.001); post-intervention 2.9% v 1.5%, p < 0.001. The supplement reached the bloodstream, so the null is not an adherence artefact - though uptake was lower than DOLAB I 3.8%, and within-trial DHA change bore no relation to outcome change | standardized | 16 weeks | business-as-usual | end-of-treatment | health |
Cited by
- Breakfast, school meals, and micronutrients — what feeding children actually buysmixedconf: mediumgc: low