On three published meta-analysis datasets, drmTMB reproduces the estimates of metafor and glmmTMB — and fits a random-effect-in-dispersion model neither can, recovered from known truth. Evidence, not assertion.
▲ The bar (Williams et al., arXiv:2604.04084): variance components to 5 dp, pooled effect to 6 dp. Every comparator slice clears it on the heterogeneity parameters; the pooled effect lands at 4–5 dp.
The headline. metafor rma.mv(~1 | study/esid), glmmTMB
equalto, and drmTMB meta_V fit the identical model on 100 effects from
17 studies. This is the design Williams et al. use to validate equalto — with drmTMB
added as a third column.
| Quantity | metafor | glmmTMB | drmTMB | Δ max |
|---|---|---|---|---|
| μ pooled effect | 0.426799 | 0.426799 | 0.426796 | 3.3e-06 |
| τ² between-study (L3) | 0.18787022 | 0.18787025 | 0.18787025 | 3.0e-08 |
| σ² within-study (L2) | 0.11199209 | 0.11199198 | 0.11199200 | 1.0e-07 |
| SE(μ) | 0.118429 | 0.118475 | 0.118475 | 4.6e-05 |
MATCH Variance components agree to < 2×10⁻⁷, pooled effect to < 5×10⁻⁶. drmTMB and glmmTMB agree on SE (both TMB-based); metafor's differs at 4×10⁻⁵ by construction.
metafor rma.mv (unstructured) vs drmTMB meta_vcov_bivariate on
Berkey et al. (1998): 5 periodontal trials, two correlated outcomes (PD, AL).
| Quantity | metafor | drmTMB | Δ |
|---|---|---|---|
| μ PD | 0.353428 | 0.353445 | 1.7e-05 |
| μ AL | -0.339215 | -0.339235 | 2.0e-05 |
| τ PD (between-study SD) | 0.10831908 | 0.10831911 | 2.9e-08 |
| τ AL (between-study SD) | 0.18069676 | 0.18069680 | 3.9e-08 |
| ρ between-study | 0.60879858 | 0.60879891 | 3.4e-07 |
MATCH Between-study SDs and correlation agree to < 10⁻⁶; pooled means to < 3×10⁻⁵ (only 5 studies). No bivariate metafor comparator existed in the package before this dossier.
Modelling the log-heterogeneity by school grade on 48 writing-to-learn studies. A genuine
three-way: metafor rma(scale=) models log(τ²); glmmTMB dispformula and
drmTMB sigma~x model log(σ). They coincide because log(τ²) = 2·log(σ) — shown on the
metafor scale below.
| log(τ²) coefficient | metafor | glmmTMB | drmTMB | Δ max |
|---|---|---|---|---|
| μ location intercept | 0.228206 | 0.228206 | 0.228178 | 2.7e-05 |
| intercept | -3.3443916 | -3.3443919 | -3.3443919 | 2.8e-07 |
| grade 2 | 1.2069189 | 1.2069205 | 1.2069192 | 1.6e-06 |
| grade 3 | 1.00277608 | 1.00277556 | 1.00277643 | 5.3e-07 |
| grade 4 | 0.08328491 | 0.08328553 | 0.08328547 | 6.2e-07 |
MATCH All coefficients agree to < 3×10⁻⁵ after the log(τ²) = 2·log(σ) reconciliation.
sigma ~ 1 + (1 | study): the heterogeneity magnitude is itself a
study-level random effect. metafor's scale= is fixed-effect only; glmmTMB's
dispformula admits no random effects. Ground truth here is simulation, and the argument
is recovery — not "trust us".
| Parameter | Truth | Mean estimate | Bias | Monte-Carlo SE |
|---|---|---|---|---|
| μ | 0.300 | 0.3042 | +0.0042 | 0.0038 |
| α₀ (log-heterogeneity intercept) | -1.386 | -1.3643 | +0.0220 | 0.0146 |
| SD of the dispersion RE | 0.500 | 0.4797 | -0.0203 | 0.0161 |
RECOVERED No detectable bias over 30 replicates (every |bias| < 3·MCSE). A single-scenario demonstration — the calibrated coverage grid is commissioned to Totoro, not claimed here.
100 replicates of a Normal–Normal meta-analysis; Wald 95% interval coverage vs nominal. This proves the simulation harness runs and coverage lands near target — the calibrated grid (4 effect measures, wide DGP range, ≥2000 reps) runs on Totoro.
SMOKE Intercept and σ land on 0.95; the slope sits ~2.7 SE low — consistent with Monte-Carlo noise across three parameters or mild Wald under-coverage. The Totoro grid resolves which. The tick marks 0.95.
A comparator on three datasets is the demonstration; this is the calibration. Four effect-measure regimes (SMD, lnRR, logOR, logIRR) crossed with study count, heterogeneity and sampling scale — 57,600 Wald 95% intervals over 48 DGP cells, each from a known truth. The question: does the interval cover at its nominal rate?
| Measure regime | μ intercept | μ slope | σ heterogeneity |
|---|---|---|---|
| SMD | 0.931 | 0.933 | 0.916 |
| lnRR | 0.931 | 0.935 | 0.906 |
| logOR | 0.929 | 0.931 | 0.934 |
| logIRR | 0.931 | 0.931 | 0.924 |
| Parameter | n = 20 | n = 40 | n = 80 | trend |
|---|---|---|---|---|
| μ intercept | 0.917 | 0.937 | 0.939 | → nominal |
| μ slope | 0.921 | 0.935 | 0.941 | → nominal |
| σ heterogeneity | 0.901 | 0.921 | 0.937 | → nominal |
HONEST Coverage runs 0.91–0.94, a few points under nominal — not a defect. This is the well-known conservatism of Wald intervals for the heterogeneity variance at small study counts (it is why metafor offers Q-profile CIs for τ²). It is uniform across all four effect measures and converges to 0.95 as studies accumulate — correct asymptotic behaviour. drmTMB's profile intervals close the finite-sample gap.
▲ This is the local broad run (400 reps/cell). The publication-grade campaign — 768 cells × 2,000 reps, adding dense known-V and the type-I lane — runs on Totoro (inst/trust-dossier/totoro/), never GitHub Actions.
Rscript run.R, seed + SHA pinned