---
title: "Glossary"
output: rmarkdown::html_vignette
vignette: >
%\VignetteIndexEntry{Glossary}
%\VignetteEngine{knitr::rmarkdown}
%\VignetteEncoding{UTF-8}
---
```{r, include = FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>")
```
A plain-language reference for the terms that recur across these articles. Each
entry defines the idea once, and the other articles link here rather than
re-explaining it. Terms are listed alphabetically. Nothing here is new: it is
the vocabulary the [*Getting started*](getting-started.html) and
[*Choosing an ICC*](choosing-an-icc.html) guides use, gathered in one place.
## Absolute agreement
One of the two `type`s of ICC. Absolute agreement asks whether raters give the
*same* score. Two raters who rank subjects identically but sit a full point apart
do **not** agree in this sense. It counts systematic rater differences (the rater
variance) as error. Contrast **consistency**. In generalizability theory the
absolute-agreement ICC is the **dependability coefficient**.
## Average-unit ICC: `ICC(*,k)`
The reliability of the **mean** of `k` raters, rather than of one rater. Averaging
cancels part of the error, so `ICC(*,k)` is always at least as high as the
single-rater `ICC(*,1)`. The `k` is the number of raters whose average you will
actually use. On incomplete data it becomes the **effective number of ratings**.
See **single-unit ICC** for its counterpart.
## Burch interval
An opt-in closed-form confidence interval for the balanced one-way design
(`ci_method = "burch"`). It gives REML-based limits with a kurtosis adjustment
(Burch 2011), so its width tracks the tail weight of the data, where the
**exact-F interval** leans on normality instead. Tracking tail weight is not the
same as widening: on the two grids this package has measured that vary only the
subject effect it came out the narrower of the two in nearly every cell. How much
narrower is conditional, and not in the direction one might guess. `"burch"`'s
width margin holds much the same up to a true ICC of 0.3 rather than shrinking as
the true ICC rises (on the larger grid; the smaller grid's margin does shrink
across its levels). The `"burch"` width margin then collapses to near parity at a
true ICC of 0.6, on the one grid reaching that value, where every cell favouring
the exact-F interval sits. It shrinks steadily as the subject count grows,
measured at 5 raters.
Burch reports the reverse for symmetric heavy-tailed data with non-normal errors,
and a third grid now measures that case.
What `"burch"` does against `"searle"` depends on what the residual is
drawn from. The three grids now measure that:
the two grids that vary only the subject effect put it narrower
nearly everywhere, while the third, which draws the residual from
the same family as the subject effect, puts it wider at every
symmetric heavy-tailed family measured (a median width ratio of
1.2963 at t(5) with 100 subjects) and narrower at every lighter-tailed one,
the normal included. So neither interval is reliably the tighter one.
That adjustment is not a remedy for heavy tails: on strongly
skewed subject effects `"burch"` under-covers about as badly as the default. See
[*When the default under-covers*](interval-methods.html#when-the-default-under-covers).
Deterministic: no resamples, no seed. See
[*Confidence-interval methods*](interval-methods.html#the-opt-in-boundary-robust-methods).
## Confidence interval vs. credible interval
A **confidence interval** (the frequentist engines' output) is a range built so
that, across many hypothetical repetitions of the study, 95% of such ranges would
contain the true ICC. A **credible interval** (the Bayesian `brms` engine's output)
is a range that holds 95% of the posterior probability. You can say directly "there
is a 95% chance the ICC lies in here, given the data and the prior." They answer
subtly different questions. See [*Confidence-interval
methods*](interval-methods.html).
## Conflated ICC
The single-level ICC you would get by **ignoring** a clustering structure
(pupils in classrooms, patients in clinics). This quantity is ten Hove et al.'s
(2022) Equation 14. It folds the between-cluster and within-cluster variation
into one "true score" and is biased for both the subject-level and cluster-level
questions. `icc()` can report it (`level = "conflated"`) purely as a
**diagnostic contrast**, to show the cost of ignoring the structure. It is never
a number to report. See [*Multilevel
designs*](multilevel-designs.html#how-much-does-ignoring-the-nesting-cost-the-conflated-icc).
## Connectedness (identification)
A design is **connected** when the raters and subjects are linked tightly enough
that the model can separate a subject effect from a rater effect. With enough missing
cells a design can split into disconnected islands, and then the variance components
are not **identified**, and no method can estimate them. `icc()` checks this and aborts
loudly rather than return a number that isn't estimable.
## Consistency
The other `type` of ICC. Consistency asks whether raters *rank* subjects the same
way, forgiving a constant offset between raters. Two raters who agree on the
ordering but differ by a fixed point still count as perfectly consistent. It leaves
the rater main effect out of the error term. Contrast **absolute agreement**.
## Credible interval
See **confidence interval vs. credible interval**.
## Dependability coefficient
The generalizability-theory name for the **absolute-agreement** ICC, written
\eqn{\Phi}. Projecting it to a different number of raters (a **D-study**) is a change
of the averaging divisor. See [*D-studies and within-cell
replicates*](d-studies-and-replicates.html).
## D-study (decision study)
A forward-looking projection: given the variance components you already estimated,
*how reliable would the mean of some other number of raters (or occasions) be?* It
reuses the existing fit, with no refitting, and answers "how many raters do I need?".
See [*D-studies and within-cell replicates*](d-studies-and-replicates.html).
## Effective number of ratings: `k_eff`
On **incomplete** data, subjects are rated different numbers of times, so there is no
single "`k`" to average over. `k_eff` is the **harmonic mean** of the per-subject
rating counts, and it is the divisor `icc()` uses for `ICC(*,k)` on ragged data. It
is always at or below the full panel size, and the report names it so the divisor is
never a black box.
## Engine
The computational backend `icc()` uses to estimate the variance components, chosen
with the `engine` argument: **glmmTMB** (the default mixed model), **lme4** (an
alternate mixed-model solver), **lavaan** (a structural-equation formulation), or
**brms** (a Bayesian fit). Some engines are just a different solver for the same
estimator. Others compute a genuinely different, though asymptotically equivalent,
estimator. See [*Estimation engines*](engines.html).
## Estimand
The true quantity you are trying to estimate: the target the ICC is aiming at.
The word matters because "the ICC" is not one number. Agreement and consistency,
single and average, subject level and cluster level are *different estimands*, and
picking the coefficient is really picking which one answers your question. See
[*Choosing an ICC*](choosing-an-icc.html).
## Exact-F interval
An opt-in closed-form confidence interval for the balanced one-way design
(`ci_method = "searle"`): it inverts the exact-F pivot of the one-way ANOVA
(Searle 1971, Ch. 9 Table 9.14; the McGraw & Wong 1996 Table 7 limits). Exact
under normality. It is best-calibrated when the data are approximately normal,
and it is deterministic. Exact is not the same as shortest: see the **Burch interval**
above for what this package measured about their relative widths. See
[*Confidence-interval methods*](interval-methods.html#the-opt-in-boundary-robust-methods).
## FIML
**Full-information maximum likelihood**: the technique the **lavaan** (SEM) engine
uses to fit **incomplete** data. Rather than dropping cases with missing cells, it
uses every observed value to estimate the model. See
[*Estimation engines*](engines.html#a-structural-equation-engine-lavaan).
## Finite-population rater variance: `θ²_r`
Raters may be treated as **fixed**, meaning the observed raters *are* the whole
population of interest. Then the "rater variance" is the spread of just those
raters, computed as a bias-corrected finite-population quantity (McGraw & Wong's
Case 3A) rather than an estimate of a wider rater universe. On balanced data it
equals the random-rater variance. Under imbalance it differs. See **fixed vs.
random raters**.
## Fixed vs. random raters
The `raters` argument. **Random** raters (the recommended default) treat the raters
you used as a sample from a larger pool, so the reliability generalizes to *new*
raters drawn from that pool. **Fixed** raters treat the observed raters as the entire
population of interest, so the reliability speaks only to *these* raters. The choice
changes the rater term from a random-sample variance to the **finite-population rater
variance**.
## Harmonic mean
An average that leans toward the smaller values: the reciprocal of the mean of the
reciprocals. It is the right average for the **effective number of ratings** because
reliability depends on the *rate* of information per subject, which the harmonic mean
captures.
## Indicator-mean estimator
How the **lavaan** (SEM) engine recovers the rater variance for **absolute
agreement**. In lavaan's formulation a rater is a single column with no random
effect, so its variance is read from the spread of the estimated column
(indicator) means (Jorgensen 2021). It is a genuinely different estimator than
the mixed model's random effect, though asymptotically equivalent to it, and can
differ modestly on small designs. See [*Estimation
engines*](engines.html#a-structural-equation-engine-lavaan).
## Modified profile likelihood
An opt-in deterministic confidence interval for the balanced, complete two-way
random absolute-agreement design (`ci_method = "mpl"`; Xiao & Liu 2013). It
profiles the likelihood in the ICC with a calibrated small-sample correction, and
returns an interval at the near-zero boundary where the Monte-Carlo default
aborts. Available at `conf_level` 0.90, 0.95, and 0.99, and deliberately
conservative. See
[*Confidence-interval methods*](interval-methods.html#the-opt-in-boundary-robust-methods).
## Monte-Carlo interval
The default confidence-interval method (`ci_method = "montecarlo"`). It draws many
parameter vectors from the fitted model's estimated covariance, on a scale that
respects the zero-variance boundary, recomputes the ICC for each draw, and takes the
2.5% and 97.5% quantiles. Fast and boundary-aware. It does assume the fitted
parameters are approximately normally distributed around the truth, and it
under-covers when the subject effects are strongly skewed or heavy-tailed. See
[*When the default
under-covers*](interval-methods.html#when-the-default-under-covers). See also
[*Confidence-interval methods*](interval-methods.html).
## Occasion (within-cell replicate)
One of several ratings the *same* rater gives the *same* subject. A design with
occasions lets `icc()` separate the [subject-by-rater
interaction](#variance-component) from pure error. `glance()` reports the
per-cell count in its `n_o` column. The printed report spells that same count
on its design line, as `N cells x N replicates`. Where `n_o` is `NA` because the
cells hold unequal counts, that slot reads `NA` too. Only a fit that splits
replicates carries the slot at all. A multilevel fit's design line reports
subjects and clusters instead, and a fit that splits none reports observations
or ratings, so read `n_o` for those. See [*D-studies and within-cell
replicates*](d-studies-and-replicates.html).
## `occasions` vs. `n_o`
Two near-identical names for different quantities. On an `icc()` fit, `tidy()`'s
`occasions` column is the **per-rater occasion divisor** that row's coefficient
applies to pure error. It reads 1 on every row that averages no occasions, which
includes every row whose error set carries no pure-error term to average. It
reads the fitted per-cell count where the row does average, and `NA` on a fit
that splits no within-cell replicates. On a `d_study()` projection the same
column reports the count each row is *projected* at instead, which `?d_study`
states. `glance()`'s `n_o` counts the **occasions observed per cell** in
the design that was fitted, and is `NA` under the condition `?icc` states for
it. So on a design with three ratings per cell, a fit reporting both settings
shows `occasions` 1 and 3 down its rows, while `n_o` is 3 for the fit as a
whole. Note that `occasions` is not a count of the ratings a coefficient
averages: an occasion-averaged `ICC(A,k)` over four raters at `occasions` 3 is
the reliability of a mean of twelve ratings. A ragged replicate fit is the case
that separates the two columns most sharply: `n_o` reads `NA` there while
`occasions` still reads 1.
## One-way vs. two-way
The `model` argument. A **two-way** design has every subject rated by the *same*
raters, so a rater main effect can be estimated and either counted as error
(agreement) or set aside (consistency). A **one-way** design has each subject rated
by possibly *different* raters, so rater identity is not modeled and only an
agreement-style `ICC(1)` / `ICC(k)` is defined.
## Parametric bootstrap
An alternative confidence-interval method (`ci_method = "bootstrap"`): simulate new
response vectors from the fitted model, refit each one, and take percentile quantiles
of the recomputed ICCs. It does not rely on the asymptotic-normal approximation the
Monte-Carlo method uses, at the cost of a full refit per resample. See
[*Confidence-interval methods*](interval-methods.html).
## Posterior mode (MAP)
The point estimate the Bayesian **brms** engine reports: the peak (mode) of the
posterior distribution of the ICC, the *maximum a posteriori* value. On a small,
right-skewed posterior it can sit below the mixed-model REML estimate. Its interval is
a **credible** interval. See [*Estimation
engines*](engines.html#a-bayesian-engine-brms).
## Prior
In the Bayesian **brms** engine, the distribution placed on each variance component
*before* seeing the data. Here that is a weakly-informative half-*t*(4, 0, 1) on
every standard deviation (ten Hove et al. 2020), the sourced prior every coverage
result depends on. Overriding it (`prior =`) voids those guarantees, so `icc()`
warns. See [*Estimation engines*](engines.html#the-prior-and-overriding-it).
## REML
**Restricted maximum likelihood**: the standard method the mixed-model engines
(glmmTMB, lme4) use to estimate variance
components. It corrects the downward bias that ordinary maximum likelihood has when
estimating variances, which matters for the small samples common in
reliability studies.
## Single-unit ICC: `ICC(*,1)`
The reliability of a **single** rater's score. Its counterpart, the **average-unit
ICC** `ICC(*,k)`, is the reliability of the mean of `k` raters and is always at least
as high.
## Subject level vs. cluster level
In a **multilevel** design (subjects nested in clusters, such as pupils in
classrooms), two reliabilities are defined. The **subject level** asks how reliably
raters distinguish subjects *within* a cluster. The **cluster level** asks how
reliably they distinguish *cluster means*. `icc()` reports both from one fit.
See [*Multilevel designs*](multilevel-designs.html#subject-level-vs.-cluster-level).
## Transformed bootstrap-*t*
An opt-in confidence-interval method for the one-way design, balanced or
unbalanced (`ci_method = "npbootstrap"`; Ukoumunne et al. 2003). It resamples
whole subjects with replacement, stabilizes the variance with a log-F transform,
studentizes, and back-transforms the endpoints, so it takes a `seed` and
`boot_samples`. Robust at the zero-variance boundary and to non-normal subject
effects. See
[*Confidence-interval methods*](interval-methods.html#the-opt-in-boundary-robust-methods).
## Variance component
A share of the total variation in the scores traced to one source. It says how
much comes from real differences between **subjects**, from some **raters**
scoring higher than others, from the **subject-by-rater** interaction, and from
residual **error**. Every ICC is a ratio built from these components: signal
variance over signal-plus-error variance.
## Zero-variance boundary
A variance component cannot be negative, so its estimate can land *exactly* at
zero, the edge of the allowed range. Ordinary interval formulas misbehave there
(they can run below zero or collapse). The Monte-Carlo default is
**boundary-aware**: it works on a scale where zero is reachable without breaking,
which is one reason glmmTMB is the recommended engine. See
[*Confidence-interval methods*](interval-methods.html).
## References
Burch, B. D. (2011). Assessing the performance of normal-based and REML-based
confidence intervals for the intraclass correlation coefficient. *Computational
Statistics and Data Analysis, 55*, 1018--1028.
Jorgensen, T. D. (2021). How to estimate absolute-error components in structural
equation models of generalizability theory. *Psych, 3*(2), 113--133.
McGraw, K. O., & Wong, S. P. (1996). Forming inferences about some intraclass
correlation coefficients. *Psychological Methods, 1*(1), 30--46.
Searle, S. R. (1971). *Linear Models*. Wiley.
ten Hove, D., Jorgensen, T. D., & van der Ark, L. A. (2020). On the usefulness of
interrater reliability coefficients. In M. Wiberg et al. (Eds.), *Quantitative
Psychology* (pp. 67--75). Springer.
ten Hove, D., Jorgensen, T. D., & van der Ark, L. A. (2022). Interrater reliability
for multilevel data: A generalizability theory approach. *Psychological Methods,
27*(4), 650--666.
Ukoumunne, O. C., Davison, A. C., Gulliford, M. C., & Chinn, S. (2003).
Non-parametric bootstrap confidence intervals for the intraclass correlation
coefficient. *Statistics in Medicine, 22*(24), 3805--3821.
Xiao, Y., & Liu, H. (2013). Modified profile likelihood approach for certain
intraclass correlation coefficients. *Computational Statistics, 28*(5),
2241--2265.