A complete Gtheory4LLM workflow for the reliability of mean continuous subjective-quality judgments.
All data below are synthetic. This tutorial uses no manuscript results, source-study data, API calls, or real LLM outputs. Its results demonstrate the software workflow; they do not establish an adequate evaluator or prompt count for a real task.
# This vignette is also installed as a runnable script:
# source(system.file("doc", "LLM-workflow.R", package = "Gtheory4LLM"))
library(Gtheory4LLM)The quantity is an item’s mean continuous quality score over evaluators, prompts, and repeated runs. We generate a complete panel of 24 items, four evaluators, three prompts, and two runs: 576 measurements. The model assumes exchangeable random effects for those populations and a Gaussian score model. The scores are simulated continuous numbers, not numeric codes assigned to ordered categories.
| Facet | Meaning in this example |
|---|---|
| Item | The object of measurement. Stable differences between items form the universe-score variance. |
| Evaluator and prompt | Both have main effects and direct item interactions. Different items can respond differently to these facets. |
| Run | Scoped to an evaluator/prompt pair. A run group labels all items. The run source is a common shift across items; remaining item-specific run variation belongs to the residual in this simulation. |
set.seed(20260912)
llm_data <- expand.grid(item = factor(seq_len(24)),
evaluator = factor(paste0("model_", seq_len(4))),
prompt = factor(paste0("prompt_", seq_len(3))),
run = factor(seq_len(2)))
llm_design <- gt_design("item", c("evaluator", "prompt", "run"),
crossed = c("evaluator", "prompt"),
nested = list(run = c("evaluator", "prompt")),
item_interactions = "additive", instrument_interactions = "additive")
llm_design$terms_requested
#> [1] "evaluator" "item" "prompt" "item:evaluator"
#> [5] "item:prompt" "evaluator:prompt:run"This declares six random sources: item, evaluator, prompt, item-by-evaluator, item-by-prompt, and evaluator-by-prompt-by-run. The run group is parent-scoped even though the coded panel is complete. Item-by-run and additional evaluator-by-prompt effects are not silently added. Omitting them is an assumption of this known simulation; a real study needs its own justification. The exact Gaussian engine currently requires a complete coded Cartesian panel with one row per full cell.
source_sd <- c(item = 1.2, evaluator = 0.4, prompt = 0.3,
"item:evaluator" = 0.5, "item:prompt" = 0.4,
"evaluator:prompt:run" = 0.25)
llm_data$quality <- 0
for (source in names(source_sd)) {
members <- strsplit(source, ":", fixed = TRUE)[[1L]]
group <- do.call(interaction, c(llm_data[members], list(drop = TRUE)))
effects <- rnorm(nlevels(group), sd = source_sd[[source]])
llm_data$quality <- llm_data$quality + effects[as.integer(group)]
}
llm_data$quality <- llm_data$quality + rnorm(nrow(llm_data), sd = 0.6)gt_preflight() reports observed counts, source
dimensions, free parameters, configured size limits, and the supported
reliability scale, without building a model matrix or running an
optimizer.
preflight <- gt_preflight(llm_data, "quality", llm_design)
print(preflight)
#> G-theory preflight: 576 rows | exact_balanced_gaussian
#> Observed levels: item=24, evaluator=4, prompt=3, run=2
#> Retained random sources: 6 | Random dimensions: 223 | Model parameters: 8
#> Structural/resource checks: PASS
#> Supported reliability scale: observed
#> Reliability of observed Gaussian-score averages.
#> Preflight checks structure and configured size limits. It does not establish identification, approximation adequacy, fit acceptance, or scientific validity.
preflight$sources
#> source observed_groups predictor_dimensions random_dimension covariance_parameters
#> 1 evaluator 4 1 4 1
#> 2 item 24 1 24 1
#> 3 prompt 3 1 3 1
#> 4 item:evaluator 96 1 96 1
#> 5 item:prompt 72 1 72 1
#> 6 evaluator:prompt:run 24 1 24 1
stopifnot(preflight$fitting_feasible)Passing this report means that the listed structural and resource checks pass. It does not prove identification, reliable estimation, numerical acceptance, or approximation adequacy. Read failed checks before changing a resource limit or simplifying a scientific model.
fit <- gt_fit(llm_data, "quality", llm_design,
control = gt_control(gaussian = list(optimizer = "CSOLNP",
extra_tries = 2L, retry_seed = 415L, threads = 1L)))
print(summary(fit))
#> G-theory fit summary: 1 outcome(s), 576 measurement rows
#> Outcomes: quality
#> Families: gaussian | Estimator: REML
#> Random sources: 6 | Instrumentation facets: 3
#> Optimizer: CSOLNP | Attempts: 1 | Selected attempt: 1
#> Optimizer completed: TRUE | Numerically accepted: TRUE
#> Likelihood approximation: exact_balanced_gaussian_likelihood
#> Boundary or nearly singular sources: none recorded
#> -2 log likelihood: 1411
#>
#> Source variances (Gaussian observed-score covariance):
#> source trait variance std_error at_boundary
#> evaluator quality 0.16915 0.15491 FALSE
#> item quality 1.95564 0.60959 FALSE
#> prompt quality 0.06368 0.07587 FALSE
#> item:evaluator quality 0.24121 0.05231 FALSE
#> item:prompt quality 0.10283 0.03178 FALSE
#> evaluator:prompt:run quality 0.04622 0.02085 FALSE
#> Residual quality 0.38950 0.02707 FALSE
#>
#> Fixed location or contrast estimates:
#> quality
#> -0.4553
#>
#> Notes:
#> - Nesting parents define instrumentation groups shared across objects; nested sibling interactions are not automatically included.
#> - No requested source was dropped for this observation family.
#>
#> Use gt_components(fit) for full covariance matrices and gt_diagnostics(fit) for details.
diagnostics <- gt_diagnostics(fit)
print(diagnostics[c("numerically_accepted", "selected_attempt",
"acceptance_failures", "approximation_adequacy")])
#> $numerically_accepted
#> [1] TRUE
#>
#> $selected_attempt
#> [1] "1"
#>
#> $acceptance_failures
#> NULL
#>
#> $approximation_adequacy
#> [1] "exact_balanced_gaussian_likelihood"
stopifnot(isTRUE(diagnostics$numerically_accepted))
gt_components(fit)
#> $evaluator
#> quality
#> quality 0.1691472
#>
#> $item
#> quality
#> quality 1.95564
#>
#> $prompt
#> quality
#> quality 0.06368045
#>
#> $`item:evaluator`
#> quality
#> quality 0.2412094
#>
#> $`item:prompt`
#> quality
#> quality 0.1028346
#>
#> $`evaluator:prompt:run`
#> quality
#> quality 0.04621689
#>
#> $Residual
#> quality
#> quality 0.3895031
#>
#> attr(,"scale")
#> [1] "Gaussian observed-score covariance"
#> attr(,"interpretation")
#> [1] "Cross-category contrasts within one categorical outcome are not separate measured outcomes."The Gaussian default estimator is REML. The trial budget allows the initial fit plus two additional attempts; it is not resampling. The chosen optimizer remains fixed across attempts. Inspect numerical acceptance and source-boundary diagnostics before interpreting the components. A completed optimizer alone is insufficient. If an installation does not support the selected optimizer or a fit is rejected, address that explicitly.
The summary table also carries an asymptotic standard error for each source variance, obtained from the numerically differentiated restricted-likelihood Hessian. A component resting on a variance boundary is flagged, because Wald theory does not apply there.
summary(fit)$variances
#> source trait variance std_error at_boundary
#> 1 evaluator quality 0.16914721 0.15491014 FALSE
#> 2 item quality 1.95564029 0.60958554 FALSE
#> 3 prompt quality 0.06368045 0.07587063 FALSE
#> 4 item:evaluator quality 0.24120936 0.05231458 FALSE
#> 5 item:prompt quality 0.10283459 0.03177643 FALSE
#> 6 evaluator:prompt:run quality 0.04621689 0.02084544 FALSE
#> 7 Residual quality 0.38950314 0.02707234 FALSEreliability <- gt_reliability(fit)
reliability
#> G-theory reliability | observed scale
#> Facet counts: evaluator=4, prompt=3, run=2
#> outcome Erho2 Erho2_se Erho2_lower Erho2_upper Phi Phi_se Phi_lower Phi_upper
#> quality 0.9464 0.01778 0.8988 0.9723 0.9173 0.03181 0.8299 0.9619
#> Erho2: relative comparisons; Phi: absolute decisions.
#> 95% intervals: delta method on the logit scale from the fitted parameter covariance.
#> Full covariance matrices and fit diagnostics remain in the returned object.
reliability$interpretation
#> [1] "Gaussian observed-score coefficient"Erho2 is the relative coefficient G: common evaluator,
prompt, and run shifts do not change relative item comparisons.
Phi is the absolute coefficient and also includes those
shifts as error. Item-dependent interactions contribute to both error
definitions. Here both coefficients refer to observed Gaussian-score
averages under the declared design. They measure consistency under that
model, not accuracy against human reference labels. A coefficient of
0.80 means that 80% of the variance in item means is universe-score
variance under this model; it is not 80% labeling accuracy.
The interval is a delta-method Wald interval on the logit scale, propagating the fitted parameter covariance matrix through the source covariances. It describes estimation uncertainty in those covariances under the declared model. It is not a guarantee about a future sample of evaluators or prompts.
grid <- expand.grid(evaluator = c(2L, 4L, 6L),
prompt = c(2L, 3L, 4L), run = c(1L, 2L))
dstudy <- gt_dstudy(fit, grid)
allocation <- dstudy$allocations
allocation$measurements_per_item <- dstudy$measurements_per_object
allocation$extrapolated <- dstudy$extrapolated
planning <- cbind(allocation,
dstudy$results[, c("outcome", "Erho2", "Erho2_lower", "Erho2_upper",
"Phi", "Phi_lower", "Phi_upper")])
planning <- planning[order(planning$measurements_per_item, -planning$Phi), ]
rownames(planning) <- NULL
print(planning)
#> evaluator prompt run measurements_per_item extrapolated outcome Erho2 Erho2_lower
#> 1 2 2 1 4 FALSE quality 0.8789244 0.7902498
#> 2 2 3 1 6 FALSE quality 0.8989630 0.8204331
#> 3 4 2 1 8 FALSE quality 0.9241948 0.8622785
#> 4 2 4 1 8 TRUE quality 0.9093288 0.8361284
#> 5 2 2 2 8 FALSE quality 0.8985872 0.8185730
#> 6 6 2 1 12 TRUE quality 0.9403393 0.8886455
#> 7 4 3 1 12 FALSE quality 0.9390021 0.8873654
#> 8 2 3 2 12 FALSE quality 0.9125791 0.8403014
#> 9 4 4 1 16 TRUE quality 0.9465851 0.9002366
#> 10 4 2 2 16 FALSE quality 0.9349508 0.8786388
#> 11 2 4 2 16 TRUE quality 0.9197397 0.8513494
#> 12 6 3 1 18 TRUE quality 0.9531530 0.9117091
#> 13 6 4 1 24 TRUE quality 0.9596917 0.9235008
#> 14 6 2 2 24 TRUE quality 0.9477350 0.8999582
#> 15 4 3 2 24 FALSE quality 0.9463767 0.8988071
#> 16 4 4 2 32 TRUE quality 0.9521950 0.9089930
#> 17 6 3 2 36 TRUE quality 0.9582059 0.9196465
#> 18 6 4 2 48 TRUE quality 0.9635285 0.9295932
#> Erho2_upper Phi Phi_lower Phi_upper
#> 1 0.9332760 0.8311242 0.6968108 0.9133367
#> 2 0.9454336 0.8543855 0.7259000 0.9285695
#> 3 0.9595798 0.8905661 0.7895704 0.9463806
#> 4 0.9517192 0.8665114 0.7404140 0.9366003
#> 5 0.9456557 0.8508181 0.7181184 0.9273660
#> 6 0.9688759 0.9123156 0.8237973 0.9586001
#> 7 0.9678246 0.9095813 0.8205098 0.9567797
#> 8 0.9539379 0.8681573 0.7404664 0.9382622
#> 9 0.9720689 0.9193968 0.8359359 0.9623144
#> 10 0.9661408 0.9017489 0.8028545 0.9538842
#> 11 0.9582099 0.8770947 0.7512927 0.9440057
#> 12 0.9756624 0.9295996 0.8559650 0.9670397
#> 13 0.9791477 0.9384895 0.8721195 0.9715377
#> 14 0.9733703 0.9201084 0.8330411 0.9637470
#> 15 0.9722742 0.9173273 0.8298542 0.9618948
#> 16 0.9754426 0.9253201 0.8430236 0.9662015
#> 17 0.9786904 0.9349788 0.8626345 0.9705244
#> 18 0.9814341 0.9425957 0.8772614 0.9741760Runs are counted within each evaluator/prompt pair, so the budget per item is evaluators x prompts x runs. A row using more levels than observed extrapolates the fitted variance model to the declared exchangeable population.
planning[planning$Phi >= 0.80, ]
#> evaluator prompt run measurements_per_item extrapolated outcome Erho2 Erho2_lower
#> 1 2 2 1 4 FALSE quality 0.8789244 0.7902498
#> 2 2 3 1 6 FALSE quality 0.8989630 0.8204331
#> 3 4 2 1 8 FALSE quality 0.9241948 0.8622785
#> 4 2 4 1 8 TRUE quality 0.9093288 0.8361284
#> 5 2 2 2 8 FALSE quality 0.8985872 0.8185730
#> 6 6 2 1 12 TRUE quality 0.9403393 0.8886455
#> 7 4 3 1 12 FALSE quality 0.9390021 0.8873654
#> 8 2 3 2 12 FALSE quality 0.9125791 0.8403014
#> 9 4 4 1 16 TRUE quality 0.9465851 0.9002366
#> 10 4 2 2 16 FALSE quality 0.9349508 0.8786388
#> 11 2 4 2 16 TRUE quality 0.9197397 0.8513494
#> 12 6 3 1 18 TRUE quality 0.9531530 0.9117091
#> 13 6 4 1 24 TRUE quality 0.9596917 0.9235008
#> 14 6 2 2 24 TRUE quality 0.9477350 0.8999582
#> 15 4 3 2 24 FALSE quality 0.9463767 0.8988071
#> 16 4 4 2 32 TRUE quality 0.9521950 0.9089930
#> 17 6 3 2 36 TRUE quality 0.9582059 0.9196465
#> 18 6 4 2 48 TRUE quality 0.9635285 0.9295932
#> Erho2_upper Phi Phi_lower Phi_upper
#> 1 0.9332760 0.8311242 0.6968108 0.9133367
#> 2 0.9454336 0.8543855 0.7259000 0.9285695
#> 3 0.9595798 0.8905661 0.7895704 0.9463806
#> 4 0.9517192 0.8665114 0.7404140 0.9366003
#> 5 0.9456557 0.8508181 0.7181184 0.9273660
#> 6 0.9688759 0.9123156 0.8237973 0.9586001
#> 7 0.9678246 0.9095813 0.8205098 0.9567797
#> 8 0.9539379 0.8681573 0.7404664 0.9382622
#> 9 0.9720689 0.9193968 0.8359359 0.9623144
#> 10 0.9661408 0.9017489 0.8028545 0.9538842
#> 11 0.9582099 0.8770947 0.7512927 0.9440057
#> 12 0.9756624 0.9295996 0.8559650 0.9670397
#> 13 0.9791477 0.9384895 0.8721195 0.9715377
#> 14 0.9733703 0.9201084 0.8330411 0.9637470
#> 15 0.9722742 0.9173273 0.8298542 0.9618948
#> 16 0.9754426 0.9253201 0.8430236 0.9662015
#> 17 0.9786904 0.9349788 0.8626345 0.9705244
#> 18 0.9814341 0.9425957 0.8772614 0.9741760A threshold of 0.80 is illustrative; choose and justify your own criterion. Read the interval columns alongside the point projection: several allocations here have overlapping intervals, so the ordering of nearby rows is not established by these data.
A facet is random when the study generalizes to a population of its levels, and fixed when the universe of generalization is exactly the levels used. LLM evaluation may target exactly a chosen temperature or prompt set, in which case those facets are fixed. Selecting levels deliberately does not by itself decide the universe of generalization.
Declare them with fixed. Following the mixed model of
Brennan (2001), the object-by-fixed-facet variance is averaged over that
facet’s levels and added to universe-score variance, and a source built
only from fixed facets shifts every item equally and leaves the
model.
fixed_prompt <- gt_reliability(fit, fixed = "prompt")
fixed_prompt
#> G-theory reliability | observed scale
#> Facet counts: evaluator=4, prompt=3, run=2
#> Fixed facets: prompt | mixed (Brennan 2001): object-by-fixed-facet variance enters the universe score
#> outcome Erho2 Erho2_se Erho2_lower Erho2_upper Phi Phi_se Phi_lower Phi_upper
#> quality 0.963 0.01261 0.9286 0.9811 0.9428 0.02464 0.8707 0.9758
#> Erho2: relative comparisons; Phi: absolute decisions.
#> 95% intervals: delta method on the logit scale from the fitted parameter covariance.
#> Full covariance matrices and fit diagnostics remain in the returned object.
fixed_prompt$source_roles
#> evaluator item
#> "absolute_error" "universe"
#> prompt item:evaluator
#> "dropped_fixed_instrumentation_constant" "relative_and_absolute_error"
#> item:prompt evaluator:prompt:run
#> "universe_after_fixed_facet_averaging" "absolute_error"
#> Residual
#> "relative_and_absolute_error"Treating the prompt set as fixed raises both coefficients, because item-by-prompt variation is now part of what the study is trying to measure rather than error. A fixed facet’s count cannot be changed, and a decision study may not project over it.
fixed acts at the reliability stage: it changes how the
already-fitted components are aggregated into a coefficient. It does not
change the G study, so it cannot repair a source whose variance was
misspecified during fitting. Observed seed disagreement can change
across temperatures even when latent variance is constant, because
category probabilities can change. If diagnostics raise doubts about
pooling temperatures, fitting within one temperature is a useful
sensitivity analysis for a conditional estimand; declaring a facet fixed
is a separate decision about generalization.
| Outcome | Current scalar reliability support |
|---|---|
| Gaussian | Observed continuous-score averages. |
| Binary or ordinal | Explicit scale = "latent" after an accepted fit. This
describes latent-response averages, not majority votes, label
proportions, or observed ordinal-score averages. |
| Unordered categorical | No implemented default scalar G/Phi. A scientific score or category-probability estimand is needed. |
Declare category order and references explicitly. Do not recode
binary, ordinal, or nominal labels as the continuous quality score used
above. Discrete fitting uses a bounded dense Laplace engine with
Gaussian latent random effects; gt_preflight() reports its
current limits, and changing the link does not remove the random-effects
distribution assumption. That engine reports point estimates only: it
computes no standard errors. Analytic reliability and D studies require
complete balanced coded panels even when discrete fitting admits missing
whole cells.
At one observation per object-by-facets cell, a discrete fit cannot identify the full-cell source that the default design requests. Declare the same design without it:
Installed help: ?gt_design, ?gt_family,
?gt_preflight, ?gt_control,
?gt_fit, ?gt_diagnostics,
?gt_reliability, and ?gt_dstudy.
Brennan, R. L. (2001). Generalizability Theory. Springer. doi:10.1007/978-1-4757-3456-0