How many evaluators, prompts, and runs?

A complete Gtheory4LLM workflow for the reliability of mean continuous subjective-quality judgments.

All data below are synthetic. This tutorial uses no manuscript results, source-study data, API calls, or real LLM outputs. Its results demonstrate the software workflow; they do not establish an adequate evaluator or prompt count for a real task.

# This vignette is also installed as a runnable script:
#   source(system.file("doc", "LLM-workflow.R", package = "Gtheory4LLM"))
library(Gtheory4LLM)

1. Define the quantity and the sampling design

The quantity is an item’s mean continuous quality score over evaluators, prompts, and repeated runs. We generate a complete panel of 24 items, four evaluators, three prompts, and two runs: 576 measurements. The model assumes exchangeable random effects for those populations and a Gaussian score model. The scores are simulated continuous numbers, not numeric codes assigned to ordered categories.

Facet Meaning in this example
Item The object of measurement. Stable differences between items form the universe-score variance.
Evaluator and prompt Both have main effects and direct item interactions. Different items can respond differently to these facets.
Run Scoped to an evaluator/prompt pair. A run group labels all items. The run source is a common shift across items; remaining item-specific run variation belongs to the residual in this simulation.
set.seed(20260912)
llm_data <- expand.grid(item = factor(seq_len(24)),
  evaluator = factor(paste0("model_", seq_len(4))),
  prompt = factor(paste0("prompt_", seq_len(3))),
  run = factor(seq_len(2)))

llm_design <- gt_design("item", c("evaluator", "prompt", "run"),
  crossed = c("evaluator", "prompt"),
  nested = list(run = c("evaluator", "prompt")),
  item_interactions = "additive", instrument_interactions = "additive")
llm_design$terms_requested
#> [1] "evaluator"            "item"                 "prompt"               "item:evaluator"      
#> [5] "item:prompt"          "evaluator:prompt:run"

This declares six random sources: item, evaluator, prompt, item-by-evaluator, item-by-prompt, and evaluator-by-prompt-by-run. The run group is parent-scoped even though the coded panel is complete. Item-by-run and additional evaluator-by-prompt effects are not silently added. Omitting them is an assumption of this known simulation; a real study needs its own justification. The exact Gaussian engine currently requires a complete coded Cartesian panel with one row per full cell.

source_sd <- c(item = 1.2, evaluator = 0.4, prompt = 0.3,
  "item:evaluator" = 0.5, "item:prompt" = 0.4,
  "evaluator:prompt:run" = 0.25)
llm_data$quality <- 0
for (source in names(source_sd)) {
  members <- strsplit(source, ":", fixed = TRUE)[[1L]]
  group <- do.call(interaction, c(llm_data[members], list(drop = TRUE)))
  effects <- rnorm(nlevels(group), sd = source_sd[[source]])
  llm_data$quality <- llm_data$quality + effects[as.integer(group)]
}
llm_data$quality <- llm_data$quality + rnorm(nrow(llm_data), sd = 0.6)

2. Inspect feasibility before fitting

gt_preflight() reports observed counts, source dimensions, free parameters, configured size limits, and the supported reliability scale, without building a model matrix or running an optimizer.

preflight <- gt_preflight(llm_data, "quality", llm_design)
print(preflight)
#> G-theory preflight: 576 rows | exact_balanced_gaussian 
#> Observed levels: item=24, evaluator=4, prompt=3, run=2 
#> Retained random sources: 6 | Random dimensions: 223 | Model parameters: 8 
#> Structural/resource checks: PASS 
#> Supported reliability scale: observed 
#> Reliability of observed Gaussian-score averages. 
#> Preflight checks structure and configured size limits. It does not establish identification, approximation adequacy, fit acceptance, or scientific validity.
preflight$sources
#>                 source observed_groups predictor_dimensions random_dimension covariance_parameters
#> 1            evaluator               4                    1                4                     1
#> 2                 item              24                    1               24                     1
#> 3               prompt               3                    1                3                     1
#> 4       item:evaluator              96                    1               96                     1
#> 5          item:prompt              72                    1               72                     1
#> 6 evaluator:prompt:run              24                    1               24                     1
stopifnot(preflight$fitting_feasible)

Passing this report means that the listed structural and resource checks pass. It does not prove identification, reliable estimation, numerical acceptance, or approximation adequacy. Read failed checks before changing a resource limit or simplifying a scientific model.

3. Fit and inspect acceptance

fit <- gt_fit(llm_data, "quality", llm_design,
  control = gt_control(gaussian = list(optimizer = "CSOLNP",
    extra_tries = 2L, retry_seed = 415L, threads = 1L)))
print(summary(fit))
#> G-theory fit summary: 1 outcome(s), 576 measurement rows
#> Outcomes: quality 
#> Families: gaussian | Estimator: REML 
#> Random sources: 6 | Instrumentation facets: 3 
#> Optimizer: CSOLNP | Attempts: 1 | Selected attempt: 1 
#> Optimizer completed: TRUE | Numerically accepted: TRUE 
#> Likelihood approximation: exact_balanced_gaussian_likelihood 
#> Boundary or nearly singular sources: none recorded 
#> -2 log likelihood: 1411 
#> 
#> Source variances (Gaussian observed-score covariance):
#>                source   trait variance std_error at_boundary
#>             evaluator quality  0.16915   0.15491       FALSE
#>                  item quality  1.95564   0.60959       FALSE
#>                prompt quality  0.06368   0.07587       FALSE
#>        item:evaluator quality  0.24121   0.05231       FALSE
#>           item:prompt quality  0.10283   0.03178       FALSE
#>  evaluator:prompt:run quality  0.04622   0.02085       FALSE
#>              Residual quality  0.38950   0.02707       FALSE
#> 
#> Fixed location or contrast estimates:
#> quality 
#> -0.4553 
#> 
#> Notes:
#> - Nesting parents define instrumentation groups shared across objects; nested sibling interactions are not automatically included.
#> - No requested source was dropped for this observation family.
#> 
#> Use gt_components(fit) for full covariance matrices and gt_diagnostics(fit) for details.
diagnostics <- gt_diagnostics(fit)
print(diagnostics[c("numerically_accepted", "selected_attempt",
                    "acceptance_failures", "approximation_adequacy")])
#> $numerically_accepted
#> [1] TRUE
#> 
#> $selected_attempt
#> [1] "1"
#> 
#> $acceptance_failures
#> NULL
#> 
#> $approximation_adequacy
#> [1] "exact_balanced_gaussian_likelihood"
stopifnot(isTRUE(diagnostics$numerically_accepted))
gt_components(fit)
#> $evaluator
#>           quality
#> quality 0.1691472
#> 
#> $item
#>         quality
#> quality 1.95564
#> 
#> $prompt
#>            quality
#> quality 0.06368045
#> 
#> $`item:evaluator`
#>           quality
#> quality 0.2412094
#> 
#> $`item:prompt`
#>           quality
#> quality 0.1028346
#> 
#> $`evaluator:prompt:run`
#>            quality
#> quality 0.04621689
#> 
#> $Residual
#>           quality
#> quality 0.3895031
#> 
#> attr(,"scale")
#> [1] "Gaussian observed-score covariance"
#> attr(,"interpretation")
#> [1] "Cross-category contrasts within one categorical outcome are not separate measured outcomes."

The Gaussian default estimator is REML. The trial budget allows the initial fit plus two additional attempts; it is not resampling. The chosen optimizer remains fixed across attempts. Inspect numerical acceptance and source-boundary diagnostics before interpreting the components. A completed optimizer alone is insufficient. If an installation does not support the selected optimizer or a fit is rejected, address that explicitly.

The summary table also carries an asymptotic standard error for each source variance, obtained from the numerically differentiated restricted-likelihood Hessian. A component resting on a variance boundary is flagged, because Wald theory does not apply there.

summary(fit)$variances
#>                 source   trait   variance  std_error at_boundary
#> 1            evaluator quality 0.16914721 0.15491014       FALSE
#> 2                 item quality 1.95564029 0.60958554       FALSE
#> 3               prompt quality 0.06368045 0.07587063       FALSE
#> 4       item:evaluator quality 0.24120936 0.05231458       FALSE
#> 5          item:prompt quality 0.10283459 0.03177643       FALSE
#> 6 evaluator:prompt:run quality 0.04621689 0.02084544       FALSE
#> 7             Residual quality 0.38950314 0.02707234       FALSE

4. Compute relative and absolute reliability

reliability <- gt_reliability(fit)
reliability
#> G-theory reliability | observed scale
#> Facet counts: evaluator=4, prompt=3, run=2 
#>  outcome  Erho2 Erho2_se Erho2_lower Erho2_upper    Phi  Phi_se Phi_lower Phi_upper
#>  quality 0.9464  0.01778      0.8988      0.9723 0.9173 0.03181    0.8299    0.9619
#> Erho2: relative comparisons; Phi: absolute decisions.
#> 95% intervals: delta method on the logit scale from the fitted parameter covariance.
#> Full covariance matrices and fit diagnostics remain in the returned object.
reliability$interpretation
#> [1] "Gaussian observed-score coefficient"

Erho2 is the relative coefficient G: common evaluator, prompt, and run shifts do not change relative item comparisons. Phi is the absolute coefficient and also includes those shifts as error. Item-dependent interactions contribute to both error definitions. Here both coefficients refer to observed Gaussian-score averages under the declared design. They measure consistency under that model, not accuracy against human reference labels. A coefficient of 0.80 means that 80% of the variance in item means is universe-score variance under this model; it is not 80% labeling accuracy.

The interval is a delta-method Wald interval on the logit scale, propagating the fitted parameter covariance matrix through the source covariances. It describes estimation uncertainty in those covariances under the declared model. It is not a guarantee about a future sample of evaluators or prompts.

5. Compare measurement budgets

grid <- expand.grid(evaluator = c(2L, 4L, 6L),
  prompt = c(2L, 3L, 4L), run = c(1L, 2L))
dstudy <- gt_dstudy(fit, grid)
allocation <- dstudy$allocations
allocation$measurements_per_item <- dstudy$measurements_per_object
allocation$extrapolated <- dstudy$extrapolated
planning <- cbind(allocation,
  dstudy$results[, c("outcome", "Erho2", "Erho2_lower", "Erho2_upper",
                     "Phi", "Phi_lower", "Phi_upper")])
planning <- planning[order(planning$measurements_per_item, -planning$Phi), ]
rownames(planning) <- NULL
print(planning)
#>    evaluator prompt run measurements_per_item extrapolated outcome     Erho2 Erho2_lower
#> 1          2      2   1                     4        FALSE quality 0.8789244   0.7902498
#> 2          2      3   1                     6        FALSE quality 0.8989630   0.8204331
#> 3          4      2   1                     8        FALSE quality 0.9241948   0.8622785
#> 4          2      4   1                     8         TRUE quality 0.9093288   0.8361284
#> 5          2      2   2                     8        FALSE quality 0.8985872   0.8185730
#> 6          6      2   1                    12         TRUE quality 0.9403393   0.8886455
#> 7          4      3   1                    12        FALSE quality 0.9390021   0.8873654
#> 8          2      3   2                    12        FALSE quality 0.9125791   0.8403014
#> 9          4      4   1                    16         TRUE quality 0.9465851   0.9002366
#> 10         4      2   2                    16        FALSE quality 0.9349508   0.8786388
#> 11         2      4   2                    16         TRUE quality 0.9197397   0.8513494
#> 12         6      3   1                    18         TRUE quality 0.9531530   0.9117091
#> 13         6      4   1                    24         TRUE quality 0.9596917   0.9235008
#> 14         6      2   2                    24         TRUE quality 0.9477350   0.8999582
#> 15         4      3   2                    24        FALSE quality 0.9463767   0.8988071
#> 16         4      4   2                    32         TRUE quality 0.9521950   0.9089930
#> 17         6      3   2                    36         TRUE quality 0.9582059   0.9196465
#> 18         6      4   2                    48         TRUE quality 0.9635285   0.9295932
#>    Erho2_upper       Phi Phi_lower Phi_upper
#> 1    0.9332760 0.8311242 0.6968108 0.9133367
#> 2    0.9454336 0.8543855 0.7259000 0.9285695
#> 3    0.9595798 0.8905661 0.7895704 0.9463806
#> 4    0.9517192 0.8665114 0.7404140 0.9366003
#> 5    0.9456557 0.8508181 0.7181184 0.9273660
#> 6    0.9688759 0.9123156 0.8237973 0.9586001
#> 7    0.9678246 0.9095813 0.8205098 0.9567797
#> 8    0.9539379 0.8681573 0.7404664 0.9382622
#> 9    0.9720689 0.9193968 0.8359359 0.9623144
#> 10   0.9661408 0.9017489 0.8028545 0.9538842
#> 11   0.9582099 0.8770947 0.7512927 0.9440057
#> 12   0.9756624 0.9295996 0.8559650 0.9670397
#> 13   0.9791477 0.9384895 0.8721195 0.9715377
#> 14   0.9733703 0.9201084 0.8330411 0.9637470
#> 15   0.9722742 0.9173273 0.8298542 0.9618948
#> 16   0.9754426 0.9253201 0.8430236 0.9662015
#> 17   0.9786904 0.9349788 0.8626345 0.9705244
#> 18   0.9814341 0.9425957 0.8772614 0.9741760

Runs are counted within each evaluator/prompt pair, so the budget per item is evaluators x prompts x runs. A row using more levels than observed extrapolates the fitted variance model to the declared exchangeable population.

planning[planning$Phi >= 0.80, ]
#>    evaluator prompt run measurements_per_item extrapolated outcome     Erho2 Erho2_lower
#> 1          2      2   1                     4        FALSE quality 0.8789244   0.7902498
#> 2          2      3   1                     6        FALSE quality 0.8989630   0.8204331
#> 3          4      2   1                     8        FALSE quality 0.9241948   0.8622785
#> 4          2      4   1                     8         TRUE quality 0.9093288   0.8361284
#> 5          2      2   2                     8        FALSE quality 0.8985872   0.8185730
#> 6          6      2   1                    12         TRUE quality 0.9403393   0.8886455
#> 7          4      3   1                    12        FALSE quality 0.9390021   0.8873654
#> 8          2      3   2                    12        FALSE quality 0.9125791   0.8403014
#> 9          4      4   1                    16         TRUE quality 0.9465851   0.9002366
#> 10         4      2   2                    16        FALSE quality 0.9349508   0.8786388
#> 11         2      4   2                    16         TRUE quality 0.9197397   0.8513494
#> 12         6      3   1                    18         TRUE quality 0.9531530   0.9117091
#> 13         6      4   1                    24         TRUE quality 0.9596917   0.9235008
#> 14         6      2   2                    24         TRUE quality 0.9477350   0.8999582
#> 15         4      3   2                    24        FALSE quality 0.9463767   0.8988071
#> 16         4      4   2                    32         TRUE quality 0.9521950   0.9089930
#> 17         6      3   2                    36         TRUE quality 0.9582059   0.9196465
#> 18         6      4   2                    48         TRUE quality 0.9635285   0.9295932
#>    Erho2_upper       Phi Phi_lower Phi_upper
#> 1    0.9332760 0.8311242 0.6968108 0.9133367
#> 2    0.9454336 0.8543855 0.7259000 0.9285695
#> 3    0.9595798 0.8905661 0.7895704 0.9463806
#> 4    0.9517192 0.8665114 0.7404140 0.9366003
#> 5    0.9456557 0.8508181 0.7181184 0.9273660
#> 6    0.9688759 0.9123156 0.8237973 0.9586001
#> 7    0.9678246 0.9095813 0.8205098 0.9567797
#> 8    0.9539379 0.8681573 0.7404664 0.9382622
#> 9    0.9720689 0.9193968 0.8359359 0.9623144
#> 10   0.9661408 0.9017489 0.8028545 0.9538842
#> 11   0.9582099 0.8770947 0.7512927 0.9440057
#> 12   0.9756624 0.9295996 0.8559650 0.9670397
#> 13   0.9791477 0.9384895 0.8721195 0.9715377
#> 14   0.9733703 0.9201084 0.8330411 0.9637470
#> 15   0.9722742 0.9173273 0.8298542 0.9618948
#> 16   0.9754426 0.9253201 0.8430236 0.9662015
#> 17   0.9786904 0.9349788 0.8626345 0.9705244
#> 18   0.9814341 0.9425957 0.8772614 0.9741760

A threshold of 0.80 is illustrative; choose and justify your own criterion. Read the interval columns alongside the point projection: several allocations here have overlapping intervals, so the ordering of nearby rows is not established by these data.

6. Fixed facets

A facet is random when the study generalizes to a population of its levels, and fixed when the universe of generalization is exactly the levels used. LLM evaluation may target exactly a chosen temperature or prompt set, in which case those facets are fixed. Selecting levels deliberately does not by itself decide the universe of generalization.

Declare them with fixed. Following the mixed model of Brennan (2001), the object-by-fixed-facet variance is averaged over that facet’s levels and added to universe-score variance, and a source built only from fixed facets shifts every item equally and leaves the model.

fixed_prompt <- gt_reliability(fit, fixed = "prompt")
fixed_prompt
#> G-theory reliability | observed scale
#> Facet counts: evaluator=4, prompt=3, run=2 
#> Fixed facets: prompt | mixed (Brennan 2001): object-by-fixed-facet variance enters the universe score 
#>  outcome Erho2 Erho2_se Erho2_lower Erho2_upper    Phi  Phi_se Phi_lower Phi_upper
#>  quality 0.963  0.01261      0.9286      0.9811 0.9428 0.02464    0.8707    0.9758
#> Erho2: relative comparisons; Phi: absolute decisions.
#> 95% intervals: delta method on the logit scale from the fitted parameter covariance.
#> Full covariance matrices and fit diagnostics remain in the returned object.
fixed_prompt$source_roles
#>                                evaluator                                     item 
#>                         "absolute_error"                               "universe" 
#>                                   prompt                           item:evaluator 
#> "dropped_fixed_instrumentation_constant"            "relative_and_absolute_error" 
#>                              item:prompt                     evaluator:prompt:run 
#>   "universe_after_fixed_facet_averaging"                         "absolute_error" 
#>                                 Residual 
#>            "relative_and_absolute_error"

Treating the prompt set as fixed raises both coefficients, because item-by-prompt variation is now part of what the study is trying to measure rather than error. A fixed facet’s count cannot be changed, and a decision study may not project over it.

fixed acts at the reliability stage: it changes how the already-fitted components are aggregated into a coefficient. It does not change the G study, so it cannot repair a source whose variance was misspecified during fitting. Observed seed disagreement can change across temperatures even when latent variance is constant, because category probabilities can change. If diagnostics raise doubts about pooling temperatures, fitting within one temperature is a useful sensitivity analysis for a conditional estimand; declaring a facet fixed is a separate decision about generalization.

7. Match other outcomes to their observation model

Outcome Current scalar reliability support
Gaussian Observed continuous-score averages.
Binary or ordinal Explicit scale = "latent" after an accepted fit. This describes latent-response averages, not majority votes, label proportions, or observed ordinal-score averages.
Unordered categorical No implemented default scalar G/Phi. A scientific score or category-probability estimand is needed.

Declare category order and references explicitly. Do not recode binary, ordinal, or nominal labels as the continuous quality score used above. Discrete fitting uses a bounded dense Laplace engine with Gaussian latent random effects; gt_preflight() reports its current limits, and changing the link does not remove the random-effects distribution assumption. That engine reports point estimates only: it computes no standard errors. Analytic reliability and D studies require complete balanced coded panels even when discrete fitting admits missing whole cells.

At one observation per object-by-facets cell, a discrete fit cannot identify the full-cell source that the default design requests. Declare the same design without it:

gt_design("item", "rater", full_cell = FALSE)$terms_requested
#> [1] "item"  "rater"