--- title: "From marginal study summaries to synthetic patients" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{From marginal study summaries to synthetic patients} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- `summary2joint` estimates a joint latent Gaussian distribution from marginal summaries in repeated independent studies of a common population. It does not identify arbitrary dependence from one set of margins and does not recover the original patient records. ## Input Use a named list specifying continuous, binary, and ordinal variables. For each study, supply its size and a named list of means/sample SDs, binary event counts, and ordered category counts. All summaries in a study refer to the same people. Unreported variables can be omitted, but every pair needs repeated joint reporting. ```{r} library(summary2joint) set.seed(31) variables <- list(age = list(type = "continuous"), response = list(type = "binary"), severity = list(type = "ordinal", levels = 3L)) studies <- lapply(seq_len(80), function(i) { z <- matrix(rnorm(300), 100, 3) z[, 2] <- 0.4 * z[, 1] + sqrt(0.84) * z[, 2] age <- 55 + 8 * z[, 1] response <- as.integer(z[, 2] > 0) severity <- findInterval(z[, 3], c(-Inf, -0.5, 0.5, Inf)) list(n = 100L, summaries = list( age = list(mean = mean(age), sd = sd(age)), response = list(events = sum(response)), severity = list(counts = tabulate(severity, 3)))) }) fit <- fit_summary_copula(studies, variables) fit fit$diagnostics[c("converged", "boundary", "inference_ok")] ``` ## Probability and generation ```{r} joint_probability(fit, lower = c(age = 60, response = 1, severity = 2)) head(simulate_summary_copula(fit, n = 100, seed = 42)) confint(fit) ``` The correlation matrix is on the latent normal scale, not generally the Pearson correlation of the observed discrete variables. Probability intervals propagate uncertainty in both the margins and dependence. Fits are ordinary serializable R objects; save them using `saveRDS()`. ## Independent groups and limitations If scientific knowledge supports independent groups, supply a partition to `fit_summary_copula_groups()`. Each group must contain at least two variables. Use `joint_probability_groups()` for the fitted grouped distribution; `simulate_summary_copula()` works for both classes. Grouping is a modeling assumption, not an automatic selection procedure. Always inspect convergence, boundaries, and inference diagnostics. A full-rank sandwich requires more studies than parameters, but that condition alone does not ensure accurate intervals. Rare events and few studies can cause undercoverage. Between-study heterogeneity can confound within-patient dependence. This package assumes a common population, aligned definitions, independent participants across studies, and reporting independent of measurements. It does not model missing patient data or rounded counts. Details and numerical controls are in `?fit_summary_copula`.